Methods and systems for fixing an error in source code
The DEEPCODE AI FIX system addresses the challenges of existing APR systems by using a language model to efficiently fix source code errors through code reduction and merging, achieving rapid and accurate fixes for complex bugs.
Patent Information
- Application Number
- PCT/IB2024/061866
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-31
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Existing automated program repair (APR) systems face challenges such as an enormous search space, continuous recompilation of candidate patches, and overfitting, making them impractical for fixing complex semantic bugs like security vulnerabilities.
The DEEPCODE AI FIX system uses a language model to fix errors in source code by modifying the original code to generate a reduced version with the error, processing this reduced code with a language model to eliminate the error, and then merging the corrected reduced code back with the original code to produce a fixed version.
This approach allows for instantaneous generation of fixes across different programs, scaling to larger edits, and significantly outperforming traditional APR systems in terms of practical applicability and accuracy.
Smart Images

Figure IB2024061866_05062025_PF_FP_ABST
Abstract
Description
6265.1015001 METHODS AND SYSTEMS FOR FIXING AN ERROR IN SOURCE CODE RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 603,945,filed on November 29, 2023, and U.S. Provisional Application No.63 / 714,448, filed on October 31, 2024, the entire teachings of which are incorporated herein by reference. BACKGROUND
[0002] Automated program repair (APR) may include automatic repair of software bugswithout intervention of a human programmer. SUMMARY
[0003] APR has been studied from both traditional software engineering and machinelearning perspectives. Traditional approaches often require test cases, crash reports, or other specifications. Given a program violating expected behavior, conventional APR systems may output a correct program fulfilling the expected behavior, for instance by performing a smart search over a space of program modifications. However, these approaches are impractical due to an enormous search space, continuous recompilation of candidate patches, and / or re-execution of tests. Some existing APR systems try to synthesize an output by solving symbolic constraints or applying learned edit patterns on abstract syntax trees (ASTs). Both traditional software engineering and machine learning approaches suffer from not generalizing, as behavior is often specific to individual programs, and / or suffer from overfitting.
[0004] Embodiments significantly differ from prior APR approaches both technically and interms of usability. Embodiments, which may be referred to herein as “DEEPCODE AI FIX,” can generalize across different programs by learning from a large dataset of diverse bugs and fixes, scale to larger edits, and generate fixes practically instantly in comparison to traditional APR systems.
[0005] Recent years have seen an increase in learning-based tools. A major hurdle for theselearning-based tools is a lack of training data. Some existing systems attempt to overcome this problem by generating synthetic fixes for specific types of bugs, such as incorrect operators or - 1 - 3924732.v36265.1015001 variable misuses. Applying various models such as long short-term memory (LSTM) networks and Transformers may increase accuracy on synthetic datasets, but conventional systems still produce mostly false positives on real-world bug distributions. As a result, traditional systems have limited practical applicability and suffer from distribution shift by learning from synthetic data. Other existing systems apply large language models (LLMs) indiscriminately to transform programs with errors into programs without errors. Still other conventional systems replace fine- tuning a model with selecting several training data samples similar to a query and using the selected training data samples as a prompt in a few-shot prediction of a larger pretrained model for code. These traditional systems are typically limited to relatively small programs that fit into a context size of a utilized model. Embodiments address the foregoing and other problems in existing methods and systems.
[0006] An example embodiment is directed to a computer-implemented method for fixing anerror in source code. The method includes modifying original source code having an error to generate reduced source code including the error. The method further includes processing the reduced source code including the error with a language model to produce error-eliminated reduced source code. The method further includes merging the error-eliminated reduced source code with the original source code to generate fixed source code.
[0007] In an example embodiment, the error may be a security vulnerability, a semanticerror, an application programming interface (API) misuse, a quality issue, or a style issue.
[0008] According to an example embodiment, the original source code may include multipleerrors. The method may further include iterating the modifying, processing, and merging for each error of the multiple errors.
[0009] In another example embodiment, the modifying may include generating a graphrepresentation of the original source code having the error. The modifying may further include iteratively (i) reducing the graph and (ii) performing a static analysis of the reduced graph to determine existence of the error in the reduced graph until a minimal reduced graph including the error is determined. In such an embodiment, the minimal reduced graph including the error may represent the generated reduced source code including the error. According to an example embodiment, in a given iteration, reducing the graph may be based on approximate provenance information provided by a static analysis report, e.g., an error or crash report. - 2 - 3924732.v36265.1015001
[0010] According to an example embodiment, the language model may be an artificialintelligence based neural network.
[0011] In another example embodiment, the method may further include training thelanguage model. According to an example embodiment, training the language model may include obtaining multiple code samples each including a respective error. Training the language model may further include obtaining multiple error-free code samples where each error-free code sample may correspond to a given code sample of the multiple code samples. Training the language model may further include reducing (i) the multiple code samples and (ii) the multiple error-free code samples and, in turn, training the language model to determine code fixes using the reduced multiple code samples and the reduced multiple error-free code samples.
[0012] According to an example embodiment, the method may further include querying thelanguage model in a zero-shot or few-shot learning manner. The querying may include obtaining multiple code samples each including a respective error. The querying may further include obtaining multiple error-free code samples where each error-free code sample may correspond to a given code sample of the multiple code samples. The querying may further include reducing (i) the multiple code samples and (ii) the multiple error-free code samples. Further still, the querying may include querying the language model with few-shot learning where the reduced multiple code samples and the reduced multiple error-free code samples may be provided as correct examples to the language model.
[0013] In another example embodiment, the merging may include comparing the reducedsource code including the error and the error-eliminated reduced source code to determine a mapping between (i) the reduced source code including the error and (ii) the error-eliminated reduced source code. The merging may further include, based on the mapping, merging the error-eliminated reduced source code with the original source code to generate the fixed source code. According to an example embodiment, in the mapping, each line in the error-eliminated reduced source code is mapped to a given line in the reduced source code including the error. In an embodiment, the merging may also include determining a mapping, e.g., a line mapping, between (i) the original source code having the error and (ii) the reduced source code including the error. This line mapping may, in turn, be utilized to perform the merging. - 3 - 3924732.v36265.1015001
[0014] Another example embodiment is directed to a computer-based system for fixing anerror in source code. The system includes a processor and a memory with computer code instructions stored or held thereon. The processor and the memory, with the computer code instructions, are configured to cause the system to implement any embodiments, or combination of embodiments, described herein.
[0015] Yet another example embodiment is directed to a non-transitory computer programproduct for fixing an error in source code. The computer program product includes a computer- readable medium with computer code instructions stored thereon. The computer code instructions are configured, when executed by a processor, to cause an apparatus associated with the processor to implement any embodiments, or combination of embodiments, described herein. As understood by one skilled in the art, one or more processors may execute the computer code instructions to cause the apparatus to implement an embodiment.
[0016] It is noted that embodiments of the method, system, and computer program productmay be configured to implement any embodiments, or combination of embodiments, described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The foregoing will be apparent from the following more particular description ofexample embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.
[0018] FIG. 1 is a simplified block diagram of example pipelines for automatic code fixingaccording to an embodiment.
[0019] FIG. 2 illustrates running an example data creation pipeline for a vulnerability fixaccording to an embodiment.
[0020] FIGs. 3A-3D are tables of example issues reported by code analysis according to anembodiment.
[0021] FIGs. 4A-4G illustrate example operations according to an embodiment.
[0022] FIG. 5 illustrates an example code reduction according to an embodiment.- 4 - 3924732.v36265.1015001
[0023] FIG. 6 illustrates example reduced code and an example predicted code fix accordingto an embodiment.
[0024] FIG. 7 is a simplified block diagram of an example overview of a data pipelineaccording to an embodiment.
[0025] FIG. 8 is a table of example evaluations of metrics for models against baselinesaccording to an embodiment.
[0026] FIGs. 9-16 illustrate example code fixes according to embodiments.
[0027] FIG. 17 is a flow diagram of a computer-implemented method for fixing an error insource code according to an embodiment.
[0028] FIG. 18 is a schematic view of an example computer network in which embodimentsmay be implemented.
[0029] FIG. 19 is a block diagram illustrating an example embodiment of a computer node inthe computer network of FIG.18. DETAILED DESCRIPTION
[0030] A description of example embodiments follows.
[0031] Introduction
[0032] The field of APR has attracted substantial interest over the years, but despitesignificant research efforts, creating a system that works well for complex semantic bugs such as security vulnerabilities has proven difficult. A promising direction to address this challenge is by leveraging language models, e.g., LLMs. Embodiments provide such functionality and are directed to methods and systems for fixing errors in source code using language models.
[0033] In this document, the effectiveness of language models for solving such code-repairtasks is investigated. It is shown that such tasks are difficult, as they may require a model to learn long-distance code relationships (i.e., relationships between code tokens that are separated by other code tokens), which in turn may employ extensive amounts of training data. At the same time, creating large, clean training or evaluation datasets for complex program bugs and their corresponding fixes may be costly and non-trivial. Embodiments address these challenges with a novel approach for querying (i.e., inference) and training (e.g., fine-tuning) language models. By leveraging program analysis to limit a language model’s attention mechanism on portions of - 5 - 3924732.v36265.1015001 code needed to perform a fix, embodiments can drastically reduce an amount of required training data and an amount of input data (i.e., number of code tokens) needed to query a model. Using a code reduction approach, for both training and inference, rather than feeding an entire program to a language model, embodiments can reduce the program’s code to a much shorter snippet that contains a reported defect together with necessary context—and use that instead.
[0034] Example evaluations show that the aforementioned code reduction approach ofembodiments can substantially improve both readily available models such as GPT-4 using few- shot learning, as well as fine-tuning models. To train and evaluate an exemplary system of embodiments, a new and comprehensive code fixing dataset was created by extensively labeling 156 exemplary non-trivial bug patterns (including 40 exemplary security rules). These exemplary patterns may be non-trivial and may require a static analyzer to perform complex interprocedural dataflow analysis to discover the patterns. An exemplary system of embodiments that employs Mixtral-8x7B can remove more than 80% of reported defects while exactly matching a human fix in between 10% and 50% of cases, outperforming baselines based on GPT-3.5 and GPT-4, or on window-based models such as TFix.
[0035] The rapid increase in the number of software systems has led to an increased demandfor methods and tools that can detect potential defects in these systems. Indeed, existing testing and bug-finding tools, ranging from test generation to dynamic and static analysis, often return thousands of high-quality findings. However, due to the sheer volume of discovered defects, most reports are not addressed by developers. Despite various efforts aiming to improve the development process and improve the situation, the most promising long-term path to addressing this issue in a systematic manner remains automatic bug fixing.
[0036] Prior to using LLMs, automated fixing tools mostly focused on a small set of bugs,trivial few-line changes, or formatting issues. As LLMs were created, it was shown that language models can also be useful in fixing lint-like issues. With even larger LLMs, a number of software security providers announced features that mostly use OpenAI® generative pretrained transformer (GPT) models in an attempt to automatically fix more complex security issues discovered by static analyzers. Other security providers announced automatic fixing tools based on custom trained models. However, outside of marketing materials that display user interfaces (UIs) of these tools, there is little understanding of the underlying accuracy of the fixes these - 6 - 3924732.v36265.1015001 tools provide. To address this issue, in an embodiment, the exemplary dataset of vulnerabilities and their fixes (described herein) was designed, by which capabilities of existing models as well as of any future proposed models can be evaluated.
[0037] Exemplary Dataset of Security and Semantic Code Fixes
[0038] To address a lack of evaluation data, an embodiment generates the exemplary dataset(described herein). To create the dataset, an embodiment performed data collection and labeled thousands of commits collected from open-source repositories. Using 380,000 exemplary candidate fix commits and tens of labelers with coding and security backgrounds, an exemplary dataset of over 5,000 labeled fix examples was assembled. This data was then split according to licenses of the code—permissively licensed data (i.e., code that is permissively licensed by the code’s developer) can be used for training or fine-tuning, and other data is used only for evaluation. Using this method can protect against train / test leakage in evaluations, because some models (e.g., StarCoder) are trained using only permissively licensed data. For example, if such a model is also evaluated using permissively licensed data, the model may perform well simply because it has encountered the same examples during training, potentially creating a skewed impression of the model’s accuracy. In contrast, if the model is only evaluated with restrictively licensed data, this helps ensure that the model is encountering the examples for the first time, thereby providing a more accurate picture of the model’s performance. According to an embodiment, the dataset is based on exemplary code scanning results of the Snyk® Code Static Application Security Testing (SAST) engine by Applicant-Assignee Snyk Limited (London, UK). The Snyk SAST engine operates on source code and is one of the fastest engines with an extensive set of security rules. Each exemplary fix in the dataset includes code where an alarm is raised by the Snyk SAST engine and corresponding code where the alarm is not raised, i.e., code with a fix. In addition, human labelers agreed that the exemplary changes correctly fix the alarms. Thus, the dataset is human-validated to confirm that the fixes were not due to false positives or false negatives of the engine.
[0039] Challenge of Difficult to Obtain Training Data
[0040] One challenge when using machine learning for program repair may be that it isdifficult to obtain a sufficiently large and clean training dataset consisting of pairs of buggy and fixed code versions for a given program issue. While histories of many open-source projects are - 7 - 3924732.v36265.1015001 available on, e.g., GitHub®, obtaining a dataset of fixes has so far only been done successfully for simple few-line local static analyses. A reason underlying this difficulty may be that, with deeper semantic issues, program changes can be observed where a defect report is present in one version of the program and absent in another, yet a core program error is not fixed. For example, a change may turn a problematic branch into dead code, some unrelated change may accidentally make a security vulnerability in code exploitable, a code change may modify many more places than a bug, or the change may affect an approximation of a static analyzer without affecting an actual presence of a reported issue. The frequency of these problems is discussed hereinbelow.
[0041] Challenge of Need to Learn Long-distance Relationships
[0042] In addition, the nature of fixing semantic program errors may be such that correctfixes must understand long-distance relationships in to-be-fixed code, yet learning such complex attention mechanisms may require large amounts of training data, which, as discussed hereinabove, may be practically infeasible to obtain. Even if such data is somehow obtained, limitation(s) in context and / or output size of a LLM may impede the LLM’s direct utilization in fine-tuning and / or few-shot learning settings. While recent models have been developed to accommodate extensive context sizes, it is emphasized that supporting such dimensions in terms of computational resources and latency may not equate to effectively learning from long-range dependencies. If not constrained by input size, these models may still encounter challenges in handling tasks with prolonged contextual requirements and may also exhibit constraints in a number of output tokens beyond a certain threshold. Overall, this may mean that a direct application of LLMs (even state-of-the-art) on the task of fixing code errors, as illustrated by pipeline 100a of FIG.1 (described hereinbelow), is unlikely to work well.
[0043] FIG. 1 is a simplified block diagram of example pipelines 100a and 100b forautomatic code fixing according to an embodiment. As shown in FIG.1, the pipeline 100a only includes LLM 102a, which is trained with unmodified data 104, e.g., git commits or other suitable known training data. The LLM 102a takes as input code 106a with a static analysis report and generates predictions in the form of fixed code 108a.
[0044] It is remarked that experiments with directly applying extensively pretrained modelssuch as GPT-4 in the pipeline 100a led to worse results on complex bugs in comparison to applying the same technique on simple few-line bugs. - 8 - 3924732.v36265.1015001
[0045] Exemplary Code-Reduced Dataset and LLM-based Code Fixing
[0046] In an embodiment, the above-described challenges may be addressed in twoexemplary steps. First, extensive labeling of open source commits may be performed and an exemplary clean dataset of issues and fixes for programs, e.g., JavaScript® programs, may be constructed. Second, learning and prediction capabilities of code models may be improved by leveraging static code analysis, thereby showing that program analysis features can play a significant role in the quality of a learned model, having the same quantitative effect as adding (an order of magnitude) more labeled data. In another example embodiment, input code may be reduced such that a static analysis report in an original program is still reported by the same static analyzer in the reduced code snippet. The reduced snippet may contain the core of a problem, meaning that the reduced snippet focuses training of an attention mechanism to very few lines of code, in turn reducing a sample length of training data—which may be difficult to obtain. The code reduction technique of embodiments may serve as a way of context extraction that allows modeling long-range dependencies and makes a repair system independent of an original file size. Using the code reduction technique of embodiments, a new cleaned dataset may be obtained, where each sample only contains code related to its particular issue.
[0047] Referring again to FIG. 1, the pipeline 100b may reflect a complete process for anexample embodiment that combines CODEREDUCE component 110 with LLM 102b. The LLM 102b may be trained with reduced data 112, e.g., modified commits, which may be generated by applying the CODEREDUCE component 110 to the unmodified training data 104 as shown in FIG. 1. Similarly, the CODEREDUCE component 110 may process input code 106b with a static analysis report to generate reduced code 114. The LLM 102b may then fix the reduced code 114 and produce output 116. In turn, MERGEBACKcomponent 120 may merge the output 116 with the original file 106b, resulting in fixed code 108b.
[0048] In an example embodiment, the CODEREDUCE component 110 may still allow forusing models pretrained on text and code while making it unnecessary to learn attention over entire files or large contexts. Because samples may be much shorter, a LLM can effectively learn to fix bugs in programs either with finetuning or few-shot learning. Once a new program needs a fix for a reported issue, its code may be reduced, and a LLM may then fix this reduced code. A LLM’s output may then be merged back into an original file by using a patch procedure. - 9 - 3924732.v36265.1015001 Interestingly, because the code reduction technique of embodiments dramatically decreases an input length fed into a model, it can enable usage of models of standard attention length that can now attend to an entire reduced code snippet with a vanilla attention mechanism.
[0049] It is noted that an automatic bug-fixing tool, such as those described herein, can beimmediately useful to developers inside their integrated development environments (IDEs), if predictions can be done in real-time. According to an example embodiment, the CODEREDUCEcomponent 110 can make possible real-time serving of predictions, because numbers of tokens to be processed and generated by a model may be significantly smaller than standard approaches. Moreover, to further improve real-time latency, embodiments achieve a speed-up over Hierarchical Delta Debugging (HDD) based on provenance information from a static analysis tool. This extension can reduce calls to an underlying static analyzer by, e.g., 30%, leading to reduced latency.
[0050] Exemplary Advantagesa) A static analysis-enabled novel code reduction technique that produces a smallrepresentation of a program, one that easily fits into an attention window of a model, yet contains the necessary information for learning a correct fix. b) An extensive evaluation of several models both with fine-tuning (e.g., MIXTRAL,STARCODERBASE, T5, etc.) and few-shot learning (e.g., GPT-3.5, GPT-4, etc.) on different exemplary datasets including, for instance, reduced code, small window, long context, and / or entire file, etc. c) Methods for reducing code while preserving a static analysis alarm and then formerging a fixed reduced code prediction back into an original source file. d) An end-to-end code fixing system that substantially improves both readilyavailable models, like GPT-4, using few-shot learning, as well as fine-tuning models based on extensive evaluations.
[0051] Overview of DEEPCODE AI FIX Embodiments
[0052] Presented hereinbelow is an overview of the DEEPCODE AI FIX system ofembodiments by way of a motivating example that contains a security vulnerability.
[0053] FIG. 2 illustrates running an example data creation pipeline 200 for a vulnerability fixaccording to an embodiment. As shown in FIG.2, a piece 206 of code, e.g., JavaScript code, - 10 - 3924732.v36265.1015001 may include a server (not shown) with a vulnerability 218, e.g., a Path Traversal vulnerability, indicated for instance by a framed line or other suitable marking known in the art.
[0054] In an example embodiment, CODEREDUCE component 210 may take the piece 206 ofcode as input and generate reduced code 214. Similarly, the CODEREDUCE 210 transformation may be applied to code fix 222—which may be a human-provided fix for this security vulnerability 218 in, e.g., a diff format—resulting in reduced code fix 224. MERGEBACKcomponent 220 may be applied to obtain fixed version 208 of the original code 206 from the piece 206 of code and a diff (not shown) of (i) the reduced code 214 and (ii) the reduced code fix 224. Shading of lines and / or ‘-’ and ‘+’ symbols may be used to represent standard diff notation. However, any other suitable known indicia may also be used.
[0055] There are static analysis tools that can discover this vulnerability 218 and report it onthe frame-highlighted line in the original code 206. To fit within the confines of FIG.2, the actual code example may be significantly simplified. The human fix 222 may not only implement removal of the vulnerability 218 via added code sequence 226, but may also include an unrelated change 228 such as doBBB()→doXYZ(), for non-limiting example.
[0056] Because a static analysis tool (not shown) may discover the vulnerability 218 (e.g., apath traversal vulnerability) in the original code 206 and may not discover it in the fixed code 222, this may be a signal that the pair 206→222 qualifies as a potential training data sample for learning. To obtain an even cleaner data sample, in an example embodiment, the DEEPCODEAI FIX system may call the CODEREDUCE component 210, which on its own may call the static analyzer multiple times until it produces the reduced code fix 224. Note that in the process of reducing the code 224, it may be discovered which lines can be used to detect the vulnerability 218. For instance, a file system library may need to be included at line 232a to keep a server handler at line 232b, and a data flow may need to be kept from a request at line 232c to a file system operation at line 232d. Note that there may be multiple trivial but wrong ways to fix the code snippet 214. For example, if any of its statements is deleted from this input, the vulnerability 218 may disappear, however a corresponding program (not shown) may also break its semantics in a significant way. This observation may highlight that relying on simple heuristics to obtain training data is potentially dangerous. It can lead to models that trivially satisfy bug removal metrics, but actually learn incorrect fixes. - 11 - 3924732.v36265.1015001
[0057] To gain a useful fix sample, in an example embodiment, any change between theoriginal code 206 and the code fix 222 may be taken such that a line of this change is also included in the reduced code 214 or adjacent to a line included in the reduced code 214. Then these changes may be applied to the reduced code 214 and a sample may be constructed, e.g., the reduced code fix 224. Next, validation may be performed to confirm that the provided fix parses and removes a static analysis alarm. This pair of code snippets 214 and 224 may be denoted REDUCEDPREand REDUCEDPOST, respectively. A static analysis report type may be denoted RULE / DESCRIPTION—e.g., in FIG.2, a rule may be PT identifying path traversal and a description may contain a short text about the vulnerability 218. Then, a LLM (not shown) can be provided with code fixing examples similar to the snippets 214 and 224 using the below exemplary prompt: fix RULE DESCRIPTION : REDUCEDPREand be fine-tuned to generate REDUCEDPOSTgiven such a prompt. Note that because the snippets 214 and 224 may be short code samples, they may be expected to fit in an attention window of a text transformer and learning can be efficient even with relatively few training samples. In another example embodiment, the same prompt structure can be used for inference as well.
[0058] In an example embodiment, a LLM may have learned to perform fixes on programslike the reduced code 214. If, e.g., the original code 206, is provided as a query, embodiments can reduce the sample 206 to the program 214, a LLM can transform the program 214 into the reduced code fix 224, and then the resulting reduced fix program 224 can be merged back into the input program 206. This step, performed by the MERGEBACK component 220, may result in the fixed version 208 of the original code 206, and may work by applying the changes observed between the code snippets 214 and 224 onto the program 206.
[0059] Other real-world fix examples generated by embodiments are shown in FIGs. 8-15,described hereinbelow.
[0060] It is a known problem that bug-fixing commits often contain patches irrelevant to theapplied fix, such as refactorings, among other examples. Note that the example procedure in FIG. 2 may also be useful for data cleaning. In an example embodiment, the first pair 206→222 may be transformed into the pair 206→208 such that the unrelated change 228 in the first pair like - 12 - 3924732.v36265.1015001 doBBB()→doXYZ() is removed. This side benefit of the CODEREDUCE component 210 can play an important role in obtaining a clean dataset.
[0061] Exemplary Components of DEEPCODE AI FIX Embodiments
[0062] Described hereinbelow are exemplary components of the DEEPCODE AI FIX system ofembodiments.
[0063] Code Analysis
[0064] In an example embodiment, the DEEPCODE AI FIX system may utilize Snyk Code—aproprietary commercial static analyzer that operates on source code—for code analysis. Further, it is noted that embodiments are not limited to utilizing Applicant-Assignee Snyk Limited tools; rather, any static analyzer known to those of skill in the art may be utilized. In this document, exemplary evaluations are performed with a JavaScript / TypeScript analysis, along with its extensions such as ReactJS and VueJS, among other examples. Other known languages and / or extensions are also suitable. An analyzer may implement, e.g., 156 different exemplary checks that are shown in tables 334a-334d of FIGs.3A-3D, respectively. Exemplary checks that may be utilized by embodiments as part of performing static analysis may be classified into five exemplary categories as follows: a) AST: Most AST checks may be ones that can be performed only on an AST.Many such rules enforce properties about the VueJS or React extensions of JavaScript / TypeScript and include checks such as missing tags, duplicate variable names, and patterns, among other examples, that usually do not depend on analyzing control or dataflow of a program. b) LOCAL: LOCAL checks may use interprocedural analysis to discover that certainincorrect values would flow into methods that would not accept them, that resources are not deallocated, or that certain expressions have no effect or always produce results with unexpected semantics, among other examples. c) FILEWIDE: FILEWIDE checks may also use interprocedural analysis, but incontrast to LOCAL checks, properties that they check may stay in different functions or methods in a program. These rules may include checks for mismatches between the signature of a method, its implementation, or its usage, among other examples. - 13 - 3924732.v36265.1015001 d) SECURITYLOCAL: Similar to LOCAL checks, SECURITYLOCAL checks may verifyincorrect API usage from a security perspective as well as other types of security misconfigurations, typically by tracking a set of all method calls on an object or by tracking values passed as parameters to method or function calls. e) SECURITYFLOW: SECURITYFLOW checks may be the most complex checkssupported by a selected static analyzer. This rule category may include all taint analysis rules that involve interprocedural dataflow analysis starting from a data source, as well as detection of various data sanitization patterns, among other examples.
[0065] Approximate Provenance Information
[0066] In addition to reports, static analyzers utilized in embodiments may provide proofs ofthe reports, such as data flow or a flow describing a discovered defect. These reports, however, may target human readability and may not be a full set of program statements that a static analyzer needs to reproduce an alarm. As a result, many existing analyzers may provide approximate provenance information that is either an over- or under-approximation of a precise provenance. Yet, even an approximation may be useful to speed-up the CODEREDUCEprocedure of embodiments. In an embodiment, a procedure may be postulated that returns an approximate set of AST nodes required for a report. A function APPROXPROVENANCENODES may be defined that takes a list of tree nodes and a static analysis report and returns a list of nodes that participate in derivation of the static analysis report.
[0067] CODEREDUCE
[0068] Given a file and a vulnerability detected by a static analyzer, in an exampleembodiment, the DEEPCODEAI FIXsystem may call the CODEREDUCEprocedure to reduce the source code to a small snippet that still reports that static analysis alarm. According to another example embodiment, the CODEREDUCE procedure may employ the cReduce tool and / or delta debugging to execute a set of transformation passes on a tree representation of input code while checking that a detected property remains present in all of its transformation phases. In yet another embodiment, all transformations may be performed on an AST and may leverage HDD.
[0069] FIGs. 4A-4G illustrate example CODEREDUCE operations 400a-400g, respectively,according to an embodiment. As shown in FIG.4A, in an example embodiment, the operation - 14 - 3924732.v36265.1015001 400a may include receiving tree representation 462 of input code including levels 464a and 464b. In another example embodiment, as shown in FIG.4B, the operation 400b may include extracting candidate nodes 466a and 466b to delete at a current level, e.g., the level 464a. Further, in yet another embodiment, the operation 400c may include iterating over the candidate nodes 466a (as shown in FIG.4B) and 466b and trying to remove their subtrees, e.g., by first attempting to remove a subtree of the candidate node 466a as shown in FIG.4C. According to an embodiment, the CODEREDUCE procedure may check if a static analysis can still detect an issue (i.e., the original issue as found in the input code without any reduction) when a subtree, e.g., the subtree of the candidate node 466a, is removed; if yes, the change may be applied, else if no, the change may be discarded. In another embodiment, as shown in FIG.4D, the operation 400d may include discarding an attempted removal of the subtree of the candidate node 466a (e.g., responsive to the removal of the subtree of node 466a as shown in FIG.4C resulting a tree that does not include the previously detected issue). As shown in FIG.4E, in yet another embodiment, the operation 400e may include applying a change to remove a subtree of the candidate node 466b. In an embodiment, as shown in FIG.4F, the operation 400f may include extracting (i.e., deleting) candidate node 466c at a next level, e.g., the level 464b. Further, according to yet another embodiment, the operation 400g may include applying a change to remove a subtree of the candidate node 466c as shown in FIG.4G. In this way, the operations 400a-400g illustrated in FIGs.4A-4G may generate a reduced tree that still includes the detected error.
[0070] In an example embodiment, a node on a higher level may become deletable only aftersome or all of its child node(s) are deleted. According to another example embodiment, the CODEREDUCEprocedure may repeat one or more iterations of a process that includes operations such as the operations 400a-400g until a fix point is reached.
[0071] Below is example pseudocode of code reduction, e.g., CODEREDUCE, functionality“Method 1” according to an embodiment: - 15 - 3924732.v36265.1015001
[0072] According to an example embodiment, HDD may itself be based on a DeltaDebugging procedure DDMIN that may be called for nodes at each level of an AST of a currently reduced code. DDMIN may take as input a set of program elements (e.g., AST nodes), a static analysis report, and code with its tree. The DDMINmay output further reduced code with a reduced tree that still keeps a static analysis alarm present in a program. DDMIN may perform multiple attempts to remove tree nodes and may call an underlying static analysis multiple times to check for the presence of a report, e.g., an error or crash report.
[0073] In an example embodiment, a CODEREDUCE procedure may be provided to accountfor approximate provenance information provided by a static analysis alarm. The above pseudocode, i.e., Method 1, illustrates an exemplary CODEREDUCE procedure according to an embodiment. This example procedure begins at a root node of an AST. At each level, the procedure filters out nodes that should be considered for deletion depending on the level in the tree (line 2). Provenance handling is shown in lines 5-9. The example procedure attempts to - 16 - 3924732.v36265.1015001 remove all nodes from the level that were not part of the approximate provenance, and if the report is not removed (line 7), the removal is accepted. Remaining lines 10-12 continue with a call to DDMINon nodes in a current level and moving to a next level. To also maintain guarantees when provenance is approximate, the example code reduction functionality may still perform the DDMIN procedure regardless of success of the provenance nodes removal step.
[0074] In an example embodiment, to make the CODEREDUCE procedure useful for theautomatic bug fixing case, the function referred to as “GETNODES” in the above pseudocode may be configured to only delete statements that are complete lines and avoid deleting lines only partially, because partially deleted lines may be harder to merge back into original code. Additionally, only nodes that delete statements and declarations but keep expressions inside statements may be considered. This may mean that, for example, statements such as “call(arg1, arg2)” will not be reduced to “call()”, “call(arg1)”, or “call(arg2)”.
[0075] When used as part of a loop until fixpoint, in an embodiment, the above pseudocodeprocedure may have 1-tree-minimality guarantees. This may mean that returned code cannot be further reduced by removing any single node in a tree, such that a static analysis report remains present. The reason for a fixpoint loop may be that a node on a higher level might become deletable only after having another node deleted on a lower level.
[0076] FIG. 5 illustrates an example code reduction according to an embodiment. As shownin FIG.5, in an example embodiment, original code 506 may be transformed into reduced code 514. Usage of variable options may remain in the reduced code 514 at line 3, but its definition from lines 4-6 of the original code 506 may be deleted. The definition of express at line 1 of the original code 506 may not be deleted because it may be an import statement related to a vulnerability, e.g., a path traversal vulnerability. The reduced code 514 may not be executable, but an analyzer can detect a vulnerability in both versions.
[0077] MERGEBACK
[0078] The CODEREDUCE procedure can bring many advantages. In an example embodiment,predictions from the CODEREDUCE procedure, just like a model input, may also be in a reduced form. According to another example embodiment, to provide end-to-end bug fixing, the DEEPCODEAI FIXsystem may merge generated code back into an original file. - 17 - 3924732.v36265.1015001
[0079] An exemplary MERGEBACK procedure is presented in the pseudocode “Method 2”below, according to an embodiment. The merging procedure starts by comparing reduced input code ^^ with predicted fix ^^. This comparison may utilize, e.g., the gitdiff tool, and may aim to compute a one-to-many mapping ^^^^→^^between lines of ^^ and ^^ (line 1). For each line ^^^^in reduced code ^^, ^^^^→^^may contain a potentially empty set of lines from prediction ^^, by which the line ^^^^should be replaced. Technically, if ^^^^is mapped to ^^^^, that may be written as (^^^^, ^^^^) ∈ ^^^^→^^and iterations over this set may return elements in sorted order. Once mappings are computed, original non-reduced file^^and the ^^ may be split into lines, andℒfixedmay be initialized to an empty sequence of strings (line 2).
[0080] To continue, the merging procedure may iterate over lines in non-reduced inputLines^^(line 3) and may check, for each line, whether the line is used in reduced code ^^ or not (line 4). If a line does not participate in reduced code ^^, this may imply that the line is not important for a bug fix and a model may not see this line while generating the prediction. Therefore, such a line may be kept as-is, so it may be appended into ℒfixed(line 7). In the other case, a prediction may decide how a line should be replaced. Therefore, a lookup may be performed in ^^^^→^^to obtain a list of prediction lines meant to replace a current line and each of - 18 - 3924732.v36265.1015001 the obtained prediction lines may be added to ℒfixed(line 5). In the end, the merge procedure may join fixed lines ℒfixedby, e.g., a newline or other suitable known character, and return a result (line 8). Finally, the example merging procedure may compute ^^^^→^^. This may be done by, e.g., doing a gitdiff and obtaining so-called “hunks” of corresponding mapped pieces of code. For each hunk, a corresponding line mapping may be calculated as visualized for example in FIG.6, described below.
[0081] FIG. 6 illustrates example reduced code 614 and example predicted code fix 616, aswell as determined mapping 636 between the reduced code 614 and the predicted fix 616, according to an embodiment. As shown in FIG.6, in an embodiment, example hunk 638a1 of the reduced code 614, which may represent a “dummy” line 1 between actual lines 1 and 2, may be mapped to hunk 638b1 of the predicted fix 616, which may represent an inserted line of code. Likewise, example hunk 638a2 of the reduced code 614, which may represent deleted line 4 and a further “dummy” line 4 between actual lines 4 and 5, may be mapped to hunk 638b2 of the predicted fix 616, which may represent two inserted lines of code.
[0082] Exemplary Data Collection
[0083] Investigated hereinbelow is the problem of collecting commit data and cleaning it upfor building a repair tool. An exemplary dataset was built from a large-scale crawl of GitHub that includes the top 500,000 exemplary repositories sorted by number of stars. From these repositories, all commits were selected that: (i) do not merge branches, (ii) do not rename files, (iii) affect JavaScript code, and (iv) both versions, i.e., pre-commit and post-commit, of the code parse correctly and have appropriate code licenses. This resulted in a dataset of 6 million exemplary commits. A static analyzer was applied on all commits to obtain around 380,000 (pre- commit, post-commit) exemplary file pairs that fixed a static analysis report—meaning that a certain report was present in a pre-version of a pair and not reported any more at a corresponding location in its post-version. Example statistics about fixes are shown per report type in Column 2 of Table 1 below.
[0084] To check data quality, 50 exemplary file pairs were randomly sampled, removingreports from each of five exemplary categories (Table 1, Column 1) and evaluated to determine if a change contains a proper fix for an issue at hand. Exemplary findings are given in Columns 3-7 of Table 1 below. Overall, Table 1 shows that a significant amount of file pairs have fixed an - 19 - 3924732.v36265.1015001 issue in a way that a model cannot and should not utilize for learning. The most frequent problem is that instead of a good fix (Table 1, Column 3), code is deleted or refactored (Table 1, Column 4).Table 1: Example fix statistics on randomly sampled commits that remove static analysis reports.
[0085] A more interesting observation, though, was that as complexity of a static analysisreport increases, a probability of observing a correct fix drops significantly. Security issues may often be left unfixed in code but may be frequently removed or moved to another location as part of a refactoring. Some cases of accidental fixes (Table 1, Column 5) were also observed, where affected code was turned into dead code, or an unrelated change removed a particular report. Besides, a static analyzer may also not be bullet-proof precision-wise. A few instances resulted in approximations that caused an analyzer to report false positives or under-approximations that caused an analysis to drop a report although no proper fix was applied (Table 1, Column 6). For some issues, incorrect fixes could be observed, but a sample could not be attributed to one of these problems (Table 1, Column 7). These results also show that as opposed to heuristic data cleaning applied in prior works such as TFix that worked on simple issues and Lint warnings, more complex static analysis reports may also need more labeling effort to account for complexities in input data. Based on this finding, in an example embodiment, a data collection pipeline, e.g., pipeline 700 of FIG.7 (described hereinbelow), may be designed accordingly and candidate file pairs may be labeled in a web application.
[0086] Exemplary Data Collection and Cleaning
[0087] In an embodiment, data may be created according to example pipeline 700 shown inFIG.7. In the pipeline 700, for each pair 742a-742n of pre- / post-commit source files (^^,^^′), static analysis 730 may be run and candidate file pairs 744a-744n may be collected where a certain issue ^^ present on a line ℓ in a pre-commit file ^^ was not reported any more in a corresponding post-commit file ^^′. Then, the candidate samples 744a-744n may be labeled 740 - 20 - 3924732.v36265.1015001 as negative or positive with respect to making a proper code change fixing an issue. Outputs of such a process may be called fix pairs 746a-746n. On cleaned fix pairs, CODEREDUCE may be applied for issue (^^,ℓ,^^). The output of performing the code reduction, apart from yielding reduced pre-commit code ^^, may allow for obtaining a reduced version ^^′ of corresponding post-commit code, in reduced pairs 748a-748n (^^, ^^′). Final data may be generated in several exemplary flavors: the FULLORIGINAL pairs 746a-746n, the CODEREDUCED pairs 748a- 748n, LONGCONTEXTpairs 752a-752n, and / or WINDOW@3 pairs 754a-754n. Given below are explanations and motivations behind each exemplary flavor.
[0088] The FULLORIGINAL dataset 746a-746n may include, e.g., pre-files scraped fromGitHub in their entirety. In post-files, only changes relevant to a fix may be retained, excluding irrelevant modifications like style fixes, refactoring, or new features, among other examples. As explained in more detail herein, CODEREDUCE may be designed to ignore such changes. To ensure a fair experimental setup in comparison to CODEREDUCE, precautions may be taken to filter noisy changes in this dataset. Without this control, noisy changes may lead to unfair comparisons in metrics dependent on direct code comparison and may also potentially harm a learning process.
[0089] The CODEREDUCED data 748a-748n may contain diffs from the FULLORIGINAL data746a-746n passed through the CODEREDUCE component. This may result in a much smaller- sized source code which, importantly, may not contain parts unrelated to an issue at hand.
[0090] The LONGCONTEXT data 752a-752n may be a result of keeping, e.g., ±50 lines ofcode around a line of a detected issue and its corresponding fix. This type of data may serve for experimenting with baseline models.
[0091] The WINDOW@3 data 754a-754n may be a result of keeping, e.g., ±3 lines of codearound a line of a detected issue and its corresponding fix. This type of data may also serve for experimenting with baseline models.
[0092] Example dataset statistics, including token size percentiles, are provided in Table 2below, according to an embodiment. Tokenization may be performed using, e.g., a tokenizer of the StarCoderBase model. It should be noted that, in an embodiment, CODEREDUCE can make data suitable even for feeding into transformers with short attention range. - 21 - 3924732.v36265.1015001Table 2: Example data statistics.
[0093] An example quantification of a ratio between token lengths of the FULLORIGINAL746a-746n and CODEREDUCED 748a-748n source code is provided in Table 3 below, according to an embodiment. Table 3: Example CODEREDUCE compression ratio for {FULLORIGINAL ↦ CODEREDUCED}.
[0094] Experimental Evaluation
[0095] Presented hereinbelow is an extensive evaluation of embodiments.
[0096] Exemplary Train / Test Split
[0097] In many domains, a conventional approach may involve randomly partitioning adataset into training and test sets. However, this methodology, commonly employed in program repair research, may be unsuitable for the APR field. Randomly splitting data may inadvertently lead to instances from the same repository or even the same file being allocated to both training and test sets, resulting in data leakage. Furthermore, using pretrained models may introduce another potential source of data leakage. Such models may have been trained on extensive datasets that encompass open-source code, creating a possibility that samples in a fine-tuning test - 22 - 3924732.v36265.1015001 dataset are also present in a pretraining dataset of an underlying model. Remarkably, these two concerns have not received widespread attention in the context of APR.
[0098] Introduced herein is an advanced train / test splitting approach that effectivelyaddresses both of the foregoing issues. In addition to utilizing code with permissive licenses, data mining may be conducted from open-source repositories governed by restrictive licenses. These restrictive licenses may prohibit utilization of code for learning but allow use of the code for testing. Consequently, a test set may be constructed from these repositories. Note that this advanced train / test splitting approach may result in utilization of entirely distinct sets of repositories for training and testing, eliminating a potential for data leakage stemming from shared repositories and files. Moreover, this approach can effectively mitigate data leakage arising from a pretrained model’s dataset. Given that repositories covered by non-permissive licenses may not be suitable for training, it can be reasonably concluded that the repositories covered by non-permissive licenses were not utilized by prior works either.
[0099] Finally, this advanced train / test splitting approach allows for having a larger set oftest samples as opposed to traditional splits where there is usually a concern with losing training samples. A test dataset may contain, e.g., 1,818 samples and they may be labeled in the same way as training samples.
[0100] Exemplary Metrics
[0101] For measuring model performance, in an example embodiment, two exemplarymetrics may be utilized: functional correctness (PASS@^^) and exact match to target code (EXACTMATCH@^^). Both metrics may be computed using an “@^^” flavor, that is, at different amounts of predictions generated by, e.g., beam search or nucleus sampling, depending on a model. To simplify the explanation, consider the following notation: for input code ^^ with an issue ^^ instance detected on line ℓ, a model may generate ^^ outputs ^^ ≡ ^^(^^, ℓ, ^^) = (^^1, … , ^^^^).
[0102] PASS@^^: In an embodiment, a code analysis engine may be capable, for a given issueinstance (^^, ℓ, ^^) and a corresponding prediction ^^, of providing a predicate DoesFix(^^, ^^, ℓ, ^^), which may capture if this prediction fixes the given issue instance. Note that line information may be important, because there can be multiple instances of the same issue on several lines. According to another example embodiment, a different predicate NoNewIssues(^^|^^) can check if ^^ introduced new issues elsewhere in code compared to ^^. Using this information, a PASS@^^ - 23 - 3924732.v36265.1015001 metric may check if any of ^^ predictions fixed a static analysis report and did not introduce any new issues, and may take an average according to exemplary Equation 1 below:
[0103] Inrequire that aprediction is parseable and / or syntactically correct code.
[0104] EXACTMATCH@^^: This metric may rely on an existence of a labeled target fix ^^′ fora given issue instance (^^, ℓ, ^^) and may determine if any predictions equal labeled target code ^^′ according to exemplary Equation 2 below:
[0105] and / ordisadvantages. PASS@^^ quality may depend on robustness of an analysis engine. For example, a trivial definition of DoesFix(…) as “an issue is not detected anymore” can be easily overfitted to—e.g., by deleting whole code. While the code analysis engine used for experiments described herein is much more advanced, it may not be free of producing false positives in certain corner cases. Defining a “suitable” metric for program repair may prove to be hard. Reporting functional correctness and exact match has an advantage of relaying a trade-off between “creativeness” and “canonicalization” of fixes, with PASS@^^ and EXACTMATCH@^^ being an upper- and lower-bound for some “true” metric.
[0106] Exemplary Training and Inference Configuration
[0107] Open-Source Models: Open-source models (StarCoderBase, T5, Mistral, Mixtral)were fine-tuned on 16 Nvidia® A100 graphics processing units (GPUs) for 60 epochs with batch size of 4 per device, gradient accumulation steps of 16, learning rate of 10-5with linear scheduler, and warm-up ratio of 0.1. DeepSpeed® with ZeRO-3 optimization was used. Further, beam search was used and various “@^^” metrics were reported. Specifically, for Mixtral, QLoRA was used to effectively fine-tune Mixtral, as the model is relatively large. Other known - 24 - 3924732.v36265.1015001 models, configurations, processors, optimizations, searches, metrics, and / or parameters are also suitable.
[0108] Models Accessible via API: GPT-3.5 and GPT-4 allow querying via APIs. Recentreleases, namely, gpt-3.5-turbo-0613 and gpt-4-0613, were employed. The models were accessed via APIs, and inference was run on a test set with few-shot learning. Other known model versions and / or inference methods are also suitable.
[0109] Few-shot Examples Choice: For each sample requesting to fix an issue ^^, few-shottraining examples were chosen randomly from a training dataset for the same issue type. For each query, 1 and 2 few-shot example(s) were provided for GPT-3.5 and GPT-4, respectively. This choice may be made based on context window limitations of each model. Seeds were preserved to keep few-shot examples consistent across experiments. Other known examples and / or seeds are also suitable.
[0110] Prompt Structure: Best practices issued for GPT-3.5 and GPT-4 by OpenAI wereused to set up a prompt and conversations. Other known prompts are also suitable. In-depth details on a shape and structure of prompts used for GPT-3.5 and GPT-4 are provided as follows.
[0111] In an example embodiment, RULE and DESCRIPTION may denote a name anddescription of an issue reported by a static analyzer in a code snippet. A model may be queried to generate a fix for this code snippet and it may be denoted by^^^^^^^^^^^^^^^^. Variable ^^ may be a number of few-shot examples provided in a prompt and a pair of code snippets (^^^^^^^^^^^^^^^^, ^^^^^^^^^^^^^^^^^^) may denote an ^^^^ℎexample fix for the issue RULE in the prompt.
[0112] GPT-3.5 and GPT-4 were designed to make conversations and best practices sharedby OpenAI were followed to build an initial conversation. According to an embodiment, a conversation can contain three different roles, namely system, user, and assistant. It may be advised to start a conversation with a system content where a role of an assistant (e.g., artificial intelligence (AI) model) is defined and instructions are given on desired output structure. After that, few-shot examples may be provided as a conversation between a user and the assistant. A conversation may be finished by a user providing vulnerable code by ^^^^^^^^^^^^^^^^so that a next turn belongs to an assistant. An assistant may complete a conversation by generating a fix to a last user query. Precisely, the following prompt may be fed into GPT models. (The last sentence of the system prompt may not be provided when a full file is fed into a model.) - 25 - 3924732.v36265.1015001
[0113] Inference Hyperparameters:number of generated tokens was set to amaximum context size of a model and a temperature of 0.2 for cases where multiple predictions were generated. While querying GPT-4 with a full file in a zero-shot setting, the temperature was 0. The instructions provided in a system prompt (described above) were sufficient to configure the model’s output. Other known token numbers and / or temperatures are also suitable.
[0114] Exemplary Comparison with Baselines
[0115] In the following exemplary experiment, bug-fixing performance of different models,including models like, e.g., GPT-4, was compared under different context extraction methods and learning setups—thus forming strong baselines to test an approach of embodiments. Text-to- text transformers pre-trained on code or natural language (StarCoderBase, T5, Mixtral) or for few-shot learning via API (GPT-3.5, GPT-4) were used. Further, three exemplary datasets were utilized: CODEREDUCED (which may represent an approach of embodiments to context extraction), LONGCONTEXT (which may represent long context), and WINDOW@3 (which may represent a baseline). As a result, the below exemplary models were evaluated: a) StarCoderBase: The StarCoderBase model, fine-tuned on CODEREDUCED dataLONGCONTEXT, may allow for comparing both approaches against each other in a fine-tuning setup. - 26 - 3924732.v36265.1015001 b) TFix: The TFix model (also named T5-WINDOW@3 for clarity) was fine-tuned onWINDOW@3 data, e.g., only 3 lines before / after an issue are taken as input to the model. Thus, the model may only utilize a linear span of context lines to be trained on. c) GPT-3.5 and GPT-4: Predictions for CODEREDUCED, LONGCONTEXT, andFULLORIGINALwere obtained via few-shot learning from the model(s). d) Mixtral-8x7B: Predictions for CODEREDUCED were obtained via fine-tuning withQLoRA.
[0116] Metrics on CODEREDUCED: During inference, after obtaining predictions, they aremerged back into full source code by means of the MERGEBACK method of embodiments, thus resulting in predicting a complete file. Metrics are then determined for full files.
[0117] Metrics on LONGCONTEXT and WINDOW@3: During inference, after obtaining such“windowed” predictions, they are pasted back into full source code, and then metrics are determined for full files.
[0118] Example results are shown in table 834 of FIG. 8, which shows an exampleevaluation of PASS@^^ and EXACTMATCH@^^ metrics for models that use the CODEREDUCE technique of embodiments (marked as ) against baselines of TFix, and various large-context models. Further example results are shown in Table 4 below. Examples of fixes produced by embodiments are shown in FIGs.9-16, described hereinbelow.
[0119] It is observed that embodiments can take advantage of “compressing” context throughreduction: embodiments outperform LONGCONTEXT, FULLORIGINAL, and WINDOW@3 baselines by a considerable margin with respect to PASS@^^ and EXACTMATCH@^^. - 27 - 3924732.v36265.1015001Table 4: Example evaluation of PASS@^^ and EXACTMATCH@^^ metrics for models that use the CODEREDUCE technique of embodiments (marked as ).
[0120] Referring again to table 834 of FIG. 8, for the GPT model family, within each pair ofCODEREDUCED-LONGCONTEXTand CODEREDUCED-FULLORIGINAL, the former leads to better PASS@^^. The gap is especially large for issue categories with complex data flow. For GPT-4 and SECURITYFLOW, the CODEREDUCE technique of embodiments improves 23%, over LONGCONTEXT, while improving 14% over FULLORIGINAL. It is also seen that GPT-4-32K- FULLORIGINAL-ZEROSHOTyields not better, but competitive performance in some categories, but querying the model in this way may have a considerable latency. Hence, it was evaluated only for ^^ = 1. To be fair, GPT-4-CODEREDUCED-ZEROSHOT was also evaluated with ^^ = 1, and the CODEREDUCE technique of embodiments yields better results in all categories, especially for SECURITYFLOW, FILEWIDE, and LOCAL. - 28 - 3924732.v36265.1015001
[0121] It is noted that there is a range of other bug-fixing models such as SequenceR,CoCoNuT, or Hoppity, among other examples, that were already shown by prior works to be noncompetitive to the T5 model used in TFix. These other bug-fixing models were omitted from evaluation, as embodiments already outperform TFix by a large margin. Also investigated were potential effects of model size between 1B and 7B and model architecture between STARCODERBASEand MISTRAL. Example results are summarized in Table 5 below, according to an embodiment.Table 5: Example effects of model size and architecture for CODEREDUCED data, with respect to PASS@^^ for ^^ = 1.
[0122] In contrast to findings observed by TFix, it may be concluded that increasing modelsize leads to significant and consistent advantages. Also, a pretrained model and its architecture can make a significant difference. Finally, examples of model predictions, fixing long-ranging issues that benefit from using the CODEREDUCE technique of embodiments, are shown in FIGs. 9-16.
[0123] FIG. 9 illustrates example fix 956 for a Sql Injection vulnerability according to anembodiment. In FIG.9, code 960 includes the vulnerability and code 961 is the corrected code. Sql Injection may be one of the most common and critical vulnerabilities.
[0124] FIG. 10 illustrates example fix 1056 for a HTTPSourceWithUncheckedTypevulnerability according to an embodiment. In FIG.10, code 1060 includes the vulnerability and code 1061 is the corrected code. The HTTPSourceWithUncheckedType vulnerability may be inside Juice-Shop, one of the intentionally vulnerable benchmark repositories.
[0125] FIG. 11 illustrates example fix 1156 for a UseCsrfForExpress vulnerability accordingto an embodiment. In FIG.11, code 1160 includes the vulnerability and the code 1161 is the corrected code. The UseCsrfForExpress vulnerability may be inside the Juice-Shop intentionally - 29 - 3924732.v36265.1015001 vulnerable benchmark repository. Note that, in an embodiment, import statement 1158a for express, definition 1158b of app, and its usage 1158c can be arbitrarily far away from each other. According to another example embodiment, the DEEPCODEAI FIXsystem can bring them all in the same range and modify several places in a file without any issue.
[0126] FIG. 12 illustrates example fix 1256 for a NoRateLimiting vulnerability according toan embodiment. In FIG.12, code 1260 includes the vulnerability and code 1161 is the corrected code. The NoRateLimiting vulnerability may be inside appsecco / dvna, another of the intentionally vulnerable benchmark repositories. This vulnerability may be hard to fix because the fix may include changes in three different locations of a file, e.g., 1258a-1258c, and some of those changes, e.g., 1258b, may involve multiple lines.
[0127] FIG. 13 illustrates example fix 1356 for a CommandInjection vulnerability accordingto an embodiment. In FIG.13, code 1360 includes the vulnerability and code 1361 is the corrected code. The CommandInjection vulnerability may be inside the appsecco / dvna intentionally vulnerable benchmark repository. In an example embodiment, it is noted (i) how the DEEPCODE AI FIX system may keep required import 1358 during reduction and (ii) the significant compression rate.
[0128] FIG. 14 illustrates example fix 1456 for a UseHelmetForExpress vulnerabilityaccording to an embodiment. In FIG.14, code 1460 includes the vulnerability and code 1461 is the corrected code. The UseHelmetForExpress vulnerability may be inside the appsecco / dvna intentionally vulnerable benchmark repository. The fix 1456 may appear simple as helmet can be added with a single line 1458a. However, without adding the right import statement 1458b, code will be broken. A great bug-fixing tool may apply imports correctly. This may make even seemingly simple fix patterns much harder, as import statements and their usages can be arbitrarily far away from each other. The rule UseHelmetForExpress may belong to category SECURITYLOCAL, but may still include changes in several different places of a file, just like other “LOCAL” rules.
[0129] FIG. 15 illustrates example fix 1556 for a cross-site scripting (XSS) vulnerabilityaccording to an embodiment. In FIG.15, code 1560 includes the vulnerability and code 1561 is the corrected code. The XSS vulnerability may be inside SirAppSec / vuln-node.js-express.js-app, yet another of the intentionally vulnerable benchmark repositories. - 30 - 3924732.v36265.1015001
[0130] FIG. 16 illustrates example fix 1656 for a HttpToHttps vulnerability according to anembodiment. In FIG.16, code 1660 includes the vulnerability and code 1661 is the corrected code. In an embodiment, the fix 1656 may include multiple changes in different locations of a file, e.g., 1658a and 1658b.
[0131] Exemplary Effects of MERGEBACK
[0132] Compared to long context or window-based baselines, using CODEREDUCEDaccording to an embodiment may contain more non-trivial steps, including MERGEBACK. Table 6 below is an example assessment of an effect of MERGEBACK on metrics, according to another embodiment. Metrics in Table 6 were calculated on full source code and reduced code, respectively. Note that this may be possible because PASS@^^ / EXACTMATCH@^^ may be computable also on reduced (before MERGEBACK) code, because it is possible to validate if a fix removes a static analysis report on reduced code.Table 6: Example influence of MERGEBACKon metrics for STARCODERBASE-CODEREDUCED.
[0133] Reduced Latency of Exemplary End-to-end Bug-fixing System
[0134] An exemplary experiment was performed to compare a CODEREDUCE method ofembodiments (Method 1) with a vanilla HDD method. To avoid differences caused by machine and analyzer speeds, a number of calls to an analyzer made by code reduction methods was evaluated. For this evaluation, 1024 exemplary programs were taken from a training dataset and code reduction speed was measured using both settings. The number of program analyzer calls executed by Method 1 averaged to 9.44 and had a geometric mean of 4.99, whereas the number - 31 - 3924732.v36265.1015001 of calls to a program analyzer performed by a HDD method averaged to 11.52 and had a geometric mean of 7.25. Accounting for geometric means that are useful for computing ratios leads to a ratio of 0.68 steps for a novel method of embodiments versus HDD, or a 30% reduction. When comparing averages, a reduction is 20%, due to several outlier samples. The most complex example in the data included 181 calls to an analyzer.
[0135] The evaluation was performed on an AWS® c5ad.8xlarge instance, runningUbuntu® 22.04; other known platforms are also suitable. In terms of latency, each call to an underlying static analyzer takes around 70 to 80^^^^. This means that an entire CODEREDUCE method of embodiments may take only, e.g., less than 1 second on average. Querying STARCODERBASE-7B takes on average 605^^^^ per prediction with an A100 GPU. Accounting for a need to check a solution after MERGEBACK and generating several candidates from a LLM, a bug-fixing system of embodiments can perform predictions within a few seconds and can be used in an IDE setting.
[0136] Exemplary Conclusions
[0137] Embodiments provide a novel learning-based system, DEEPCODE AI FIX, forautomatically fixing non-trivial coding errors and security vulnerabilities. Further, embodiments may utilize a code reduction mechanism that incorporates program analysis into a machine learning pipeline. Embodiments can learn to fix non-trivial coding errors accurately because a code reduction technique of embodiments extracts essential information needed for a fix, allowing a learning process of embodiments to avoid having to uncover arbitrarily long-range dependencies and focusing the process on attending to crucial parts of code.
[0138] An exemplary analysis of open-source commits reveals insights about a difficulty ofcollecting large amounts of data for program repair. As a solution, an exemplary high-quality dataset was constructed based on bugs and fixes extracted from millions of real-world GitHub commits.
[0139] An approach of embodiments was evaluated on multiple models—StarCoderBase,Mixtral, T5, GPT-3.5, and GPT-4—and demonstrated that directly using a model that tries to learn complex long-range dependencies needed to produce a correct fix, performs generally worse than using a context extracted by reducing code according to an embodiment. Overall, embodiments can outperform a previous state-of-the-art baseline TFix and help multiple models - 32 - 3924732.v36265.1015001 in different setups to perform better. Embodiments may utilize program analysis to deal with long-range dependencies and data flows. Such functionality greatly simplifies an attention learning task.
[0140] Embodiments not only offer innovations in program repair, but also enableincorporating program analysis into machine learning and clearly highlight a need for code analysis, even when using powerful LLMs.
[0141] Exemplary Method Embodiment
[0142] FIG. 17 is a flow diagram of a computer-implemented method 1700 for fixing anerror in source code according to an embodiment. The method 1700 begins at step 1701 by modifying original source code having an error, e.g., the source code 106b (FIG.1) or 206 (FIG. 2), to generate reduced source code including the error, e.g., the source code 114 (FIG.1), 214 (FIG.2), or 514 (FIG.5). At step 1702, the method 1700 processes the reduced source code including the error with a language model, e.g., the LLM 102b (FIG.1), to produce error- eliminated reduced source code, e.g., the output 116 (FIG.1) or 616 (FIG.6). Then, at step 1703, the method 1700 merges the error-eliminated reduced source code with the original source code to generate fixed source code, e.g., the fixed code 108b (FIG.1).
[0143] As noted, the method 1700 is computer-implemented and, as such, the functionalityand effective operations, e.g., the modifying (1701), processing (1702), and merging (1703), are automatically implemented by one or more digital processors. Moreover, the method 1700 can be implemented using any computer device or combination of computing devices known in the art. Among other examples, the method 1700 can be implemented using computer(s) / device(s) 50 and / or 60 described hereinbelow in relation to FIGs.18 and 19.
[0144] In an example embodiment of the method 1700, the error may be a securityvulnerability, e.g., the vulnerability 218 (FIG.2), a semantic error, an API misuse, a quality issue, or a style issue, among other examples.
[0145] According to an example embodiment of the method 1700, the original source codemay include multiple errors, e.g., the errors of code statements 1658a and 1658b (FIG.16). The method 1700 may further include iterating the modifying (1701), processing (1702), and merging (1703) for each error of the multiple errors. - 33 - 3924732.v36265.1015001
[0146] In another example embodiment of the method 1700, the modifying (1701) mayinclude generating a graph representation of the original source code having the error. The modifying (1701) may further include iteratively reducing the graph and performing a static analysis of the reduced graph to determine existence of the error in the reduced graph until a minimal reduced graph including the error is determined. An example of such functionality is described hereinabove in relation to FIGs.4A-4G. The minimal reduced graph including the error may represent the generated reduced source code including the error. According to an example embodiment of the method 1700, in a given iteration, reducing the graph may be based on approximate provenance information provided by a static analysis report.
[0147] According to an example embodiment of the method 1700, the language model maybe an artificial intelligence based neural network.
[0148] In another example embodiment, the method 1700 may further include training thelanguage model. According to an example embodiment of the method 1700, training the language model may include obtaining multiple code samples each including a respective error, e.g., multiple code samples of the training data 104 (FIG.1). Training the language model may further include obtaining multiple error-free code samples, e.g., error-free code samples of the training data 104. Each error-free code sample may correspond to a given code sample of the multiple code samples. Training the language model may further include reducing the multiple code samples and the multiple error-free code samples. Training the language model may further include training the language model to determine code fixes using the reduced multiple code samples and the reduced multiple error-free code samples, e.g., reduced multiple code samples and reduced multiple error-free code samples of the reduced training data 112 (FIG.1).
[0149] According to an example embodiment, the method 1700 may further include queryingthe language model in a zero-shot or few-shot learning manner. The querying may include obtaining multiple code samples each including a respective error. The querying may further include obtaining multiple error-free code samples. Each error-free code sample may correspond to a given code sample of the multiple code samples. The querying may further include reducing the multiple code samples and the multiple error-free code samples. The querying may further include querying the language model with few-shot learning. The reduced multiple code samples - 34 - 3924732.v36265.1015001 and the reduced multiple error-free code samples may be provided as correct examples to the language model.
[0150] In another example embodiment of the method 1700, the merging (1703) may includecomparing the reduced source code including the error and the error-eliminated reduced source code to determine a mapping, e.g., the mapping 636 (FIG.6), between the reduced source code including the error and the error-eliminated reduced source code. The merging (1703) may further include, based on the mapping, merging the error-eliminated reduced source code with the original source code to generate the fixed source code. According to an example embodiment of the method 1700, in the mapping, each line in the error-eliminated reduced source code may be mapped to a given line in the reduced source code including the error.
[0151] Computer Support
[0152] FIG. 18 is a schematic view of an example computer network in which embodimentsmay be implemented. Client computer(s) / devices 50 and server computer(s) 60 provide processing, storage, and input / output (I / O) devices executing application programs and the like. Client computer(s) / device(s) 50 can also be linked through communications network 70 to other computing devices, including other client device(s) / processor(s) 50 and server computer(s) 60. The communications network 70 can be part of a remote access network, a global network (e.g., the Internet), cloud computing servers or service, a worldwide collection of computers, local area or wide area networks, and gateways that currently use respective protocols (e.g., TCP / IP, Bluetooth®, etc.) to communicate with one another. Other electronic device / computer network architectures are also suitable.
[0153] FIG. 19 is a block diagram illustrating an example embodiment of a computer node(e.g., client processor(s) / device(s) 50 or server computer(s) 60) in the computer network 70 of FIG.18. Each computer node 50, 60 contains system bus 79, where a bus is a set of hardware lines used for data transfer among components of a computer or processing system. The system bus 79 is essentially a shared conduit that connects different elements of a computer system (e.g., processor, disk storage, memory, I / O ports, network ports, etc.) that enables transfer of information between the elements. Attached to the system bus 79 is an I / O devices interface 82 for connecting various input and output devices (e.g., keyboard, mouse, display(s), printer(s), speaker(s), etc.) to the computer node 50, 60. A network interface 86 allows the computer node - 35 - 3924732.v36265.1015001 to connect to various other devices attached to a network (e.g., the network 70 of FIG.18). A memory 90 provides volatile storage for computer software instructions 92a and data 94a used to implement an embodiment of the present disclosure (e.g., the method 1700 of FIG.17). A disk storage 95 provides non-volatile storage for the computer software instructions 92b and data 94b used to implement an embodiment of the present disclosure. A central processor unit 84 is also attached to the system bus 79 and provides for execution of computer instructions.
[0154] In one embodiment, the processor routines 92a-92b and data 94a-94b are a computerprogram product (generally referenced as 92), including a non-transitory, computer readable medium (e.g., a removable storage medium such as DVD-ROM(s), CD-ROM(s), diskette(s), tape(s), etc.) that provides at least a portion of the software instructions for the disclosure system. The computer program product 92 can be installed by any suitable software installation procedure, as is well known in the art. In another embodiment, at least a portion of the software instructions may also be downloaded over a cable, communication, and / or wireless connection. In other embodiments, the disclosure programs are a computer program propagated signal product embodied on a propagated signal on a propagation medium (e.g., a radio wave, an infrared wave, a laser wave, a sound wave, or an electrical wave propagated over a global network such as the Internet, or other network(s)). Such carrier medium or signals provide at least a portion of the software instructions for the present disclosure routines / program 92.
[0155] In alternative embodiments, the propagated signal is an analog carrier wave or digitalsignal carried on the propagated medium. For example, the propagated signal may be a digitized signal propagated over a global network (e.g., the Internet), a telecommunications network, or other networks (such as the network 70 of FIG.18). In one embodiment, the propagated signal is a signal that is transmitted over the propagation medium over a period of time, such as the instructions for a software application sent in packets over a network over a period of milliseconds, seconds, minutes, or longer. In another embodiment, the computer readable medium of the computer program product 92 is a propagation medium that the computer system 50 may receive and read, such as by receiving the propagation medium and identifying a propagated signal embodied in the propagation medium, as described above for computer program propagated signal product.
[0156] Generally speaking, the term “carrier medium” or transient carrier encompasses the- 36 - 3924732.v36265.1015001 foregoing transient signals, propagated signals, propagated medium, storage medium, and the like.
[0157] In other embodiments, the program product 92 may be implemented as a so-calledSoftware as a Service (SaaS), or other installation or communication supporting end-users.
[0158] Example Code, Datasets, Instructions, and Configurations for ImplementingEmbodiments
[0159] Provided hereinbelow as Appendices A-R are example configurations and code thatmay be used to implement embodiments. Further provided hereinbelow as Appendices S and T are example datasets that may be utilized for implementing embodiments.
[0160] The Appendices are referred to as follows:a) Appendix A (requirements.txt)i. A person having ordinary skill in the art can recognize that the file“requirements.txt” (Appendix A) can be placed in a directory named “deepcode_ai_fix”. b) Appendix B (data_schemas.py)i. A person having ordinary skill in the art can recognize that the file“data_schemas.py” (Appendix B) can be moved to a directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. c) Appendix C (train_autofix.sh)i. A person having ordinary skill in the art can recognize that the file“train_autofix.sh” (Appendix C) can be moved to a directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. d) Appendix D (predict_autofix.sh)i. A person having ordinary skill in the art can recognize that the file“predict_autofix.sh” (Appendix D) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. e) Appendix E (args.py)i. A person having ordinary skill in the art can recognize that the file“args.py” (Appendix E) can be moved to the directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. - 37 - 3924732.v36265.1015001 f) Appendix F (predict_llm.py)i. A person having ordinary skill in the art can recognize that the file“predict_llm.py” (Appendix F) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. g) Appendix G (predict_llm.sh)i. A person having ordinary skill in the art can recognize that the file“predict_llm.sh” (Appendix G) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. h) Appendix H (data_processor.py)i. A person having ordinary skill in the art can recognize that the file“data_processor.py” (Appendix H) can be moved to the directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. i) Appendix I (metrics_aggregator.py)i. A person having ordinary skill in the art can recognize that the file“metrics_aggregator.py” (Appendix I) can be moved to the directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. j) Appendix J (metrics_impl.py)i. A person having ordinary skill in the art can recognize that the file“metrics_impl.py” (Appendix J) can be moved to the directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. k) Appendix K (trainer.py)i. A person having ordinary skill in the art can recognize that the file“trainer.py” (Appendix K) can be moved to the directory path “autofix / ml / lib” under the “deepcode_ai_fix” directory. l) Appendix L (predict_autofix.py)i. A person having ordinary skill in the art can recognize that the file“predict_autofix.py” (Appendix L) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. m) Appendix M (train_autofix.py)- 38 - 3924732.v36265.1015001 i. A person having ordinary skill in the art can recognize that the file“train_autofix.py” (Appendix M) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. n) Appendix N (deepspeed_config.json)i. A person having ordinary skill in the art can recognize that the file“deepspeed_config.json” (Appendix N) can be moved to the directory path “autofix / ml / bin” under the “deepcode_ai_fix” directory. o) Appendix O (loading.py)i. A person having ordinary skill in the art can recognize that the file“loading.py” (Appendix O) can be moved to a directory path “autofix / ml / utils” under the “deepcode_ai_fix” directory. p) Appendix P (inference_pipeline.py)i. A person having ordinary skill in the art can recognize that the file“inference_pipeline.py” (Appendix P) can be moved to a directory path “autofix / ml_inference” under the “deepcode_ai_fix” directory. q) Appendix Q (distributed_training.py)i. A person having ordinary skill in the art can recognize that the file“distributed_training.py” (Appendix Q) can be moved to a directory “ml_utils” under the “deepcode_ai_fix” directory. r) Appendix R (model_type.py)i. A person having ordinary skill in the art can recognize that the file“model_type.py” (Appendix R) can be moved to the directory “ml_utils” under the “deepcode_ai_fix” directory. s) Appendix S (train_example.json)i. A person having ordinary skill in the art can recognize that the file“train_example.json” (Appendix S) can be moved to the “deepcode_ai_fix” directory. t) Appendix T (test_example.json)- 39 - 3924732.v36265.1015001 i. A person having ordinary skill in the art can recognize that the file“test_example.json” (Appendix T) can be moved to the “deepcode_ai_fix” directory.
[0161] Provided hereinbelow are example instructions and configurations that may beutilized for implementing embodiments in conjunction with the provided Appendices.
[0162] Overview
[0163] Embodiments provide state-of-the-art functionality for automatically fixing securityvulnerabilities and coding errors in software systems. Embodiments can utilize LLMs and leverage program analysis to limit an LLM’s attention mechanism to the portions of code needed to perform the fix. In an example embodiment, such an approach drastically reduces the amount of required training data. Concretely, for both training and inference, rather than feeding an entire program to an LLM, an embodiment reduces the program’s code to a much shorter snippet that contains a reported defect together with necessary context—and uses that instead.
[0164] Utilizing Embodiments
[0165] Users can utilize embodiments in modern IDEs. In an implementation, Snyk IDEExtension can be installed and AI Fix suggestions can be enabled. Further documentation is available on Applicant-Assignee’s (Snyk Limited) website as follows: a) “Setting up Snyk Code in IDE” (docs.snyk.io / scm-ide-and-ci-cd-integrations / snyk-ide-plugins-and-extensions) b) “DeepCode AI Fix” (docs.snyk.io / scan-using-snyk / snyk-code / manage-code-vulnerabilities / fix-code-vulnerabilities-automatically)
[0166] Scientific Paper
[0167] Further details regarding example embodiments, results, and findings can be found inthe following scientific paper: Berabi, B., et al., “DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models,” arXiv preprint arXiv:2402.13291 (2024) (arxiv.org / abs / 2402.13291), which is herein incorporated by reference in its entirety. It should be noted that the paper presents findings from early 2023. Further, it is noted that while the paper focuses on training and evaluation using JavaScript, embodiments can support any language known to those of skill in the art, including, e.g., Java, Python, C / C++, Go, and Apex.
[0168] Example Configuration- 40 - 3924732.v36265.1015001
[0169] To implement a non-limiting example, a system with Python 3 is installed on thesystem being used. Then, the below steps are utilized to create a virtual environment and install the necessary dependencies: cd deepcode_ai_fix python3 -m venv dc_ai_fix_venv source dc_ai_fix_venv / bin / activate pip install -r requirements.txt
[0170] If any errors are encountered during the installation of the requirements (AppendixA), it may be due to certain libraries needing system-wide packages. Error messages can be reviewed carefully and / or required system-wide dependencies can be installed according to the system’s specifications. For instance, it may be necessary to install mpi packages as follows: sudo apt install libmpich-dev libopenmpi-dev
[0171] Additionally, it may be necessary to install PyTorch®, Nvidia drivers, and / or CUDAbased on GPU requirements.
[0172] Example Dataset and Models
[0173] Exemplary dataset schemas are defined in the file “autofix / ml / lib / data_schemas.py”(Appendix B). Below are provided detailed explanations for each schema and field.
[0174] LabelledDataSchema (used for fine-tuning)Field Descriptionrule name of static analysis rule, e.g., SQLi (SQL injection)message message returned from a static analyzer describing a reportedissue line_number line number where a static analysis report startsline_end line number where a static analysis report endscol_begin column number where a static analysis report startscol_end column number where a static analysis report endspre_file content of an entire file in pre-version (before fix)post_file content of an entire file in post-version (after fix)repo <org_id> / <repo_name> in GitHubpre_filename file name in pre-version (before fix)- 41 - 3924732.v36265.1015001 post_filename file name in post-version (after fix)pre_sha commit Secure Hash Algorithm (SHA) hash of pre-versionpost_sha commit SHA hash of post-versionpre_reduced Reduced code snippet containing static analysis reportpost_reduced Reduced fixed code snippetreduction_line_map Line mapping between pre_reduced and pre_filereduction_match_line_num Line number of a static analysis report in pre_reducedlanguage programming language
[0175] The files “train_example.json” (Appendix S) and “test_example.json” (Appendix T)are examples of training and test files, respectively, implementing LabelledDataSchema.
[0176] PredictionSchema
[0177] Everything under LabelledDataSchema and:Field Descriptionpredictions a list of strings containing predictions
[0178] EvaluationSchema
[0179] Everything under PredictionSchema and:Field Descriptiontrue_fix a list of booleans indicating for each prediction (corresponding indices) whetherthe prediction passed static analysis checks eval_status a list of strings for each prediction (corresponding indices) summarizing staticanalysis evaluation. Passed or error message in case of failure exact_match a list of booleans indicating for each prediction (corresponding indices) whetherthe prediction exactly matches a target fix (post_reduced or post_file depending on the experiment)
[0180] Example Fine-tuning and Obtaining Predictions
[0181] Convenient scripts are provided for both training and inference. The training script“autofix / ml / bin / train_autofix.sh” (Appendix C) and inference script “autofix / ml / bin / predict_autofix.sh” (Appendix D) should be reviewed carefully.
[0182] In an example embodiment, parameters are preset to the values primarily used in theabove referenced paper. However, depending on the desired experiment, it may be necessary to - 42 - 3924732.v36265.1015001 adjust or add a few parameters. For example, if training the Mixtral8x7B model, it may be necessary to enable parameter-efficient fine-tuning (LoRA). Detailed training information for each model is available in the above referenced paper, and available arguments can be viewed in the “autofix / ml / lib / args.py” file (Appendix E).
[0183] To access some models, the license agreement for the HuggingFace® (HF) UI mayneed to be accepted. If model loading crashes due to this reason, there may be useful error messages. The instructions for HF can be followed. Further, an account on HF can be created and a HF token to be used can be exported as follows: export HUGGING_FACE_HUB_TOKEN="<your_token>"
[0184] Training Exampleenv MODEL_NAME="bigcode / starcoderbase-3b" NUM_EPOCHS=60 INPUT_MAX_NUM_TOKENS=512 . / autofix / ml / bin / train_autofix.sh
[0185] Inference Exampleenv MODEL_NAME="path_to_best_model_dir_from_training_script" MAX_NUM_TOKENS=512 BATCH_SIZE=1. / autofix / ml / bin / predict_autofix.sh
[0186] Running Experiments Against Third-party LLMs
[0187] In the above referenced paper, embodiments are evaluated in a few-shot learningsetting using LLMs that are accessible only via API (such as GPT-4). The code to run these experiments is provided in the Appendices. However, there is one caveat: these experiments were conducted using a private API endpoint. If access to such an endpoint is available, a uniform resource locator (URL) for the endpoint can easily be passed as a command-line argument. If not, a slight modification to the code may be needed to use the OpenAI library directly instead of the requests library, as implemented in the “autofix / ml / bin / predict_llm.py” file (Appendix F).
[0188] All other details, including parameters, prompt construction, and selection of few-shot examples, can be found in the code provided in the Appendices. A convenient script “autofix / ml / bin / predict_llm.sh” (Appendix G) is also provided to run example experiments. The instructions in the script can be followed and the parameter combinations used in the above referenced paper can be reviewed, depending on the desired model.
[0189] Running Evaluations- 43 - 3924732.v36265.1015001
[0190] The above referenced paper provides a detailed description of the evaluation process.This section offers guidance on running Snyk Code on users’ own predictions. Users may need to replicate the evaluation process themselves.
[0191] To evaluate predictions, a static analyzer can be used as a third party by utilizing theanalyzer’s API. Applicant-Assignee Snyk Limited offers a command-line interface (CLI) that can be used to run the analyzer. The steps described in the Snyk Limited official documentation (docs.snyk.io / snyk-cli / scan-and-maintain-projects-using-the-cli / snyk-cli-for-snyk-code) can be followed.
[0192] Once the analyzer is set up, it is recommended to create a Python script to automateinvoking CLI commands on predictions.
[0193] Embodiments or aspects thereof may be implemented in the form of hardwareincluding but not limited to hardware circuitry, firmware, or software. If implemented in software, the software may be stored on any non-transient computer readable medium that is configured to enable a processor to load the software or subsets of instructions thereof. The processor then executes the instructions and is configured to operate or cause an apparatus to operate in a manner as described herein.
[0194] Further, hardware, firmware, software, routines, or instructions may be describedherein as performing certain actions and / or functions of the data processors. However, it should be appreciated that such descriptions contained herein are merely for convenience and that such actions in fact result from computing devices, processors, controllers, or other devices executing the firmware, software, routines, instructions, etc.
[0195] It should be understood that the flow diagrams, block diagrams, and networkdiagrams may include more or fewer elements, be arranged differently, or be represented differently. But it further should be understood that certain implementations may dictate the block and network diagrams and the number of block and network diagrams illustrating the execution of the embodiments be implemented in a particular way.
[0196] Accordingly, further embodiments may also be implemented in a variety of computerarchitectures, physical, virtual, cloud computers, and / or some combination thereof, and, thus, the data processors described herein are intended for purposes of illustration only and not as a limitation of the embodiments. - 44 - 3924732.v36265.1015001
[0197] The teachings of all patents, published applications, and references cited herein areincorporated by reference in their entirety.
[0198] While example embodiments have been particularly shown and described, it will beunderstood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the embodiments encompassed by the appended claims.
[0199] For example, the foregoing description and details of embodiments referenceApplicant-Assignee (Snyk Limited), software, tools, and platforms, for purposes of illustration and not limitation. Other similar software, tools, and platforms are also suitable. - 45 - 3924732.v36265.1015001 APPENDIX A – “requirements.txt” accelerate~=0.19 bitsandbytes~=0.9 black~=24.3.0 datasets~=2.14 python-dateutil~=2.8 deepspeed<0.12 einops~=0.6 fairscale~=0.4 fastparquet~=2023.8 fsspec~=2023.6 https: / / download.pytorch.org / whl / cu118 / torch-2.1.2%2Bcu118-cp310-cp310- linux_x86_64.whl https: / / download.pytorch.org / whl / triton-2.1.0-0-cp310-cp310- manylinux2014_x86_64.manylinux_2_17_x86_64.whl https: / / github.com / Dao-AILab / flash- attention / releases / download / v2.5.2 / flash_attn- 2.5.2+cu118torch2.1cxx11abiFALSE-cp310-cp310-linux_x86_64.whl huggingface-hub>=0.19.3 ipython~=8.10 isort~=5.10 numpy~=1.24 pandas-stubs~=1.5.0 pandas~=1.5.0 pandera~=0.19.3 parquet~=1.3 peft~=0.7 pyarrow==15.0.2 requests scikit-learn~=1.5.0 sentencepiece~=0.1 setuptools~=69.5 simple-parsing~=0.1 structlog~=22.1 tqdm~=4.62 transformers>=4.39.3 thriftpy2~=0.5.0 mpi4py~=4.0.0 - 46 - 3924732.v36265.1015001 APPENDIX B – “data_schemas.py” """We use strict dataframe schema validations. Use `pandera` for data validations. See below for examples. """ import numpy as np import pandera as pa from pandera.typing import Int32 from pandera.typing import Object from pandera.typing import Series class LabelledDataSchema(pa.SchemaModel): rule: Series[str] message: Series[str] line_number: Series[Int32] line_end: Series[Int32] col_begin: Series[Int32] col_end: Series[Int32] severity: Series[Int32] event_id: Series[Int32] pre_file: Series[str] post_file: Series[str] repo: Series[str] pre_filename: Series[str] post_filename: Series[str] pre_sha: Series[str] post_sha: Series[str] pre_reduced: Series[str] post_reduced: Series[str] reduction_line_map: Series[Object] reduction_match_line_num: Series[Int32] language: Series[str] jj: Series[str] @pa.check("reduction_line_map") def is_list_of_int32(cls, s: Series) -> Series[bool]: # `.pipe()` is needed to reassure checker return s.apply( lambda elem: isinstance(elem, (list, tuple, np.ndarray)) and (len(elem) == 0 or isinstance(elem[0], (int, np.int32))) ).pipe(Series[bool]) class PredictionSchema(LabelledDataSchema): predictions: Series[Object] # List[str], checked at runtime @pa.check("predictions") def is_list_of_strings(cls, s: Series) -> Series[bool]: # `.pipe()` is needed to reassure checker return s.apply( lambda elem: isinstance(elem, (list, tuple, np.ndarray)) and (len(elem) == 0 or isinstance(elem[0], (str, np.str_))) ).pipe(Series[bool]) class EvaluationSchema(PredictionSchema): true_fix: Series[Object] # List[bool], checked at runtime eval_status: Series[Object] # List[str], checked at runtime exact_match: Series[Object] # List[bool], checked at runtime @pa.check("predictions", "eval_status") - 47 - 3924732.v36265.1015001 def is_list_of_strings(cls, s: Series) -> Series[bool]: return super().is_list_of_strings(s) @pa.check("true_fix", "exact_match") def is_list_of_bool(cls, s: Series) -> Series[bool]: # `.pipe()` is needed to reassure checker return s.apply( lambda elem: isinstance(elem, (list, tuple, np.ndarray)) and (len(elem) == 0 or isinstance(elem[0], (bool, np.bool_))) ).pipe(Series[bool]) - 48 - 3924732.v36265.1015001 APPENDIX C – “train_autofix.sh” #! / bin / bash if [[ -z "${MODEL_NAME}" ]]; then echo "please provide a model name in the MODEL_NAME env variable" exit 1 fi LAST_MODEL_NAME_COMPONENT=$(echo "$MODEL_NAME" | awk -F / '{print $NF}') if [[ -z "${NUM_EPOCHS}" ]]; then echo "please provide the number of epochs in the NUM_EPOCHS env variable" exit 1 fi if [[ -z "${INPUT_MAX_NUM_TOKENS}" ]]; then echo "please provide the tokenizer max length in the INPUT_MAX_NUM_TOKENS env variable" exit 1 fi OUTPUT_DIR=" / tmp / train_autofix" # These disable the annoying warning messages from HF. export TRANSFORMERS_NO_ADVISORY_WARNINGS="True" export TOKENIZERS_PARALLELISM="False" # dc_ai_fix_venv / bin / python3 autofix / ml / bin / train_autofix.py \ dc_ai_fix_venv / bin / deepspeed autofix / ml / bin / train_autofix.py \ --output_dir="${OUTPUT_DIR}" \ --model_name="${MODEL_NAME}" \ --model_id="${MODEL_NAME}" \ --num_train_epochs="${NUM_EPOCHS}" \ --learning_rate="1e-5" \ --warmup_ratio=0.1 \ --save_strategy="epoch" \ --save_total_limit=1 \ --evaluation_strategy="epoch" \ --load_best_model_at_end="True" \ --greater_is_better="False" \ --metric_for_best_model="eval_loss" \ --ddp_find_unused_parameters="False" \ --remove_unused_columns="False" \ --logging_steps=1 \ --seed=42 \ --data_id="train.parquet" \ --tokenizer_max_length="${INPUT_MAX_NUM_TOKENS}" \ --max_new_tokens="${INPUT_MAX_NUM_TOKENS}" \ --per_device_train_batch_size=4 \ --per_device_eval_batch_size=16 \ --gradient_accumulation_steps=16 \ --model_artifact="llm-trials / ${LAST_MODEL_NAME_COMPONENT}" \ --bf16="True" \ --preprocessing_batch_size=1000 \ --torch_compile_mode="max-autotune" \ --deepspeed="autofix / ml / bin / deepspeed_config.json" \ --dataloader_num_workers=10 - 49 - 3924732.v36265.1015001 APPENDIX D – “predict_autofix.sh” #! / bin / bash if [[ -z "${MODEL_NAME}" ]]; then echo "please provide a model name in the MODEL_NAME env variable" exit 1 fi LAST_MODEL_NAME_COMPONENT=$(echo "$MODEL_NAME" | awk -F / '{print $NF}') if [[ -z "${MAX_NUM_TOKENS}" ]]; then echo "please provide the tokenizer max length in the MAX_NUM_TOKENS env variable" exit 1 fi if [[ -z "${BATCH_SIZE}" ]]; then echo "please provide the batch size in the BATCH_SIZE env variable" exit 1 fi NUM_BEAMS=5 export TRANSFORMERS_NO_ADVISORY_WARNINGS="True" dc_ai_fix_venv / bin / python3 autofix / ml / bin / predict_autofix.py \ --output_dir=" / tmp / predict_autofix" \ --per_device_eval_batch_size="${BATCH_SIZE}" \ --data_id="test.parquet" \ --model_id="t5-small" \ --model_name="${LAST_MODEL_NAME_COMPONENT}" \ --model_id="${MODEL_NAME}" \ --beam_size="${NUM_BEAMS}" \ --num_return_seqs="${NUM_BEAMS}" \ --tokenizer_max_length="${MAX_NUM_TOKENS}" \ --max_new_tokens="${MAX_NUM_TOKENS}" \ --seed=42 \ --dataloader_num_workers=10 \ --bf16_full_eval="True" \ --torch_compile_mode="max-autotune" \ --preprocessing_batch_size=1000 \ - 50 - 3924732.v36265.1015001 APPENDIX E – “args.py” from dataclasses import dataclass from dataclasses import field from typing import cast import transformers from transformers.hf_argparser import DataClassType @dataclass class DatasetArgs: dataset_seed: int = field( default=42, metadata={"help": "The seed used for all random operations."} ) train_size: float = field( "Relative train split size as float in range str = field( existent", "Repeatable data id of the artifact."}, int = field( "help": "The batch size to use during dataset preprocessing.see: https: / / huggingface.co / docs / datasets / v2.14.5 / en / package_reference / main_cla sses#datasets.Dataset.map.batch_size" }, ) keep_in_memory: bool = field( default=True, metadata={ "help": "Whether the `dataset.map` results should be kept in memory or written to a HuggingFace cache file. For more details, see: https: / / huggingface.co / docs / datasets / v2.14.5 / en / package_reference / main_cla sses#datasets.Dataset.map.keep_in_memory" }, ) @dataclass class ModelArgs: model_name: str = field( metadata={"help": "Model name to be used, e.g. `t5-large`."} ) tokenizer_max_length: int = field( metadata={"help": "Maximum length for tokenized samples"} ) model_id: str = field( metadata={ "help": ( "The model identifier" ) - 51 - 3924732.v36265.1015001 }, ) use_bits_and_bytes: bool = field( default=False, metadata={ "help": ( "Whether to use bits and bytes config to load and fine- tune the model." ) }, ) use_peft: bool = field( default=False, metadata={ "help": ( "Whether to use parameter efficient fine-tuning. If use_bits_and_bytes is True, this must be se to True as well." ) }, ) use_flash_attention_2: bool = field( default=False, metadata={ "help": ( "Whether to use flash attention 2. Needs batch size of 1 to work properly." ) }, ) @dataclass class InferenceArgs: max_new_tokens: int = field( metadata={ "help": "The number of new tokens to generate. It must be set to limit the maximum number of tokens generated by the model." } ) beam_size: int = field( default=1, metadata={"help": "Beam size during validation and prediction"} ) num_return_seqs: int = field( default=1, metadata={ "help": "The number of returned sequences for each sample during prediction" }, ) early_stopping: bool = field( default=False, metadata={ - 52 - 3924732.v36265.1015001 "help": "Whether to stop generation once <num_return_seqs> many sequences are done during beam search or continue with it untill all sequences are finished." }, ) @dataclass class AutofixArgs: model_args: ModelArgs dataset_args: DatasetArgs training_args: transformers.TrainingArguments inference_args: InferenceArgs def parse_autofix_args_from_cmd() -> AutofixArgs: argument_types = [ ModelArgs, DatasetArgs, transformers.TrainingArguments, InferenceArgs, ] dataclass_types = [cast(DataClassType, arg_type) for arg_type in argument_types] hf_parser = transformers.HfArgumentParser(dataclass_types=dataclass_types) ma: ModelArgs da: DatasetArgs ta: transformers.TrainingArguments ia: InferenceArgs ma, da, ta, ia = hf_parser.parse_args_into_dataclasses( return_remaining_strings=True )[: len(argument_types)] return AutofixArgs( model_args=ma, dataset_args=da, training_args=ta, inference_args=ia, ) - 53 - 3924732.v36265.1015001 APPENDIX F – “predict_llm.py” import os import tempfile import time from dataclasses import dataclass from pathlib import Path import pandas as pd import requests import simple_parsing import tqdm import autofix.ml.lib.data_processor as preprocessor from autofix.ml.lib.data_schemas import PredictionSchema from ml_utils.distributed_training import MultiProcessLogger logger = MultiProcessLogger() @dataclass(frozen=True) class LLMInferenceArgs: end_point: str model_name: str model_artifact_id: str data_id: str num_predictions: int temperature: float num_shots: int max_new_tokens: int @staticmethod def parse() -> "LLMInferenceArgs": return simple_parsing.parse(LLMInferenceArgs) def predict_llm() -> None: args = LLMInferenceArgs.parse() data_id = args.data_id all_test_data = pd.read_parquet("test.parquet") all_train_data = pd.read_parquet("train.parquet") logger.info( "Starting inference with settings", data_id=data_id, model_name=args.model_name, num_return_seqs=args.num_predictions, max_new_tokens=args.max_new_tokens, temperature=args.temperature, num_shots=args.num_shots, num_train_samples=len(all_train_data), num_test_samples=len(all_test_data), ) checkpoint_dir = f"{args.model_name}_{data_id}_predictions_checkpoint" Path(checkpoint_dir).mkdir(parents=True, exist_ok=True) checkpointed_rules = set(os.listdir(checkpoint_dir)) grouped_by_rule = all_test_data.groupby(PredictionSchema.rule) for rule_id, rule_df in grouped_by_rule: if f"{rule_id}.parquet" in checkpointed_rules: logger.warning("Found cached results. Skipping", rule=rule_id) continue logger.info("running for", rule=rule_id) - 54 - 3924732.v36265.1015001 all_rule_predictions = [] for _, sample in tqdm.tqdm(rule_df.iterrows(), total=len(rule_df)): conversation = preprocessor.create_conversation_with_examples( query=sample, few_shot_examples=all_train_data, num_shots=args.num_shots ) inference_not_done = True max_tries = 3 current_tries = 0 while inference_not_done: try: input_data = { "model_input": { "temperature": args.temperature, "num_generations": args.num_predictions, "max_tokens": args.max_new_tokens, "context": "Assistant is a code assistant designed to fix issues in given code snippets.\nInstructions:\n-Do not generate additional text or code. Output only the fixed code snippet\n-Do not generate explanations, comments, notes. Note that the code we provide is incomplete, it is intentionally reduced to a smaller snippet, do not try to complete it in anyway. Leave evertything as it is and just apply the changes related to the fix.", "messages": conversation, }, "provider": "openai", "model": args.model_name, } response = requests.post( args.end_point, headers={"Authorization": "Bearer " + os.environ["TOKEN"]}, json=input_data, ) inference_not_done = response.status_code != 200 if inference_not_done: logger.warning( "A non 200 response occured", status=response.status_code ) raise RuntimeError("A non 200 response occured") except Exception as e: current_tries += 1 if current_tries == max_tries: logger.error("Maximum number of tries exceeded") break logger.warning("caught an exception, will try again soon", error=e) time.sleep(60) try: predictions: list[str] = [ - 55 - 3924732.v36265.1015001 message["content"] for message in response.json()["messages"] ] except Exception as e: logger.error("caught an exception in accessing predictions", error=e) predictions = [ "<FAILED_INFERENCE>" for _ in range(args.num_predictions) ] all_rule_predictions.append(predictions) # type ignore rule_df["predictions"] = all_rule_predictions rule_df.to_parquet(os.path.join(checkpoint_dir, f"{rule_id}.parquet")) logger.info("done with generating predictions") all_preds_df = pd.read_parquet(checkpoint_dir) logger.info("read predictions from checkpoints", num_samples=len(all_preds_df)) with tempfile.TemporaryDirectory(prefix="autofix_llm_predictions-") as save_dir: prediction_file_path = Path( os.path.join(save_dir, "predictions", "predictions.parquet") ) prediction_file_path.parent.mkdir(parents=True, exist_ok=True) all_preds_df.to_parquet( str(prediction_file_path), index=False, ) logger.info( "Done", data_id=data_id, model_id=args.model_artifact_id ) if __name__ == "__main__": predict_llm() - 56 - 3924732.v36265.1015001 APPENDIX G – “predict_llm.sh” #! / bin / bash if [[ -z "${END_POINT}" ]]; then echo "please provide an url to query in the END_POINT env variable" exit 1 fi if [[ -z "${MODEL_NAME}" ]]; then echo "please provide a model name in the MODEL_NAME env variable" exit 1 fi if [[ -z "${DATA_ID}" ]]; then echo "please provide the data id in the DATA_ID env variable" exit 1 fi if [[ -z "${NUM_PREDICTIONS}" ]]; then echo "please provide the number of predictions to generate in the NUM_PREDICTIONS env variable" exit 1 fi if [[ -z "${TEMPERATURE}" ]]; then echo "please provide the temperature in the TEMPERATURE env variable" exit 1 fi if [[ -z "${NUM_SHOTS}" ]]; then echo "please provide the num_shots in the NUM_SHOTS env variable" exit 1 fi if [[ -z "${MAX_NEW_TOKENS}" ]]; then echo "please provide the max_new_tokens in the MAX_NEW_TOKENS env variable" exit 1 fi dc_ai_fix_venv / bin / activate / python3 autofix / ml / bin / predict_llm.py \ --end_point="${END_POINT}" \ --model_name="${MODEL_NAME}" \ --data_id="${DATA_ID}" \ --num_predictions="${NUM_PREDICTIONS}" \ --temperature="${TEMPERATURE}" \ --num_shots="${NUM_SHOTS}" \ --- 57 - 3924732.v36265.1015001 APPENDIX H – “data_processor.py” from typing import List, Optional, TypeVar, Union import pandas as pd import tokenizers from pandera.typing import DataFrame from transformers import PreTrainedTokenizer from transformers.tokenization_utils_base import BatchEncoding from autofix.ml.lib.data_schemas import LabelledDataSchema from ml_utils.model_type import ModelType TDataFrame = TypeVar("TDataFrame", bound=pd.DataFrame) # These are special tokens that will be added to the tokenizer. _BEGIN_BUGGY_CODE = "<beginbuggycode>" _END_BUGGY_CODE = "<endbuggycode>" _BEGIN_FIXED_CODE = "<beginfixedcode>" _END_FIXED_CODE = "<endfixedcode>" def make_special_tokens() -> list[tokenizers.AddedToken]: return [ tokenizers.AddedToken(_BEGIN_BUGGY_CODE, normalized=False), tokenizers.AddedToken(_END_BUGGY_CODE, normalized=False), tokenizers.AddedToken(_BEGIN_FIXED_CODE, normalized=False), tokenizers.AddedToken(_END_FIXED_CODE, normalized=False), ] def make_prompt_for_seq2seq_model( rule_key: str, pre_code: str, ) -> str: return f"fix {rule_key}\n{_BEGIN_BUGGY_CODE}\n{pre_code}\n{_END_BUGGY_CODE}\n{_BEGIN_FI XED_CODE}\n" def wrap_fixed_code_for_seq2seq_model(post_code: str) -> str: return f"{post_code}{_END_FIXED_CODE}" def make_prompt_for_starcoder( tokenizer: PreTrainedTokenizer, rule_key: str, rule_message: str, pre_code: str, post_code: Optional[str] = None, ) -> str: prompt = f"Fix {rule_key}\n\n{rule_message}\n{_BEGIN_BUGGY_CODE}\n{pre_code}\n{_END_BUGGY _CODE}\n{_BEGIN_FIXED_CODE}\n" if post_code is not None: prompt += f"{post_code}{_END_FIXED_CODE}" return prompt def make_prompt_for_chat( tokenizer: PreTrainedTokenizer, rule_key: str, rule_message: str, pre_code: str, post_code: Optional[str] = None, ) -> str: chat = [ - 58 - 3924732.v36265.1015001 { "role": "user", "content": f"Fix the bug {rule_key}: {rule_message} in the code {pre_code}", }, ] if post_code is not None: chat.append({"role": "assistant", "content": f"{post_code}"}) return tokenizer.apply_chat_template( chat, tokenize=False, add_generation_prompt=True ) # type: ignore def make_prompt_for_causal_model( rule_key: str, rule_message: str, pre_code: str, tokenizer: PreTrainedTokenizer, post_code: Optional[str] = None, ) -> str: if tokenizer.chat_template is not None: return make_prompt_for_chat( tokenizer=tokenizer, rule_key=rule_key, rule_message=rule_message, pre_code=pre_code, post_code=post_code, ) else: return make_prompt_for_starcoder( tokenizer=tokenizer, rule_key=rule_key, rule_message=rule_message, pre_code=pre_code, post_code=post_code, ) def make_prompt_for_llm_user(rule_key: str, rule_message: str, pre_code: str) -> str: p = f"generate the fixed code for the bug {rule_key} with the error message {rule_message}\n{pre_code}\n" return p def create_conversation_with_examples( query, few_shot_examples: DataFrame[LabelledDataSchema], num_shots: int ) -> list[dict[str, str]]: shots = few_shot_examples[ few_shot_examples[LabelledDataSchema.rule] == query[LabelledDataSchema.rule] ] shots = shots.sample(n=num_shots, random_state=42) example_conversation = [] for _, example in shots.iterrows(): shot_messages = [ { "role": "USER", - 59 - 3924732.v36265.1015001 "content": make_prompt_for_llm_user( example[LabelledDataSchema.rule], example[LabelledDataSchema.message], example[LabelledDataSchema.pre_reduced], ), }, { "role": "ASSISTANT", "content": f"{example[LabelledDataSchema.post_reduced]}\n", }, ] example_conversation += shot_messages return example_conversation + [ { "role": "USER", "content": make_prompt_for_llm_user( query[LabelledDataSchema.rule], query[LabelledDataSchema.message], query[LabelledDataSchema.pre_reduced], ), } ] def make_prompts( batch: DataFrame[LabelledDataSchema], model_type: ModelType, tokenizer: PreTrainedTokenizer, include_post: bool, ) -> List[str]: if model_type == ModelType.SEQ2SEQ: return [ make_prompt_for_seq2seq_model(rule_key, pre_code) for rule_key, pre_code in zip( batch[LabelledDataSchema.rule], batch[LabelledDataSchema.pre_reduced], ) ] if model_type == ModelType.CAUSAL: return [ make_prompt_for_causal_model( rule_key=rule_key, rule_message=rule_message, pre_code=pre_code, tokenizer=tokenizer, post_code=post_code if include_post else None, ) for rule_key, rule_message, pre_code, post_code in zip( batch[LabelledDataSchema.rule], batch[LabelledDataSchema.message], batch[LabelledDataSchema.pre_reduced], batch[LabelledDataSchema.post_reduced], ) ] - 60 - 3924732.v36265.1015001 else: raise ValueError(f"{model_type} is not yet handled!") def tokenize_prompts( prompts: Union[str, List[str]], labels: Optional[Union[str, List[str]]], tokenizer: PreTrainedTokenizer, truncation: bool, max_length: Optional[int], model_type: ModelType, return_tensors: Optional[str] = None, ) -> BatchEncoding: if model_type == ModelType.SEQ2SEQ: if labels is not None: if isinstance(labels, str): labels = wrap_fixed_code_for_seq2seq_model(labels) else: labels = [wrap_fixed_code_for_seq2seq_model(label) for label in labels] return tokenizer( text=prompts, text_target=labels, # type: ignore truncation=truncation, max_length=max_length, padding=False, return_tensors=return_tensors, ) if model_type == ModelType.CAUSAL: return tokenizer( text=prompts, truncation=truncation, max_length=max_length, padding=False, return_tensors=return_tensors, ) else: raise ValueError(f"{model_type} is not yet handled!") - 61 - 3924732.v36265.1015001 APPENDIX I – “metrics_aggregator.py” from dataclasses import dataclass import pandas as pd from autofix.ml.lib.data_schemas import EvaluationSchema from autofix.ml.lib.metrics_impl import ExactMatchAtKAccuracy from autofix.ml.lib.metrics_impl import FloatMetric from autofix.ml.lib.metrics_impl import PassAtKAccuracy from autofix.ml.lib.metrics_impl import SampleCount @dataclass class AutofixMetricAggregator: """This class aggregates the metrics once they are in EvaluationSchema """ def aggregate_all_rules_metrics(self, predictions: pd.DataFrame) -> pd.DataFrame: """Takes the checked predictions, scores and computes various metrics based on them. Output `DataFrame` has a column `rule_id` and number of columns corresponding to metrics. """ metrics: list[FloatMetric] = [ PassAtKAccuracy(), ExactMatchAtKAccuracy(), SampleCount(), ] existing_rules = list(predictions[EvaluationSchema.rule].unique()) rule_to_predictions: dict[str, Any] = { # type: ignore rule_id: predictions[predictions[EvaluationSchema.rule] == rule_id] for rule_id in existing_rules } # Create an empty metrics dataframe, containing only `rule_id`s. Metric columns will be # appended on the right. result: pd.DataFrame = pd.DataFrame({"rule_id": existing_rules}) for m in metrics: metric_results = result.apply( lambda row: m.compute_metric_for_predictions( rule_to_predictions[row.rule_id] ), axis=1, ) result = pd.concat([result, metric_results], axis=1) percentage_cols = result.select_dtypes(include="number").columns.difference( ["samples"] ) result[percentage_cols] = result[percentage_cols].mul(100).round(2) return result - 62 - 3924732.v36265.1015001 APPENDIX J – “metrics_impl.py” """Metrics for the metric aggregator.""" from abc import ABCMeta from abc import abstractmethod from typing import List, Union, final import numpy as np import pandas as pd import pandera as pa from pandera.typing import DataFrame from autofix.ml.lib.data_schemas import EvaluationSchema class Metric(metaclass=ABCMeta): """Generic metric class.""" @abstractmethod def compute_metric_for_predictions( self, df: DataFrame[EvaluationSchema] ) -> pd.Series: """Takes evaluations and computes a metric. Args: df: `DataFrame` with the schema EvaluationSchema """ pass @abstractmethod def get_string_name(self) -> Union[str, List[str]]: pass class FloatMetric(Metric): @abstractmethod def compute_metric_for_predictions( self, df: DataFrame[EvaluationSchema] ) -> pd.Series: pass @final class PassAtKAccuracy(FloatMetric): """Average % of datapoints that have at least one "pass" in `k` predictions (`pass@k`).""" step_size = 2 def get_string_name(self) -> str: return "pass@k accuracy" @pa.check_types def compute_metric_for_predictions( self, df: DataFrame[EvaluationSchema] ) -> pd.Series: # `row["true_fix"]` is `List[bool]`` of length `k` containing if the `i`-th prediction has # fixed the problem. res = pd.Series(dtype="float64") assert df[EvaluationSchema.predictions] is not None num_predictions = len(df[EvaluationSchema.predictions].values[0]) # type: ignore for i in range(1, num_predictions, self.step_size): res[f"pass@{i}"] = np.float64( df.apply(lambda row: any(row["true_fix"][:i]), axis=1).sum() / len(df) - 63 - 3924732.v36265.1015001 ) res[f"pass@{num_predictions}"] = np.float64( df.apply(lambda row: any(row["true_fix"][:num_predictions]), axis=1).sum() / len(df) ) return res @final class ExactMatchAtKAccuracy(FloatMetric): """Evaluates exact match accuracy for k predictions.""" step_size = 2 def get_string_name(self) -> str: return "exact@k accuracy" @pa.check_types def compute_metric_for_predictions( self, df: DataFrame[EvaluationSchema] ) -> pd.Series: res = pd.Series(dtype="float64") assert df[EvaluationSchema.predictions] is not None num_predictions = len(df[EvaluationSchema.predictions].values[0]) # type: ignore for i in range(1, num_predictions, self.step_size): res[f"exact@{i}"] = np.float64( df.apply(lambda row: any(row["exact_match"][:i]), axis=1).sum() / len(df) ) res[f"exact@{num_predictions}"] = np.float64( df.apply( lambda row: any(row["exact_match"][:num_predictions]), axis=1 ).sum() / len(df) ) return res @final class SampleCount(FloatMetric): """Counts the number of predictions.""" def get_string_name(self) -> str: return "samples" @pa.check_types def compute_metric_for_predictions( self, df: DataFrame[EvaluationSchema] ) -> pd.Series: res = pd.Series(dtype="float64") assert df[EvaluationSchema.predictions] is not None res[self.get_string_name()] = len(df) return res - 64 - 3924732.v36265.1015001 APPENDIX K – “trainer.py” from typing import Union, cast import torch import torch.utils.data.dataset import transformers from datasets.arrow_dataset import Dataset from transformers import PreTrainedModel from transformers import PreTrainedTokenizer from transformers import PreTrainedTokenizerFast from transformers.tokenization_utils_base import PreTrainedTokenizerBase from autofix.ml.lib.args import AutofixArgs from ml_utils.distributed_training import MultiProcessLogger from ml_utils.model_type import ModelType logger = MultiProcessLogger() _TokenizerT = Union[ PreTrainedTokenizer, PreTrainedTokenizerFast, PreTrainedTokenizerBase ] class AutofixTrainer(transformers.Trainer): def __init__( self, model: PreTrainedModel, tokenizer: _TokenizerT, args: AutofixArgs, train_ds: Dataset, val_ds: Dataset, ): model_type = ModelType.infer_from_model(model) if model_type == ModelType.CAUSAL: data_collator = transformers.DataCollatorForLanguageModeling( tokenizer=tokenizer, mlm=False, pad_to_multiple_of=8, return_tensors="pt", ) elif model_type == ModelType.SEQ2SEQ: data_collator = transformers.DataCollatorForSeq2Seq( tokenizer=tokenizer, model=model, padding="longest", pad_to_multiple_of=8, max_length=args.model_args.tokenizer_max_length, return_tensors="pt", ) else: raise ValueError(f"unsupported ModelType: {model_type}") super().__init__( model=model, train_dataset=cast(torch.utils.data.dataset.Dataset, train_ds), eval_dataset=cast(torch.utils.data.dataset.Dataset, val_ds), tokenizer=tokenizer, args=args.training_args, - 65 - 3924732.v36265.1015001 data_collator=data_collator, ) - 66 - 3924732.v36265.1015001 APPENDIX L – “predict_autofix.py” import os from typing import cast import pandas as pd import structlog import torch import tqdm import transformers from datasets.arrow_dataset import Dataset from pandera.typing import DataFrame from transformers.pipelines.pt_utils import KeyDataset import autofix.ml.lib.args as arguments import autofix.ml.lib.data_processor as data_processor import autofix.ml.utils.loading as ld from autofix.ml.lib.data_schemas import LabelledDataSchema from autofix.ml.lib.data_schemas import PredictionSchema from autofix.ml_inference.inference_pipeline import make_generation_config from autofix.ml_inference.inference_pipeline import make_inference_pipeline from ml_utils.model_type import ModelType logger = structlog.get_logger() transformers.logging.set_verbosity_info() def prediction_worker( args: arguments.AutofixArgs, data_id: str, single_test_data: DataFrame[LabelledDataSchema], pipeline: transformers.Pipeline, model: transformers.PreTrainedModel, tokenizer: transformers.PreTrainedTokenizerBase, ) -> list[str]: # Step 1 - preprocess the dataset test_ds = Dataset.from_pandas(single_test_data) logger.info( f"Preprocessing the test dataset", num_samples=len(single_test_data), data_id=data_id, ) PROMPT_COLUMN_NAME = "prompt" model_type = ModelType.infer_from_model(model) def process_batch(batch: DataFrame[LabelledDataSchema]) -> dict[str, list[str]]: return { PROMPT_COLUMN_NAME: data_processor.make_prompts( batch, model_type, tokenizer, include_post=False # type: ignore ) } test_ds = test_ds.map( process_batch, remove_columns=test_ds.column_names, keep_in_memory=args.dataset_args.keep_in_memory, batched=True, - 67 - 3924732.v36265.1015001 batch_size=args.dataset_args.preprocessing_batch_size, desc=f"preprocessing the test dataset {data_id}.", ) # Step 2 - generate the model predictions. logger.info( "Generating predictions...", num_samples=len(test_ds), maxlen=args.model_args.tokenizer_max_length, num_beams=args.inference_args.beam_size, batch_size=args.training_args.per_device_eval_batch_size, ) test_ds_wrapper = KeyDataset( cast(torch.utils.data.Dataset, test_ds), key=PROMPT_COLUMN_NAME ) pipeline_iterator = pipeline( test_ds_wrapper, num_workers=args.training_args.dataloader_num_workers, batch_size=args.training_args.per_device_eval_batch_size, **make_generation_config(args.inference_args, tokenizer), ) # `predictions` is a list (with size `num_samples`) of lists (with size `num_return_seqs`) that contains # the beam search predictions for every sample. predictions = [ [out["generated_text"] for out in cast(list[dict], predictions_for_one_sample)] # `predictions_for_one_sample` is a list that contains the `num_return_seqs` predictions for a sample. for predictions_for_one_sample in tqdm.tqdm( pipeline_iterator, desc=f"Generating predictions {data_id}.", ) ] return predictions # type: ignore def predict_autofix() -> None: # Step 0 - setup. args = arguments.parse_autofix_args_from_cmd() if not args.model_args.model_id: raise RuntimeError("model_id can not be empty or None.") os.makedirs(args.training_args.output_dir, exist_ok=True) device = ( torch.device(f"cuda:0") if torch.cuda.is_available() else torch.device("cpu") ) # Step 1 - fetch the data. single_test_data = DataFrame[LabelledDataSchema](pd.read_parquet(args.dataset_args.data_id)) logger.info("Finished fetching data") # Step 2 - load the model and the tokenizer. torch_dtype = torch.float32 if args.training_args.bf16 or args.training_args.bf16_full_eval: torch_dtype = torch.bfloat16 - 68 - 3924732.v36265.1015001 elif args.training_args.fp16 or args.training_args.fp16_full_eval: torch_dtype = torch.float16 logger.info( "Loading the model", torch_dtype=torch_dtype, device=device ) model_path = str(args.model_args.model_id) if not args.model_args.model_id: raise RuntimeError("model_id can not be empty or None.") tokenizer = ld.load_suitable_tokenizer(model_path) model_loading_kwargs: dict = { "torch_dtype": torch_dtype, "device_map": str(device), "use_flash_attention_2": args.model_args.use_flash_attention_2, } if args.model_args.use_bits_and_bytes: bnb_config = transformers.BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model_loading_kwargs.update({"quantization_config": bnb_config}) model = ld.load_suitable_model( model_path, **model_loading_kwargs ) gpu_memory_used = ( torch.cuda.max_memory_allocated(device=device) if torch.cuda.is_available() else 0 ) logger.info( "Done loading model", gpu_memory_used=gpu_memory_used ) USE_TRUNCATION = True MAX_INPUT_SIZE = args.model_args.tokenizer_max_length logger.info( "Creating the inference pipeline", ) pipeline =model, tokenizer, truncate_input=USE_TRUNCATION, )- predictions = prediction_worker( args, args.dataset_args.data_id, single_test_data, pipeline, model, tokenizer # type: ignore ) # Step 4 - add the "predictions" column and save the final parquet. - 69 - 3924732.v36265.1015001 assert ( predictions is not None ), "Predictions in main process must be returned." single_test_data[PredictionSchema.predictions] = predictions single_test_data.to_parquet( "predictions.parquet", index=False, ) logger.info("Done generating predictions") if __name__ == "__main__": predict_autofix() - 70 - 3924732.v36265.1015001 APPENDIX M – “train_autofix.py” import os import tempfile from typing import cast import pandas as pd import peft import torch import transformers from sklearn.model_selection import train_test_split from datasets.arrow_dataset import Dataset from pandera.typing import DataFrame from peft.utils.other import prepare_model_for_kbit_training from tokenizers import AddedToken from transformers import PreTrainedModel from transformers import PreTrainedTokenizerBase from transformers import T5TokenizerFast from transformers.tokenization_utils_base import BatchEncoding import autofix.ml.lib.args as arguments import autofix.ml.lib.data_processor as data_processor import autofix.ml.utils.loading as ld from autofix.ml.lib.data_schemas import LabelledDataSchema from autofix.ml.lib.trainer import AutofixTrainer from ml_utils.distributed_training import MultiProcessLogger from ml_utils.distributed_training import is_main_process from ml_utils.model_type import ModelType logger = MultiProcessLogger() transformers.logging.set_verbosity_info() def _patch_token_ids(model: PreTrainedModel, tokenizer: PreTrainedTokenizerBase): if isinstance(tokenizer, T5TokenizerFast): logger.info( "Adding code-related tokens to the T5 tokenizer", model_name_or_path=model.config.name_or_path, ) tokenizer.add_tokens(["{", "}", "<", ">", "\\", "^", "`", "~"]) tokenizer.add_tokens(AddedToken("\n", normalized=False)) # type: ignore tokenizer.add_tokens(AddedToken("\t", normalized=False)) # type: ignore tokenizer.add_tokens(AddedToken("\r\n", normalized=False)) # type: ignore model.resize_token_embeddings(len(tokenizer)) if model.config.model_type in [ "gpt2", "gpt_bigcode", "mosaic_gpt", "llama", "mpt", "mixtral", "stablelm_epoch", "starcoder2", ]: - 71 - 3924732.v36265.1015001 # These models (e.g. GPT2) don't use padding tokens. To make it work with the # padding data collators, we set the padding token to unknown token, if defined, or else to the end of sequence token. if tokenizer.pad_token_id is None: if tokenizer.unk_token_id is not None: tokenizer.pad_token = tokenizer.unk_token tokenizer.pad_token_id = tokenizer.unk_token_id else: tokenizer.pad_token = tokenizer.eos_token tokenizer.pad_token_id = tokenizer.eos_token_id model.pad_token_id = tokenizer.pad_token_id # type: ignore model.config.pad_token_id = tokenizer.pad_token_id logger.info( "Setting the pad_token_id for a decoder model", new_pad_token_id=tokenizer.pad_token_id, new_pad_token=tokenizer.pad_token, model_pad_token_id=model.pad_token_id, model_config_pad_token_id=model.config.pad_token_id, model_name_or_path=model.config.name_or_path, ) assert tokenizer.pad_token_id is not None assert ( model.pad_token_id == tokenizer.pad_token_id ), "The model's pad token id and the tokenizer's pad token id are not the same" # Handle the chat template. if tokenizer.chat_template is None: old_tokenizer_len = len(tokenizer) tokenizer.add_tokens(data_processor.make_special_tokens(), special_tokens=True) # type: ignore logger.info( "Adding new special tokens", num_new_special_tokens=len(tokenizer) - old_tokenizer_len, new_num_tokens=len(tokenizer), ) model.resize_token_embeddings(len(tokenizer)) def train_autofix() -> None: # Step 0 - setup. args: arguments.AutofixArgs = arguments.parse_autofix_args_from_cmd() transformers.trainer_utils.set_seed(args.training_args.seed) logger.info("Set transformers seed.", seed=args.training_args.seed) if not args.model_args.model_id: raise RuntimeError("model_id can not be empty or None.") best_model_path = tempfile.TemporaryDirectory() # Step 1 - fetch the data. train_size = args.dataset_args.train_size with args.training_args.main_process_first(): train_df= pd.read_parquet(args.dataset_args.data_id) stratify = train_df[LabelledDataSchema.rule] n_unique_rules = train_df[LabelledDataSchema.rule].nunique() if train_size < 1 and len(train_df) * (1.0 - train_size) < n_unique_rules: - 72 - 3924732.v36265.1015001 stratify = None logger.warning( "Dataset is too sparse to split across rules. Use random split instead.", data_size=len(train_df), unique_rules=n_unique_rules, ) train_df, val_df = train_test_split( train_df, train_size=train_size, random_state=args.dataset_args.dataset_seed, shuffle=True, stratify=stratify, ) logger.info( "Read and split the data successfully", train_size=len(train_df), val_size=len(val_df), ) train_ds = Dataset.from_pandas(train_df) logger.info("Training size", num=len(train_ds)) val_ds = Dataset.from_pandas(val_df) logger.info("Validation size", num=len(val_ds)) # Step 2 - fetch the model and the tokenizer model_loading_kwargs: dict = { "torch_dtype": torch.bfloat16, "use_flash_attention_2": args.model_args.use_flash_attention_2, } if args.model_args.use_bits_and_bytes: bnb_config = transformers.BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model_loading_kwargs.update({"quantization_config": bnb_config}) if args.training_args.deepspeed is None: logger.info("Training without deepspeed") model_loading_kwargs.update({"device_map": "auto"}) model = ld.load_suitable_model( model_identifier=args.model_args.model_name, **model_loading_kwargs, ) if args.model_args.use_peft: model = prepare_model_for_kbit_training( model, use_gradient_checkpointing=args.training_args.gradient_checkpointing ) lora_config = peft.tuners.lora.LoraConfig( r=16, lora_alpha=16, target_modules=[ "self_attn.q_proj", - 73 - 3924732.v36265.1015001 "self_attn.k_proj", "self_attn.v_proj", "self_attn.o_proj", "lm_head", ], lora_dropout=0.05, bias="all", task_type="CAUSAL_LM", inference_mode=False, ) model = peft.mapping.get_peft_model(model, lora_config) logger.info( "Created peft the model", config=lora_config ) model.print_trainable_parameters() tokenizer = ld.load_suitable_tokenizer(args.model_args.model_name) _patch_token_ids(model, tokenizer) # type: ignore # Step 3 - preprocess the datasets. model_type = ModelType.infer_from_model(model) # type: ignore def process_batch(batch: DataFrame[LabelledDataSchema]) -> BatchEncoding: prompts = data_processor.make_prompts( batch, model_type, tokenizer, include_post=True ) return data_processor.tokenize_prompts( prompts=prompts, labels=cast(list[str], batch[LabelledDataSchema.post_reduced]), tokenizer=tokenizer, truncation=True, max_length=args.model_args.tokenizer_max_length, model_type=model_type, ) def process_dataset(ds: Dataset, ds_name: str) -> Dataset: return ds.map( process_batch, batched=True, batch_size=args.dataset_args.preprocessing_batch_size, keep_in_memory=args.dataset_args.keep_in_memory, desc=f"Processing the {ds_name} dataset", remove_columns=ds.column_names, ) train_ds = process_dataset(train_ds, "training") val_ds = process_dataset(val_ds, "validation") # Step 4 - train the model. logger.info("Starting with autofix training...") trainer = AutofixTrainer( model=model, # type: ignore tokenizer=tokenizer, args=args, train_ds=train_ds, val_ds=val_ds, - 74 - 3924732.v36265.1015001 ) trainer.train() # Step 5 - save the trained model and tokenizer logger.info("saving the best model", best_model_path=best_model_path) trainer.save_model(str(best_model_path)) if trainer.is_deepspeed_enabled: # If we're training with deepspeed and ZERO-3, then we additionally need to convert # the bf16 / fp16 weights from deepspeed's checkpoints into a fp32 torch state_dict, # that we can later use for inference. # `trainer.deepspeed` is actually the model object (wrapped into a deepspeed wrapper) assert trainer.deepspeed is not None # This saves the current deepspeed checkpoint into `best_checkpoint_dir`. # If `--load_best_model_at_end` is specified, this checkpoint will correspond # to the best checkpoint. Otherwise, it will correspond to the current state of the model. checkpoint_tag = "export-checkpoint" logger.info( "saving the best deepspeed checkpoint", path=args.training_args.output_dir, tag=checkpoint_tag, ) trainer.deepspeed.save_checkpoint(args.training_args.output_dir, tag=checkpoint_tag) # type: ignore if is_main_process(): out_path = os.path.join(str(best_model_path), "pytorch_model.bin") logger.info("converting the best weights to fp32", out_path=out_path) from deepspeed.utils.zero_to_fp32 import ( convert_zero_checkpoint_to_fp32_state_dict, ) convert_zero_checkpoint_to_fp32_state_dict( args.training_args.output_dir, output_file=out_path, tag=checkpoint_tag, ) if __name__ == "__main__": train_autofix() - 75 - 3924732.v36265.1015001 APPENDIX N – “deepspeed_config.json” { "optimizer": { "type": "AdamW", "params": { "lr": "auto", "betas": "auto", "eps": "auto", "weight_decay": "auto" } }, "scheduler": { "type": "WarmupDecayLR", "params": { "warmup_min_lr": 0, "warmup_max_lr": "auto", "warmup_num_steps": "auto", "warmup_type": "linear", "total_num_steps": "auto" } }, "zero_optimization": { "stage": 3, "offload_optimizer": { "device": "cpu", "pin_memory": true }, "overlap_comm": true, "reduce_bucket_size": 1e8, "contiguous_gradients": true, "stage3_gather_16bit_weights_on_model_save": false }, "bf16": { "enabled": "auto" }, "gradient_accumulation_steps": "auto", "gradient_clipping": "auto", "steps_per_print": 2000, "train_batch_size": "auto", "train_micro_batch_size_per_gpu": "auto", "wall_clock_breakdown": false } - 76 - 3924732.v36265.1015001 APPENDIX O – “loading.py” import os from typing import Any import transformers from ml_utils.distributed_training import MultiProcessLogger logger = MultiProcessLogger() def load_suitable_model( model_identifier: str, **from_pretrained_kwargs, ) -> Any: # DO NOT CHANGE THIS TYPE ANNOTATION """Loads the model depending on model model_identifier. Args: model_identifier: either local model directory or an HuggingFace ID from_pretrained_kwargs: all the keyword arguments that transformers.from_pretrained accepts can be passed here. """ loaders = [ transformers.AutoModelForCausalLM, transformers.AutoModelForSeq2SeqLM, ] for loader in loaders: try: model = loader.from_pretrained( model_identifier, trust_remote_code=True, token=os.environ.get("HUGGING_FACE_HUB_TOKEN"), **from_pretrained_kwargs, ) logger.info( "Loaded the model", model_id=model_identifier, dtype=model.dtype, model_type=loader.__name__, ) return model except Exception as e: logger.warning( "Failed to load the model. Will try the next loader", failed_loader=loader.__name__, error=e, ) raise ValueError("Failed loading the model") def load_suitable_tokenizer( model_identifier: str, ) -> transformers.PreTrainedTokenizerBase: """Loads the tokenizer depending on model_identifier. Args: model_identifier: can be either local model path directory or an HuggingFace ID """ tokenizer = transformers.AutoTokenizer.from_pretrained( - 77 - 3924732.v36265.1015001 model_identifier, trust_remote_code=True ) logger.info( "Loaded the tokenizer.", path=model_identifier, model_type=type(tokenizer), num_tokens=len(tokenizer), ) return tokenizer - 78 - 3924732.v36265.1015001 APPENDIX P – “inference_pipeline.py” from typing import Any, Optional, Union import torch import torch.utils.data.dataset import transformers from transformers import PreTrainedModel from transformers import PreTrainedTokenizer from transformers import PreTrainedTokenizerFast from transformers.tokenization_utils_base import PreTrainedTokenizerBase import autofix.ml.lib.data_processor as data_processor from autofix.ml.lib.args import InferenceArgs from ml_utils.distributed_training import MultiProcessLogger from ml_utils.model_type import ModelType logger = MultiProcessLogger() _TokenizerT = Union[ PreTrainedTokenizer, PreTrainedTokenizerFast, PreTrainedTokenizerBase ] class _AutofixCausalInferencePipeline(transformers.TextGenerationPipeline): def __init__( self, truncate_input: bool, max_input_size: Optional[int], log_metrics: bool, *args, **kwargs, ): super().__init__(*args, **kwargs) self.truncate_input = truncate_input self.max_input_size = max_input_size self.log_metrics = log_metrics def preprocess(self, inputs, *args, **kwargs): # `self.tokenizer` is initialized in the base class constructor. assert self.tokenizer is not None tokenized_input = data_processor.tokenize_prompts( inputs, labels=None, tokenizer=self.tokenizer, truncation=self.truncate_input, max_length=self.max_input_size, model_type=ModelType.CAUSAL, return_tensors="pt", ) # Setting this is required to make the rest of `transformers.TextGenerationPipeline` to work. tokenized_input["prompt_text"] = inputs return tokenized_input def __call__(self, *args, **kwargs): # `return_full_text=True` makes the pipeline return only newly generated text # (instead of the input + newly generated text). kwargs["return_full_text"] = False - 79 - 3924732.v36265.1015001 return super().__call__(*args, **kwargs) class _AutofixSeq2SeqPipeline(transformers.Text2TextGenerationPipeline): def __init__( self, truncate_input: bool, max_input_size: Optional[int], log_metrics: bool, *args, **kwargs, ): super().__init__(*args, **kwargs) self.truncate_input = truncate_input self.max_input_size = max_input_size self.log_metrics = log_metrics def preprocess(self, inputs, *args, **kwargs): # `self.tokenizer` is initialized in the base class constructor. assert self.tokenizer is not None tokenized_input = data_processor.tokenize_prompts( inputs, labels=None, tokenizer=self.tokenizer, truncation=self.truncate_input, max_length=self.max_input_size, model_type=ModelType.SEQ2SEQ, # the rest of `TextGenerationPipeline` expects tensors, not lists. return_tensors="pt", ) return tokenized_input def make_inference_pipeline( model: PreTrainedModel, tokenizer: _TokenizerT, truncate_input: bool, max_input_size: int, log_metrics: bool = False, device: Optional[torch.device] = None, ) -> transformers.Pipeline: model_type = ModelType.infer_from_model(model) model.eval() if model_type == ModelType.SEQ2SEQ: return _AutofixSeq2SeqPipeline( tokenizer=tokenizer, model=model, truncate_input=truncate_input, max_input_size=max_input_size, log_metrics=log_metrics, ) if model_type == ModelType.CAUSAL: if tokenizer.padding_side == "right": logger.warning( "switching the tokenizer padding side from 'right' to 'left' for a causal LM" ) - 80 - 3924732.v36265.1015001 tokenizer.padding_side = "left" kwargs = {} if device is not None: kwargs["device"] = device return _AutofixCausalInferencePipeline( tokenizer=tokenizer, model=model, truncate_input=truncate_input, max_input_size=max_input_size, log_metrics=log_metrics, **kwargs, ) else: raise ValueError(f"unsupported ModelType: {model_type}") def make_generation_config( inference_args: InferenceArgs, tokenizer: _TokenizerT ) -> dict[str, Any]: if inference_args.max_new_tokens is None: raise ValueError( "please set the --max_new_tokens argument to limit the max output size" ) eos_token_ids = [tokenizer.eos_token_id] if tokenizer.chat_template is None: # Sanity check and extend eos_token_ids end_of_fixed_code_str: str = data_processor._END_FIXED_CODE end_of_fixed_code_token_id: int = tokenizer.convert_tokens_to_ids(end_of_fixed_code_str) # type: ignore if end_of_fixed_code_str != tokenizer.convert_ids_to_tokens(end_of_fixed_code_token_id): # type: ignore # A sanity check that `end_of_fixed_code_str` is used as an independent token in the tokenizer. raise Exception( f"the '{end_of_fixed_code_str}' is not present in the tokenizer's vocabulary (the tokenizer returned token_id={end_of_fixed_code_token_id} instead" ) eos_token_ids = [tokenizer.eos_token_id, end_of_fixed_code_token_id] if tokenizer.pad_token_id is None: if tokenizer.unk_token_id is not None: tokenizer.pad_token_id = tokenizer.unk_token_id else: tokenizer.pad_token_id = tokenizer.eos_token_id return { "num_return_sequences": inference_args.num_return_seqs, "early_stopping": inference_args.early_stopping, # We set the `max_length` to None explicitly to avoid the noisy HF warnings. "max_length": None, "max_new_tokens": inference_args.max_new_tokens, "num_beams": inference_args.beam_size, - 81 - 3924732.v36265.1015001 # Setting the padding token id is needed to silence the warnings that HF throws at every call to `predict`. "pad_token_id": tokenizer.pad_token_id, # We use `end_of_fixed_code_token_id` as an "end of sequence" token to stop predicting at `end_of_fixed_code_token_id`. "eos_token_id": eos_token_ids, } - 82 - 3924732.v36265.1015001 APPENDIX Q – “distributed_training.py” import os import structlog class ConditionalLogger: def __init__(self, should_log: bool): self.logger = structlog.get_logger() self.should_log = should_log def debug(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.debug(msg, *posargs, **kwargs) def info(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.info(msg, *posargs, **kwargs) def warning(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.warning(msg, *posargs, **kwargs) def error(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.error(msg, *posargs, **kwargs) def fatal(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.fatal(msg, *posargs, **kwargs) def exception(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.exception(msg, *posargs, **kwargs) def critical(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.critical(msg, *posargs, **kwargs) def msg(self, msg: str, *posargs, **kwargs) -> None: if self.should_log: self.logger.msg(msg, *posargs, **kwargs) class MultiProcessLogger(ConditionalLogger): def __init__(self): super().__init__(should_log=is_main_process()) def is_multiprocessing() -> bool: # Both torch DDP and deepspeed set this env variable to the local rank of the process. # If the variable is not set, then we conclude that we're not in multiprocessing environment, # and hence the process is the main (and only) process. return os.environ.get("LOCAL_RANK") is not None def get_process_local_rank() -> int: return int(os.environ.get("LOCAL_RANK", 0)) def is_main_process() -> bool: return get_process_local_rank() == 0 - 83 - 3924732.v36265.1015001 APPENDIX R – “model_type.py” from enum import Enum import transformers class ModelType(Enum): CAUSAL = "causal" SEQ2SEQ = "seq2seq" @staticmethod def infer_from_model(model: transformers.PreTrainedModel) -> "ModelType": t: str = model.config.model_type if t in ["t5", "codet5p"]: return ModelType.SEQ2SEQ if t in [ "gpt2", "gpt_bigcode", "mosaic_gpt", "llama", "mpt", "mixtral", "stablelm_epoch", "starcoder2", ]: return ModelType.CAUSAL raise ValueError(f"unknown model type for the '{t}' architecture") - 84 - 3924732.v36265.1015001 APPENDIX S – “train_example.json” {"rule":"AmbiguousConditional","message":"Ambiguous conditional expression: is the sum intended as a condition? Consider adding parentheses to improve readability.","line_number":499,"line_end":505,"col_begin":16,"col_end":18 ,"severity":0,"event_id":4209,"pre_file":"var C = require('..\ / constants.js'),\n engine = require('engine'),\n Element = require('.\ / element.js'),\n Clip = Element,\n Brush = require('..\ / graphics\ / brush.js'),\n provideEvents = require('..\ / events.js').provideEvents,\n AnimationError = require('..\ / errors.js').AnimationError,\n Errors = require('..\ / loc.js').Errors,\n ResMan = require('..\ / resource_manager.js'),\n FontDetector = require('..\ / ..\ / vendor\ / font_detector.js'),\n utils = require('..\ / utils.js'),\n is = utils.is,\n iter = utils.iter;\n\n\n\ / * X_ERROR, X_FOCUS, X_RESIZE, X_SELECT, touch events *\ / \n\nvar DOM_TO_EVT_MAP = {\n 'mouseup': C.X_MUP,\n 'mousedown': C.X_MDOWN,\n 'mousemove': C.X_MMOVE,\n 'mouseover': C.X_MOVER,\n 'mouseout': C.X_MOUT,\n 'click': C.X_MCLICK,\n 'dblclick': C.X_MDCLICK,\n 'keyup': C.X_KUP,\n 'keydown': C.X_KDOWN,\n 'keypress': C.X_KPRESS\n};\n\n\ / \ / Animation\n\ / \ / ---------------------- ------------------------------------------------ anm.Animation\n *\n * Create an Animation.\n *\n tree, an id-to-element map, background fill,It also may render itself to any context with {@link anm.Animation#render}\n * method.\n *\n * @constructor\n *\ / \nfunction Animation() {\n this.id = utils.guid();\n this.tree = [];\n this.hash = {};\n this.name = '';\n this.duration = undefined;\n this.bgfill = null;\n this.width = undefined;\n this.height = undefined;\n this.zoom = 1.0;\n this.speed = 1.0;\n this.repeat = false;\n this.meta = {};\n \ / \ / this.fps = undefined;\n this.__informEnabled = true;\n this._laters = [];\n this._initHandlers(); \ / \ / TODO: make automatic\n}\n\nAnimation.DEFAULT_DURATION = 10;\n\n\ / \ / mouse\ / keyboard events are assigned in L.loadAnimation\n\ / * TODO: move them into animation *\ / \nprovideEvents(Animation, [ C.X_MCLICK, C.X_MDCLICK, C.X_MUP,anm.Element());`\n * * `anim.add([new anm.Element(), new anm.Element()]);`\n * * `anim.add(function(ctx) {...}, function(t) { ... });`\n * * `anim.add(function(ctx) {...}, function(t) { ... },\n * function(ctx, prev(ctx)) { ... });`\n *\n * @param {anm.Element|anm.Clip|Array[Element]} subject Any number of Elements to add\n *\n * @return {anm.Element} The Element was appended.\n *\n *\ / \nAnimation.prototype.add = function(arg1, arg2, arg3) {\n \ / \ / this method only adds an element to a top-level\n \ / \ / FIXME: allow to add - 85 - 3924732.v36265.1015001 elements deeper or rename this\n \ / \ / method to avoid confusion?\n if (arg2) { \ / \ / element by functions mode\n var elm = new Element(arg1, arg2);\n if (arg3) elm.changeTransform(arg3);\n this.addToTree(elm);\n \ / \ / return elm;\n } else if (is.arr(arg1)) { \ / \ / elements array mode\n var clip = new Clip();\n clip.add(arg1);\n this.addToTree(_clip);\n \ / \ / return clip;\n } else { \ / \ / element object mode\n this.addToTree(arg1);\n }\n return this;\n}\n\ / * addS allowed to add static element before, such as image, may be return it in some form? *\ / \n\ / **\n * @method remove\n * @chainable\n *\n * Remove (unregister) element from this animation.\n *\n * @param {anm.Element} element\n *\ / \nAnimation.prototype.remove = function(elm) {\n \ / \ / error will be thrown in _unregister method\n \ / \ / if (!this.hash[elm.id]) throw new AnimErr(Errors.A.ELEMENT_IS_NOT_REGISTERED);\n if (elm.parent) {\n \ / \ / it will unregister element inside\n elm.parent.remove(elm);\n } else {\n this._unregister(elm);\n }\n return this;\n}\n\ / \ / > Animation.prototype.clear % ()\n\ / * Animation.prototype.clear = function() {\n this.hash = {};\n this.tree = [];\n this.duration = 0;\n var hash = this.hash;\n this.hash = {};\n for (var elmId in hash) {\n hash[elm.id]._unbind(); \ / \ / unsafe, because calls unregistering\n }\n} *\ / \n\ / **\n * @method traverse\n * @chainable\n *\n * Visit every element in a tree, no matter how deep it is.\n *\n * @param {Function} visitor\n * @param {anm.Element} visitor.element\n * @param {Object} [data]\n *\ / \n\ / \ / visitElems\nAnimation.prototype.traverse = function(visitor, data) {\n for (var elmId in this.hash) {\n visitor(this.hash[elmId], data);\n }\n return this;\n}\n\ / **\n * @method each\n * @chainable\n *\n * Visit every root element (direct Animation child) in a tree.\n *\n * @param {Function} visitor\n * @param {anm.Element} visitor.child\n * @param {Object} [data]\n *\ / \nAnimation.prototype.each = function(visitor, data) {\n for (var i = 0, tlen = this.tree.length; i < tlen; i++) {\n visitor(this.tree[i], data);\n }\n return this;\n}\n\ / **\n * @method iter\n * @chainable\n *\n * Iterate through every root (direct Animation child) element in a tree.\n *\n * @param {Function} iterator\n * @param {anm.Element} iterator.child\n * @param {Boolean} iterator.return `false`, if this element should be removed\n *\ / \nAnimation.prototype.iter = function(func, rfunc) {\n iter(this.tree).each(func, rfunc);\n return this;\n}\n\ / **\n * @method render\n *\n * Render the Animation for given context at given time.\n *\n * @param {Canvas2DContext} context\n * @param {Number} time\n * @param {Number} [dt] The difference in time between current frame and previous one\n *\ / \nAnimation.prototype.render = function(ctx, time, dt) {\n ctx.save();\n var zoom = this.zoom;\n try {\n if (zoom != 1) {\n ctx.scale(zoom, zoom);\n }\n if (this.bgfill) {\n if (!this.bgfill instanceof Brush) this.bgfill = Brush.fill(this.bgfill);\n ctx.fillStyle = this.bgfill.apply(ctx);\n ctx.fillRect(0, 0, this.width, this.height);\n }\n this.each(function(child) {\n child.render(ctx, time, dt);\n });\n } finally { ctx.restore(); }\n this.fire(C.X_DRAW,ctx);\n}\nAnimation.prototype.handle__x = function(type, evt) {\n this.traverse(function(elm) {\n elm.fire(type, evt);\n });\n return true;\n}\n\ / \ / TODO: test\n\ / **\n * @method getFittingDuration\n *\n * Get the duration where - 86 - 3924732.v36265.1015001 all child elements' bands fit.\n *\n * @return {Number} The calculated duration\n *\ / \nAnimation.prototype.getFittingDuration = function() {\n var max_pos = -Infinity;\n var me = this;\n this.each(function(child) {\n var elm_tpos = child._max_tpos();\n if (elm_tpos > max_pos) max_pos = elm_tpos;\n });\n return max_pos;\n}\n\ / **\n * @method reset\n * @chainable\n *\n * Reset all elements.\n =();\n });\n return this;\n}\n\ / **\n * @method dispose\n * @chainable\n *\n * Remove every possible allocated data to either never use this animation again or\n * start using it from scratch as if it never was used before.\n *\ / \nAnimation.prototype.dispose = function() {\n this.disposeHandlers();\n var me = this;\n \ / * FIXME: unregistering removes from tree, ensure it is safe *\ / \n this.iter(function(child) {\n me._unregister_no_rm(child);\n child.dispose();\n return false;\n });\n return this;\n}\n\ / **\n * @method isEmpty\n *\n * Does Animation has any Elements inside.\n *\n * @return {Boolean} `true` if no Elements, `false` if there are some.\n *\ / \nAnimation.prototype.isEmpty = function() {\n return this.tree.length == 0;\n}\n\ / **\n * @method toString\n *\n * Get a pretty description of this Animation\n *\n * @return {String} pretty string\n *\ / \nAnimation.prototype.toString = function() {\n return \"[ Animation \"+(this.name ? \"'\"+this.name+\"'\" : \"\")+\"]\";\n}\n\ / **\n * @method subscribeEvents\n * @private\n *\n * @param {Canvas} canvas\n *\ / \nAnimation.prototype.subscribeEvents = function(canvas) {\n engine.subscribeAnimationToEvents(canvas, this, DOM_TO_EVT_MAP);\n}\n\ / **\n * @method unsubscribeEvents\n * @private\n *\n * @param {Canvas} canvas\n *\ / \nAnimation.prototype.unsubscribeEvents = function(canvas) {\n engine.unsubscribeAnimationFromEvents(canvas, this);\n}\n\ / **\n * @method addToTree\n * @private\n *\n * @param {anm.Element} element\n *\ / \nAnimation.prototype.addToTree = function(elm) {\n if (!elm.children) {\n throw new AnimationError('It appears that it is not a clip object or element that you pass');\n }\n this._register(elm);\n \ / *if (elm.children) this._addElems(elm.children);*\ / \n this.tree.push(elm);\n}\n\ / *Animation.prototype._addElems = function(elems) {\n for (var ei = 0; ei < elems.length; ei++) {\n var _elm = elems[ei];\n this._register(_elm);\n }\n}*\ / \nAnimation.prototype._register = function(elm) {\n if (this.hash[elm.id]) throw new AnimationError(Errors.A.ELEMENT_IS_REGISTERED);\n elm.registered = true;\n elm.anim = this;\n this.hash[elm.id] = elm;\n var me = this;\n elm.each(function(child) {\n me._register(child);\n });\n}\nAnimation.prototype._unregister_no_rm = function(elm) {\n this._unregister(elm, true);\n}\nAnimation.prototype._unregister = function(elm, save_in_tree) { \ / \ / save_in_tree is optional and false by default\n if (!elm.registered) throw new AnimationError(Errors.A.ELEMENT_IS_NOT_REGISTERED);\n var me = this;\n elm.each(function(child) {\n me._unregister(child);\n });\n var pos = -1;\n if (!save_in_tree) {\n while ((pos = this.tree.indexOf(elm)) >= 0) {\n this.tree.splice(pos, 1); \ / \ / FIXME: why it does not goes deeply in the tree?\n }\n }\n - 87 - 3924732.v36265.1015001 delete this.hash[elm.id];\n elm.registered = false;\n elm.anim = null;\n \ / \ / elm.parent = null;\n}\nAnimation.prototype._collectRemoteResources = function(player) {\n var remotes = [],\n anim = this;\n this.traverse(function(elm) {\n if (elm._hasRemoteResources(anim, player)) {\n remotes = remotes.concat(elm._collectRemoteResources(anim, player)\ / * || []*\ / );\n }\n });\n if(this.fonts && this.fonts.length) {\n remotes = remotes.concat(this.fonts.map(function(f){return f.url;}));\n }\n return remotes;\n}\nAnimation.prototype._loadRemoteResources = function(player) {\n var anim = this;\n this.traverse(function(elm) {\n if (elm._hasRemoteResources(anim, player)) {\n elm._loadRemoteResources(anim, player);\n }\n });\n anim.loadFonts(player);\n}\nAnimation.prototype.__ensureHasMaskCanvas = function(lvl) {\n if (this.__maskCvs && this.__backCvs &&\n this.__maskCvs[lvl] && this.__backCvs[lvl]) return;\n if (!this.__maskCvs) { this.__maskCvs = []; this.__maskCtx = []; }\n if (!this.__backCvs) { this.__backCvs = []; this.__backCtx = []; }\n this.__maskCvs[lvl] = engine.createCanvas(1, 1);\n this.__maskCtx[lvl] = engine.getContext(this.__maskCvs[lvl], '2d');\n this.__backCvs[lvl] = engine.createCanvas(1, 1);\n this.__backCtx[lvl] = engine.getContext(this.__backCvs[lvl], '2d');\n}\nAnimation.prototype.__removeMaskCanvases = function() {\n if (!this.__maskCvs && !this.__backCvs) return;\n if (this.__maskCvs) {\n for (var i = 0, il = this.__maskCvs.length; i < il; i++) {\n if (this.__maskCvs[i]) { \ / \ / use `continue`?\n engine.disposeElement(this.__maskCvs[i]);\n this.__maskCvs[i] = null; \ / \ / is it required?\n this.__maskCtx[i] = null; \ / \ / is it required?\n }\n }\n this.__maskCvs = null;\n this.__maskCtx = null;\n }\n if (this.__backCvs) {\n for (var i = 0, il = this.__backCvs.length; i < il; i++) {\n if (this.__backCvs[i]) { \ / \ / use `continue`?\n engine.disposeElement(this.__backCvs[i]);\n this.__backCvs[i] = null; \ / \ / is it required?\n this.__backCtx[i] = null; \ / \ / is it required?\n }\n }\n this.__backCvs = null;\n this.__backCtx = null;\n }\n}\n\ / **\n * @method find\n *\n * Searches for {@link anm.Element elements} by name inside another\n * {@link anm.Element element} or inside the whole Animation itself, if no other\n * element was provided.\n *\n * @param {String} name Name of the element(s) to find\n * @param {anm.Element} [where] Where to search elements for; if omitted, searches in Animation\n *\n * @return {Array} An array of found elements\n *\ / \nAnimation.prototype.find = function(name, where) {\n var where = where || this;\n var found = [];\n if (where.name == name) found.push(name);\n where.traverse(function(elm) {\n if (elm.name == name) found.push(elm);\n });\n return found;\n}\n\ / **\n * @method findById\n *\n * Searches for {@link anm.Element elements} by ID inside another inside the\n * Animation. Actually, just gets it from hash map, so O(1).\n *\n * @param {String} id ID of the element to find\n * @return {anm.Element|Null} An element you've searched for, or null\n *\ / \nAnimation.prototype.findById = function(id) {\n return this.hash[id];\n}\n\ / *\n * @method invokeAllLaters\n * @private\n *\ / \nAnimation.prototype.invokeAllLaters = function() {\n for (var i = - 88 - 3924732.v36265.1015001 0; i < this._laters.length; i++) {\n this._laters[i].call(this);\n };\n}\n\ / *\n * @method clearAllLaters\n * @private\n *\ / \nAnimation.prototype.clearAllLaters = function() {\n this._laters = [];\n}\n\ / *\n * @method invokeLater\n * @private\n *\ / \nAnimation.prototype.invokeLater = function(f) {\n this._laters.push(f);\n}\n\nvar FONT_LOAD_TIMEOUT = 10000; \ / \ / in ms\n\ / *\n * @method loadFonts\n * @private\n *\ / \nAnimation.prototype.loadFonts = function(player) {\n if (!this.fonts || !this.fonts.length) {\n return;\n }\n\n var fonts = this.fonts,\n style = engine.createStyle(),\n css = '',\n fontsToLoad = [],\n detector = new FontDetector();\n style.type = 'text\ / css';\n\n for (var i = 0; i < fonts.length; i++) {\n var font = fonts[i];\n if (!font.url || !font.face || detector.detect(font.face)) {\n \ / \ / no font name or url || font already available\n continue;\n }\n fontsToLoad.push(font);\n css += '@font-face {' +\n 'font-family: \"' + font.face + '\"; ' +\n 'src:' + font.woff ? 'url(\"'+font.woff+'\") format(\"woff\"), ' : '' +\n 'url(\"'+font.url+'\");' +\n (font.style ? 'style: ' + font.style +'; ' : '') +\n (font.weight ? 'weight: ' + font.weight + '; ' : '') +\n '}\\n';\n }\n\n if (fontsToLoad.length == 0) {\n return;\n };\n\n style.innerHTML = css;\n document.head.appendChild(style); \ / \ / FIXME: should use engine\n\n for (var i = 0; i < fontsToLoad.length; i++) {\n \ / \ / FIXME: should not require a player (probably)\n ResMan.loadOrGet(player.id, fontsToLoad[i].url, function(success) {\n var face = fontsToLoad[i].face,\n interval = 100,\n counter = 0,\n intervalId,\n checkLoaded = function() {\n counter += interval;\n var loaded = detector.detect(face);\n if (loaded || counter > FONT_LOAD_TIMEOUT) {\n \ / \ / after 10 seconds, we'll just assume the font has been loaded\n \ / \ / and carry on. this should help when the font could not be\n \ / \ / reached for whatever reason.\n clearInterval(intervalId);\n success();\n }\n };\n intervalId = setInterval(checkLoaded, interval)\n });\n }\n\n};\n\nmodule.exports = Animation;\n","post_file":"var C = require('..\ / constants.js'),\n engine = require('engine'),\n Element = require('.\ / element.js'),\n Clip = Element,\n Brush = require('..\ / graphics\ / brush.js'),\n provideEvents = require('..\ / events.js').provideEvents,\n AnimationError = require('..\ / errors.js').AnimationError,\n Errors = require('..\ / loc.js').Errors,\n ResMan = require('..\ / resource_manager.js'),\n FontDetector = require('..\ / ..\ / vendor\ / font_detector.js'),\n utils = require('..\ / utils.js'),\n is = utils.is,\n iter = utils.iter;\n\n\n\ / * X_ERROR, X_FOCUS, X_RESIZE, X_SELECT, touch events *\ / \n\nvar DOM_TO_EVT_MAP = {\n 'mouseup': C.X_MUP,\n 'mousedown': C.X_MDOWN,\n 'mousemove': C.X_MMOVE,\n 'mouseover': C.X_MOVER,\n 'mouseout': C.X_MOUT,\n 'click': C.X_MCLICK,\n 'dblclick': C.X_MDCLICK,\n 'keyup': C.X_KUP,\n 'keydown': C.X_KDOWN,\n 'keypress': C.X_KPRESS\n};\n\n\ / \ / -- ------------------------------------------------- 89 - 3924732.v36265.1015001 anm.Animation\n *\n * Create an Animation.\n *\n * It holds an elements tree, an id-to-element map, background fill, zoom and\n * repeat option. It also may render itself to any context with {@link anm.Animation#render}\n * method.\n *\n * @constructor\n *\ / \nfunction Animation() {\n this.id = utils.guid();\n this.tree = [];\n this.hash = {};\n this.name = '';\n this.duration = undefined;\n this.bgfill = null;\n this.width = undefined;\n this.height = undefined;\n this.zoom = 1.0;\n this.speed = 1.0;\n this.repeat = false;\n this.meta = {};\n \ / \ / this.fps = undefined;\n this.__informEnabled = true;\n this._laters = [];\n this._initHandlers(); \ / \ / TODO: make automatic\n}\n\nAnimation.DEFAULT_DURATION = 10;\n\n\ / \ / mouse\ / keyboard events are assigned in L.loadAnimation\n\ / * TODO: move them into animation *\ / \nprovideEvents(Animation, [ C.X_MCLICK, C.X_MDCLICK, C.X_MUP, C.X_MDOWN,\n C.X_MMOVE, C.X_MOVER, C.X_MOUT,\n, new anm.Element()]);`\n * * `anim.add(function(ctx) {...}, function(t) { ... });`\n * * `anim.add(function(ctx) {...}, function(t) { ... },\n * function(ctx, prev(ctx)) { ... });`\n *\n * @param {anm.Element|anm.Clip|Array[Element]} subject Any number of Elements to add\n *\n * @return {anm.Element} The Element was appended.\n *\n *\ / \nAnimation.prototype.add = function(arg1, arg2, arg3) {\n \ / \ / this method only adds an element to a top-level\n \ / \ / FIXME: allow to add elements deeper or rename this\n \ / \ / method to avoid confusion?\n if (arg2) { \ / \ / element by functions mode\n var elm = new Element(arg1, arg2);\n if (arg3) elm.changeTransform(arg3);\n this.addToTree(elm);\n \ / \ / return elm;\n } else if (is.arr(arg1)) { \ / \ / elements array mode\n var clip = new Clip();\n clip.add(arg1);\n this.addToTree(_clip);\n \ / \ / return clip;\n } else { \ / \ / element object mode\n this.addToTree(arg1);\n }\n return this;\n}\n\ / * addS allowed to add static element before, such as image, may be return it in some form? *\ / \n\ / **\n * @method remove\n * @chainable\n *\n * Remove (unregister) element from this animation.\n *\n * @param {anm.Element} element\n *\ / \nAnimation.prototype.remove = function(elm) {\n \ / \ / error will be thrown in _unregister method\n \ / \ / if (!this.hash[elm.id]) throw new AnimErr(Errors.A.ELEMENT_IS_NOT_REGISTERED);\n if (elm.parent) {\n \ / \ / it will unregister element inside\n elm.parent.remove(elm);\n } else {\n this._unregister(elm);\n }\n return this;\n}\n\ / \ / > Animation.prototype.clear % ()\n\ / * Animation.prototype.clear = function() {\n this.hash = {};\n this.tree = [];\n this.duration = 0;\n var hash = this.hash;\n this.hash = {};\n for (var elmId in hash) {\n hash[elm.id]._unbind(); \ / \ / unsafe, because calls unregistering\n }\n} *\ / \n\ / **\n * @method traverse\n * @chainable\n *\n * Visit every element in a tree, no matter how deep it is.\n *\n * @param {Function} visitor\n * @param {anm.Element} visitor.element\n * - 90 - 3924732.v36265.1015001 @param {Object} [data]\n *\ / \n\ / \ / visitElems\nAnimation.prototype.traverse = function(visitor, data) {\n for (var elmId in this.hash) {\n visitor(this.hash[elmId], data);\n }\n return this;\n}\n\ / **\n * @method each\n * @chainable\n *\n * Visit every root element (direct Animation child) in a tree.\n *\n * @param {Function} visitor\n * @param {anm.Element} visitor.child\n * @param {Object} [data]\n *\ / \nAnimation.prototype.each = function(visitor, data) {\n for (var i = 0, tlen = this.tree.length; i < tlen; i++) {\n visitor(this.tree[i], data);\n }\n return this;\n}\n\ / **\n * @method iter\n * @chainable\n *\n * Iterate through every root (direct Animation child) element in a tree.\n *\n * @param {Function} iterator\n * @param {anm.Element} iterator.child\n * @param {Boolean} iterator.return `false`, if this element should be removed\n *\ / \nAnimation.prototype.iter = function(func, rfunc) {\n iter(this.tree).each(func, rfunc);\n return this;\n}\n\ / **\n * @method render\n *\n * Render the Animation for given context at given time.\n *\n * @param {Canvas2DContext} context\n * @param {Number} time\n * @param {Number} [dt] The difference in time between current frame and previous one\n *\ / \nAnimation.prototype.render = function(ctx, time, dt) {\n ctx.save();\n var zoom = this.zoom;\n try {\n if (zoom != 1) {\n ctx.scale(zoom, zoom);\n }\n if (this.bgfill) {\n if (!this.bgfill instanceof Brush) this.bgfill = Brush.fill(this.bgfill);\n ctx.fillStyle = this.bgfill.apply(ctx);\n ctx.fillRect(0, 0, this.width, this.height);\n }\n this.each(function(child) {\n child.render(ctx, time, dt);\n });\n } finally { ctx.restore(); }\n this.fire(C.X_DRAW,ctx);\n}\nAnimation.prototype.handle__x = function(type, evt) {\n this.traverse(function(elm) {\n elm.fire(type, evt);\n });\n return true;\n}\n\ / \ / TODO: test\n\ / **\n * @method getFittingDuration\n *\n * Get the duration where all child elements' bands fit.\n *\n * @return {Number} The calculated duration\n *\ / \nAnimation.prototype.getFittingDuration = function() {\n var max_pos = -Infinity;\n var me = this;\n this.each(function(child) {\n var elm_tpos = child._max_tpos();\n if (elm_tpos > max_pos) max_pos = elm_tpos;\n });\n return * @method reset\n * @chainable\n *\n * Reset all elements.\n= true;\n this.each(function(child) {\n child.reset();\n });\n return this;\n}\n\ / **\n * @method dispose\n * @chainable\n *\n * Remove every possible allocated data to either never use this animation again or\n * start using it from scratch as if it never was used before.\n *\ / \nAnimation.prototype.dispose = function() {\n this.disposeHandlers();\n var me = this;\n \ / * FIXME: unregistering removes from tree, ensure it is safe *\ / \n this.iter(function(child) {\n me._unregister_no_rm(child);\n child.dispose();\n return false;\n });\n return this;\n}\n\ / **\n * @method isEmpty\n *\n * Does Animation has any Elements inside.\n *\n * @return {Boolean} `true` if no Elements, `false` if there are some.\n *\ / \nAnimation.prototype.isEmpty = function() {\n return this.tree.length == 0;\n}\n\ / **\n * @method toString\n *\n * Get a pretty description of this Animation\n *\n * @return {String} pretty string\n *\ / \nAnimation.prototype.toString = function() {\n return \"[ Animation \"+(this.name ? \"'\"+this.name+\"'\" : \"\")+\"]\";\n}\n\ / **\n * @method - 91 - 3924732.v36265.1015001 subscribeEvents\n * @private\n *\n * @param {Canvas} canvas\n *\ / \nAnimation.prototype.subscribeEvents = function(canvas) {\n engine.subscribeAnimationToEvents(canvas, this, DOM_TO_EVT_MAP);\n}\n\ / **\n * @method unsubscribeEvents\n * @private\n *\n * @param {Canvas} canvas\n *\ / \nAnimation.prototype.unsubscribeEvents = function(canvas) {\n engine.unsubscribeAnimationFromEvents(canvas, this);\n}\n\ / **\n * @method addToTree\n * @private\n *\n * @param {anm.Element} element\n *\ / \nAnimation.prototype.addToTree = function(elm) {\n if (!elm.children) {\n throw new AnimationError('It appears that it is not a clip object or element that you pass');\n }\n this._register(elm);\n \ / *if (elm.children) this._addElems(elm.children);*\ / \n this.tree.push(elm);\n}\n\ / *Animation.prototype._addElems = function(elems) {\n for (var ei = 0; ei < elems.length; ei++) {\n var _elm = elems[ei];\n this._register(_elm);\n }\n}*\ / \nAnimation.prototype._register = function(elm) {\n if (this.hash[elm.id]) throw new AnimationError(Errors.A.ELEMENT_IS_REGISTERED);\n elm.registered = true;\n elm.anim = this;\n this.hash[elm.id] = elm;\n var me = this;\n elm.each(function(child) {\n me._register(child);\n });\n}\nAnimation.prototype._unregister_no_rm = function(elm) {\n this._unregister(elm, true);\n}\nAnimation.prototype._unregister = function(elm, save_in_tree) { \ / \ / save_in_tree is optional and false by default\n if (!elm.registered) throw new AnimationError(Errors.A.ELEMENT_IS_NOT_REGISTERED);\n var me = this;\n elm.each(function(child) {\n me._unregister(child);\n });\n var pos = -1;\n if (!save_in_tree) {\n while ((pos = this.tree.indexOf(elm)) >= 0) {\n this.tree.splice(pos, 1); \ / \ / FIXME: why it does not goes deeply in the tree?\n }\n }\n delete this.hash[elm.id];\n elm.registered = false;\n elm.anim = null;\n \ / \ / elm.parent = null;\n}\nAnimation.prototype._collectRemoteResources = function(player) {\n var remotes = [],\n anim = this;\n this.traverse(function(elm) {\n if (elm._hasRemoteResources(anim, player)) {\n remotes = remotes.concat(elm._collectRemoteResources(anim, player)\ / * || []*\ / );\n }\n });\n if(this.fonts && this.fonts.length) {\n remotes = remotes.concat(this.fonts.map(function(f){return f.url;}));\n }\n return remotes;\n}\nAnimation.prototype._loadRemoteResources = function(player) {\n var anim = this;\n this.traverse(function(elm) {\n if (elm._hasRemoteResources(anim, player)) {\n elm._loadRemoteResources(anim, player);\n }\n });\n anim.loadFonts(player);\n}\nAnimation.prototype.__ensureHasMaskCanvas = function(lvl) {\n if (this.__maskCvs && this.__backCvs &&\n this.__maskCvs[lvl] && this.__backCvs[lvl]) return;\n if (!this.__maskCvs) { this.__maskCvs = []; this.__maskCtx = []; }\n if (!this.__backCvs) { this.__backCvs = []; this.__backCtx = []; }\n this.__maskCvs[lvl] = engine.createCanvas(1, 1);\n this.__maskCtx[lvl] = engine.getContext(this.__maskCvs[lvl], '2d');\n this.__backCvs[lvl] = engine.createCanvas(1, 1);\n this.__backCtx[lvl] = engine.getContext(this.__backCvs[lvl], '2d');\n}\nAnimation.prototype.__removeMaskCanvases = function() {\n if (!this.__maskCvs && !this.__backCvs) return;\n if (this.__maskCvs) {\n - 92 - 3924732.v36265.1015001 for (var i = 0, il = this.__maskCvs.length; i < il; i++) {\n if (this.__maskCvs[i]) { \ / \ / use `continue`?\n engine.disposeElement(this.__maskCvs[i]);\n this.__maskCvs[i] = null; \ / \ / is it required?\n this.__maskCtx[i] = null; \ / \ / is it required?\n }\n }\n this.__maskCvs = null;\n this.__maskCtx = null;\n }\n if (this.__backCvs) {\n for (var i = 0, il = this.__backCvs.length; i < il; i++) {\n if (this.__backCvs[i]) { \ / \ / use `continue`?\n engine.disposeElement(this.__backCvs[i]);\n this.__backCvs[i] = null; \ / \ / is it required?\n this.__backCtx[i] = null; \ / \ / is it required?\n }\n }\n this.__backCvs = null;\n this.__backCtx = null;\n }\n}\n\ / **\n * @method find\n *\n * Searches for {@link anm.Element elements} by name inside another\n * {@link anm.Element element} or inside the whole Animation itself, if no other\n * element was provided.\n *\n * @param {String} name Name of the element(s) to find\n * @param {anm.Element} [where] Where to search elements for; if omitted, searches in Animation\n *\n * @return {Array} An array of found elements\n *\ / \nAnimation.prototype.find = function(name, where) {\n var where = where || this;\n var found = [];\n if (where.name == name) found.push(name);\n where.traverse(function(elm) {\n if (elm.name == name) found.push(elm);\n });\n return found;\n}\n\ / **\n * @method findById\n *\n * Searches for {@link anm.Element elements} by ID inside another inside the\n * Animation. Actually, just gets it from hash map, so O(1).\n *\n * @param {String} id ID of the element to find\n * @return {anm.Element|Null} An element you've searched for, or null\n *\ / \nAnimation.prototype.findById = function(id) {\n return this.hash[id];\n}\n\ / *\n * @method invokeAllLaters\n * @private\n *\ / \nAnimation.prototype.invokeAllLaters = function() {\n for (var i = 0; i < this._laters.length; i++) {\n this._laters[i].call(this);\n };\n}\n\ / *\n * @method clearAllLaters\n * @private\n *\ / \nAnimation.prototype.clearAllLaters = function() {\n this._laters = [];\n}\n\ / *\n * @method invokeLater\n * @private\n *\ / \nAnimation.prototype.invokeLater = function(f) {\n this._laters.push(f);\n}\n\nvar FONT_LOAD_TIMEOUT = 10000; \ / \ / in ms\n\ / *\n * @method loadFonts\n * @private\n *\ / \nAnimation.prototype.loadFonts = function(player) {\n if (!this.fonts || !this.fonts.length) {\n return;\n }\n\n var fonts = this.fonts,\n style = engine.createStyle(),\n css = '',\n fontsToLoad = [],\n detector = new FontDetector();\n style.type = 'text\ / css';\n\n for (var i = 0; i < fonts.length; i++) {\n var font = fonts[i];\n if (!font.url || !font.face || detector.detect(font.face)) {\n \ / \ / no font name or url || font already available\n continue;\n }\n fontsToLoad.push(font);\n css += '@font-face {' +\n 'font-family: \"' + font.face + '\"; ' +\n 'src:' + (font.woff ? 'url(\"'+font.woff+'\") format(\"woff\"), ' : '') +\n 'url(\"'+font.url+'\");' +\n (font.style ? 'style: ' + font.style +'; ' : '') +\n (font.weight ? 'weight: ' + font.weight + '; ' : '') +\n '}\\n';\n }\n\n if (fontsToLoad.length == 0) {\n return;\n };\n\n style.innerHTML = css;\n document.head.appendChild(style); \ / \ / FIXME: should use engine\n\n for (var i = 0; i < fontsToLoad.length; i++) {\n - 93 - 3924732.v36265.1015001 \ / \ / FIXME: should not require a player (probably)\n ResMan.loadOrGet(player.id, fontsToLoad[i].url, function(success) {\n var face = fontsToLoad[i].face,\n interval = 100,\n counter = 0,\n intervalId,\n checkLoaded = function() {\n counter += interval;\n var loaded = detector.detect(face);\n if (loaded || counter > FONT_LOAD_TIMEOUT) {\n \ / \ / after 10 seconds, we'll just assume the font has been loaded\n \ / \ / and carry on. this should help when the font could not be\n \ / \ / reached for whatever reason.\n clearInterval(intervalId);\n success();\n }\n };\n intervalId = setInterval(checkLoaded, interval)\n });\n }\n\n};\n\nmodule.exports = Animation;\n","repo":"Animatron\ / player","pre_filename":"src\ / anm\ / animati on\ / animation.js","post_filename":"src\ / anm\ / animation\ / animation.js","pre _sha":"44cf839ed28814eddef222d7a1ce3229f936b849","post_sha":"c07e64a1331b2 ef2b8eb66fc5e86d22d8d3d00f2","pre_reduced":"Animation.prototype.loadFonts = function(player) {\n for (var i = 0; i < fonts.length; i++) {\n css += '@font-face {' +\n 'font-family: \"' + font.face + '\"; ' +\n 'src:' + font.woff ? 'url(\"'+font.woff+'\") format(\"woff\"), ' : '' +\n 'url(\"'+font.url+'\");' +\n (font.style ? 'style: ' + font.style +'; ' : '') +\n (font.weight ? 'weight: ' + font.weight + '; ' : '') +\n '}\\n';\n }\n};","post_reduced":"Animation.prototype.loadFonts = function(player) {\n for (var i = 0; i < fonts.length; i++) {\n css += '@font-face {' +\n 'font-family: \"' + font.face + '\"; ' +\n 'src:' + (font.woff ? 'url(\"'+font.woff+'\") format(\"woff\"), ' : '') +\n 'url(\"'+font.url+'\");' +\n (font.style ? 'style: ' + font.style +'; ' : '') +\n (font.weight ? 'weight: ' + font.weight + '; ' : '') +\n '}\\n';\n }\n};\n","reduction_line_map":[479,491,498,499,500,501,502,503,504,505,536 ],"reduction_match_line_num":3,"language":"javascript","jj":""} - 94 - 3924732.v36265.1015001 APPENDIX T – “test_example.json” {"rule":"AmbiguousConditional","message":"Ambiguous conditional expression: is the sum intended as a condition? Consider adding parentheses to improve readability.","line_number":510,"line_end":510,"col_begin":28,"col_end":83 ,"severity":0,"event_id":10324,"pre_file":"function AddonHierarchical_Lesson_Report_create() {\n var presenter = function () {};\n var presentationController;\n var pageIndex = 0;\n var totalChecks = 0;\n var totalErrors = 0;\n var totalMistakes = 0;\n var totalPoints = 0;\n var totalMaxScore = 0;\n\n var isPreview = false;\n\n presenter.ERROR_MESSAGES = {\n EXPAND_DEPTH_NOT_NUMERIC: \"Depth of expand is not proper\",\n\n C01: \"Wrong classes name format\",\n C02: \"Class names has to be separated by new line\",\n\n D01: \"Values in Disable score on pages property should be numeric and non empty\",\n D02: \"Values in Disable score on pages property should be greater than 0\",\n D03: \"Values in Disable score on pages property should be unique\"\n };\n\n function returnErrorObject(ec) { return { isValid: false, errorCode: ec }; }\n\n function returnCorrectObject(v) { return { isValid: true, value: v }; }\n\n presenter.showErrorMessage = function (message, substitutions) {\n var errorContainer;\n if (typeof(substitutions) == 'undefined') {\n errorContainer = '' + message + '<\ / p>';\n } else {\n var messageSubst = message;\n for (var key in substitutions) {\n messageSubst = messageSubst.replace('%' + key + '%', substitutions[key]);\n }\n errorContainer = '' + messageSubst + '<\ / p>';\n }\n presenter.$view.html(errorContainer);\n };\n\n presenter.setPlayerController = function (controller) {\n presentationController = controller;\n };\n\n presenter.run = function (view, model) {\n isPreview = false;\n presenter.initialize(view, model);\n };\n\n presenter.createPreview = function (view, model) {\n isPreview = true;\n presenter.initialize(view, model);\n };\n\n function addHeader() {\n var headerHTML = \" \" + presenter.configuration.titleLabel + \"<\ / td>\";\n if (presenter.configuration.showResults) headerHTML += \" \" + presenter.configuration.resultsLabel + \"<\ / td>\";\n if (presenter.configuration.showChecks) headerHTML += \" \" + presenter.configuration.checksLabel + \"<\ / td>\";\n if (presenter.configuration.showMistakes) headerHTML += \" \" + presenter.configuration.mistakesLabel + \"<\ / td>\";\n if (presenter.configuration.showErrors) headerHTML += \" \" + presenter.configuration.errorsLabel + \"<\ / td>\";\n if (presenter.configuration.showPageScore) headerHTML += \" <\ / td>\";\n if (presenter.configuration.showMaxScoreField) headerHTML += \"<\ / td>\";\n $(\"<\ / tr>\").prependTo($(\"#\" + presenter.treeID).find('table')).addClass(\"hier_report- header\").html(headerHTML);\n }\n\n function addFooter() {\n - 95 - 3924732.v36265.1015001 var row = document.createElement('tr');\n $(row).appendTo($(\"#\" + presenter.treeID).find('table'));\n $(row).addClass(\"hier_report- footer\");\n\n $(\"<\ / td>\").appendTo($(row)).html(presenter.configuration.totalLabel );\n\n if (presenter.configuration.showResults) {\n var score = resetScore();\n\n if (!isPreview) {\n var playerUtils = new PlayerUtils({});\n playerUtils.scoreService = presentationController.getScore();\n var totalScore = playerUtils.getPresentationScore(presentationController.getPresentation()) ;\n score = {\n score: totalScore.scaledScore,\n count: 1\n }\n }\n createProgressCell(row, score);\n }\n\n if (presenter.configuration.showChecks) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- checks\").html(totalChecks);\n }\n\n if (presenter.configuration.showMistakes) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- mistakes\").html(totalMistakes);\n }\n\n if (presenter.configuration.showErrors) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- errors\").html(totalErrors);\n }\n\n if (presenter.configuration.showPageScore) {\n var content = totalPoints + \"\ / <\ / span>\" + totalMaxScore;\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-page- score\").html(content);\n }\n\n if (presenter.configuration.showMaxScoreField) {\n $(\"<\ / td>\").appendTo($(row));\n }\n }\n\n function createRow(index, parentIndex, isChapter) {\n var row = document.createElement('tr');\n\n $(row).appendTo($(\"#\" + presenter.treeID).find('table'));\n $(row).addClass(\"treegrid-\" + index);\n $(row).addClass(presenter.configuration.classes[index % presenter.configuration.classes.length]);\n\n if (parentIndex != null) {\n \t$(row).addClass(\"treegrid-parent-\" + parentIndex);\n }\n\n if (isChapter) {\n \t$(row).addClass(\"hier_report- chapter\");\n } else {\n $(row).addClass(index % 2 > 0 ? \"hier_report-odd\" : \"hier_report-even\");\n }\n\n return row;\n }\n\n function createProgressCell(row, score, index, isChapter) {\n var progressCell = document.createElement('td');\n $(progressCell).appendTo($(row)).addClass(\"hier_report-progress\");\n\n var progressbar = document.createElement('div');\n $(progressbar).appendTo($(progressCell));\n $(progressbar).attr(\"id\", \"progressbar-\" + index);\n $(progressbar).addClass(\"hier_report-progressbar\");\n\n var percent = Math.floor(score.score \ / score.count * 100);\n\n var progressInfo = document.createElement('div');\n $(progressInfo).appendTo($(progressCell)).attr(\"style\", \"float: right\").html(percent + \"%\");\n\n if (!isChapter) {\n $(progressbar).progressbar({\n value: Math.floor(score.score * 100),\n max: 100\n });\n }\n }\n\n function createScoreCells(row, pageId, index, isChapter) {\n var isScoreEnable = - 96 - 3924732.v36265.1015001 presenter.configuration.disabledScorePages.indexOf(index) === -1;\n var score = resetScore();\n if (!isPreview) score = presentationController.getScore().getPageScoreById(pageId);\n var points = 0;\n\n if (!isChapter) {\n points = score.score;\n \tscore.count = 1;\n \tscore.score = score.maxScore !== 0 ? score.score \ / score.maxScore : 0;\n }\n\n if (isScoreEnable) {\n if (presenter.configuration.showResults) {\n createProgressCell(row, score, index, isChapter);\n }\n\n if (presenter.configuration.showChecks) {\n var checksCell = document.createElement('td');\n $(checksCell).appendTo($(row))\n .addClass(\"hier_report-checks\")\n .html(score.checkCount);\n totalChecks += score.checkCount;\n }\n\n if (presenter.configuration.showMistakes) {\n var mistakesCell = document.createElement('td');\n $(mistakesCell).appendTo($(row))\n .addClass(\"hier_report-mistakes\")\n .html(score.mistakeCount);\n totalMistakes += score.mistakeCount;\n }\n\n if (presenter.configuration.showErrors) {\n var errorsCell = document.createElement('td');\n $(errorsCell).appendTo($(row))\n .addClass(\"hier_report-errors\")\n .html(score.errorCount);\n totalErrors += score.errorCount;\n }\n\n if (presenter.configuration.showPageScore) {\n $(\"<\ / td>\").appendTo($(row))\n .addClass(\"hier_report-page-score\")\n .html(points + \"\ / <\ / span>\" + score.maxScore);\n totalPoints += points;\n totalMaxScore += score.maxScore;\n }\n\n if (presenter.configuration.showMaxScoreField) {\n var className = (points === score.maxScore && score.maxScore !== 0 ? \"page-max-score\" : \"page-non-max-score\");\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-\" + className);\n }\n } else {\n var c = presenter.configuration;\n var columns = [c.showResults, c.showChecks, c.showMistakes, c.showErrors, c.showPageScore, c.showMaxScoreField].filter(function(a) { return a }).length;\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-score-disabled- row\");\n }\n\n }\n\n function generatePageLinks(text, isChapter, pageId) {\n var $element = $(document.createElement('td')),\n $link = $(\"<\ / a>\").text(text).attr('href', '#').attr('data-page-id', pageId);\n\n $element.append($('<div class=\"text- wrapper\">').html(isChapter ? text : $link));\n\n return $element;\n }\n\n function addRow(name, index, parrentIndex, isChapter, pageId) {\n var row = createRow(index, parrentIndex, isChapter);\n\n var nameCell = generatePageLinks(name, isChapter, pageId);\n $(nameCell).appendTo($(row));\n\n createScoreCells(row, pageId, index, isChapter);\n }\n\n function updateRow(pageIndex, pageScore, isEmptyChapter) {\n var row = - 97 - 3924732.v36265.1015001 $(\".treegrid-\" + pageIndex);\n\n if (presenter.configuration.showResults) {\n if (isEmptyChapter) {\n var progresscell = $(row).find(\".hier_report- progress\");\n $(progresscell).children().remove();\n $(progresscell).html(\"-\");\n } else {\n var percent = (Math.floor((pageScore.score \ / pageScore.count) * 100));\n var progressbar = $(row).find(\"#progressbar-\" + pageIndex);\n $(progressbar).progressbar({value: Math.floor((pageScore.score \ / pageScore.count) * 100), max: 100});\n $(progressbar).closest(\"div\").next().html(percent + \"%\");\n }\n }\n\n if (presenter.configuration.showChecks) {\n $(row).find(\".hier_report-checks\").html(isEmptyChapter ? \"-\" : pageScore.checkCount);\n }\n\n if (presenter.configuration.showMistakes) {\n$(row).find(\".hier_report-mistakes\").html ? \"-\" : pageScore.mistakeCount);\n }\n\n if (presenter.configuration.showErrors) {\n$(row).find(\".hier_report-errors\").html(isEmptyChapter ? : pageScore.errorCount);\n }\n }\n\n function updateScore(score, update) {\n score.score? update.score : update.score\ / update.maxScore;\n += update.errorCount;\n score.checkCount += update.checkCount;\n score.mistakeCount += update.mistakeCount;\n score.count += update.count;\n return score;\n }\n\n function resetScore() {\n return {\n score: 0,\n maxScore: 0,\n errorCount: 0,\n checkCount: 0,\n mistakeCount: 0,\n count: 0\n };\n }\n\n presenter.createPreviewTree = function() {\n var pagesMockup = [\n {name : \"Page1\", parent : null},\n {name : \"Unit1\", parent : null},\n {name : \"Page2\", parent : 1},\n {name : \"Chapter1\", parent : 1},\n {name : \"Page3\", parent : 3},\n {name : \"Page4\", parent : 3},\n {name : \"Chapter2\", parent : 1},\n {name : \"Page5\", parent : 6},\n {name : \"Page6\", parent : 1},\n {name : \"Page7\", parent : null},\n {name : \"Page8\", parent : null},\n {name : \"Page9\", parent : null},\n {name : \"Page10\", parent : null},\n {name : \"Page11\", parent : null}\n ];\n\n var chapterScore = resetScore();\n for (var i = 0; i < pagesMockup.length; i++) {\n addRow(pagesMockup[i].name, i, pagesMockup[i].parent, false, \"some_id\");\n }\n return chapterScore;\n };\n\n presenter.createTree = function (root, parrentIndex, pageCount) {\n var chapterIndex = 0,\n chapterScore = resetScore(),\n pageScore = resetScore(),\n isEmpty = true,\n values = {};\n\n for (var i = 0; i < pageCount; i++) {\n var isChapter = (root.get(i).type == \"chapter\");\n if (!isChapter && !root.get(i).isReportable()) continue;\n if (!isChapter && root.get(i).isReportable()) {\n \tisEmpty = false;\n }\n var pageId = \"chapter\";\n if (!isChapter) {\n \tpageId = root.get(i).getId();\n }\n addRow(root.get(i).getName(), pageIndex, parrentIndex, isChapter, pageId);\n pageScore = presentationController.getScore().getPageScoreById(pageId);\n pageScore.count = 1;\n pageIndex++;\n if (isChapter) - 98 - 3924732.v36265.1015001 {\n \tchapterIndex = pageIndex - 1;\n \tvalues = presenter.createTree(root.get(i), chapterIndex, root.get(i).size());\n \tupdateRow(chapterIndex, values.pagesScore, values.isEmpty);\n \tpageScore = values.pagesScore;\n }\n \tchapterScore = updateScore(chapterScore, pageScore);\n }\n\n return { pagesScore: chapterScore, isEmpty: isEmpty };\n };\n\n function handleMouseClickActions() {\n var commander = presentationController.getCommands(),\n $report = presenter.$view.find('.hier_report tr');\n\n $report.find('td a').each(function () {\n $(this).click(function (event) {\n event.preventDefault();\n event.stopPropagation();\n commander.gotoPageId($(this).attr('data-page-id'));\n });\n });\n\n $report.find('.treegrid-expander').each(function () {\n $(this).click(function (event) {\n event.preventDefault();\n event.stopPropagation();\n });\n });\n }\n\n function expandTree(level) {\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n if ($(this).treegrid('getDepth') < level) {\n $(this).treegrid('expand');\n }\n });\n }\n\n function saveTreeState() {\n var state = [];\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n state.push($(this).treegrid('isExpanded'))\n });\n return state;\n }\n\n function restoreTreeState(state) {\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n $(this).treegrid(state[$(this).treegrid('getNodeId')] ? 'expand' : 'collapse');\n });\n }\n\n presenter.getState = function () {\n return JSON.stringify({\n 'treeState': saveTreeState(),\n 'isVisible': presenter.configuration.isVisible\n });\n };\n\n presenter.setState = function (stateString) {\n var state = JSON.parse(stateString);\n\n restoreTreeState(state.treeState);\n\n presenter.setVisibility(state.isVisible);\n presenter.configuration.isVisible = state.isVisible;\n };\n\n function parseClasses(classes_text) {\n function isValidClassName(class_name) {\n return \ / ^[a-z_-][a-z\\d_- ]*$\ / i.test(class_name);\n }\n\n if (ModelValidationUtils.isStringEmpty(classes_text)) {\n return returnCorrectObject([]);\n }\n\n var classes = classes_text.split('\\n');\n for (var i=0; i<classes.length; i++) {\n if (classes[i].indexOf(' ') !== -1) {\n return returnErrorObject(\"C02\");\n }\n\n if (!isValidClassName(classes[i])) {\n return returnErrorObject(\"C01\");\n }\n }\n\n return returnCorrectObject(classes);\n }\n\n function parseScoreDisable(pages_text) {\n if (ModelValidationUtils.isStringEmpty(pages_text)) {\n return returnCorrectObject([]);\n }\n\n var i;\n\n var pages = pages_text.split(';');\n for (i=0; i<pages.length; i++) {\n var numberObject = ModelValidationUtils.validateInteger(pages[i]);\n if (!numberObject.isValid) {\n return - 99 - 3924732.v36265.1015001 returnErrorObject(\"D01\");\n }\n\n pages[i] = numberObject.value - 1; \ / \ / indexing from 0\n\n if (pages[i] < 0) {\n return returnErrorObject(\"D02\");\n }\n }\n\n for (i=1; i<pages.length; i++) {\n if (pages.sort()[i] === pages.sort()[i-1]) {\n return returnErrorObject(\"D03\");\n }\n }\n\n return returnCorrectObject(pages.sort());\n }\n\n presenter.validateModel = function (model) {\n var expandDepth = returnCorrectObject(0);\n\n if (model['expandDepth'].length > 0) {\n expandDepth = ModelValidationUtils.validateInteger(model['expandDepth']);\n if (!expandDepth.isValid) {\n return returnErrorObject('EXPAND_DEPTH_NOT_NUMERIC');\n }\n }\n\n var validatedClasses = parseClasses(model[\"classes\"]);\n if (!validatedClasses.isValid) {\n return returnErrorObject(validatedClasses.errorCode);\n }\n\n var validatedDisabledScorePages = parseScoreDisable(model[\"scoredisabled\"]);\n if (!validatedDisabledScorePages.isValid) {\n return returnErrorObject(validatedDisabledScorePages.errorCode);\n }\n\n return {\n ID: model.ID,\n isValid: true,\n width: parseInt(model[\"Width\"], 10),\n height: parseInt(model[\"Height\"], 10),\n isVisible: ModelValidationUtils.validateBoolean(model[\"Is Visible\"]),\n\n showResults: ModelValidationUtils.validateBoolean showErrors: ModelValidationUtils.validateBoolean showChecks: ModelValidationUtils.validateBoolean showMistakes: ModelValidationUtils.validateBoolean ),\n resultsLabel: model['resultsLabel'],\nmodel['errorsLabel'],\n checksLabel: model['checksLabel'],\n mistakesLabel: model['mistakesLabel'],\n showTotal: ModelValidationUtils.validateBoolean(model[\"total\"]),\n totalLabel: model['totalLabel'],\n titleLabel: model['titleLabel'],\n expandDepth: expandDepth.value,\n classes: validatedClasses.value,\n showPageScore: ModelValidationUtils.validateBoolean(model[\"showpagescore\"]),\n showMaxScoreField: ModelValidationUtils.validateBoolean(model[\"showmaxscorefield\"]),\n disabledScorePages: validatedDisabledScorePages.value\n };\n };\n\n presenter.setVisibility = function (isVisible) {\n presenter.$view.css(\"visibility\", isVisible ? \"visible\" : \"hidden\");\n };\n\n presenter.initialize = function (view, model) {\n presenter.$view = $(view);\n\n presenter.configuration = presenter.validateModel(model);\n if (!presenter.configuration.isValid) {\n presenter.showErrorMessage(presenter.ERROR_MESSAGES[presenter.configuratio n.errorCode]);\n return;\n }\n\n $('.hier_report').attr(\"style\", \"height: \" + presenter.configuration.height + \"px\");\n presenter.treeID = presenter.configuration.ID + isPreview ? \"Preview\" : \"\";\n presenter.$view.find(\"div\").first().attr('id', presenter.treeID);\n\n presenter.setVisibility(presenter.configuration.isVisible);\n\n addHeader();\n if (isPreview) {\n presenter.createPreviewTree();\n } else {\n var - 100 - 3924732.v36265.1015001 presentation = presentationController.getPresentation();\n presenter.createTree(presentation.getTableOfContents(), null, presentation.getTableOfContents().size());\n }\n\n if (presenter.configuration.showTotal) {\n addFooter();\n }\n\n $(\"#\" + presenter.treeID).find('table').not('.hier_report- header').not('.hier_report-footer').treegrid({\n 'initialState': 'collapsed',\n 'expanderTemplate': '<div class=\"treegrid-expander\"><\ / div>'\n });\n\n expandTree(presenter.configuration.expandDepth);\n if (!isPreview) {\n handleMouseClickActions();\n }\n };\n\n return presenter;\n}","post_file":"function AddonHierarchical_Lesson_Report_create() {\n var presenter = function () {};\n var presentationController;\n var pageIndex = 0;\n var totalChecks = 0;\n var totalErrors = 0;\n var totalMistakes = 0;\n var totalPoints = 0;\n var totalMaxScore = 0;\n\n var isPreview = false;\n\n presenter.ERROR_MESSAGES = {\n EXPAND_DEPTH_NOT_NUMERIC: \"Depth of expand is not proper\",\n\n C01: \"Wrong classes name format\",\n C02: \"Class names has to be separated by new line\",\n\n D01: \"Values in Disable score on pages property should be numeric and non empty\",\n D02: \"Values in Disable score on pages property should be greater than 0\",\n D03: \"Values in Disable score on pages property should be unique\"\n };\n\n function returnErrorObject(ec) { return { isValid: false, errorCode: ec }; }\n\n function returnCorrectObject(v) { return { isValid: true, value: v }; }\n\n presenter.showErrorMessage = function (message, substitutions) {\n var errorContainer;\n if (typeof(substitutions) == 'undefined') {\n errorContainer = '' + message + '<\ / p>';\n } else {\n var messageSubst = message;\n for (var key in substitutions) {\n messageSubst = messageSubst.replace('%' + key + '%', substitutions[key]);\n }\n errorContainer = '' + messageSubst + '<\ / p>';\n }\n presenter.$view.html(errorContainer);\n };\n\n presenter.setPlayerController = function (controller) {\n presentationController = controller;\n };\n\n presenter.run = function (view, model) {\n isPreview = false;\n presenter.initialize(view, model);\n };\n\n presenter.createPreview = function (view, model) {\n isPreview = true;\n presenter.initialize(view, model);\n };\n\n function addHeader() {\n var headerHTML = \" \" + presenter.configuration.titleLabel + \"<\ / td>\";\n if (presenter.configuration.showResults) headerHTML += \" \" + presenter.configuration.resultsLabel + \"<\ / td>\";\n if (presenter.configuration.showChecks) headerHTML += \" \" + presenter.configuration.checksLabel + \"<\ / td>\";\n if (presenter.configuration.showMistakes) headerHTML += \" \" + presenter.configuration.mistakesLabel + \"<\ / td>\";\n if (presenter.configuration.showErrors) headerHTML += \" \" + presenter.configuration.errorsLabel + \"<\ / td>\";\n if (presenter.configuration.showPageScore) headerHTML += \" <\ / td>\";\n if (presenter.configuration.showMaxScoreField) headerHTML += - 101 - 3924732.v36265.1015001 \"<\ / td>\";\n $(\"<\ / tr>\").prependTo($(\"#\" + presenter.treeID).find('table')).addClass(\"hier_report- header\").html(headerHTML);\n }\n\n function addFooter() {\n var row = document.createElement('tr');\n $(row).appendTo($(\"#\" + presenter.treeID).find('table'));\n $(row).addClass(\"hier_report- footer\");\n\n $(\"<\ / td>\").appendTo($(row)).html(presenter.configuration.totalLabel );\n\n if (presenter.configuration.showResults) {\n var score = resetScore();\n\n if (!isPreview) {\n var playerUtils = new PlayerUtils({});\n playerUtils.scoreService = presentationController.getScore();\n var totalScore = playerUtils.getPresentationScore(presentationController.getPresentation()) ;\n score = {\n score: totalScore.scaledScore,\n count: 1\n }\n }\n createProgressCell(row, score);\n }\n\n if (presenter.configuration.showChecks) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- checks\").html(totalChecks);\n }\n\n if (presenter.configuration.showMistakes) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- mistakes\").html(totalMistakes);\n }\n\n if (presenter.configuration.showErrors) {\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report- errors\").html(totalErrors);\n }\n\n if (presenter.configuration.showPageScore) {\n var content = totalPoints + \"\ / <\ / span>\" + totalMaxScore;\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-page- score\").html(content);\n }\n\n if (presenter.configuration.showMaxScoreField) {\n $(\"<\ / td>\").appendTo($(row));\n }\n }\n\n function createRow(index, parentIndex, isChapter) {\n var row = document.createElement('tr');\n\n $(row).appendTo($(\"#\" + presenter.treeID).find('table'));\n $(row).addClass(\"treegrid-\" + index);\n $(row).addClass(presenter.configuration.classes[index % presenter.configuration.classes.length]);\n\n if (parentIndex != null) {\n \t$(row).addClass(\"treegrid-parent-\" + parentIndex);\n }\n\n if (isChapter) {\n \t$(row).addClass(\"hier_report- chapter\");\n } else {\n $(row).addClass(index % 2 > 0 ? \"hier_report-odd\" : \"hier_report-even\");\n }\n\n return row;\n }\n\n function createProgressCell(row, score, index, isChapter) {\n var progressCell = document.createElement('td');\n $(progressCell).appendTo($(row)).addClass(\"hier_report-progress\");\n\n var progressbar = document.createElement('div');\n $(progressbar).appendTo($(progressCell));\n $(progressbar).attr(\"id\", \"progressbar-\" + index);\n $(progressbar).addClass(\"hier_report-progressbar\");\n\n var percent = Math.floor(score.score \ / score.count * 100);\n\n var progressInfo = document.createElement('div');\n $(progressInfo).appendTo($(progressCell)).attr(\"style\", \"float: right\").html(percent + \"%\");\n\n if (!isChapter) {\n $(progressbar).progressbar({\n value: - 102 - 3924732.v36265.1015001 Math.floor(score.score * 100),\n max: 100\n });\n }\n }\n\n function createScoreCells(row, pageId, index, isChapter) {\n var isScoreEnable = presenter.configuration.disabledScorePages.indexOf(index) === -1;\n var score = resetScore();\n if (!isPreview) score = presentationController.getScore().getPageScoreById(pageId);\n var points = 0;\n\n if (!isChapter) {\n points = score.score;\n \tscore.count = 1;\n \tscore.score = score.maxScore !== 0 ? score.score \ / score.maxScore : 0;\n }\n\n if (isScoreEnable) {\n if (presenter.configuration.showResults) {\n createProgressCell(row, score, index, isChapter);\n }\n\n if (presenter.configuration.showChecks) {\n var checksCell = document.createElement('td');\n $(checksCell).appendTo($(row))\n .addClass(\"hier_report-checks\")\n .html(score.checkCount);\n totalChecks += score.checkCount;\n }\n\n if (presenter.configuration.showMistakes) {\n var mistakesCell = document.createElement('td');\n $(mistakesCell).appendTo($(row))\n .addClass(\"hier_report-mistakes\")\n .html(score.mistakeCount);\n totalMistakes += score.mistakeCount;\n }\n\n if (presenter.configuration.showErrors) {\n var errorsCell = document.createElement('td');\n $(errorsCell).appendTo($(row))\n .addClass(\"hier_report-errors\")\n .html(score.errorCount);\n totalErrors += score.errorCount;\n }\n\n if (presenter.configuration.showPageScore) {\n $(\"<\ / td>\").appendTo($(row))\n .addClass(\"hier_report-page-score\")\n .html(points + \"\ / <\ / span>\" + score.maxScore);\n totalPoints += points;\n totalMaxScore += score.maxScore;\n }\n\n if (presenter.configuration.showMaxScoreField) {\n var className = (points === score.maxScore && score.maxScore !== 0 ? \"page-max-score\" : \"page-non-max-score\");\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-\" + className);\n }\n } else {\n var c = presenter.configuration;\n var columns = [c.showResults, c.showChecks, c.showMistakes, c.showErrors, c.showPageScore, c.showMaxScoreField].filter(function(a) { return a }).length;\n $(\"<\ / td>\").appendTo($(row)).addClass(\"hier_report-score-disabled- row\");\n }\n\n }\n\n function generatePageLinks(text, isChapter, pageId) {\n var $element = $(document.createElement('td')),\n $link = $(\"<\ / a>\").text(text).attr('href', '#').attr('data-page-id', pageId);\n\n $element.append($('<div class=\"text- wrapper\">').html(isChapter ? text : $link));\n\n return $element;\n }\n\n function addRow(name, index, parrentIndex, isChapter, pageId) {\n var row = createRow(index, parrentIndex, isChapter);\n\n var nameCell = generatePageLinks(name, isChapter, - 103 - 3924732.v36265.1015001 pageId);\n $(nameCell).appendTo($(row));\n\n createScoreCells(row, pageId, index, isChapter);\n }\n\n function updateRow(pageIndex, pageScore, isEmptyChapter) {\n var row = $(\".treegrid-\" + pageIndex);\n\n if (presenter.configuration.showResults) {\n if (isEmptyChapter) {\n var progresscell = $(row).find(\".hier_report- progress\");\n $(progresscell).children().remove();\n $(progresscell).html(\"-\");\n } else {\n var percent = (Math.floor((pageScore.score \ / pageScore.count) * 100));\n var progressbar = $(row).find(\"#progressbar-\" + pageIndex);\n $(progressbar).progressbar({value: Math.floor((pageScore.score \ / pageScore.count) * 100), max: 100});\n $(progressbar).closest(\"div\").next().html(percent + \"%\");\n }\n }\n\n if (presenter.configuration.showChecks) {\n $(row).find(\".hier_report-checks\").html(isEmptyChapter ? \"-\" : pageScore.checkCount);\n }\n\n if (presenter.configuration.showMistakes) {\n$(row).find(\".hier_report-mistakes\").html ? \"-\" : pageScore.mistakeCount);\n }\n\n if (presenter.configuration.showErrors) {\n$(row).find(\".hier_report-errors\").html(isEmptyChapter ? : pageScore.errorCount);\n }\n }\n\n function updateScore(score, update) {\n score.score? update.score : update.score\ / update.maxScore;\n score.errorCount += update.errorCount;\n score.checkCount += update.checkCount;\n score.mistakeCount += update.mistakeCount;\n score.count += update.count;\n return score;\n }\n\n function resetScore() {\n return {\n score: 0,\n maxScore: 0,\n errorCount: 0,\n checkCount: 0,\n mistakeCount: 0,\n count: 0\n };\n }\n\n presenter.createPreviewTree = function() {\n var pagesMockup = [\n {name : \"Page1\", parent : null},\n {name : \"Unit1\", parent : null},\n {name : \"Page2\", parent : 1},\n {name : \"Chapter1\", parent : 1},\n {name : \"Page3\", parent : 3},\n {name : \"Page4\", parent : 3},\n {name : \"Chapter2\", parent : 1},\n {name : \"Page5\", parent : 6},\n {name : \"Page6\", parent : 1},\n {name : \"Page7\", parent : null},\n {name : \"Page8\", parent : null},\n {name : \"Page9\", parent : null},\n {name : \"Page10\", parent : null},\n {name : \"Page11\", parent : null}\n ];\n\n var chapterScore = resetScore();\n for (var i = 0; i < pagesMockup.length; i++) {\n addRow(pagesMockup[i].name, i, pagesMockup[i].parent, false, \"some_id\");\n }\n return chapterScore;\n };\n\n presenter.createTree = function (root, parrentIndex, pageCount) {\n var chapterIndex = 0,\n chapterScore = resetScore(),\n pageScore = resetScore(),\n isEmpty = true,\n values = {};\n\n for (var i = 0; i < pageCount; i++) {\n var isChapter = (root.get(i).type == \"chapter\");\n if (!isChapter && !root.get(i).isReportable()) continue;\n if (!isChapter && root.get(i).isReportable()) {\n \tisEmpty = false;\n }\n var pageId = \"chapter\";\n if (!isChapter) {\n \tpageId = root.get(i).getId();\n }\n addRow(root.get(i).getName(), pageIndex, parrentIndex, isChapter, - 104 - 3924732.v36265.1015001 pageId);\n pageScore = presentationController.getScore().getPageScoreById(pageId);\n pageScore.count = 1;\n pageIndex++;\n if (isChapter) {\n \tchapterIndex = pageIndex - 1;\n \tvalues = presenter.createTree(root.get(i), chapterIndex, root.get(i).size());\n \tupdateRow(chapterIndex, values.pagesScore, values.isEmpty);\n \tpageScore = values.pagesScore;\n }\n \tchapterScore = updateScore(chapterScore, pageScore);\n }\n\n return { pagesScore: chapterScore, isEmpty: isEmpty };\n };\n\n function handleMouseClickActions() {\n var commander = presentationController.getCommands(),\n $report = presenter.$view.find('.hier_report tr');\n\n $report.find('td a').each(function () {\n $(this).click(function (event) {\n event.preventDefault();\n event.stopPropagation();\n commander.gotoPageId($(this).attr('data-page-id'));\n });\n });\n\n $report.find('.treegrid-expander').each(function () {\n $(this).click(function (event) {\n event.preventDefault();\n event.stopPropagation();\n });\n });\n }\n\n function expandTree(level) {\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n if ($(this).treegrid('getDepth') < level) {\n $(this).treegrid('expand');\n }\n });\n }\n\n function saveTreeState() {\n var state = [];\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n state.push($(this).treegrid('isExpanded'))\n });\n return state;\n }\n\n function restoreTreeState(state) {\n $('.hier_report table').find('tr').not('.hier_report- header').not('.hier_report-footer').each(function () {\n $(this).treegrid(state[$(this).treegrid('getNodeId')] ? 'expand' : 'collapse');\n });\n }\n\n presenter.getState = function () {\n return JSON.stringify({\n 'treeState': saveTreeState(),\n 'isVisible': presenter.configuration.isVisible\n });\n };\n\n presenter.setState = function (stateString) {\n var state = JSON.parse(stateString);\n\n restoreTreeState(state.treeState);\n\n presenter.setVisibility(state.isVisible);\n presenter.configuration.isVisible = state.isVisible;\n };\n\n function parseClasses(classes_text) {\n function isValidClassName(class_name) {\n return \ / ^[a-z_-][a-z\\d_- ]*$\ / i.test(class_name);\n }\n\n if (ModelValidationUtils.isStringEmpty(classes_text)) {\n return returnCorrectObject([]);\n }\n\n var classes = classes_text.split('\\n');\n for (var i=0; i<classes.length; i++) {\n if (classes[i].indexOf(' ') !== -1) {\n return returnErrorObject(\"C02\");\n }\n\n if (!isValidClassName(classes[i])) {\n return returnErrorObject(\"C01\");\n }\n }\n\n return returnCorrectObject(classes);\n }\n\n function parseScoreDisable(pages_text) {\n if (ModelValidationUtils.isStringEmpty(pages_text)) {\n return returnCorrectObject([]);\n }\n\n var i;\n\n var pages - 105 - 3924732.v36265.1015001 = pages_text.split(';');\n for (i=0; i<pages.length; i++) {\n var numberObject = ModelValidationUtils.validateInteger(pages[i]);\n if (!numberObject.isValid) {\n return returnErrorObject(\"D01\");\n }\n\n pages[i] = numberObject.value - 1; \ / \ / indexing from 0\n\n if (pages[i] < 0) {\n return returnErrorObject(\"D02\");\n }\n }\n\n for (i=1; i<pages.length; i++) {\n if (pages.sort()[i] === pages.sort()[i-1]) {\n return returnErrorObject(\"D03\");\n }\n }\n\n return returnCorrectObject(pages.sort());\n }\n\n presenter.validateModel = function (model) {\n var expandDepth = returnCorrectObject(0);\n\n if (model['expandDepth'].length > 0) {\n expandDepth = ModelValidationUtils.validateInteger(model['expandDepth']);\n if (!expandDepth.isValid) {\n return returnErrorObject('EXPAND_DEPTH_NOT_NUMERIC');\n }\n }\n\n var validatedClasses = parseClasses(model[\"classes\"]);\n if (!validatedClasses.isValid) {\n return returnErrorObject(validatedClasses.errorCode);\n }\n\n var validatedDisabledScorePages = parseScoreDisable(model[\"scoredisabled\"]);\n if (!validatedDisabledScorePages.isValid) {\n return returnErrorObject(validatedDisabledScorePages.errorCode);\n }\n\n return {\n ID: model.ID,\n isValid: true,\n width: parseInt(model[\"Width\"], 10),\n height: parseInt(model[\"Height\"], 10),\n isVisible: ModelValidationUtils.validateBoolean(model[\"Is Visible\"]),\n\n showResults: ModelValidationUtils.validateBoolean showErrors: ModelValidationUtils.validateBoolean showChecks: ModelValidationUtils.validateBoolean showMistakes: ModelValidationUtils.validateBoolean ),\n resultsLabel: model['resultsLabel'],\nmodel['errorsLabel'],\n checksLabel: model['checksLabel'],\n mistakesLabel: model['mistakesLabel'],\n showTotal: ModelValidationUtils.validateBoolean(model[\"total\"]),\n totalLabel: model['totalLabel'],\n titleLabel: model['titleLabel'],\n expandDepth: expandDepth.value,\n classes: validatedClasses.value,\n showPageScore: ModelValidationUtils.validateBoolean(model[\"showpagescore\"]),\n showMaxScoreField: ModelValidationUtils.validateBoolean(model[\"showmaxscorefield\"]),\n disabledScorePages: validatedDisabledScorePages.value\n };\n };\n\n presenter.setVisibility = function (isVisible) {\n presenter.$view.css(\"visibility\", isVisible ? \"visible\" : \"hidden\");\n };\n\n presenter.initialize = function (view, model) {\n presenter.$view = $(view);\n\n presenter.configuration = presenter.validateModel(model);\n if (!presenter.configuration.isValid) {\n presenter.showErrorMessage(presenter.ERROR_MESSAGES[presenter.configuratio n.errorCode]);\n return;\n }\n\n $('.hier_report').attr(\"style\", \"height: \" + presenter.configuration.height + \"px\");\n presenter.treeID = presenter.configuration.ID + (isPreview ? \"Preview\" : \"\");\n presenter.$view.find(\"div\").first().attr('id', presenter.treeID);\n\n - 106 - 3924732.v36265.1015001 presenter.setVisibility(presenter.configuration.isVisible);\n\n addHeader();\n if (isPreview) {\n presenter.createPreviewTree();\n } else {\n var presentation = presentationController.getPresentation();\n presenter.createTree(presentation.getTableOfContents(), null, presentation.getTableOfContents().size());\n }\n\n if (presenter.configuration.showTotal) {\n addFooter();\n }\n\n $(\"#\" + presenter.treeID).find('table').not('.hier_report- header').not('.hier_report-footer').treegrid({\n 'initialState': 'collapsed',\n 'expanderTemplate': '<div class=\"treegrid-expander\"><\ / div>'\n });\n\n expandTree(presenter.configuration.expandDepth);\n if (!isPreview) {\n handleMouseClickActions();\n }\n };\n\n return presenter;\n}","repo":"icplayer\ / icplayer","pre_filename":"addons\ / Hierarc e_reduced":"function presenter.initialize = function (view, model) {\n = presenter.configuration.ID + isPreview ? \"Preview\"};\n return presenter;\n}","post_reduced":"function AddonHierarchical_Lesson_Report_create() {\n presenter.initialize = function (view, model) {\n presenter.treeID = presenter.configuration.ID + (isPreview ? \"Preview\" : \"\");\n };\n return presenter;\n}\n","reduction_line_map":[0,499,509,535,537,538],"reduction_m atch_line_num":3,"language":"javascript","jj":""} - 107 - 3924732.v3
Claims
6265.1015001 CLAIMS What is claimed is:
1. A computer-implemented method for fixing an error in source code, the computer-implemented method comprising: modifying original source code having an error to generate reduced source code including the error; processing the reduced source code including the error with a language model to produce error-eliminated reduced source code; and merging the error-eliminated reduced source code with the original source code to generate fixed source code.
2. The computer-implemented method of Claim 1, wherein the error is at least one of: asecurity vulnerability, a semantic error, an application programming interface (API) misuse, a quality issue, and a style issue.
3. The computer-implemented method of Claim 1, wherein the original source codeincludes multiple errors and the computer-implemented method further comprises: iterating the modifying, processing, and merging for each error of the multiple errors.
4. The computer-implemented method of Claim 1, wherein modifying the original sourcecode to generate the reduced source code including the error includes: generating a graph representation of the original source code having the error; iteratively: (i) reducing the graph and (ii) performing a static analysis of the reduced graph to determine existence of the error in the reduced graph until a minimal reduced graph including the error is determined; and modifying the original source code in accordance with the minimal reduced graph including the error to generate the reduced source code including the error. - 108 - 3924732.v36265.10150015. The computer-implemented method of Claim 4, wherein in a given iteration, reducing thegraph is based on approximate provenance information provided by a static analysis report.
6. The computer-implemented method of Claim 1, wherein the language model is anartificial intelligence based neural network.
7. The computer-implemented method of Claim 1, further comprising:training the language model.
8. The computer-implemented method of Claim 7, wherein training the language modelincludes: obtaining multiple code samples each including a respective error; obtaining multiple error-free code samples, each error-free code sample corresponding to a given code sample of the multiple code samples; reducing (i) the multiple code samples and (ii) the multiple error-free code samples; and training the language model to determine code fixes using the reduced multiple code samples and the reduced multiple error-free code samples.
9. The computer-implemented method of Claim 1, further comprising querying thelanguage model in a zero-shot or few-shot learning manner, wherein the querying includes: obtaining multiple code samples each including a respective error; obtaining multiple error-free code samples, each error-free code sample corresponding to a given code sample of the multiple code samples; reducing (i) the multiple code samples and (ii) the multiple error-free code samples; and - 109 - 3924732.v36265.1015001 querying the language model with few-shot learning where the reduced multiple code samples and the reduced multiple error-free code samples are provided as correct examples to the language model.
10. The computer-implemented method of Claim 1, wherein the merging includes:comparing the reduced source code including the error and the error-eliminated reduced source code to determine a mapping between (i) the reduced source code including the error and (ii) the error-eliminated reduced source code; and based on the mapping, merging the error-eliminated reduced source code with the original source code to generate the fixed source code.
11. The computer-implemented method of Claim 10, wherein, in the mapping, each line inthe error-eliminated reduced source code is mapped to a given line in the reduced source code including the error.
12. A computer-based system for fixing an error in source code, the computer-based systemcomprising: a processor; and a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the computer- based system to: modify original source code having an error to generate reduced source code including the error; process the reduced source code including the error with a language model to produce error-eliminated reduced source code; and merge the error-eliminated reduced source code with the original source code to generate fixed source code. - 110 - 3924732.v36265.101500113. The computer-based system of Claim 12, wherein the original source code includesmultiple errors and wherein the processor and the memory, with the computer code instructions, are further configured to cause the computer-based system to: iterate the modifying, processing, and merging for each error of the multiple errors.
14. The computer-based system of Claim 12 where, in modifying the original source code togenerate the reduced source code including the error, the processor and the memory, with the computer code instructions, are configured to cause the computer-based system to: generate a graph representation of the original source code having the error; and iteratively: (i) reduce the graph and (ii) perform a static analysis of the reduced graph to determine existence of the error in the reduced graph until a minimal reduced graph including the error is determined, wherein the minimal reduced graph including the error represents the generated reduced source code including the error.
15. The computer-based system of Claim 14 where, in reducing the graph, the processor andthe memory, with the computer code instructions, are further configured to cause the computer-based system to, in a given iteration, reduce the graph based on approximate provenance information provided by a static analysis report.
16. The computer-based system of Claim 12, wherein the language model is an artificialintelligence based neural network and where the processor and the memory, with the computer code instructions, are further configured to cause the computer-based system to: train the language model.
17. The computer-based system of Claim 16 where, in training the language model, theprocessor and the memory, with the computer code instructions, are further configured to cause the computer-based system to: obtain multiple code samples each including a respective error; - 111 - 3924732.v36265.1015001 obtain multiple error-free code samples, each error-free code sample corresponding to a given code sample of the multiple code samples; reduce (i) the multiple code samples and (ii) the multiple error-free code samples; and train the language model to determine code fixes using the reduced multiple code samples and the reduced multiple error-free code samples.
18. The computer-based system of Claim 12, wherein the processor and the memory, withthe computer code instructions, are further configured to cause the computer-based system to query the language model in a zero-shot or few-shot learning manner, wherein the querying includes: obtaining multiple code samples each including a respective error; obtaining multiple error-free code samples, each error-free code sample corresponding to a given code sample of the multiple code samples; reducing (i) the multiple code samples and (ii) the multiple error-free code samples; and querying the language model with few-shot learning where the reduced multiple code samples and the reduced multiple error-free code sample are provided as correct examples to the language model.
19. The computer-based system of Claim 12, wherein the processor and the memory, withthe computer code instructions, are further configured to cause the computer-based system to: compare the reduced source code including the error and the error-eliminated reduced source code to determine a mapping between (i) the reduced source code including the error and (ii) the error-eliminated reduced source code; and based on the mapping, merge the error-eliminated reduced source code with the original source code to generate the fixed source code. - 112 - 3924732.v36265.101500120. A non-transitory computer program product for fixing an error in source code, the non-transitory computer program product comprising a non-transitory computer-readable medium with computer code instructions stored thereon, the computer code instructions being configured, when executed by a processor, to cause an apparatus associated with the processor to: modify original source code having an error to generate reduced source code including the error; process the reduced source code including the error with a language model to produce error-eliminated reduced source code; and merge the error-eliminated reduced source code with the original source code to generate fixed source code. - 113 - 3924732.v3