Systems and methods for automated software code generation

CodeValet addresses the limitations of existing generative coding assistants by using predictive models to generate reliable, testable software code, achieving robust and efficient automated code generation.

WO2025099638A1PCT designated stage expired Publication Date: 2025-05-15CODEVALET INC

Patent Information

Application Number
PCT/IB2024/061052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-08
Filing Date
2024-11-07
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

Existing generative coding assistants struggle to produce reliable, working code due to context size-related issues and generative hallucinations, limiting their effectiveness in software development.

Method used

The CodeValet system employs predictive models, specifically generative models trained on large source code corpora, to generate accurate, testable software code. It includes features like fine-tuning on domain-specific codebases, automated testing pipelines, and continuous self-improvement.

Benefits of technology

CodeValet achieves robust automated code generation by producing complete, compilable code that is formally testable, reducing hallucinations, and improving efficiency through modular and scalable implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061052_15052025_PF_FP_ABST
    Figure IB2024061052_15052025_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for automatic software code generation utilizing one or more predictive model trained on code-related data, said code-related data involving obtained features and aspects from arbitrary code-related data. At least one instruction, such as natural language description(s), of at least one of desired code components and features can be processed by at least one of the at least one predictive models to produce corresponding code that matches at least one intent behind the at least one instruction. The generated code is analyzed and tested for correctness. At least one portion of the at least one trained predictive model leverages techniques such as attention mechanisms and transfer learning to improve contextual code generation. The present technology can be at least enabled with version control and codebases to ensure seamless integration with existing workflows. Automated code generation improves programmer productivity and enables rapid software prototyping.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEMS AND METHODS FOR AUTOMATED SOFTWARE CODE GENERATION

[0002] Cross Reference to Related Applications

[0003] This application claims the benefit of US. Provisional Patent Application 63 / 547,652, filed on November 8, 2023, which is incorporated by reference herein in its entirety.

[0004] Field

[0005] The present disclosure relates generally to artificial intelligence and machine learning technologies, and more specifically to systems and methods for automatic software code generation using neural networks

[0006] Background

[0007] In recent years, artificial intelligence (Al) has seen rapid advancements, especially in the field of natural language processing (NLP). Large neural network models such as GPT-3 and GPT-4, as well as other open source models such as Llama 2, have demonstrated the ability to generate remarkably human-like text for a variety of applications. There is growing interest in leveraging these powerful generative models for automating software development tasks such as code generation. However, existing generative coding assistants have limitations in producing reliable, working code due to multiple issues, including context size-related ones and / or generative hallucinations. Therefore, there is a need for systems and methods leveraging generative model(s) tailored for software development, typically and mainly code generation, that can produce accurate, testable software code. In conjunction to generating testable code, such methods and systems can further be utilized, in one or more embodiments, to generate software documentation and / or software code enabling the testing of software code. Summary

[0008] The present disclosure proposes methods and systems, herein referred as CodeValet, that uses one or more predictive models of one or more types, such as generative models, to automatically generate software-related data, mainly code, based on one or more types of one or more prompts, such as natural language prompts. A generative model leveraged by CodeValet is typically trained and / or has been previously trained on a large corpus of source code, aiming at learning common patterns and programming constructs. Given a natural language description of a desired code aspect, such as a feature and / or a component, CodeValet, after performing several processing actions, can produce corresponding code that matches the intent of the description. In association with a task of generating new code, CodeValet would typically at least one of analyze and test the generated code to ensure correctness and functionality and refer to crystalized knowledge enabling a validating of one or more portions of the new code. It will be understood that CodeValet’s leveraging such techniques enables uniquely robust automated code generation capabilities.

[0009] Some technical features and advantages of CodeValet can include: a. Autonomously obtaining, such as by generating, features and aspects from arbitrary data containing an arbitrary codebase, or representation thereof, aiming at obtaining crystalized knowledge associated therewith, enabling further aspects of CodeValet, such as finetuning of one or more predictive models and on- demand accessing of crystalized knowledge upon generation of, for instance, new code according to a user input. b. Leveraging fine-tuning of one or more generative models, such as a large language model, on domain-specific codebase(s) to improve code generation accuracy and reduce hallucinations. The models can access and incorporate structured knowledge about a given codebase. c. Generates complete, compilable code rather than just suggestions or snippets. The code is formally testable against requirements. d. Includes automated testing pipelines to validate at least one aspect, such as functionality, security, performance, code structure, etc., of generated code before deployment. Makes Al-generated code more robust. e. Modular, scalable implementation that allows distributed training and parallelized code generation. Improves efficiency. f. Can work from a minimal starting codebase or boilerplate; not strictly reliant on large existing code. g. Can continuously improve over time by learning from user feedback, new data, such as one or more portions of one or more CodeValet’s outputs, and / or version control. Has self-recursive and self-improving capabilities. h. Integrates with developer workflows like repositories and IDEs for seamless adoption. In embodiments, augments rather than displaces developers. i. Complementary techniques like crystallized knowledge libraries and encoder-decoder models to further enhance accuracy of natural language to code mappings. j. In summary, CodeValet uniquely combines aspects and features generation, codebase-specific fine-tuning, formal testing pipelines, modular implementation and continuous self-improvement to produce accurate and reliable Al- generated code tailored to the application domain. In embodiments, the emphasis is on integrating with existing developer workflows rather than fully automating them, enabling proper validation of generated code before use.

[0010] Brief Description of the Drawings

[0011] Figure 1 is a flow diagram illustrating an embodiment of a method for processing an input;

[0012] Figure 2 is a flow diagram illustrating an embodiment of a method for processing at least one first indication of at least one of at least one codebase and at least one codebase-related data;

[0013] Figure 3 is a system architecture diagram;

[0014] Figure 4 is a schematic diagram illustrating an electronic device which can be used in accordance with one or more non-limiting embodiments of the present technology; Figure 5 is a flow diagram illustrating an embodiment of at least one aspect of a method for processing at least one first indication of at least one of at least one codebase and at least one codebase-related data; and

[0015] Figure 6 if a flow diagram illustrating an embodiment of at least one aspect of a method for processing an input, in conjunction with the embodiment of Figure 5.

[0016] Description

[0017] Overview

[0018] Many Al coding assistants have been released this year, but they mostly generate untested, unreliable code or code suggestions to inspire developers. On the other hand, CodeValet is a coding autopilot aiming at creating solid and accurate code which it at least one of tests, debugs and deploys for the developer or architect. It also has self-improving and self-recursing capabilities, enabling CodeValet to learn from the codebase and its user to optimize performance and precision with each iteration.

[0019] While CodeValet initially ingest and acquire a deep level of understanding of an arbitrary codebase through various means, which generally requires compute resources increasingly higher in correlation with the arbitrary data size, such as the arbitrary codebase size, once such steps are made, new generations made by CodeValet would typically require less compute resources as CodeValet acquires knowledge related to the user’s input, such as knowledge obtained from user feedback(s), interaction(s), and content of the arbitrary data.

[0020] CodeValet uses natural language processing (NLP) instructions from the user to direct it to add to, change, or improve a given codebase with a combination of generative and heuristic methods that result in new and / or altered data, mainly code, that at least one of is formally testable against and aiming at being compliant with requirement(s) obtained in relation with one or more input of the user’s. Input(s) provided to CodeValet can also be, In embodiments, at least one indication of at least one output obtained from a one or more, typically a plurality of, operations previously performed by one or more instances of one or more implementations of CodeValet. Features and aspects obtaining

[0021] CodeValet’s first step is obtaining, such as by generating, features and aspects from arbitrary data containing an arbitrary codebase, or representation thereof. The first step of obtaining can typically involve performing a plurality of operations associated with the arbitrary data, such as the arbitrary codebase.

[0022] In embodiments, at least one subset of the operations is organized in the form of modular script(s). In embodiments, the script(s) can be added, removed and / or edited according to a user’s need and / or according to at least one other criterion. In embodiments, script(s) can be organized in folders. In embodiments, script(s) can be associated with one or more config files. In embodiments, one or more script(s) use at least one predictive model in their execution, such as a predictive model aiming at rendering one or more input unique compared to one or more previously generated output of the present technology according to one or more similar features and aspects obtaining step. In embodiments, an enabling of one or more script(s) can be done according to at least one criterion, such as the kind of data currently being processed. For example, script(s) aiming at establishing a call graph of programming elements like functions, methods, etc. would not necessarily be triggered in a processing of documentation(s). The present technology is not limited by the below-listed operations, and that various other operations can be provided, possibly influencing at least one result of at least one part of the operations.

[0023] Table 1. Example operations

[0024] It will thus be understood that the above-mentioned exemplary operations are presented by way of example, and that, typically, an operation can be added and / or removed from a given set of operation(s) and / or updated in a given set of operation(s).

[0025] In embodiments, at least one machine learning algorithm, at least one deep learning algorithm and / or the like can be used to at least one of analyze, determine and obtain at least one indication of at least one part of at least one result of at least one part of a given operation, such as paraphrasing at least one obtained question associated with one or more determined portions of the current processing input. In embodiments, at least one result of one or more of the operations is processed in such a way that renders it directly usable and suitable for training one or more given generative models. In embodiments, at least one indication of one or more results of at least one machine- readable instruction associated with one or more operations can be stored within one or more storage mediums.

[0026] Training methodology and datasets

[0027] CodeValet employs robust training methodology(ies) and dataset(s) to produce accurate code generation model(s).

[0028] In embodiments, the model training pipeline begins, in accordance with above-mentioned steps, by obtaining features / aspects of the arbitrary codebase and, In embodiments, related input arbitrary data. Training-usable features / aspects can be used in training techniques such as LoRA (Low Rank Adapters) and / or other fine-tuning techniques suitable to increase at least one metric / aspect of one or more given generative models, which, In embodiments, involves transformer-based model(s) leveraging attention-related mechanism(s). In embodiments, a plurality of models can be involved, aiming at creating models considered experts at their specific task(s) and / or in specific aspect(s) / field(s) related to the arbitrary data. In embodiments, at least one classification model can be trained in association with at least one generative model, aiming at enabling a classification, for instance, of one or more tasks a given user interacting with CodeValet, thereby enabling further actions by CodeValet, such as the determining of structured plan(s) to accomplish user requirements.

[0029] In embodiments, one or more code predictive models can be trained in accordance with backpropagation techniques over multiple epochs. In embodiments, at least one early stopping criterion, dropout(s), weight decay and / or regularization method(s) can be involved in the training of one or more predictive models. The present technology is not limited by these training-related techniques, and that various other techniques can be used, possibly influencing one or more metrics / aspects of one or more predictive models trained in association therewith. In addition to programming-related dataset(s), such as dataset(s) generated according to the above-mentioned steps, CodeValet, In embodiments, trains one or more generative models on natural language. In embodiments, one or more pretrained NLP models can be leveraged to enrich linguistic capability. In one or more embodiments, transfer learning techniques from large foundation models can be leveraged. In one or more embodiments, joint training techniques on, for instance, paired code and text examples are leveraged.

[0030] The result is one or more predictive models, comprising generative model(s) and / or other predictive models(s) enabling one or more action of CodeValet, such as determining the requirement(s) and, for instance, its type(s), and / or the like, enabling generation of code-related data according to an utilizing of Al-related techniques and methodologies, optimized training data, general purpose and domain- related knowledge to achieve superior results. Learned representations and correlations enable CodeValet to reliably provide functional code from input(s), such as human- readable descriptions.

[0031] Neural network model architecture, training, and NLP

[0032] CodeValet leverages advanced techniques, such as training and NLP techniques, and network architectures to ensure robust comprehension of instructions and descriptions, such as human-generated ones, for code generation. CodeValet can leverage neural network architecture(s) optimized for at least one aspect of software code generation. In embodiments, CodeValet leverages transformer-based language model(s), such as model(s) employing attention mechanisms like self-attention, to perform one or more aspects of code-related generation associated with one or more previously determined arbitrary codebases. Attention-related mechanisms can allow a given model to build contextual representations of word sequence(s) by relating element(s) to other element(s) in a passage. Attention-related mechanisms provides CodeValet with a deeper, more nuanced understanding of natural language compared to simpler, non-attention-related neural networks. While other model architectures can be used in one or more implementations of CodeValet, according to the current state-of- the-art, transformer-based models, such as Large Language Models (LLMs), typically outperform other generative model architectures. CodeValet aims at being modular and independent of a specific version of a specific model.

[0033] For efficiency reasons and due to it’s modularity, in one or more embodiments, CodeValet can leverage transfer learning from large pre-trained language model(s) with billions of parameters. Starting with robust general-purpose, typically code generation capable, model(s) that learned from massive datasets saves significant training time. Model(s) can further be fine-tuned on programming-related dataset(s) to improve one or more aspects, such as knowledge-related and understanding-related metrics, associated therewith. In one or more embodiments, model(s) that have already been fine-tuned at least once prior any CodeValet’s fine-tuning are preferred over least domain-specialized model(s). For instance, in one or more embodiments, model(s) determined as having better expertise in one or more programming languages, such as programming language(s) at least enabling implementation(s) of one or more aspects of CodeValet, are preferred.

[0034] In embodiments, CodeValet can also be able to leverage multi-task training, aiming at concurrently optimizing model aspect(s) / metric(s) on aligned text and code examples for, In embodiments, align the representations across modalities to improve mapping of specifications to implementations. It will be appreciated that techniques like knowledge distillation typically aim to transfer learned relationships from one model to another.

[0035] A given model generation skills can be enhanced through training on diverse dialogue, narrative and conversational dataset(s). This augments the linguistic breadth of the system beyond sterile technical prose. In embodiments, at least one portion of such enhancements / augmentations can be performed as part of the training of an obtained model, such as an open source model. To handle specialized vocabulary, CodeValet’s vocabulary embeddings can be supplemented with program-specific word vectors such as code documentation and API definitions.

[0036] In an embodiment, advanced techniques, such as Conditional Crystalized Knowledge (CCK) can be leveraged by CodeValet. Conditional Crystalized Knowledge comprises, In embodiments, technique(s) for embedding crystalized knowledge inside a predictive model in such a way that the predictive model can consider additional fed knowledge while in the process of generating new code-related data. CCK, In embodiments, can involve first determining if one or more sequences of characters / tokens can be present in a given input to the predictive model. Such knowledge can be a representation thereof, such as by leveraging encoder / decoder network(s), such as autoencoder network(s), to obtain one or more indications, such as condensed representations, of one or more code-related knowledge. In an embodiment, the representations can be at least one of injected in and trained within the predictive model. In an embodiment, CCK can involve adapter(s) in attention layer(s) of the predictive model, wherein said adapter(s) determine if at least one portion of the one or more representations should be considered by the predictive model at the current state of said predictive model in an active prediction process. In embodiments, the effective use of CCK require performing at least one training of a newly-CCK-altered predictive model. In an embodiment, CCK can further involve at least one classification mechanism, such as at least one gate, to determine at least one relevant portion of at least one part of the one or more representations for the predictive model to consider. In an embodiment, CCK is used in conjunction with other techniques mentioned herein.

[0037] Code generation process flow

[0038] The code generation process begins when instructions are provided, such as user provided natural language specification(s) describing the desired functionality(ies) or component(s) to be generated. This input, typically in a textual format, can be preprocessed with optimizing in mind, such as by leveraging normalization to resolve syntactic variations. The preprocessing, In embodiments, typically involves performing several transformations, such as at least one paraphrasing, at least one of at least one formatting and at least one structuring, at least one at least identifying of at least one potential concern preventing an enabling of a successful processing of the current input, at least one organizing and / or at least one summarizing. In embodiments, one or more transformations can be enabled by at least one criterion, such as an evaluation of a result of a previous transformation in order to determine the relevancy of further transformation(s). In embodiments, one or more transformation(s) can involve one or more previously-trained predictive models. For instance, In embodiments, one or more classification models can be leveraged in such a way that enables higher accuracy of the different types of requirement(s) comprised in the input.

[0039] CodeValet thereby performs several other transformations on the current input by leveraging previously crystalized knowledge, let it be in the internal layers of one or more given models or external / accessed, in order to obtain a clear, step by step plan that would enable a completion and implementation of identified and clarified requirements in the original input. A continuous refinement of such plan can be performed, let it be by leveraging voting-related mechanisms involving multiple predictive models aiming at determining any concern / issue / unclear aspect associated with the (original and / or current) input and / or the requirements and eliminating such by performing actions such as fetching and querying external data sources, such as documentations, obtaining user validations, and / or the like. In one or more embodiments, the refining of the plan can involve iterating through various files and folders, according to the arbitrary codebase(s), and fetching summaries associated therewith to determine at least one relevancy aspect of one or more portions of the plan. A visual representation of such refinement(s) can be found in several portions of Figure 1.

[0040] In embodiments, CodeValet can leverage several times in the refining process structured abstract syntax tree (AST) to further enhance the plan.

[0041] In an embodiment, the code portions of the plan can be refined in accordance with determined guidelines and rules, such as PEP8-related guidelines. In embodiments, rules, such as a maximum number of lines per function / other coderelated component, one comment prior each for loop / function / other code-related component, using snake case, and / or the like can be enforced and of great relevancy, aiming at, and thereby helping, at least one generative model to stick to a determined coding structure. In embodiments, rules such as not going above a determined number of actions per function can be enforced, enabling a higher degree of parallelism and better test(s) generation. In embodiments, one or more rules / guidelines can be detected from at least one part of provided arbitrary codebase(s). Upon successfully defining one or more clear and granular plans, CodeValet can thereby start executing such. In embodiments, at least one portion of the executing is enabled by at least one criterion, such as at least one user validation / approval.

[0042] In an embodiment, the generating of code, as well as associated test(s), associated with the one or more plans can be distributed across agent(s), i.e. , predictive model(s)-enabled code generators and testers, typically, when possible, multiple agents, thereby enabling parallelism. By generating not only code to meet the requirements associated with the input but also code to test such is beneficial in several ways and enables CodeValet to make sure of the robustness of new requirement- oriented code. In embodiments using it, multi-layered testing enables a validating of correctness, security, robustness and other quality attributes. In embodiments, the test- oriented code can be generated according to at least one criterion, such as one or more parts, such as the parameters, of generated requirements-oriented code, context(s) obtained from the target arbitrary codebase, at least one portion of the plan associated with such generating and / or generating of the requirement-oriented code, determined, such as user-determined, parameters and rules in CodeValet, and / or the like. In embodiments, the testing process can follow a progression of unit, integration, system, stress and other tests. If any issues are uncovered during testing, new code generation can thereby be enabled, as appropriate, and according to at least one criterion, such as the one or more failed tests, and / or the nature of identified issue(s). For instance, In embodiments, issues can have a varying degree of graveness, and such degree of graveness can directly be correlated or enable which part(s) of the code, let it be the requirement-oriented code and / or the test-oriented code, needs to be (re)generated. In embodiments, one or more aspects and / or portions of such issues-derived execution can be provided as feedback to incrementally improve CodeValet's model(s). In embodiments, code can also be patched based on test insights before the final output is presented to the user.

[0043] Additional documentation such as function comments and in-line annotations can be automatically generated to produce comprehensive, production- ready code. The validated code is finally returned to the user along with any generated documentation, tests, and metadata.

[0044] Testing and validation procedures

[0045] An advantage of CodeValet is the automated testing and validation of generated code before deployment. This helps ensure the correctness and functionality of the Al-produced code. Some testing and validation techniques employed by one or more embodiments of CodeValet could include:

[0046] • Unit Testing - CodeValet applies unit tests to exercise and validate the behavior of individual code components such as functions or methods. Unit tests are generated based on analysis of the component's logic and interface.

[0047] • Integration Testing - Groups of unit-tested components are combined and tested together to verify interoperability. Stubbing and mocking can be used to simulate dependencies where needed.

[0048] • Regression Testing - Existing test suites are re-run against the updated codebase to detect any new bugs or regressions introduced by changes.

[0049] • Static Analysis - The source code is analyzed without execution to detect code quality issues, security vulnerabilities, etc.

[0050] • Dynamic Analysis - The running code is monitored and profiled to capture runtime metrics for optimization.

[0051] • Fuzz Testing - Random or unexpected inputs are fed to methods to ensure robustness and account for edge cases.

[0052] • Code Review - Human code reviewers can be used to analyze the style, design and documentation of generated code for maintainability.

[0053] • Beta Testing - The code can deployed to a small set of users before full release to garner feedback.

[0054] In embodiments, one or more aspects of these tests involving human feedback can be implemented to leverage Al agents, i.e., predictive model(s)-enabled machine readable instructions, instead and / or in compliment to human(s).

[0055] The results of the testing and validation procedures can be used to further improve CodeValet's models and code generation skill. This incremental learning approach leads to higher quality code over time. We will now discuss a few of these topics in greater detail.

[0056] Dynamic Analysis

[0057] In embodiments, CodeValet utilizes various forms of dynamic analysis to capture runtime metrics and validate the real-world performance of the Al-generated code. One technique employed can be performance profiling, where CodeValet instruments identify sections of the code to add timers and gauges. This allows CodeValet to measure metrics like time spent in specific functions, database queries, API calls and other segments. By identifying the most time-consuming portions of the code, optimization opportunities can be uncovered.

[0058] In embodiments, in addition to timing data, CodeValet can also track memory allocation and object lifetimes in the running software via memory profiling. Monitoring how memory is utilized at runtime enables CodeValet to detect potential leaks and inefficient usage patterns. This ensures that the generated code is optimized for memory usage. In embodiments, to understand behavior under load, the instrumented code can be exercised with concurrent simulated users and request volumes at test scale. This load testing pinpoints limitations related to threads, I / O or caching bottlenecks. In embodiments, critical sections that need to be concurrently refactored can be revealed.

[0059] Stress testing is another validation technique where computing resources like CPU, memory and network bandwidth can be artificially constrained during execution to simulate workload-related stress. This verifies the stability and recovery capabilities of the generated code when subjected to extreme conditions. For distributed systems, In embodiments, CodeValet implements distributed tracing by propagating unique transaction IDs across process boundaries. This tracks flows across services and helps isolate latency issues in complex environments.

[0060] In an embodiment, additional logging statements can be inserted into the code to capture runtime debug variables, environment information and other contextual data. These logs enhance visibility for diagnosing production issues. In embodiments, key application metrics can also be tracked in real-time and stored in time-series databases for analysis. Statistical monitoring of these telemetry streams allows CodeValet to detect anomalies and emerging performance regressions.

[0061] In embodiments, any unhandled exceptions raised in the software can be reported back to CodeValet for diagnosis. This error telemetry enables previously undetected bugs that escaped testing to be identified and fixed. Leveraging such runtime validation techniques thereby typically enables CodeValet to continuously refine and enhance the performance, stability and reliability of the code produced by its generative models. The insights from dynamic analysis can also inform incremental improvements to the Al system itself.

[0062] Enhancements such as version control integration

[0063] In an embodiment, CodeValet scans an arbitrary code repository to construct an inventory of the files, directories, classes, functions and other artifacts. This mapping of the codebase provides crucial context. When a code generation request is received, In embodiments, CodeValet consults the repo mapping to determine the optimal placement for injecting the new code as part of the plan-determining process above-described. In an embodiment, the repository is processed according to the above-mentioned features and aspects obtaining steps.

[0064] In an embodiment, CodeValet generates changesets containing the added / modified code files that can be automatically committed and pushed to the repo. Commit messages summarize the implemented features based on the original request and would, In embodiments, be generated in association with one or more predictive models. For larger changes, in embodiments, CodeValet would stage at least one portion of the changes on a feature branch and / or one or more determined types of branches. This allows incremental refinement before merge. CodeValet manages the pull request workflow to solicit code reviews. In embodiments, such would enable CodeValet to analyze version histories to identify prior relevant changes that can inform the current generation task. Code ancestry helps determining knowledge.

[0065] In an embodiment, CodeValet can also ingest specs for changes to existing methods or components. Such, in an embodiment, would involve to locate target artifacts via the repo map, checks out the specific file versions, modifiying the code, then submitting the deltas. In embodiments, unit tests can be generated for the new code and integrated into the project's testing hierarchies. In embodiments, the tests would also be versioned to support continuous validation.

[0066] In embodiments, usage patterns of generated code can be passively collected. In embodiments, low utility artifacts can be refactored or consolidated to optimize the architecture. In embodiments, the repo itself evolves dynamically. In embodiments, by integrating with developers' native version control workflows, CodeValet would aim to feel like a natural extension of the team rather than an external disruption, improving adoption while still enabling transformative automation of coding tasks.

[0067] In embodiments, at least one predictive model can be encapsulated within at least one first container for ease of deployment, and further wherein at least one portion of the at least one first container includes at least one of at least one necessary dependency and at least one configuration enabling an executing of its associated at least one predictive model. In embodiments, the at least one first container involves at least one Docker container. The use of Docker container(s) enhances at least one aspect, such as deployment aspect(s), in comparison with non-containerized model(s). In embodiments, the at least one first container can include one or more machine- readable instructions, such as scripts and / or commands, which enhances at least one aspect, such as at least one facility aspect, of the execution of the associated predictive model(s). For example, in embodiments, the one or more machine-readable instructions include at least one instruction to load at least one required dependency, configure at least one target deployment environment, run the predictive model and / or the like. In embodiments, the one or more machine-readable instructions further involves handling input and / or output data, such as fetching data from and / or writing data to at least one database, processing data prior providing at least one related, such as derived, input to the predictive model(s) associated therewith, processing data obtained from the predictive model(s) associated therewith prior performing further processing and / or the like.

[0068] In embodiments, at least one of the at least one first container can be implemented using a cloud-based infrastructure, such as Amazon Web Services (AWS), Microsoft Azure or Google Cloud Platform (GCP). In embodiments, at least one of the at least one first container can be configured to communicate with one or more other components associated with one or more steps of the present technology, such as at least one data storage module and / or at least one user interface module. For instance, a given container can include APIs and / or messaging protocols that allow it to provide to and / or obtain data from one or more other modules. In embodiments, a given container can also be configured to receive updates or new versions of the predictive model, such as leveraging at least one version control system, allowing the technology to adapt to changing conditions or improve its performance over time.

[0069] With reference to Figure 2, a flow diagram illustrating processing at least one first indication of at least one of at least one codebase and at least one codebaserelated data is shown. As shown, an input, such as documents, raw codebase or previous storage information or data, such as newly generated or implemented code by Code Valet is examined. CodeValet will determine if a transformation is required at this moment in time and will then proceed to a subsequent step. For example, if CodeValet determines that a transformation is not required, then the input is transmitted to storage for later use. If CodeValet determines that transformation is required, then CodeValet proceeds with the transformation. As shown, transformation can include, but is not limited to graph generation, summarizing, vectorizing, paraphrasing, compression, organizstion, formatting, structuring, data augmentation, tokenization, data normalization, norice reduction, filtering compiling, combining, etc. Upon completion of the first transformation, CodeValet further determines if subsequent transformations are required. If no, then the input is transmitted to storage and if yes, then transformation is initiated.

[0070] During transformations, checks, such as uniqueness checks can be performed according to criteria that can be established by the user.

[0071] If the imput is to be stored, then the input is stored in a manner that permits the input to be fetched at a later date. For example, input can be stored as formatted commands. In embodiments, some data can be stored in different ways other than data, such as summaries of files, folders, call graphs, etc. which can enable easeier fetching or validation when generating new code upon user input. Exemplary electronic device

[0072] As shown in Figure 3, in an embodiment, a system architecture can comprise a network that interconnects a user device and remote CodeValet execution machines and databases. As shown, a user can use the user device to remotely access the CodeValet execution machine and database via the internet and request execution outputs or generation of code based upon code-related data that the user inputs into the CodeValet execution machine.

[0073] With reference to Figure 4, a schematic diagram illustrating an electronic device 1000 which can be used in accordance with one or more non-limiting embodiments of the present technology is shown.

[0074] Referring to Figure 4, there is shown an electronic device 1000 suitable for use with at least one implementation of the present technology, the electronic device 1000 comprising various hardware components including one or more single or multicore processors collectively represented by processor 1002, one or more graphics processing units (GPU) 1004, one or more storage drives such as one or more solid- state drives 1006, one or more random-access memory 1008, one or more display interfaces 1010, and one or more input / output interfaces 1012.

[0075] Communication between the various components of the electronic device 1000 can be enabled by one or more internal and / or external buses 1014, such as a PCI bus, universal serial bus, IEEE 1394 "Firewire" bus, SCSI bus, Serial-ATA bus, and / or the like, to which the various hardware components can be electronically coupled.

[0076] At least one of the one or more input / output interfaces 1012 can be coupled to one or more touchscreens 1016 and / or to one or more internal and / or external buses 1014. A given touchscreen 1016 can be portion of at least one given display. In embodiments, the touchscreen(s) 1016 is the display. The touchscreen(s) 1016 can equally be referred to as screen(s) 1016. In embodiments, the touchscreen(s) 1016 comprises touch hardware 1018 (e.g., pressure-sensitive cells embedded in a layer of a display allowing detection of a physical interaction between a user and a given display) and one or more touch input / output controllers 1020 allowing communication with the display interface(s) 1010 and / or the one or more internal and / or external buses 1014. In embodiments, the input / output interface 1012 can be connected to a keyboard (not shown), a mouse (not shown) and / or a trackpad (not shown) allowing the user to interact with the electronic device 1000 in addition or in replacement of the touchscreen 1016. In embodiments, a plurality of electronic devices 1000 can be used in conjunction to enable further scalability, such as parallelism.

[0077] According to implementations of the present technology, the solid-state drive(s) 1006 stores machine-readable instructions suitable for being loaded into the random-access memory 1008 and executed by the processor(s) 1002 and / or the GPU(s) 1004.

[0078] The electronic device 1000 can be implemented as a server, a desktop computer, a laptop computer, a tablet, a smartphone, a given, or a plurality thereof, device comprising and / or being a device capable of executing computer interpretable instructions, such as an ARDUINO®, a personal digital assistant and / or any device that can be configured to implement the present technology, as it can be understood by a person skilled in the art.

[0079] Referring now to Figure 5, the flow diagram illustrates a two-phase training process for a system designed to generate software code. In Phase 1 , generalized training is performed on various codebases using an encoder / decoder network pattern. This phase aims to capture general programming constructs and patterns across different languages and frameworks. The encoder processes input data, such as Python or other code files, and generates condensed embeddings representing the essential features of the input. These embeddings are typically stored in long-term memory for later use.

[0080] In Phase 2, fine-tuning is applied to specific codebases for individual files or other related data. During this phase, the system uses the previously generated embeddings from Phase 1 and further refines them to adapt to the particularities of a specific project or domain. This fine-tuning process ensures that the model can generate more accurate and context-relevant code based on domain-specific knowledge. The output of this phase is a set of condensed embeddings that are suitable for use in downstream tasks such as code generation or analysis. Turning now to Figure 6, this flow diagram illustrates how all files or coderelated data are encoded into latent spaces and stored for future use. After completing the two-phase training process described in Figure 5, each file or piece of code-related data is processed by the system’s encoder. The encoder converts the input into latent space representations — compact, high-dimensional vectors that capture the underlying semantics of the code. These latent space representations are then stored in long-term memory for later retrieval during tasks such as predictive modeling or generating new code based on user input. By storing these representations, the system can efficiently access pre-learned knowledge about various codebases without needing to reprocess large amounts of raw data each time it generates new outputs.

[0081] Figures 5 and 6 depict how the system leverages both generalized training across multiple codebases and fine-tuning on specific projects to create a robust foundation for automated software code generation. The use of latent space representations allows for efficient storage and retrieval of knowledge, enabling faster and more accurate responses to user queries or prompts for new code generation.

Claims

Claims:1 . A system for automated software code generation, comprising: at least one processor; and at least one memory device storing instructions that, when executed, cause the system to: obtain features and aspects from identified code-related data, receive user input comprising instructions for desired code, process the input using a predictive model to generate a structured plan for implementing requirements identified in the user input, generate code corresponding to the desired code, and analyze and test the generated code for correctness.

2. The system of claim 1 , further comprising instructions that, when executed, cause the system to: train the predictive model using the obtained features and aspects; and employ a Conditional Crystallized Knowledge (CCK) process to embed domain-specific knowledge into the predictive model.

3. The system of claim 1 , further comprising instructions that, when executed, cause the system to present the structured plan to the user for approval, and receive an indication of user approval of the structured plan before generating code corresponding to the desired code.

4. The system of claim 3, further comprising instructions that, when executed, cause the system to output validated code after analyzing and testing the generated code for correctness.

5. The system of claim 1 , wherein the predictive model and structured plan are used to generate code corresponding to the desired code.

6. The system of claim 2, wherein the CCK process comprises: generating condensed representations of code-related knowledge using encoder-decoder networks; embedding the condensed representations within the predictive model; implementing conditional access mechanisms to allow the predictive model to selectively utilize the embedded knowledge during code generation.

7. The system of claim 6, wherein the conditional access mechanisms comprise adapters in attention layers of the predictive model, the adapters determining which portions of the embedded representations should be considered by the predictive model during code generation.

8. The system of claim 6, wherein implementing the CCK process further comprises training the predictive model after embedding the condensed representations to enable effective utilization of the embedded knowledge.

9. The system of claim 6, wherein implementing the CCK process further comprises implementing at least one classification mechanism to determine relevant portions of the embedded representations for the predictive model to consider during code generation.

10. The system of claim 9, wherein the at least one classification mechanism comprises at least one gate to control access to different portions of the embedded representations based on the current code generation task.11 . The system of claim 1 , wherein analyzing and testing the generated code comprises performing at least one of: unit testing, integration testing, regression testing, static analysis, dynamic analysis, and fuzz testing.

12. The system of claim 1 , wherein the instructions further cause the system to create test code corresponding to the generated code.

13. The system of claim 12, wherein the act of creating test code corresponding to the generated code comprises: analyzing parameters and structure of the generated code; determining context from a target codebase associated with the generated code; generating unit tests based on the analysis of the generated code and the determined context; generating integration tests to verify interoperability of the generated code with existing code components; and implementing at least one of: regression tests, static analysis, dynamic analysis, and fuzz testing; and integrating the generated test code into existing testing hierarchies of the target codebase.

14. The system of claim 1 , wherein the instructions further cause the system to integrate the generated code with a version control system.

15. The system of claim 1 , wherein obtaining features and aspects from code-related data comprises performing a plurality of operations organized as modular scripts.

16. The system of claim 15, wherein the plurality of operations comprises at least one of: obtaining indications of function, method, or class calls from a given file; obtaining indications of dependencies and associated declarations from a given file; obtaining indications of coding style and error handling from function, method, or class declarations;obtaining indications of file and folder structures representing implemented features; obtaining indications of graphs associated with programming elements within files; and obtaining indications of code lines to generate "TODO" summaries.

17. The system of claim 16, wherein the instructions further cause the system to: organize the modular scripts in folders; associate the modular scripts with one or more configuration files; and enable or disable one or more of the modular scripts based on at least one criterion related to the type of data being processed.

18. The system of claim 15, wherein at least one of the modular scripts uses at least one predictive model in its execution to generate unique outputs compared to previously generated outputs for similar features and aspects.

19. The system of claim 15, wherein obtaining features and aspects further comprises: processing results of the plurality of operations to render them suitable for training the predictive model; and storing indications of results of the plurality of operations within one or more storage mediums.

20. The system of claim 1 , wherein the instructions further cause the system to preprocess the input by performing at least one of: paraphrasing, formatting, structuring, identifying potential concerns, organizing, and summarizing.21 . The system of claim 20, wherein preprocessing the input comprises preprocessing the input to normalize syntactic variations; classifying requirements within the input; generating a structured plan to implement the classified requirements; and refining the plan iteratively.

22. The system of claim 1 , wherein the instructions further cause the system to encapsulate at least one of the one or more predictive models within a container for deployment, wherein the container is a Docker container configured to execute on a cloud-based infrastructure.

23. The system of claim 1 , wherein obtaining features and aspects from identified code-related data comprises analyzing function calls, method declarations, and class structures within the codebase; extracting dependency information and coding styles; and generating a call graph of programming elements.

24. The system of claim 2, wherein training the predictive model comprises fine-tuning a pre-trained large language model on the obtained features and aspects; and employing a transfer learning process to leverage knowledge from the pretrained model.

25. The system of claim 1 , further comprising: integrating the generated and tested code with a version control system; automatically creating code commits and pull requests; and managing code placement within an existing repository structure.

26. The system of claim 1 , wherein the instructions further cause the system to continuously improve the predictive model based on feedback from testing results and user interactions.

27. The system of claim 1 , wherein analyzing and testing the generated code for correctness comprises: generating unit tests for individual code components; performing integration testing on groups of components; conducting static and dynamic code analysis; executing fuzz testing with random inputs; and running regression tests against existing test suites.

28. A method for automated software code generation, comprising: training a predictive model; receiving user input comprising instructions for desired code; embedding domain-specific knowledge into the predictive model; processing the input using the predictive model to generate a structured plan for implementing requirements identified in the user input; presenting the structured plan for approval; receiving an indication of user approval of the structured plan; using the predictive model and structured plan to generate code corresponding to the desired code; testing the generated code for correctness; and outputting validated code.

29. The method of claim 28, wherein a Conditional Crystallized Knowledge (CCK) process is employed to embed domain-specific knowledge into the predictive model.

30. The method of claim 29, wherein the CCK process comprises: generating condensed representations of code-related knowledge using encoder-decoder networks; embedding the condensed representations within the predictive model; and implementing conditional access mechanisms to allow the predictive model to selectively utilize the embedded knowledge during code generation.31 . The method of claim 30, wherein: the conditional access mechanisms comprise adapters in attention layers of the predictive model, the adapters determining which portions of the embedded representations should be considered by the predictive model during code generation; andthe CCK process further comprises training the predictive model after embedding the condensed representations to enable effective utilization of the embedded knowledge, and implementing at least one classification mechanism to determine relevant portions of the embedded representations for the predictive model to consider during code generation, wherein the at least one classification mechanism comprises at least one gate to control access to different portions of the embedded representations based on the current code generation task.

32. The method of claim 28, wherein the predictive model comprises a transformer-based neural network employing attention mechanisms.

33. The method of claim 28, wherein analyzing and testing the generated code comprises performing at least one of: unit testing, integration testing, regression testing, static analysis, dynamic analysis, and fuzz testing.

34. The method of claim 28, further comprising the following steps to create test code corresponding to the generated code: analyzing parameters and structure of the generated code; determining context from a target codebase associated with the generated code; generating unit tests based on the analysis of the generated code and the determined context; generating integration tests to verify interoperability of the generated code with existing code components; and implementing at least one of: regression tests, static analysis, dynamic analysis, and fuzz testing; and integrating the generated test code into existing testing hierarchies of the target codebase.

35. The method of claim 28, wherein the instructions further cause the system to integrate the generated code with a version control system.

36. The method of claim 28, wherein obtaining features and aspects from code-related data comprises performing a plurality of operations organized as modular scripts, wherein the plurality of operations comprises at least one of: obtaining indications of function, method, or class calls from a given file; obtaining indications of dependencies and associated declarations from a given file; obtaining indications of coding style and error handling from function, method, or class declarations; obtaining indications of file and folder structures representing implemented features; obtaining indications of graphs associated with programming elements within files; and obtaining indications of code lines to generate "TODO" summaries.

37. The method of claim 36, wherein the instructions further cause the system to: organize the modular scripts in folders; associate the modular scripts with one or more configuration files; and enable or disable one or more of the modular scripts based on at least one criterion related to the type of data being processed.

38. The method of claim 36, wherein at least one of the modular scripts uses at least one predictive model in its execution to generate unique outputs compared to previously generated outputs for similar features and aspects.

39. The method of claim 36, wherein obtaining features and aspects further comprises: processing results of the plurality of operations to render them suitable for training the predictive model; and storing indications of results of the plurality of operations within one or more storage mediums.

40. The method of claim 28, wherein the instructions further cause the system to preprocess the input by performing at least one of: paraphrasing, formatting, structuring, identifying potential concerns, organizing, and summarizing.41 . The method of claim 40, wherein preprocessing the input comprises: preprocessing the input to normalize syntactic variations; classifying requirements within the input; generating a structured plan to implement the classified requirements; and refining the plan iteratively.

42. The method of claim 28, wherein the instructions further cause the system to encapsulate at least one of the one or more predictive models within a container for deployment, wherein the container is a Docker container configured to execute on a cloud-based infrastructure.

43. The method of claim 28, wherein obtaining features and aspects from identified code-related data comprises analyzing function calls, method declarations, and class structures within the codebase; extracting dependency information and coding styles; and generating a call graph of programming elements.

44. The method of claim 28, wherein training the predictive model comprises fine-tuning a pre-trained large language model on the obtained features and aspects; and employing a transfer learning process to leverage knowledge from the pretrained model.

45. The method of claim 28, further comprising: integrating the generated and tested code with a version control system; automatically creating code commits and pull requests; and managing code placement within an existing repository structure.

46. The method of claim 28, wherein the instructions further cause the system to continuously improve the predictive model based on feedback from testing results and user interactions.

47. The method of claim 28, wherein analyzing and testing the generated code for correctness comprises: generating unit tests for individual code components; performing integration testing on groups of components; conducting static and dynamic code analysis; executing fuzz testing with random inputs; and running regression tests against existing test suites.

Citation Information

Patent Citations

  • Integrated system and method for the management of a complete end-to-end software delivery process

    US20030046681A1

  • Artificial intelligence and machine learning infrastructure

    US20200125941A1

  • Efficient bundling and delivery of client-side scripts

    US20200150939A1

  • Utilizing machine learning models for automated software code modification

    US20220244937A1

  • Automated code generation based on pseudo-code

    US20230041718A1

Cited By

  • Code automatic generation and optimization system based on multiple modes

    CN120276718A