Method and system for large language model-based lint tools

WO2026199286A1PCT designated stage Publication Date: 2026-10-01MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/085244
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure CN2025085244_01102026_PF_FP_ABST
    Figure CN2025085244_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Dependency documents are identified for a project and are stored in a dependency knowledge base. The dependency documents are also converted into document dependency vectors. A lint tool query event for source code triggers the generation of a prompt for an LLM. The prompt includes source code referenced by the lint tool and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code. The prompt is submitted to the LLM with the source code and further augmented with one or more selected dependency documents. At least a portion of the response to the prompt from the LLM is presented with the source code.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR LARGE LANGUAGE MODEL-BASED LINT TOOLSBACKGROUND

[0001] In the realm of software development, lint tools have become indispensable for aiding programmers in identifying and rectifying potential errors in their code. Lint tools analyze source code to flag programming errors, bugs, stylistic errors, and suspicious constructs and provide the developer with selectable options for fixing, analyzing and developing their source code. For example, a lint tool may be configured to highlight suspicious code and / or provide functionality for selecting code to trigger the presentation of a pop-up window or menu with options to evaluate or modify the highlighted and / or selected code with potential fixes or suggested considerations. The primary function of lint tools is to improve code quality and ensure adherence to coding standards.

[0002] One limitation associated with current lint tools is that they access the suggested fixes and other information provided to the developers from leveraged models based on the lint prompts that essentially comprise only the referenced source code. This can be a problem because many models are trained with aged or historical knowledge, which can make it difficult for them to correlate the referenced source code with the most up-to-date programming knowledge. As a result, the suggestions and fixes provided by these tools may not reflect the latest advancements in programming languages, frameworks, and best practices.

[0003] Several problems can result when a model is provided a source code prompt from a lint tool that utilizes conventions that the model is not trained on. For instance, the model may generate suggestions based on outdated programming practices which, if incorporated, can resulting in inferior programming quality and programmer frustration. It can also result in an inability of the lint tools to correctly handle new programming languages, libraries, or frameworks, leading to missed opportunities for optimization and error correction.

[0004] To address these challenges, there is a need for improved systems and methods that incorporate the latest programming knowledge into lint tools.

[0005] The subject matter claimed herein is not limited to embodiments that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is only provided to illustrate one example technology area where some embodiments described herein may be practiced.SUMMARY OF THE INVENTION

[0006] Systems and methods are provided for facilitating large language model (LLM) –based lint tools and, even more particularly, for facilitating access to and incorporation of updated dependency documentation utilized by the lint tools.

[0007] In some aspects, the techniques described herein relate to methods for facilitating large language model (LLM) -based lint tools, the methods include act for identifying a project; identifying dependencies of the project; identifying repositories for dependency documents corresponding to the dependencies; crawling the repositories for the dependency documents and storing a copy of the dependency documents in a dependency knowledge base; converting the dependency documents into document dependency vectors; detecting a lint tool query event for source code associated with the project; and creating a prompt for a LLM in response to a query event.

[0008] In some aspects, the prompt includes at least a portion of the aforementioned source code and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents include a threshold similarity match to the vector of the source code.

[0009] In some aspects, the techniques described herein relate to a method, wherein the method further includes submitting the prompt with the source code and one or more selected dependency documents to the LLM; and presenting the lint tool with a response to the prompt from the LLM for presentation with the source code.

[0010] In some aspects, the techniques described herein relate to a method, wherein the method further includes causing the lint tool to present at least a portion of the response with the source code.

[0011] In some aspects, the techniques described herein relate to a method, wherein the method further includes storing the document dependency vectors in a dependency vector database. This may include splitting the dependency documents into a plurality of snippets and vectorizing each of the snippets as a separate vector that is stored in the dependency vector database.

[0012] In some aspects, the techniques described herein relate to a method, wherein the query event that triggers the prompt generation includes a scan of source code. The event may also be triggered additionally, or separately, in response to a user selection of a portion of the source code. The selected portion of the source code may include a highlighted portion of the source code highlighted by the lint tool, for example.

[0013] In some aspects, the techniques described herein relate to a method, wherein determining the document dependency vectors corresponding to the one or more selected dependency documents include a threshold similarity match to the vector of the source code includes selecting a set of top N document dependency vectors having a closest similarity to the vector of the source code, wherein N is a predetermined number.

[0014] In some aspects, computing systems are configured with one or more hardware processors and stored computer-executable instructions that are executable by the one or more processors to implement the disclosed methods.

[0015] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to describe the manner in which the advantages and features of the systems and methods described herein can be obtained, a more particular description of the embodiments briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the appended drawings. Understanding that these drawings depict only typical embodiments of the systems and methods described herein, and are not therefore to be considered to be limiting of their scope, certain systems and methods will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:

[0017] FIG. 1 illustrates a computing environment that may be utilized to implement the disclosed methods for facilitating identification and masking of secrets in processed text.

[0018] FIG. 2 illustrates an example of an improved lint tool processing flow.

[0019] FIG. 3 illustrates a flowchart of acts associated with methods for facilitating large language model (LLM) –based lint tools.DETAILED DESCRIPTION OF THE INVENTION

[0020] As disclosed herein, systems and methods are provided for facilitating large language model (LLM) –based lint tools and, even more particularly, for facilitating access to and incorporation of updated dependency documentation utilized by the lint tools.

[0021] Beneficially, the disclosed embodiments augment lint tool prompts containing source code with dependency documentation identified based on matching similarities of the vectors of current project dependency documents with the vectors of the project source code. In this manner, the LLMs receive augmented prompts that enable the LLMs to better identify the most current and contextually relevant response for presentation by the lint tool.

[0022] When the dependency documentation is processed with the source code by the models in this manner, it effectively operates as updated training data that the models are applied to. In this manner, the models leveraged by the lint tools can provide more accurate, relevant, and effective suggestions, thereby enhancing the overall quality of software development.

[0023] Attention is now directed to FIG. 1, which illustrates a computing system that can comprise or be utilized to facilitate the disclosed embodiments. As shown, the computing system 100 is in communication with one or more user system (s) 110 and / or third-party system (s) 120. In some implementations, the user system (s) 110 and third-party system (s) 120 are remotely located from the computing system 100 and are independently controlled computing systems. In other implementations, the user system (s) 110 and / or third-party system (s) 120 comprise distributed components of the computing system 100, such that they share storage and processing capabilities.

[0024] Computing system 100 is connected to the other user system (s) 110 and third-party system (s) 120 through a network of wired and / or wireless connections, such as currently represented as the cloud.

[0025] Each of the illustrated systems includes input and output devices (I / O devices 130) for receiving inputs and rendering outputs, respectively, even though they are only explicitly shown for computing system 100. Non-limiting examples of input devices include microphones, keyboards, mouse devices, touch pads, and camera sensors. Non-limiting examples of output devices include speakers, desktop display screens, mixed-reality display devices, and haptic feedback devices.

[0026] The disclosed systems also include one or more storage system (s) 140 of volatile and / or non-volatile storage and one or more hardware processor (s) 150 configured to execute computer-executable instructions 160 stored in the storage system (s) 140 to cause the computing system 100 to implement the methods and functionality disclosed herein.

[0027] The storage system (s) 140 also store, as described in more detail below, the referenced dependency knowledge base 165 and vector database 170 (which store the referenced source code vectors and dependency document vectors and which may be incorporated into the dependency knowledge base) , as well as the referenced project files and source code (180) .

[0028] The storage system (s) 140 also stores, in some instances, training data (not shown) used to train the machine learning model (s) 180 to generate responses to the lint prompts and to perform the vector matching processes described herein.

[0029] Interfaces (180) , such as lint tool interfaces, and other interfaces are also provided to facilitate interfacing with the model (s) 185 and remote user and third-party system (s) that store dependency documentation, as well as with other system components and to present the processed prompt responses through the lint tool (s) to a developer or other end-user.

[0030] As noted above, computing system 100 can be utilized to implement the disclosed methods, including the methods associated with the improved lint tool processing flow of FIG. 2, and the corresponding acts illustrated in flowchart 300 of FIG. 3 for facilitating the processing of the lint tools, particularly for facilitating the processing of LLM-based lint tools.

[0031] As shown, in FIG. 2, the disclosed improved lint tool processing flow includes four basic processes, namely, project scanning, dependency processing, vector database processing and Retrieval Augmented Generation (RAG) processing for the open file (s) .

[0032] The project scanning generally relates to the evaluation of a project and the files of the project, as well as the identification of the dependencies within the project. The referenced dependency processing generally relates to the identification of documents associated with the dependencies and storing of those documents in a dependency knowledge base. The vector database processing generally relates to the embedding or vectorization of the dependency documents, as composite vectors and / or as discrete vectors of dependency document snippets. The vector database processing can also include the vectorizing of source code and matching of the source code vectors with the dependency document vectors to identify the corresponding dependency documents most contextually relevant to the source code. Finally, the RAG processing for the open file (s) generally relates to the generation of prompts for a lint tool that includes referenced source code and dependency documents that correspond to the source code (which were selected based on the vector matching process referenced above) .

[0033] These processes are described in more detail with reference to the acts shown in flowchart 300 of FIG. 3 and which may be implemented by the computing system of FIG. 1. As shown, these acts include the identification of a project, dependencies of the project, and repositories for dependency documents corresponding to the dependencies of the project (310) , the storing dependency documents in a dependency knowledge base and dependency document vectors in a dependency vector database (act 320) , the creation and submission of a prompt for an LLM with project source code and further augmented one or more dependency documents selected from the dependency knowledge base based on similarity of the corresponding dependency document vectors with the vector (s) of the project source code (act 330) , and obtaining and presenting a response to the prompt from the LLM with the source code (act 340) .

[0034] Each of these acts will now be described in more detail.

[0035] Initially, it is noted that the act of identifying a project, dependencies in the project, and the repositories for the dependency documents (act 310) can be performed in different manners according to different needs and preferences, as well as based on the different types of lint tools and programming languages that are being used.

[0036] In some instances, a project corresponds to one or more program files of source code being developed. In some instances, the project is associated with a folder that stores the source code program files and other project files, such as dependency files.

[0037] An integrated development environment (IDE) which may include the lint tools referenced herein, will also include, in some instances, controls and menus for selecting a project and for identifying all files associated with a project.

[0038] Depending on the type of programming language being used, the project files that identify or include the dependencies can vary. For instance, JavaScript and TypeScript projects can include package. jason, package-lock. json, and / or yarn. lock dependency file (s) . A Python project can include requirements. txt, Pipfile and / or pyproject. toml dependency file (s) . A Java project can include pom. xml and / or build. gradle dependency file (s) . C# (. NET) project dependency files can include . csproj, . sln, and / or packages. config. Go project dependency files can include go. mod and / or go. sum. Rust project dependency files can include Cargo. toml and / or Cargo. lock dependency files. PHP projects can include composer. json and / or composer. lock files.

[0039] The foregoing listing of project types and dependency files are only illustrative and non-exhaustive. That said, the disclosed methods include referencing an index that identifies different types of dependency files and searching the project files for those dependency files. The methods also include scanning the source code files and file metadata for declared dependencies.

[0040] Once the dependencies are identified, the methods also include identifying the different local and / or third-party repositories corresponding to the different dependencies. In some instances, the repositories are the same third-party repositories. In other instances, different language-based dependencies correspond to different third-party repositories.

[0041] In some instances, by way of non-limiting example, the JavaScript and TypeScript languages correspond to the industry known npm registry, whereas the Python projects correspond to the industry known PyPI (Python Package Index) repository, the Java projects correspond to industry known Maven central repository, C# (. NET) projects correspond to the industry known NuGet repository, the Go projects correspond to the industry known Go Proxy repository, the Rust projects correspond to the industry known Crates. io repository, and PHP projects correspond to the industry known Packagist repository.

[0042] Notwithstanding the foregoing examples, the systems may query an index, or an LLM or other tool to identify existing repositories for documentation corresponding to the identified dependencies.

[0043] Once the dependencies are identified, the system crawls the documents in the repository for the identified dependencies. The documentation being crawled for includes coding samples and examples of use, API descriptions, problem and solution descriptions, issue discussions, and other information related to the dependencies referenced.

[0044] The system downloads a text format copy of the dependency documentation and stores those copies in the aforementioned dependency knowledge base. The system also creates and stores dependency document vectors for each of the dependency documents in a vector database, which may be part of the dependency knowledge base (act 320) .

[0045] The system may make a single vector per document and / or a plurality of vectors for each document. For instance, the system may segment the documents into different snippets that are each vectorized and associated with the different portions of the documents they correspond to.

[0046] An LLM or other model with embedding functionality can be used to convert the text snippets and / or whole documents into vectors. A vector index correlates the vectors to the different documents and document segments.

[0047] The next acts correspond to the RAG processing and corresponding query processes used by the lint tool.

[0048] For instance, when working on developing code, a lint tool can be used to identify source code and that is provided to a LLM in a query. However, before the query is submitted, according to the present embodiments, it is augmented with dependency documentation from one or more of the dependency documents stored in the dependency knowledge base.

[0049] The selection of the dependency documents to augment the query with is based on the similarity between the vectors of the dependency documents and a newly generated vector corresponding to the referenced source code.

[0050] Accordingly, in some instances, the system identifies the referenced source code from the lint tool and generates a vector embedding, using the same vector embedding technology used for generating the dependency documentation vectors. Then, a comparison is made between the vector of the source code and the indexed dependency documentation vectors.

[0051] The system determines which document dependency vectors comprise a threshold similarity match to the vector of the source code and then selects the corresponding dependency documents to be used for augmenting the source code in the query to the LLM. This similarity match can be a percentage match or another type of predetermined match, such as, but not limited to, the cosine similarity match.

[0052] According to some instances, the similarity match is based on the closest top N matches of the indexed dependency documentation vectors to the vector of the source code. Once the dependency documentation vectors matching the vector of the source code within the predetermined threshold are identified, their corresponding dependency documentation is selected for augmenting the prompt with the source code.

[0053] The prompt may be augmented with the totality of the selected dependency documents and / or only a portion of the selected dependency documents to accommodate different needs and preferences (e.g., a limit in prompt size and / or to omit redundant or cumulative information) .

[0054] The prompt may also be augmented with commands to instructions to find a fix or example, for instance, to satisfy the functionality provided by the lint tool.

[0055] The prompt is then submitted to the LLM to obtain a corresponding response that is provided to the lint tool. The system then causes the lint tool to present the response, or at least a portion of the response, to the end-user. In some instances, this includes presenting the response (e.g., fix, suggestion, example (s) , snippets of code, contextual information, links to third-party content) with the reference source code that was selected or highlighted to trigger the query event and prompt generation / submission process.

[0056] In some instances, the query event comprises a scan of the source code by the lint tool when it is loaded into an IDE for development. In some instances, the query event comprises a user selection of a portion of the source code concurrently with or separate from the scan of the source code by the lint tool. In some instances, the query event is triggered by a user selecting source code highlighted by the lint tool after scanning the source code to highlight suspicious code or code that matches a user query.

[0057] As noted previously, the foregoing functionality can be implemented by the computing system 100 of FIG. 1, which may take various different forms and which may include and / or be in communication with one or more user systems and third-party system.

[0058] For example, computing system 100 may be embodied as a tablet, a desktop, a laptop, a mobile device, or a standalone device, such as those described throughout this disclosure. Computing system 100 may also be a distributed system that includes one or more connected computing components / devices that are in communication with computing system 100.

[0059] In its most basic configuration, computing system 100 includes various different components, such as the referenced processors. Without limitation, illustrative types of hardware logic components / processors that can be used include Field-Programmable Gate Arrays ( “FPGA” ) , Program-Specific or Application-Specific Integrated Circuits ( “ASIC” ) , Program-Specific Standard Products ( “ASSP” ) , System-On-A-Chip Systems ( “SOC” ) , Complex Programmable Logic Devices ( “CPLD” ) , Central Processing Units ( “CPU” ) , Graphical Processing Units ( “GPU” ) , or any other type of programmable hardware.

[0060] As used herein, the terms “executable module, ” “executable component, ” “component, ” “module, ” “service, ” or “engine” can refer to hardware processing units or to software objects, routines, or methods that may be executed on computing system 100. The different components, modules, engines, and services described herein may be implemented as objects or processors that execute on computing system 100 (e.g. as separate threads) .

[0061] The referenced storage system (s) may be physical system memory, which may be volatile, non-volatile, or some combination of the two. The term “memory” may also be used herein to refer to non-volatile mass storage such as physical storage media. If computing system 100 is distributed, the processing, memory, and / or storage capability may be distributed as well.

[0062] The storage system (s) may include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general-purpose or special-purpose computer system. Computer-readable media that store computer-executable instructions in the form of data are “physical computer storage media” or a “hardware storage device. ” Furthermore, computer-readable storage media, which includes physical computer storage media and hardware storage devices, exclude signals, carrier waves, and propagating signals. On the other hand, computer-readable media that carry computer-executable instructions are “transmission media” and include signals, carrier waves, and propagating signals. Thus, by way of example and not limitation, the current embodiments can comprise at least two distinctly different kinds of computer-readable media: computer storage media and transmission media.

[0063] Computer storage media (aka “hardware storage device” ) are computer-readable hardware storage devices, such as RAM, ROM, EEPROM, CD-ROM, solid state drives ( “SSD” ) that are based on RAM, Flash memory, phase-change memory ( “PCM” ) , or other types of memory, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions, data, or data structures and that can be accessed by a general-purpose or special-purpose computer.

[0064] Computer system 100 may also be connected (via a wired or wireless connection) to external sensors (e.g., one or more remote cameras) or devices via a network defined as one or more data links and / or data switches that enable the transport of electronic data between computer systems, modules, and / or other electronic devices. For example, computer system 100 can communicate with any number of devices or cloud services to obtain or process data. In some cases, the network may itself be a cloud network (e.g., the cloud shown in FIG. 1) . Furthermore, computer system 100 may also be connected through one or more wired or wireless networks to remote / separate computer systems (s) that are configured to perform any of the processing described with regard to computer system 100.

[0065] When information is transferred, or provided, over a network (either hardwired, wireless, or a combination of hardwired and wireless) to a computer, the computer properly views the connection as a transmission medium. Computer system 100 will include one or more communication channels that are used to communicate with the network. Transmissions media include a network that can be used to carry data or desired program code means in the form of computer-executable instructions or in the form of data structures. Further, these computer-executable instructions can be accessed by a general-purpose or special-purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0066] Upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to computer storage media (or vice versa) . For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a network interface card or “NIC” ) and then eventually transferred to computer system RAM and / or to less volatile computer storage media at a computer system. Thus, it should be understood that computer storage media can be included in computer system components that also (or even primarily) utilize transmission media.

[0067] Computer-executable (or computer-interpretable) instructions comprise, for example, instructions that cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a certain function or group of functions. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0068] Those skilled in the art will appreciate that the embodiments may be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, pagers, routers, switches, and the like. The embodiments may also be practiced in distributed system environments where local and remote computer systems that are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network each perform tasks (e.g. cloud computing, cloud services and the like) . In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0069] The present invention may be embodied in other specific forms without departing from its characteristics. The embodiments described are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0070] The present invention can also be described in accordance with the following numbered clauses.

[0071] Clause 1. A method for facilitating large language model (LLM) -based lint tools, the method comprising: identifying a project; identifying dependencies of the project; identifying repositories for dependency documents corresponding to the dependencies; crawling the repositories for the dependency documents and storing a copy of the dependency documents in a dependency knowledge base; converting the dependency documents into document dependency vectors; detecting a lint tool query event for source code associated with the project; creating a prompt for a LLM in response to the query event, the prompt including the source code and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code; submitting the prompt with the source code and one or more selected dependency documents to the LLM; and presenting the lint tool with a response to the prompt from the LLM for presentation with the source code.

[0072] Clause 2. The method of clause 1, wherein the method further includes causing the lint tool to present at least a portion of the response with the source code.

[0073] Clause 3. The method of clause 1, wherein the method further includes identifying the repositories from an index that identifies different repositories corresponding to different types of programming languages.

[0074] Clause 4. The method of clause 1, wherein the method further includes storing the document dependency vectors in a dependency vector database.

[0075] Clause 5. The method of clause 4, wherein converting the dependency documents into document dependency vectors by at least splitting the dependency documents into a plurality of snippets and vectorizing each of the snippets as a separate vector that is stored in the dependency vector database.

[0076] Clause 6. The method of clause 1, wherein the query event comprises a scan of the source code.

[0077] Clause 7. The method of clause 1, wherein the query event comprises a scan of the source code coupled with a user selection of a portion of the source code.

[0078] Clause 8. The method of clause 7, wherein the selected portion of the source code comprises a highlighted portion of the source code highlighted by the lint tool.

[0079] Clause 9. The method of clause 1, wherein determining the document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code comprises selecting a set of top N document dependency vectors having a closest similarity to the vector of the source code, wherein N is a predetermined number.

[0080] Clause 10. A computing system comprising: one or more hardware processor; and one or more hardware storage device having stored computer-executable instructions that are executable by the one or more hardware processor to cause the computing system to implement a method for facilitating large language model (LLM) -based lint tools, the method comprising: identifying a project; identifying dependencies of the project; identifying repositories for dependency documents corresponding to the dependencies; crawling the repositories for the dependency documents and storing a copy of the dependency documents in a dependency knowledge base; converting the dependency documents into document dependency vectors; detecting a lint tool query event for source code associated with the project; creating a prompt for a LLM in response to the query event, the prompt including the source code and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code; submitting the prompt with the source code and one or more selected dependency documents to the LLM; and presenting the lint tool with a response to the prompt from the LLM for presentation with the source code.

[0081] Clause 11. The computing system of clause 10, wherein the method further includes causing the lint tool to present at least a portion of the response with the source code.

[0082] Clause 12. The computing system of clause 10, wherein determining the document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code comprises selecting a set of top N document dependency vectors having a closest similarity to the vector of the source code, wherein N is a predetermined number.

[0083] Clause 13. The computing system of clause 10, wherein the method further includes storing the document dependency vectors in a dependency vector database.

[0084] Clause 14. The computing system of clause 13, wherein converting the dependency documents into document dependency vectors by at least splitting the dependency documents into a plurality of snippets and vectorizing each of the snippets as a separate vector that is stored in the dependency vector database.

[0085] Clause 15. The computing system of clause 10, wherein the query event comprises a scan of the source code.

[0086] Clause 16. The computing system of clause 10, wherein the query event comprises a scan of the source code coupled with a user selection of a portion of the source code.

[0087] Clause 17. The computing system of clause 16, wherein the selected portion of the source code comprises a highlighted portion of the source code highlighted by the lint tool.

[0088] Clause 18. The computing system of clause 10, wherein the method further includes identifying the repositories from an index that identifies different repositories corresponding to different types of programming languages.

Claims

1.A method for facilitating large language model (LLM) -based lint tools, the method comprising:identifying a project;identifying dependencies of the project;identifying repositories for dependency documents corresponding to the dependencies;crawling the repositories for the dependency documents and storing a copy of the dependency documents in a dependency knowledge base;converting the dependency documents into document dependency vectors;detecting a lint tool query event for source code associated with the project;creating a prompt for a LLM in response to the query event, the prompt including the source code and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code;submitting the prompt with the source code and one or more selected dependency documents to the LLM; andpresenting the lint tool with a response to the prompt from the LLM for presentation with the source code.2.The method of claim 1, wherein the method further includes causing the lint tool to present at least a portion of the response with the source code.3.The method of claim 1, wherein the method further includes identifying the repositories from an index that identifies different repositories corresponding to different types of programming languages.4.The method of claim 1, wherein the method further includes storing the document dependency vectors in a dependency vector database.5.The method of claim 4, wherein converting the dependency documents into document dependency vectors by at least splitting the dependency documents into a plurality of snippets and vectorizing each of the snippets as a separate vector that is stored in the dependency vector database.6.The method of claim 1, wherein the query event comprises a scan of the source code.7.The method of claim 1, wherein the query event comprises a scan of the source code coupled with a user selection of a portion of the source code.8.The method of claim 7, wherein the selected portion of the source code comprises a highlighted portion of the source code highlighted by the lint tool.9.The method of claim 1, wherein determining the document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code comprises selecting a set of top N document dependency vectors having a closest similarity to the vector of the source code, wherein N is a predetermined number.10.A computing system comprising:one or more hardware processor; andone or more hardware storage device having stored computer-executable instructions that are executable by the one or more hardware processor to cause the computing system to implement a method for facilitating large language model (LLM) -based lint tools, the method comprising:identifying a project;identifying dependencies of the project;identifying repositories for dependency documents corresponding to the dependencies;crawling the repositories for the dependency documents and storing a copy of the dependency documents in a dependency knowledge base;converting the dependency documents into document dependency vectors;detecting a lint tool query event for source code associated with the project;creating a prompt for a LLM in response to the query event, the prompt including the source code and one or more selected dependency documents from the dependency knowledge base, wherein the one or more selected dependency documents are selected in response to creating a vector of the source code and determining document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code;submitting the prompt with the source code and one or more selected dependency documents to the LLM; andpresenting the lint tool with a response to the prompt from the LLM for presentation with the source code.11.The computing system of claim 10, wherein the method further includes causing the lint tool to present at least a portion of the response with the source code.12.The computing system of claim 10, wherein determining the document dependency vectors corresponding to the one or more selected dependency documents comprise a threshold similarity match to the vector of the source code comprises selecting a set of top N document dependency vectors having a closest similarity to the vector of the source code, wherein N is a predetermined number.13.The computing system of claim 10, wherein the method further includes storing the document dependency vectors in a dependency vector database.14.The computing system of claim 13, wherein converting the dependency documents into document dependency vectors by at least splitting the dependency documents into a plurality of snippets and vectorizing each of the snippets as a separate vector that is stored in the dependency vector database.15.The computing system of claim 10, wherein the query event comprises a scan of the source code.16.The computing system of claim 10, wherein the query event comprises a scan of the source code coupled with a user selection of a portion of the source code.17.The computing system of claim 16, wherein the selected portion of the source code comprises a highlighted portion of the source code highlighted by the lint tool.18.The computing system of claim 10, wherein the method further includes identifying the repositories from an index that identifies different repositories corresponding to different types of programming languages.