Code adoption rate determination method and device, equipment, storage medium and product
By building a code line information database and inverted index, the problem of difficult to quantify the contribution of large language models in the software development process is solved, and the accurate evaluation of the code adoption rate is achieved, and the development efficiency and code quality are improved.
Patent Information
- Application Number
- CN202510979810.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-16
AI Technical Summary
The actual contribution of large language models in the software development process is difficult to quantify and evaluate, making it difficult for developers to track and evaluate their code adoption.
Build a code line information library and establish an inverted index. By matching the auxiliary code generated by the large language model and the new code added to the target code repository, determine the code adoption rate.
This has achieved quantitative evaluation of the contribution of large language models in the software development process, and improved the determination efficiency and accuracy of code adoption rate.
Smart Images

Figure CN120492310A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of software development technology, and in particular to a method, apparatus, device, storage medium, and product for determining a code adoption rate. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, large language models (LLMs) have become a promising new auxiliary tool in software development. They can assist developers with a variety of programming tasks, including code generation, code completion, code interpretation, bug fixes, and test case generation, significantly improving development efficiency and code quality. Developers are increasingly interacting with these large language models through various channels, including integrated development environment (IDE) plugins, web interfaces, and command-line tools.
[0003] However, while LLMs can generate a large number of coding suggestions, it is often difficult to accurately track whether and to what extent developers adopt these suggestions. This makes it difficult to quantify the actual contribution of large language models in the software development process.
[0004] Regarding related technologies, the problem that the actual contribution of large language models in the software development process is difficult to quantify and evaluate has not yet been effectively solved. Summary of the Invention
[0005] The present application provides a method, apparatus, device, storage medium, and product for determining code adoption rate, to at least address the problem in related technologies that the actual contribution of large language models in the software development process is difficult to quantify and evaluate.
[0006] The present application provides a method for determining a code adoption rate, comprising: constructing a code line information library including a first code, and constructing an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, and newly added code in a target code warehouse within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate a correspondence between a first word element included in the first code and position information of a code line including the first word element in the first code in the code line information library; matching a second code with the first code through the code line information library and the inverted index to determine a code adoption rate corresponding to the auxiliary code, wherein, if the first code is the auxiliary code, the second code is the newly added code, and if the first code is the newly added code, the second code is the auxiliary code.
[0007] The present application also provides a device for determining a code adoption rate, comprising: a construction module for constructing a code line information library including a first code, and constructing an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, and newly added code in a target code warehouse within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate the correspondence between a first word element included in the first code and position information of a code line including the first word element in the first code in the code line information library; a matching module for matching a second code with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein, if the first code is the auxiliary code, the second code is the newly added code, and if the first code is the newly added code, the second code is the auxiliary code.
[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for determining a code adoption rate when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for determining the code adoption rate are implemented.
[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for determining code adoption rate when the computer program is executed by a processor.
[0011] Through the present application, a code line information library including a first code is constructed, and an inverted index corresponding to the first code is constructed based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, newly added code in the target code warehouse within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate the correspondence between the first word included in the first code and the position information of the code line including the first word in the first code in the code line information library; the second code is matched with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code. Therefore, the technical problem in the related art that the actual contribution of the large language model in the software development process is difficult to quantify and evaluate can be solved, and the technical effect of determining the actual contribution of the large language model based on the code adoption rate is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 This is a hardware structure block diagram of a computer terminal for a method for determining a code adoption rate according to an embodiment of the present application;
[0014] Figure 2 is a flowchart of a method for determining a code adoption rate according to an embodiment of the present application;
[0015] Figure 3 This is the architecture of a system for determining a code adoption rate according to an embodiment of the present application;
[0016] Figure 4 4 is a structural block diagram of a device for determining a code adoption rate according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0020] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the method for determining the code adoption rate depends, the specific application environment architecture or specific hardware architecture is described herein.
[0021] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for determining a code adoption rate according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microcontroller unit (MCU) or a field-programmable gate array (FPGA) and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the code adoption rate in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] Figure 2 Flowchart of the method for determining the code adoption rate according to an embodiment of the present application. Figure 2 As shown, the process includes the following steps:
[0025] Step S202: construct a code line information library including a first code, and construct an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, or newly added code in a target code repository within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate a correspondence between a first word element included in the first code and position information of a code line in the first code including the first word element in the code line information library;
[0026] Step S204: Match the second code with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein, when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code.
[0027] It should be noted that, in the embodiments of the present application, the auxiliary code generated by the large language model is generally used as the first code to construct the inverted index and the code line information library, and the newly added code of the target code repository is used as the second code to calculate the code adoption rate.
[0028] Through the above steps, a code line information library including a first code is constructed, and an inverted index corresponding to the first code is constructed based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, newly added code in the target code warehouse within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate the correspondence between the first word element included in the first code and the position information of the code line including the first word element in the first code in the code line information library; the second code is matched with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein, when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code. Therefore, the technical problem in the related art that the actual contribution of the large language model in the software development process is difficult to quantify and evaluate can be solved, and the technical effect of determining the actual contribution of the large language model based on the code adoption rate is achieved.
[0029] An embodiment of the present application provides a method for determining a code adoption rate, and the method is described in detail in conjunction with the execution flow of the method for determining a code adoption rate.
[0030] In an exemplary embodiment, constructing a code line information library including a first code includes: determining a valid code area in the first code and parsing a programming language type to which the first code belongs; extracting each line of code in the valid code area, and storing each line of code and code information of each line of code in the code line information library to construct the code line information library, wherein the code information includes: the first code and the programming language type.
[0031] Optionally, determining a valid code area in the first code includes: determining a text format of the first code, and locating a target symbol in the first code based on the text format; and determining a code block surrounded by the target symbol in the first code as the valid code area.
[0032] For example, in a text containing Markdown format, locate and extract the code block surrounded by specific symbols (equivalent to the target symbols, such as three backquotes \`\`\`) as the valid code area.
[0033] Each line of code may be normalized and then stored in a code line information library. The code information of each line of code also includes: information of the user who requested the generation of the first code, the creation time of the first code, and a response identifier (ID) of the first code.
[0034] Through this application, by identifying valid code areas and parsing programming languages, the accuracy and efficiency of code analysis are improved, the possibility of cross-language matching is promoted, and the data storage and retrieval process is optimized, providing strong support for efficient and reliable code adoption evaluation.
[0035] In an exemplary embodiment, constructing an inverted index corresponding to the first code based on the code line information library includes: determining a second word in each line of code, wherein the first word includes: the second word; creating the inverted index according to the second word and a unique identifier corresponding to each line of code in the code line information library, wherein the unique identifier is used to indicate the location information of each line of code in the code line information library.
[0036] Furthermore, determining the second word in each line of code includes one of the following: determining the second word in each line of code through a lexical analyzer according to the grammatical rules corresponding to the programming language type, wherein the lexical analyzer is a dedicated lexical analyzer corresponding to the programming language type; determining the second word in each line of code through a preset matching method, wherein the preset matching method includes: regular expression.
[0037] Prioritize using a lexical analyzer specifically tailored to the language type identified for each line of code. This analyzer (for example, by calling a function like `lexer.get_tokens(code)`, where `lexer` is the lexical analysis engine for the specific programming language and `code` is the text of the line of code to be analyzed) accurately decomposes the line of code into a series of lexical units (i.e., tokens) and their types (for example, distinguishing between variable names, keywords, and comments) according to the strict grammatical rules of the language. This precise parsing based on the structure of the programming language is key to ensuring the quality of token recognition and subsequent indexing.
[0038] If you can't find a language-specific analyzer during processing, or if a dedicated lexical analyzer has difficulty analyzing a line of code (for example, if the code snippet is incomplete or contains minor grammatical errors), you can enable a predefined matching method (such as regular expressions) as a fallback token extraction strategy. While predefined matching methods don't have the same deep understanding of code structure as a dedicated analyzer, they can still effectively extract the vast majority of meaningful text fragments, such as potential identifiers and comments, from the text, thereby preserving the most information.
[0039] Furthermore, the inverted index is created based on the second term and the unique identifier corresponding to each line of code in the code line information library, including: if an entry corresponding to the second term already exists in the index library corresponding to the first code, adding the unique identifier to the first position list corresponding to the entry to create the inverted index; if an entry corresponding to the second term does not exist in the index library corresponding to the first code, creating a new entry in the index library based on the second term, and adding the unique identifier to the second position list corresponding to the new entry to create the inverted index. The second term is each valid term extracted from each line of code.
[0040] This application constructs an inverted index based on a code line information library, which not only greatly improves the efficiency and accuracy of code search and matching, but also provides an efficient, accurate and scalable technical means for evaluating the adoption of artificial intelligence (AI) generated code through its flexibility, low storage requirements and powerful analytical capabilities.
[0041] In an exemplary embodiment, matching the second code with the first code using the code line information library and the inverted index to determine a code adoption rate corresponding to the auxiliary code includes: determining a third token in the second code; determining a third code in the first code that matches the third token using the code line information library and the inverted index; and determining the code adoption rate using the third code. The third token is each token in the second code.
[0042] Furthermore, determining the third code in the first code that matches the third word element through the code line information library and the inverted index includes: traversing the inverted index through the third word element to determine a third position list corresponding to the third word element in the index library where the inverted index is located, wherein the third position list is a list corresponding to entries created based on the third word element in the index library; determining the third code from the code line information library through the third position list, wherein the third position list records the unique identifier of the third code.
[0043] That is, the inverted index is used to determine the third position list corresponding to each word (i.e., the third word) in the second code in the index library, and then the third position list is used to determine all third codes in the first code that include the third word.
[0044] Through this application, compared with full text comparison, word-based inverted index retrieval greatly reduces the amount of calculation, and reduces resource consumption and processing time in the code matching process.
[0045] Furthermore, determining the code adoption rate through the third code includes: determining the text similarity of the fourth code and the fifth code, wherein the fourth code is any code line in the second code that includes the third word, and the fifth code is any code line in the third code; and adjusting the text similarity through type information corresponding to the fourth code and the fifth code respectively, wherein the type information includes: programming language type and code type; and determining the code adoption rate through the adjusted text similarity.
[0046] Optionally, determining the text similarity of the fourth code and the fifth code includes: respectively disassembling the fourth code and the fifth code to obtain a first word set and a second word set; determining the intersection and difference of the first word set and the second word set; and determining the text similarity through the intersection and the difference.
[0047] That is, the present application adopts a comparison strategy based on token sets to calculate text similarity. This strategy first decomposes the fourth code and the fifth code into sets of their constituent tokens. Subsequently, the intersection and difference of the two token sets are analyzed, and an algorithm based on edit distance (such as Levenshtein distance) is used to quantify the similarity between these set components, and finally a comprehensive assessment of the similarity between the two contents is made. Optionally, the text similarity between the fourth code and the fifth code can also be determined by calculating the Jaccard similarity coefficient of the intersection and the union, where the union is the union of the first token set and the second token set.
[0048] Optionally, the text similarity is adjusted according to the type information corresponding to the fourth code and the fifth code, including: comparing the programming language types corresponding to the fourth code and the fifth code, respectively, to obtain a first result; and determining whether the code types corresponding to the fourth code and the fifth code, respectively, are target types, to obtain a second result; and adjusting the text similarity according to the first result and the second result.
[0049] If the first comparison result indicates that the programming language types corresponding to the fourth code and the fifth code are inconsistent and do not belong to the predefined "related languages", a penalty is imposed on the text similarity and marked as "language mismatch".
[0050] Determine whether the code types corresponding to the fourth code and the fifth code are target types, wherein the target type is annotation content; if at least one of the code types of the fourth code and the fifth code is annotation content, the text similarity can be lowered. It should also be noted that if at least one of the code types of the fourth code and the fifth code is annotation content, the similarity threshold can also be increased. Exemplarily, the current fourth code is code 1, and there are 3 fifth codes, namely code 2, code 3, and code 4, which are used to calculate the text similarity with the current fourth code. If the code with the highest text similarity with code 1 (and the text similarity is higher than the similarity threshold) among codes 2, 3, and 4 is code 3, then code 3 and the current fourth code are the best matches to each other.
[0051] Optionally, determining the code adoption rate through the adjusted text similarity includes: determining a sixth code in the second code that matches the first code through the adjusted text similarity; and determining the code adoption rate through the number of code lines of the sixth code and the number of code lines of the second code.
[0052] A ratio of the number of code lines of the sixth code to the number of code lines of the second code is calculated, and the ratio is used as a code adoption rate.
[0053] In an exemplary embodiment, after matching the second code with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, the method further includes: determining each line of code corresponding to the first code, and obtaining a target inverted index pre-constructed for the second code, wherein each line of code is each line of code in the valid code area of the first code; determining a first code line list in the second code that matches each line of code through the target inverted index, wherein the target inverted index is used to indicate the correspondence between the third word included in the second code and the position information of the code line in the second code including the third word; determining the missed code lines in the second code through the first code line list.
[0054] Optionally, a target inverted index and a target code line information library may be established for the second code in the same manner as establishing an inverted index and a code line information library for the first code. The target inverted index is used to indicate the correspondence between the third token included in the second code and the position information of the code line in the second code including the third token in the target code line information library.
[0055] Furthermore, the first code line list that matches each line of code in the second code is determined through the target inverted index, including: determining each word in each line of code; determining a fourth position list through each word and the target inverted index, wherein the fourth position list is a list corresponding to entries created based on each word in the target index library where the target inverted index is located, and the fourth position list records a target unique identifier, wherein the target unique identifier is used to indicate the position information of the code line in the second code including each word; determining the first code line list through the fourth position list.
[0056] Furthermore, determining the missed code lines in the second code through the first code line list includes: filtering out a second code line list from the first code line list through the metadata of each line of code, wherein the metadata includes: user information, generation timestamp; determining the missed code lines through each line of code and the second code line list.
[0057] That is, the code lines in the first code line list that have consistent user information for each line of code and corresponding generation timestamps are filtered out to obtain a second code line list. The generation timestamps correspond, meaning that the generation timestamps of the code lines allowed to be filtered out in the first code line list must be later than the generation timestamps of each line of code, and the interval between the two generation timestamps must be a fixed length (e.g., 7 days).
[0058] In an exemplary embodiment, before constructing a code line information library including the first code, the method further includes: obtaining a model call request for the large language model initiated by a user through different access portals; obtaining response content generated by the large language model in response to the model call request; uniformly formatting the response content, and determining the first code based on the uniformly formatted response content.
[0059] In order to better understand the process of the above-mentioned method for determining the code adoption rate, the implementation process of the above-mentioned method for determining the code adoption rate is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of the present application.
[0060] In the related technologies, the practice of large language model-assisted programming faces the following technical challenges that need to be addressed urgently:
[0061] 1) Difficulty collecting and analyzing data during LLM interactions: Developers access LLM services through different portals, and their interaction data (such as questions, model responses, and contextual information) is scattered and formatted differently. The lack of a unified, systematic logging mechanism makes it difficult to comprehensively record and structure this interaction data, hindering in-depth analysis and optimization of LLM usage, developer habits, and model performance.
[0062] 2) The actual adoption of LLM-generated code is difficult to quantify: Although LLM can generate numerous coding suggestions, it is often difficult to accurately track whether and to what extent developers have adopted these suggestions. Traditional code audits and version comparisons are inefficient and costly in accurately identifying and quantifying the actual contribution of LLM-generated content to project code. This makes it very difficult to assess the true value and return on investment (ROI) of LLM in software development.
[0063] 3) The accuracy of code adoption assessments needs to be improved: When developers adopt code generated by LLM, they often make adaptive modifications based on their needs, such as renaming variables, adjusting code structure, and adding or removing comments. These modifications can prevent simple text matching algorithms from effectively identifying code homology, leading to false negatives (FNs, a type II error in statistics and machine learning) and underestimating the actual contribution of LLM.
[0064] Therefore, there is an urgent need for a technical solution that can systematically and automatically record the LLM interaction process, capture the change history of the code repository, and use efficient and accurate analysis methods to evaluate the adoption of LLM-generated code and its actual impact on the software development process. The embodiments of this application aim to provide such a solution to overcome the above-mentioned problems existing in the related art.
[0065] This embodiment of the present application provides a comprehensive code generation traceability and adoption evaluation system (hereinafter referred to as the system, which is used to implement the above-mentioned code adoption rate determination method) specifically for large language model (LLM)-assisted programming scenarios. Through integrated log collection, code version analysis, efficient text indexing, and intelligent matching mechanisms, the system fully records the interaction process between developers and LLM, accurately captures new code added to Git code repositories, and automatically evaluates the adoption of LLM-generated content in actual projects.
[0066] It's important to note that Git is a distributed version control system. In software development, version control systems like Git are standard tools for managing code changes, tracking project history, and collaborating. Every code change submitted by a developer records the project's evolution.
[0067] The core functions of the system include:
[0068] 1) Comprehensive collection and management of LLM interaction logs: Through reverse proxy and asynchronous processing, LLM question-and-answer logs from multiple portals, such as IDE plug-ins and web pages, are uniformly captured, and key information such as user questions, model responses, and call duration is structured and stored.
[0069] 2) Accurate tracking and extraction of Git commit changes: Automatically analyze the commit history of a specified code repository within a specific time period, accurately extracting all newly added code lines and their associated commit metadata (such as committer, time, and file path).
[0070] 3) Efficient inverted index construction for code text: The code generated by the LLM (equivalent to the auxiliary code generated by the large language model in the above embodiment) and the newly added code in Git (equivalent to the newly added code in the target code repository within the target time period after the auxiliary code was generated by the large language model in the above embodiment) are normalized and tokenized. An inverted index is constructed from tokens to code line locations, and a detailed code line information library is established to provide a basis for fast retrieval and similarity comparison.
[0071] 4) Intelligent Code Adoption Assessment Based on Inverted Indexes: Utilizing the constructed inverted index and code line information repository, this system efficiently matches newly added Git code with LLM-generated code snippets. By calculating text similarity and dynamically adjusting thresholds based on factors such as language consistency and annotation content, it identifies AI-generated code actually adopted by developers and quantifies the code adoption rate.
[0072] 5) False negative detection and calibration mechanism: A reverse query-based cross-validation strategy is introduced. Starting from the code lines generated by LLM, potential matches are found in the newly added Git code. This can identify and correct possible false negatives in the main evaluation process, improving the accuracy and recall of the evaluation.
[0073] Through the collaborative work of the above components, the system can provide reliable data support and technical means for code generation traceability, AI-assisted programming effect evaluation, developer behavior analysis, etc. The ultimate goal is to deeply understand and optimize the application efficiency of large language models in the software development life cycle.
[0074] like Figure 3 As shown, the system includes: an LLM log collection and persistence management component (used to obtain the auxiliary code generated by the large language model in the above embodiment), a Git change capture and extraction component (used to obtain the newly added code of the target code repository in the above embodiment), a code inverted index construction component, a code retrieval and evaluation adoption component, and a false negative detection component.
[0075] like Figure 3 As shown, the LLM Log Collection and Persistence Management component: This component targets various usage entry points in Large Language Model (LLM)-assisted programming scenarios. It provides a reverse proxy and asynchronous processing mechanism for question-and-answer log collection and structured storage, fully recording the interaction between developers and models. Its core goal is to capture key information such as question content, model responses, and call duration, without changing user habits or tool configuration, and store it in a structured manner, providing reliable data support for subsequent code generation traceability and adoption evaluation.
[0076] The LLM log collection and persistence management component allows the following functions to be performed:
[0077] (1) Unified collection of multiple access paths. This system supports log access capabilities for two major types of models using channels:
[0078] Integrated development environment access path: such as Visual Studio Code (VS Code, a code editor) and JetBrains series tools, users can directly call the locally deployed language model service through built-in plug-ins.
[0079] Interactive web access path: For example, accessing pages such as OpenWebUI through a browser to ask natural language questions and generate code.
[0080] At the access layer, these various call methods are uniformly forwarded through the reverse proxy system deployed by the LLM log collection and persistence management component. The system automatically identifies the source type (such as a development tool or browser page) based on the access path and marks it as the access source category field for subsequent data management and analysis.
[0081] (2) Request information parsing and context capture mechanism.
[0082] When a user initiates a model call request, the system automatically identifies and extracts the following information:
[0083] User identity information: including basic identification fields such as user name and user Internet Protocol (IP) address;
[0084] Request content structure: Extract the latest round of user input in the conversation context as the question content, while retaining all context information as the original request details;
[0085] Call scenario type: Marks the call scenarios as general conversation, code writing, test case generation, etc. to facilitate classification analysis;
[0086] Call start time: The system records the time point when the request enters the processing flow, providing a basis for subsequent calculation of execution time.
[0087] The above-mentioned user identity information and other information are collected during the request phase and are continuously mounted in the system context until the model (i.e., large language model) response phase ends.
[0088] (3) Model response content extraction and execution time statistics. When the model returns the results, the system automatically captures the following elements:
[0089] Generate response content: Identify the complete output text of the model and support the extraction of natural language answers and code snippets, especially accurately extracting multi-line code blocks wrapped in Markdown structures;
[0090] Execution time calculation: Based on the timestamps of the request start and response end, the time consumed by the current conversation generation is calculated for performance analysis and user behavior modeling.
[0091] Structural processing of call return results: The above information is uniformly formatted as log entity content and enters the subsequent storage process.
[0092] In other words, the process performs multi-layer judgment and fault-tolerant processing on the return body to ensure that the response data can be retained to the maximum extent possible in the event of structural anomalies or intermediate network fluctuations.
[0093] (4) Asynchronous log storage and fault-tolerant writing mechanism.
[0094] To ensure that the collection component has no performance impact on the main inference process, the system uses an asynchronous timed processing mechanism to achieve log persistence:
[0095] Structured write fields: Each log contains key fields such as user information, call source, conversation type, request text, complete context, model response content, execution time, and collection time;
[0096] Asynchronous scheduling triggering: All log writing operations are triggered through the event queue and executed asynchronously by independent threads to avoid blocking the main request processing flow;
[0097] Abnormal retry mechanism: If the write operation fails due to network interruption, database abnormality, etc., the system automatically performs exponential backoff retries until the task is completed or the upper limit is reached;
[0098] Batch submission optimization: For request-response pairs generated in a short period of time, the system can aggregate them and write them uniformly, improving storage efficiency and reducing resource consumption.
[0099] It should be noted that the log storage process is strictly limited to the local intranet database system to ensure that all interactive data is not transferred through the external network and meet information security requirements.
[0100] (5) Configurable and scalable design.
[0101] New channel access expansion: The system supports access to more access paths through configuration expansion, such as command line interfaces, continuous integration tools, and mobile applications. New paths only require adding corresponding rules and reusing the log collection logic.
[0102] Response structure adaptation: When different model service providers return inconsistent formats, the system can automatically enable the corresponding parsing strategy based on the response header or content structure to maintain universal adaptability;
[0103] Dynamic start / stop and grayscale control: The system supports configuring log collection switches for specific users, specific channels, or specific application levels. It can also be connected to the configuration center or cache platform for dynamic management and control.
[0104] Log security isolation mechanism: All log data is used only for traceability and identification purposes and does not directly participate in business generation. The system supports desensitization processing and access control authorization for sensitive fields to meet data compliance requirements.
[0105] like Figure 3As shown in the figure, the Git Change Capture and Extraction component is used to automatically capture the change history of project code from the version control system and accurately extract all newly added code or text lines. It mainly consists of two sub-functional modules that work together: the commit information acquisition module and the new line extraction module.
[0106] The optional commit information acquisition module is responsible for interacting with the specified code repository to obtain all commit records and detailed code changes within a specific time period. The operation process of the commit information acquisition module includes:
[0107] First, the submission information acquisition module receives the user-specified code repository path, start time (`since`), and end time (`until`) as input parameters.
[0108] The commit information acquisition module then calls the version control system's logging function. This function is configured with specific parameters to retrieve the required information: 1) limiting the time range of commit records; 2) formatting the output to ensure that each commit includes a unique commit identifier and commit timestamp; 3) generating detailed code change comparison information, showing the line-level details of the additions, deletions, and modifications introduced by each commit.
[0109] After execution, the commit information acquisition module captures the standard output of the version control system log. This output contains the original commit metadata and detailed code change difference information.
[0110] Finally, the commit information acquisition module returns a raw text string containing all eligible commit records and their code change details. If an error occurs during the process (for example, the version control system command fails to execute), the error is recorded and a null value is returned.
[0111] Optionally, a new line extraction module is used to receive the original text output generated by the submission information acquisition module and parse it to identify and extract all added text lines. The operation process of the new line extraction module includes:
[0112] The module scans the input changelog text line by line.
[0113] Commit identification: When the module identifies the start marker of a new code commit record, it extracts the unique identifier of the commit and prepares to record the new lines related to this commit. At the same time, it also finds and parses the time information of the commit.
[0114] File change identification: When the module detects a specific tag indicating that a file has changed, it parses the changed file name or path.
[0115] New Line Identification: The core logic is to identify text lines that represent newly added content. In the detailed record of the code change, new text lines are usually marked with a special indicator (for example, starting with a specific symbol such as a plus sign `+`, but excluding those symbol combinations used to mark file difference header information).
[0116] Information aggregation: For each line of text identified as newly added, the module extracts its actual content (stripping leading special indicators and spaces) and associates it with the currently parsed commit identifier, file path, and commit time.
[0117] The module maintains contextual information during the processing, such as the submission ID, submission time, project name, and file name currently being processed, to ensure that each new line can be correctly classified and recorded.
[0118] Output: This module ultimately outputs a structured list of data, where each element represents a newly extracted line. Each element typically contains the following information: the unique identifier of the commit to which the line belongs, the specific text content of the newly added line, the file path where the line is located, and the timestamp of the commit.
[0119] Through the collaborative work of the commit information acquisition module and the new line extraction module, the Git change capture and extraction component can effectively and automatically track code evolution from the code repository and focus on quantitative analysis of new contributions.
[0120] like Figure 3 As shown, the inverted code index construction component: The core function of this component is to deeply analyze and structure code text resources, constructing an efficient mapping relationship between vocabulary units and their precise locations in the code (i.e., an inverted index), supplemented by a detailed code line information library. This mechanism is designed to provide core technical support for subsequent rapid code content retrieval, similarity comparison, and other advanced code analysis functions, and is a key link in achieving code content understanding and traceability.
[0121] (1) Core processing flow and pre-dependencies.
[0122] The effective operation of this component depends on two key pre-processing steps: normalization of code line content and accurate extraction of key information units (lexers) in the code.
[0123] Pre-processing step 1: Normalize the code line content (data cleaning).
[0124] Before building the index, all input raw code lines (equivalent to the first code in the above embodiment) must go through a normalization process to eliminate non-semantic differences in the code text, such as removing extra whitespace characters at the beginning and end of the line, replacing multiple consecutive spaces or tabs with a single space, and removing semicolons at the end of the line that do not affect the execution logic. Importantly, this process will carefully retain the comment information in the code, because comments often contain important context and developer intentions. The goal of normalization is to ensure that logically equivalent or highly similar code snippets can also show consistency at the text level, thereby significantly improving the quality of subsequent index construction and the accuracy and recall of retrieval.
[0125] Pre-processing step 2: Accurate extraction of code words.
[0126] After normalization, the lines of code are further broken down into a series of minimal units with independent semantics, called "lexical elements." This process prioritizes the use of advanced lexical analysis techniques. The lexical analyzer intelligently identifies and extracts the various basic elements that make up the code, such as variable names, function names, class names, keywords that control the flow of the code, specific numerical values or strings (literals) used in the program, and code comments, based on the grammatical rules of the specific programming language used. Each identified lexical element carries certain semantic information, forming the basis for the subsequent inverted index.
[0127] (2) Inverted index construction mechanism.
[0128] The construction of the inverted index is the core function of this component. The detailed process is as follows:
[0129] Step 31: Input Receiving: The component's input is a series of structured data records, each of which represents a response (i.e., auxiliary code) generated by a program (e.g., a large language model). These records contain metadata such as the response's unique identifier, user information, the specific code text, and the timestamp of the code creation.
[0130] Step 32: Initialization operation. Before processing begins, the system initializes two core data structures:
[0131] The first core data structure initialized is the inverted term index. The inverted term index is a mapping structure (logically similar to a multi-level lookup table). Its primary function is to store every unique term that appears in the code, along with the specific location information for all code lines where that term appears. Simply put, a term is the "key" for search, and its corresponding "value" is a list of the unique internal numbers of all code lines containing that term.
[0132] The second core data structure initialized is the code line information repository (equivalent to the code line information repository in the above embodiment, referred to as the repository). The code line information repository is an ordered collection (which can be understood as an information registry) that centrally stores the complete details of each processed line of code. For each line of code, the repository records its original unprocessed text, the normalized text, the programming language to which the line of code belongs (if identifiable), associated metadata (such as the response it originated from and the creation time), and a unique internal identifier within the repository (for example, its serial number in the registry).
[0133] Step 33: Divide the code text into blocks and parse it line by line.
[0134] Code Region Identification: The component first analyzes the input code text (i.e., the first code) to accurately identify valid code regions. For example, in text containing Markdown formatting, it can locate and extract code blocks surrounded by specific symbols (such as three backticks). While identifying code blocks, it also attempts to parse the programming language type declared in the code block (e.g., `python`, `java`, etc.). This language information is crucial for subsequent token extraction.
[0135] Careful line-by-line processing: Within the identified code block, the component will traverse and process line by line. For each non-empty line of code with actual content:
[0136] Step 33-1: Perform normalization: Call the aforementioned "code line content normalization" process to clean and standardize the current code line (i.e. each line of code).
[0137] Step 33-2: Store in the repository: The normalized line of code, along with its original text, identified programming language, and metadata inherited from the input data, such as the response ID, user information, and creation time, is stored in the "Line of Code Information Repository." The system also assigns a unique internal row identifier (i.e., a unique identifier) to this line of code within the repository.
[0138] Step 33-3: Extract tokens: Perform “code token accurate extraction” on the normalized code lines.
[0139] Optional, regarding word extraction, including the preferred strategy and the backup strategy.
[0140] The preferred strategy is lexical analysis based on language characteristics: the system will prioritize using a dedicated lexical analyzer for the identified programming language for that line of code. This analyzer (for example, by calling a function like `lexer.get_tokens(code)`, where `lexer` is the lexical analysis engine for the specific programming language and `code` is the code line to be analyzed) can accurately decompose the code line into a series of lexical units (i.e., tokens) and their types (for example, distinguishing between variable names, keywords, and comments) according to the strict grammatical rules of the language. This precise parsing based on the structure of the programming language is key to ensuring the quality of token recognition and subsequent indexing.
[0141] Among them, the backup strategy is: General pattern matching: If a language-specific analyzer cannot be found during processing, or if a dedicated analyzer encounters difficulty analyzing a line of code (for example, if the code snippet is incomplete or contains minor grammatical errors), the system will automatically use a backup token extraction strategy based on general text pattern matching (such as using carefully designed regular expressions). Although this strategy does not have the deep understanding of code structure that a dedicated analyzer does, it can still effectively extract the vast majority of meaningful text fragments from the text, such as potential identifiers and comments, thereby maximizing information preservation.
[0142] Step 33-4: Update the inverted index: For each valid word (i.e., the second word) extracted from the current code line: Check whether the word already exists in the "word-word inverted index library" (equivalent to the index library in the above embodiment). If the word appears for the first time, use this word as the new "key" to create a new entry in the index library (equivalent to the newly added entry in the above embodiment), and initialize its "value" (i.e., the code line position list) to empty. Then, append the unique internal row identifier of the current code line in the "code line information repository" to the value list corresponding to the word (equivalent to the second position list in the above embodiment), thereby recording that the word has appeared in this line.
[0143] (3) Output results and core role.
[0144] After this component is executed, it will generate and output the following two important results:
[0145] The completed inverted index of terms: This is a structured index that allows us to quickly find the exact location of all code lines containing any term.
[0146] A comprehensive codeline database that stores each line of code in its original form, processed form, and rich contextual metadata.
[0147] like Figure 3 As shown, the Code Retrieval and Adoption Assessment component includes a code retrieval and adoption assessment component. Building on the inverted index and code line information repository constructed by the aforementioned "Code Inverted Index Construction Component," this component focuses on enabling efficient retrieval and matching between newly added code in version control systems (such as Git) and code snippets generated by the Large Language Model (LLM). The core goal is to automatically identify AI-generated or assisted code that developers actually adopt, quantify the code adoption rate, and provide data support for analyzing the actual application effects of AI in programming.
[0148] (1) Overall business process of code adoption assessment. The code adoption assessment component implements a comprehensive analysis of code changes and AI responses within a specified time period through a collaborative workflow, including:
[0149] Step 41: Analysis cycle setting and data initialization.
[0150] The system uses a specific date or time period (for example, daily) as the unit of analysis. After the user specifies the analysis date, the system determines the precise start and end timestamps.
[0151] Initialize necessary service modules, including Git version analyzer, database manager, and core code matcher.
[0152] Step 42: Comprehensive collection and index construction of AI-generated code.
[0153] First, we retrieve all LLM response records from the database within the specified analysis period. These records contain the original AI-generated code text and related metadata (such as user, timestamp, etc.).
[0154] The inverted index building function in the "Code Inverted Index Building Component" is called to process all collected LLM response text. The processing steps include: 1) Parsing formats such as Markdown to extract pure code blocks and their declared language. 2) Performing code line normalization on each valid line of code. 3) Storing the normalized code line along with its original text, language, source AI response ID, user information, etc. in the "Code Line Information Repository." 4) Performing tokenization on the normalized code line and associating each token with the identifier of the code line it resides in (indexed in the repository) to build a global "term inverted index." This index is key to subsequent efficient retrieval.
[0155] Step 43: Perform incremental code analysis and matching by project module.
[0156] The system iterates through the pre-defined code modules (projects / repositories). For each module, the following operations are performed:
[0157] Step 43-1: Code base update: Try to pull the latest code from the remote Git repository to ensure that the analysis is based on the latest version.
[0158] Step 43-2: Obtaining commit records: Use the module commit record obtaining function provided by the Git analyzer to obtain all Git commit records of the module within a specified time period, including detailed code change differences.
[0159] Step 43-3: Extract newly added lines of code: From the diff information of the commit record, accurately identify all newly added lines of code by extracting the newly added lines of code. At the same time, record the commit hash, file path, and commit time of these newly added lines of code.
[0160] Step 43-4: Similar code retrieval and matching: Call the core code matching logic to compare the extracted Git new code lines with the pre-built AI code inverted index and line library to find similar code pairs.
[0161] Step 43-5: Results Statistics and Persistence: Count the total number of new lines of code for the current module and the number of matched AI-generated lines of code. Calculate the code adoption rate (for example, number of matching lines divided by total number of new lines). Save the module's analysis results (including specific similar code pair details, statistical data, analysis date, etc.) to the database through transactional operations to ensure data consistency.
[0162] Step 44: Display and summarize the results.
[0163] The system can display analytical statistics for each module, such as total new lines, AI-matched lines, and adoption rate, and can list specific similar code pairs in descending order of similarity, including Git-side code, AI-side code, both languages, commit information, AI response information, and similarity scores.
[0164] (2) An efficient code retrieval and similarity matching mechanism based on inverted index. This mechanism is the core of the component and is responsible for quickly locating and verifying fragments similar to Git-added code in a large amount of AI-generated code. This is achieved through the following process:
[0165] Step 51: Input and Preparation.
[0166] Git adds a new code line list: including code text, commit, file path, and commit time.
[0167] Global token inverted index: built from all lines of code generated by AI.
[0168] Global code line information repository: stores detailed information about all AI code lines.
[0169] Similarity threshold: The preset basic similarity judgment standard.
[0170] Step 52: Execute the matching process for each line of Git newly added code.
[0171] Step 52-1: Git code line preprocessing:
[0172] Perform the same code line normalization process on the newly added Git code lines as when building the AI code line index to eliminate formatting differences. Detect the programming language of the Git code line.
[0173] The language is determined by the file name suffix first. If it cannot be determined, the language is guessed by the code content.
[0174] Perform tokenization on the normalized Git code lines to extract the token sets they contain. This process also takes language characteristics into consideration, prioritizing the use of a language-based lexical analyzer and falling back to general regular expression matching if this fails.
[0175] Step 52-2: Quickly retrieve candidate AI code lines (using inverted index).
[0176] Iterate over each token extracted from the Git code line (i.e., the third token).
[0177] Using each token as a key, the global "Token Inverted Index" is searched. The inverted index returns a list (the third position list) containing the unique index (i.e., unique identifier) of all AI code lines that appear in the "Code Line Information Repository" where the token appears.
[0178] All found AI code line indexes are gathered together to form a non-repeating "candidate AI code line set" (equivalent to the third code in the above embodiment). This set is the target for subsequent detailed comparison. The inverted index greatly narrows the search scope and avoids a full comparison.
[0179] Step 52-3: Compare candidate AI code lines one by one in detail.
[0180] Traverse each AI code line index in the "Candidate AI Code Line Set" and obtain the complete information of the AI code line (original text, normalized text, language, source ID, etc.) from the "Code Line Information Repository".
[0181] Perform similarity calculations, including:
[0182] 1) A token-based comparison strategy is used to calculate text similarity. This strategy first decomposes the normalized Git code line and the candidate AI code line into their constituent token sets. It then analyzes the intersection and difference of the two token sets and applies algorithms based on edit distance (such as the Levenshtein distance) to quantify the similarity between these sets, ultimately providing a comprehensive assessment of the similarity between the two content. This approach effectively addresses situations such as reordering or repetition of code elements, focusing on determining the degree of overlap between the two in core semantic units, thereby generating a quantitative similarity score.
[0183] 2) Language consistency consideration: Compare the programming languages of Git code lines and AI code lines. If the languages are different and do not belong to the predefined "related languages", a penalty may be imposed on the similarity score and marked as "language mismatch".
[0184] 3) Apply logic that dynamically adjusts the similarity threshold: Determine whether a Git line is primarily comments. Adjust the base similarity threshold based on whether it is a comment and whether there is a language mismatch. Generally, comment lines may require a higher similarity score to be considered a valid match.
[0185] 4) Best Match Selection: If the calculated similarity score exceeds a dynamically adjusted threshold and is the highest score among all candidate AI code lines for the current Git code line, this AI code line is recorded as the best match for the current Git line. This record contains rich metadata such as Git line information, AI line information, similarity, whether it is annotated, whether there is a language mismatch, and the number of shared tokens.
[0186] Step 53: Output the matching result.
[0187] After processing all newly added Git code lines, the component outputs a list of "similar code pairs." Each element in the list details a pair of Git code lines and AI code lines that are considered similar, along with all relevant contextual information.
[0188] Through the above process, this component can not only efficiently retrieve potential similar code fragments from massive code data, but also combine multiple heuristic rules to perform accurate similarity assessment, thereby providing solid technical support for code adoption analysis.
[0189] like Figure 3As shown in the figure, the false negative detection component is included. This component aims to establish a continuous monitoring and calibration mechanism to evaluate and improve the recognition accuracy of the "Code Retrieval and Adoption Evaluation Component." The core goal is to proactively detect and correct false negatives that may occur when the previous component identifies the matching relationship between Git code commits and LLM-generated code through complementary detection strategies, thereby ensuring the authenticity and reliability of code adoption data.
[0190] The false negative detection mechanism in the false negative detection component is based on cross-validation using reverse queries. A false negative refers to a situation where lines of code in the Git repository that are actually derived from or significantly influenced by LLM-assisted generation are not identified by the Code Retrieval and Adoption Evaluation Component, which executes the main adoption evaluation process. This mechanism uses a reverse query cross-validation strategy to identify potential false negatives. This mechanism does not create new data, but rather, by shifting the matching perspective, discovers connections in existing static datasets that might be overlooked by the main process.
[0191] (1) The core idea of the false negative detection component is explained through the following three core data sets:
[0192] Set A (LLM output): The collection of all lines of code generated and recorded by LLM during a specific analysis cycle. Each line includes metadata such as the original text, normalized text, tokens, generation timestamp, and user information. This data primarily comes from LLM logs collected by the "LLM Log Collection and Persistence Component" and processed and stored by the "Code Inverted Index Construction Component."
[0193] Set B (Git New Lines): This is the collection of all new lines of code captured from each project's Git repository during the same analysis cycle. Each line contains the original text, normalized text, tokens, commit hash, file path, commit timestamp, user information, and more. This data comes from the Git Change Capture and Extraction Component.
[0194] Set C (Identified Matches): The set of correspondences between Git-added lines and LLM-generated lines identified by the Code Retrieval and Adoption Evaluation Component through its primary matching logic (Git lines dominate, looking for similarities in LLM output).
[0195] False negatives (FN) are those matching pairs that actually exist in the "intersection of set A and set B" (that is, the LLM generated content actually contained in Git) but are not recorded in "set C".
[0196] (2) Review and potential limitations of forward matching logic:
[0197] The core process of the "Code Retrieval and Adoption Evaluation Component," also known as "forward matching," is to traverse each Git-added line (element `b` in Set B) and then search for the most similar candidate `a` in the set of code lines generated by the LLM (Set A). If the similarity exceeds a threshold, `(b, a)` is recorded in Set C.
[0198] A limitation of this method is that some LLM-generated code may undergo manual adaptive modifications before being applied by developers to Git commits, such as variable renaming, code structure fine-tuning, and comment additions and deletions. These modifications may cause its normalized form or token set to differ from the original LLM output, causing positive matches based on specific similarity algorithms (such as the Levenshtein distance) to fail to meet the threshold, resulting in false negatives.
[0199] (3) Reverse matching process (LLM-led verification). To make up for the potential blind spots of forward matching, this mechanism introduces a reverse matching process, including:
[0200] Step 61: Iterate LLM code lines: traverse each line of LLM-generated code in set A (denoted as `ai_line`, equivalent to each line of code in the above embodiment).
[0201] Step 62: Circle candidate Git code lines:
[0202] Based on `ai_line`'s metadata (such as user information and generation timestamp), relevant Git-added code lines (denoted as `git_line`, equivalent to the second code line list in the above embodiment) are filtered from Set B. The filtering criteria typically include: the submitter of `git_line` is the same as the user who generated `ai_line`, and the submission time of `git_line` is within a reasonable window after the generation time of `ai_line` (for example, within 7 days). This is because developers typically wait some time after receiving LLM assistance to apply it to actual projects.
[0203] Optionally, the specific process of identifying candidate Git code lines includes:
[0204] Step 62-1: First, the code line normalization and lemma processing similar to that when the "code inverted index building component" processes the LLM output is performed on the currently traversed ai_line to extract the lemma set it contains.
[0205] Step 62-2: Secondly, this step relies on a pre-built inverted index for Set B (Git new lines) (equivalent to the target inverted index in the above embodiment). This index is constructed in the same way as the "Code Inverted Index Construction Component" constructs the index for Set A, namely, normalizing and tokenizing each Git new line and storing a mapping from the token to the Git line it belongs to (e.g., line ID or unique identifier).
[0206] Step 62-3: Using the set of tokens extracted from ai_line, query the inverted index for the newly added Git line. By searching each token in ai_line, a preliminary list of candidate git_lines (equivalent to the first code line list in the above embodiment) containing at least one of these tokens can be quickly obtained from the inverted index. This list is a preliminary selection of Git code lines that are potentially relevant in content, identified through token matching.
[0207] Step 62-4: Finally, apply the original metadata filtering conditions to this preliminary candidate git_line list for precise filtering: that is, ensure that the submitter of the candidate git_line is consistent with the user who generated ai_line, and that the submission time of git_line falls within a reasonable time window after the generation time of ai_line (for example, within 7 days).
[0208] Step 63: Perform similarity comparison:
[0209] Perform the same code line normalization and tokenization process as in "Code Retrieval and Adoption Evaluation" on `ai_line` and each candidate `git_line`.
[0210] Compute a text similarity score between the two (e.g., using the Levenshtein distance algorithm based on a set of tokens).
[0211] Step 64: Identify potential false negatives:
[0212] If `similarity(ai_line,git_line)` is above a preset validation threshold (this threshold may be different from the positive match threshold, for example, it can be appropriately relaxed to capture more potential associations), and `git_line` is not marked as matched in set C (or the specific matching relationship `(ai_line,git_line)` is not in C), then `(ai_line,git_line)` is marked as a potential false negative.
[0213] The significance of cross-validation: Forward matching is "Git-driven, searching for the best AI source," while reverse matching is "AI-driven, observing whether it is adopted in some form." These two perspectives differ in their tolerance for fuzzy code similarity and the coverage of candidate sets. Even if the data itself is static, changing the matching logic and verification direction, like "trying to open the same lock with a different key," can effectively discover real connections that a single matching strategy may miss, thereby improving overall recall.
[0214] Combine Figure 3 In general, 1) The LLM log collection and persistence component implements unified interception and log collection of requests from multiple entry points (such as IDE plug-ins and web interfaces) by deploying a reverse proxy in the LLM service access path. During implementation, for each captured request, the system parses and extracts user identity information, request content (especially the latest round of questions and complete context), call scenario, and timestamp. For LLM responses, the complete output text (paying special attention to the extraction of code blocks in Markdown format) and execution time are extracted. After all collected information is structured, it is batch-written to the local database through asynchronous mechanisms (such as message queues and independent processing threads), and includes exception retry and exponential backoff strategies to ensure minimal impact on the main process performance.
[0215] The Git Change Capture and Extraction component is implemented by invoking the Git log functionality of the target repository. To perform the analysis, specify the repository path and the time interval to be analyzed (starting with `since` and ending with `until`). Configure the Git command to output a log containing commit hashes, timestamps, and full file diffs. After receiving the raw log, the component scans the log line by line, specifically identifying and extracting new lines marked with a specific prefix (such as `+`, excluding diff header information). Each extracted line of new code is associated with its commit hash, file path, and commit time.
[0216] The implementation process for building a code inverted index component first normalizes the input code text (primarily from LLM-generated content), including removing excess whitespace, unifying the representation, and retaining comments. The normalized code lines are then tokenized, preferably using a lexical analyzer based on the code language's characteristics (such as a Lexer generated by ANTLR or similar tools) to decompose them into tokens. If there is no specific language analyzer or the analysis fails, it falls back to general token extraction based on regular expressions. The inverted index construction process includes: (a) receiving LLM response records as input, each record containing a response ID, code text, etc. (b) initializing a token inverted index library (mapping tokens to a list of code line IDs where the tokens appear) and a code line information repository (storing detailed information about each line of code, including the original text, normalized text, language, source metadata, and unique line ID). (c) Identify code blocks in the input text (e.g., `python...` in Markdown) and determine the programming language. (d) Process the code line by line: perform normalization; store the normalized code and metadata in a repository to obtain a unique row ID; extract tokens from this line; and for each token, update the inverted index and add its row ID to the token list.
[0217] The implementation of the inverted index-based code retrieval and adoption evaluation component revolves around a periodic (e.g., daily) evaluation process: (a) Collect all LLM-generated code within a specified period and call the aforementioned "Code Text Inverted Index Construction Component" to build or update the global AI code inverted index and line repository. (b) Traverse each configured code project / module. For each module: First, pull the latest code from the Git repository; then call the "Git Commit Change Capture and New Line Extraction Component" to obtain all new code lines added during the period. (c) For each new Git code line: Perform the same normalization, language detection (prioritizing file name suffixes, and if not, guessing the content), and tokenization as when building the AI code index. (d) Use the tokens extracted from the Git lines to query the AI code inverted index library and quickly filter out a set of candidate AI code lines that contain any of the same tokens. (e) Perform a detailed comparison of each candidate AI code line: Complete information is retrieved from the repository, and a similarity score is calculated with the current Git line using a similarity algorithm based on a set of tokens (e.g., Levenshtein distance). This process considers programming language consistency and dynamically adjusts the similarity threshold based on factors such as whether the code line is a comment. (f) The best matching AI code line is selected, and information such as matching pairs and similarity is recorded. The module's statistical results (total new lines, matching lines, and adoption rate) are stored in the database.
[0218] To improve the recall rate of adoption evaluation, the false negative detection component implements a reverse query cross-validation mechanism: (a) It traverses all LLM-generated code lines (`ai_line`) processed by the "Code Text Inverted Index Construction Component." (b) It tokenizes the current `ai_line`. These tokens are used to query an inverted index pre-built for Git newly added code lines, thereby obtaining a preliminary set of candidate `git_line`s. This candidate set is then filtered based on `ai_line` metadata (such as user information and generation timestamp), selecting `git_line`s with the same committer and a commit time within a reasonable window after the `ai_line` generation time. For the detailed description of tokenization, inverted index construction, and filtering logic, please refer to the "Reverse Matching Process" under "False Negative Detection Component" in the Summary of the Invention. (c) The same normalization and tokenization process is performed on `ai_line` and each candidate `git_line` selected in step (b), and the textual similarity between them is calculated (for example, using the Levenshtein distance algorithm based on token sets). (d) If the similarity is higher than a specific verification threshold (this threshold can be different from the threshold of the main evaluation process and can be appropriately relaxed), and the `(ai_line,git_line)` pair is not identified as a match in the main evaluation process (i.e., step (f) in "Implementation of code snippet retrieval and adoption evaluation component based on inverted index"), it is marked as a potential false negative.
[0219] Combining all the above components, the system of the present application completes the code adoption evaluation through the collaborative work of the above components. First, the LLM log collection and persistence component and the Git change capture component independently and continuously collect raw data. The code inverted index construction component processes the LLM log data regularly or on demand to build the data structure required for efficient retrieval. The inverted index-based code retrieval and adoption evaluation component uses these indexes to match and evaluate the similarity between the new code in Git and the code generated by LLM, and output preliminary adoption results. Finally, the underreporting detection component supplements the verification of the evaluation results to discover and correct potential underreporting, further improving the accuracy and completeness of the evaluation. The data and results generated by all processes are structured and stored to support subsequent statistical analysis and display.
[0220] In summary, this application provides an advanced and practical solution for large language model-assisted programming adoption assessment, which not only effectively solves the key problems existing in current technologies and improves the accuracy and efficiency of assessment, but also provides strong data support and technical means for in-depth understanding and optimization of LLM applications in software development. Specifically:
[0221] 1) Comprehensive and objective data support and interaction traceability: The LLM log collection and persistence management component comprehensively and uniformly records the entire interaction process between developers and the large language model, including key information such as questions, context, model responses, and time consumption. This addresses the issue of fragmented interaction data, varying formats, and difficulty in unified management. Furthermore, combined with the precise code change history captured by the Git commit change capture component, this provides a solid and objective data foundation for subsequent adoption evaluation and LLM application effectiveness analysis, and enhances the traceability of the code generation process.
[0222] 2) Accurate and Efficient Quantitative Evaluation of Code Adoption: To address the difficulty in quantifying the actual adoption of LLM-generated code, this application utilizes an inverted index of code text, an efficient retrieval component, and an adoption evaluation component based on this. This allows for efficient retrieval and matching of newly added Git commits with LLM-generated code snippets within massive amounts of code. By accurately calculating similarity and incorporating contextual information, this enables automated and quantitative evaluation of the adoption rate of AI-generated code, making it possible to measure the actual contribution and return on investment of LLM in software development.
[0223] 3) Significantly improve the accuracy and recall of adoption evaluation: Traditional text matching methods struggle to cope with developers' adaptive modifications to LLM-generated code and are prone to false negatives. This application not only uses a similarity algorithm based on word sets and a dynamic threshold adjustment strategy in forward matching, but also innovatively introduces a false negative detection component based on reverse queries (and also leverages the efficiency advantages of inverted indexes when screening candidate Git code lines). By starting with LLM-generated code and reverse-verifying its representation in Git commits, it effectively identifies matching pairs that may have been missed by the main evaluation process, significantly reducing the false negative rate and greatly improving the overall accuracy and recall of code adoption evaluation.
[0224] 4) Efficiency Improvements and Cost Reductions through Automated Processes: This application automates multiple processes, including LLM log collection, code change capture, index building, code matching, and adoption evaluation. This significantly reduces the need for manual auditing and analysis, significantly improves evaluation efficiency, and reduces associated costs. The application of technologies such as asynchronous processing and inverted indexing ensures efficient system operation.
[0225] 5) Data Insights to Drive LLM Application Optimization: In-depth analysis of collected and evaluated data reveals developer LLM usage patterns and preferences, LLM effectiveness in different scenarios, and which types of code suggestions are most likely to be adopted. These insights provide valuable basis for optimizing LLM models, improving the functionality of auxiliary programming tools, and guiding developers to more effectively utilize LLM.
[0226] 6) Enhanced System Scalability and Compliance: This application's component-based design supports flexible configuration and expansion, adapting to new LLM channels or the integration of different model service providers. Furthermore, log data is processed locally within the intranet, and sensitive information is masked, meeting enterprise requirements for data security and compliance.
[0227] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0228] This embodiment also provides a device for determining a code adoption rate, which is used to implement the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0229] Figure 4 is a structural block diagram of a device for determining a code adoption rate according to an embodiment of the present application, such as Figure 4 As shown, the device includes:
[0230] A construction module 42 is configured to construct a code line information library including a first code, and construct an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, or newly added code in a target code repository within a target time period after the large language model generates the auxiliary code, and the inverted index is configured to indicate a correspondence between a first word element included in the first code and position information of a code line in the first code including the first word element in the code line information library;
[0231] The matching module 44 is used to match the second code with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code.
[0232] Through the above-mentioned device, a code line information library including a first code is constructed, and an inverted index corresponding to the first code is constructed based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, newly added code in the target code warehouse within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate the correspondence between the first word element included in the first code and the position information of the code line including the first word element in the first code in the code line information library; the second code is matched with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein, when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code. Therefore, the technical problem in the related art that the actual contribution of the large language model in the software development process is difficult to quantify and evaluate can be solved, and the technical effect of determining the actual contribution of the large language model based on the code adoption rate is achieved.
[0233] In an exemplary embodiment, the construction module is further used to determine a valid code area in the first code and parse the programming language type to which the first code belongs; extract each line of code in the valid code area, and store each line of code and the code information of each line of code in the code line information library to construct the code line information library, wherein the code information includes: the first code and the programming language type.
[0234] In an exemplary embodiment, the construction module is further used to determine the text format of the first code, and locate the target symbol in the first code based on the text format; and determine the code block surrounded by the target symbol in the first code as the valid code area.
[0235] In an exemplary embodiment, the construction module is further used to determine the second word in each line of code, wherein the first word includes: the second word; and create the inverted index based on the second word and the corresponding unique identifier of each line of code in the code line information library, wherein the unique identifier is used to indicate the location information of each line of code in the code line information library.
[0236] In an exemplary embodiment, the construction module is further used for one of the following: determining the second word in each line of code through a lexical analyzer according to the grammatical rules corresponding to the programming language type, wherein the lexical analyzer is a dedicated lexical analyzer corresponding to the programming language type; determining the second word in each line of code through a preset matching method, wherein the preset matching method includes: regular expression.
[0237] In an exemplary embodiment, the construction module is also used to, when an entry corresponding to the second word already exists in the index library corresponding to the first code, add the unique identifier to the first position list corresponding to the entry to create the inverted index; when an entry corresponding to the second word does not exist in the index library corresponding to the first code, create a new entry in the index library based on the second word, and add the unique identifier to the second position list corresponding to the new entry to create the inverted index.
[0238] In an exemplary embodiment, the matching module is further used to determine a third word in the second code; determine a third code in the first code that matches the third word through the code line information library and the inverted index; and determine the code adoption rate through the third code.
[0239] In an exemplary embodiment, the matching module is further used to traverse the inverted index through the third word to determine a third position list corresponding to the third word in the index library where the inverted index is located, wherein the third position list is a list corresponding to entries created based on the third word in the index library; and determine the third code from the code line information library through the third position list, wherein the third position list records a unique identifier of the third code.
[0240] In an exemplary embodiment, the matching module is further used to determine the text similarity between a fourth code and a fifth code, wherein the fourth code is any code line in the second code that includes the third word, and the fifth code is any code line in the third code; and to adjust the text similarity according to type information corresponding to the fourth code and the fifth code respectively, wherein the type information includes: programming language type and code type; and to determine the code adoption rate according to the adjusted text similarity.
[0241] In an exemplary embodiment, the matching module is further used to decompose the fourth code and the fifth code respectively to obtain a first word set and a second word set; determine the intersection and difference of the first word set and the second word set; and determine the text similarity through the intersection and the difference.
[0242] In an exemplary embodiment, the matching module is further used to compare the programming language types corresponding to the fourth code and the fifth code respectively to obtain a first result; and determine whether the code types corresponding to the fourth code and the fifth code respectively are target types to obtain a second result; and adjust the text similarity according to the first result and the second result.
[0243] In an exemplary embodiment, the matching module is further configured to determine a sixth code in the second code that matches the first code based on the adjusted text similarity; and determine the code adoption rate based on the number of code lines of the sixth code and the number of code lines of the second code.
[0244] In an exemplary embodiment, the device also includes a missed detection module, which is used to determine each line of code corresponding to the first code and obtain a target inverted index pre-constructed for the second code, wherein each line of code is each line of code in the valid code area of the first code; determine a first code line list in the second code that matches each line of code through the target inverted index, wherein the target inverted index is used to indicate the correspondence between the third word included in the second code and the position information of the code line in the second code including the third word; determine the missed code line in the second code through the first code line list.
[0245] In an exemplary embodiment, the missed detection module is also used to determine each word in each line of code; determine a fourth position list through each word and the target inverted index, wherein the fourth position list is a list corresponding to entries created based on each word in the target index library where the target inverted index is located, and the fourth position list records a target unique identifier, wherein the target unique identifier is used to indicate the position information of the code line including each word in the second code; determine the first code line list through the fourth position list.
[0246] In an exemplary embodiment, the missed code detection module is further used to filter out a second code line list from the first code line list through the metadata of each line of code, wherein the metadata includes: user information, generation timestamp; and determine the missed code line through each line of code and the second code line list.
[0247] In an exemplary embodiment, the device also includes a determination module for obtaining a model call request for the large language model initiated by a user through different access portals; obtaining response content generated by the large language model in response to the model call request; uniformly formatting the response content, and determining the first code through the uniformly formatted response content.
[0248] For the description of the features in the embodiment corresponding to the device for determining the code adoption rate, please refer to the relevant description of the embodiment corresponding to the method for determining the code adoption rate, which will not be repeated here.
[0249] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for determining the code adoption rate.
[0250] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned code adoption rate determination method embodiments when running.
[0251] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0252] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the computer program implements the steps of any of the above-mentioned code adoption rate determination method embodiments.
[0253] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned code adoption rate determination method embodiments are implemented.
[0254] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0255] The above is a detailed introduction to the determination of a code adoption rate provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for determining a code adoption rate, characterized in that: include: Constructing a code line information library including a first code, and constructing an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, or newly added code in a target code repository within a target time period after the large language model generates the auxiliary code, and the inverted index is used to indicate a correspondence between a first word element included in the first code and position information of a code line in the first code including the first word element in the code line information library; The second code is matched with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein, when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code.
2. The method for determining the code adoption rate according to claim 1, wherein: Build a code line information library including the first code, including: Determining a valid code region in the first code, and parsing a programming language type to which the first code belongs; Each line of code in the valid code area is extracted, and each line of code and code information of each line of code are stored in the code line information library to construct the code line information library, wherein the code information includes: the first code and the programming language type.
3. The method for determining the code adoption rate according to claim 2, wherein: Determining a valid code area in the first code includes: determining a text format of the first code, and locating a target symbol in the first code based on the text format; A code block surrounded by the target symbol in the first code is determined as the valid code area.
4. The method for determining the code adoption rate according to claim 2, wherein: Constructing an inverted index corresponding to the first code based on the code line information library includes: Determining a second word in each line of code, wherein the first word includes: the second word; The inverted index is created according to the second word and the unique identifier corresponding to each line of code in the code line information library, wherein the unique identifier is used to indicate the position information of each line of code in the code line information library.
5. The method for determining the code adoption rate according to claim 4, wherein: Determining the second token in each line of code includes one of the following: Determining the second word in each line of code by a lexical analyzer according to grammatical rules corresponding to the programming language type, wherein the lexical analyzer is a dedicated lexical analyzer corresponding to the programming language type; The second word in each line of code is determined by a preset matching method, wherein the preset matching method includes: a regular expression.
6. The method for determining the code adoption rate according to claim 4, wherein: Creating the inverted index according to the second word and the unique identifier corresponding to each line of code in the code line information library includes: If an entry corresponding to the second word already exists in the index library corresponding to the first code, adding the unique identifier to the first position list corresponding to the entry to create the inverted index; If there is no entry corresponding to the second word in the index library corresponding to the first code, a new entry is created in the index library based on the second word, and the unique identifier is added to the second position list corresponding to the new entry to create the inverted index.
7. The method for determining code adoption rate according to claim 1, wherein: Matching the second code with the first code by using the code line information library and the inverted index to determine a code adoption rate corresponding to the auxiliary code includes: determining a third word in the second code; Determine, by using the code line information library and the inverted index, a third code in the first code that matches the third word; The code adoption rate is determined by the third code.
8. The method for determining the code adoption rate according to claim 7, wherein: Determining, by using the code line information library and the inverted index, a third code in the first code that matches the third word, includes: Traversing the inverted index using the third word to determine a third position list corresponding to the third word in the index library where the inverted index is located, wherein the third position list is a list corresponding to entries created based on the third word in the index library; The third code is determined from the code line information library through the third position list, wherein the third position list records a unique identifier of the third code.
9. The method for determining the code adoption rate according to claim 7, wherein: Determining the code adoption rate by using the third code includes: determining text similarity between a fourth code and a fifth code, wherein the fourth code is any code line in the second code including the third word-gram, and the fifth code is any code line in the third code; and Adjusting the text similarity according to type information corresponding to the fourth code and the fifth code, wherein the type information includes: programming language type and code type; The code adoption rate is determined by the adjusted text similarity.
10. The method for determining the code adoption rate according to claim 9, wherein: Determining the textual similarity between the fourth code and the fifth code, including: Decomposing the fourth code and the fifth code respectively to obtain a first word unit set and a second word unit set; Determining the intersection and difference of the first word-gram set and the second word-gram set; The text similarity is determined by the intersection and the difference.
11. The method for determining code adoption rate according to claim 9, wherein: Adjusting the text similarity according to the type information corresponding to the fourth code and the fifth code respectively includes: Comparing the programming language types corresponding to the fourth code and the fifth code, respectively, to obtain a first result; and determining whether the code types respectively corresponding to the fourth code and the fifth code are target types, to obtain a second result; The text similarity is adjusted based on the first result and the second result.
12. The method for determining code adoption rate according to claim 9, wherein: The code adoption rate is determined by adjusting the text similarity, including: Determining a sixth code in the second code that matches the first code based on the adjusted text similarity; The code adoption rate is determined by the number of code lines of the sixth code and the number of code lines of the second code.
13. The method for determining code adoption rate according to claim 1, wherein: After matching the second code with the first code using the code line information library and the inverted index to determine a code adoption rate corresponding to the auxiliary code, the method further includes: Determine each line of code corresponding to the first code, and obtain a target inverted index pre-constructed for the second code, wherein each line of code is each line of code in a valid code region of the first code; Determining, by using the target inverted index, a list of first code lines in the second code that match each line of code, wherein the target inverted index is used to indicate a correspondence between a third word element included in the second code and position information of a code line in the second code including the third word element; The missing code lines in the second code are determined through the first code line list.
14. The method for determining code adoption rate according to claim 13, wherein: Determining, by using the target inverted index, a list of first code lines in the second code that matches each line of code, includes: Determine each token in each line of code; Determining a fourth position list using each word and the target inverted index, wherein the fourth position list is a list corresponding to entries created based on each word in a target index library where the target inverted index is located, and the fourth position list includes a target unique identifier, wherein the target unique identifier is used to indicate position information of a code line in the second code that includes each word; The first code line list is determined by a fourth position list.
15. The method for determining code adoption rate according to claim 13, wherein: Determining the missing code lines in the second code by using the first code line list includes: Filtering a second code line list from the first code line list using metadata of each code line, wherein the metadata includes: user information and a generation timestamp; The missed code line is determined by using each line of code and the second code line list.
16. The method for determining code adoption rate according to claim 1, wherein: Before building the code line information library including the first code, the method further includes: Obtaining model call requests for the large language model initiated by users through different access portals; Obtaining response content generated by the large language model in response to the model call request; The response content is uniformly formatted, and the first code is determined based on the uniformly formatted response content.
17. A device for determining a code adoption rate, characterized in that: include: a construction module, configured to construct a code line information library including a first code, and construct an inverted index corresponding to the first code based on the code line information library, wherein the first code includes one of the following: auxiliary code generated by a large language model, or newly added code in a target code repository within a target time period after the large language model generates the auxiliary code, and the inverted index is configured to indicate a correspondence between a first word element included in the first code and position information of a code line in the first code including the first word element in the code line information library; A matching module is used to match the second code with the first code through the code line information library and the inverted index to determine the code adoption rate corresponding to the auxiliary code, wherein when the first code is the auxiliary code, the second code is the newly added code, and when the first code is the newly added code, the second code is the auxiliary code.
18. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for determining the code adoption rate according to any one of claims 1 to 16 when executing the computer program.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for determining the code adoption rate according to any one of claims 1 to 16 are implemented.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for determining the code adoption rate according to any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
Method and device for calculating contribution degree of source code
CN108932198A
Code contribution statistical method and device
CN111367529A
Method, device and equipment for determining code coverage rate
CN112148590A
Incremental code coverage rate testing method and device, storage medium and electronic equipment
CN112799939A
System and method for evaluating code contributions of software developers
CN114365095A
Cited By
Code adoption rate determination method and device, medium, electronic equipment and product
CN121050713A
Code adoption rate determination method and device, medium, electronic equipment and product
CN121050713B