Code adoption condition determination method and device, electronic equipment and storage medium
By generating code vectors and using vector databases and large language model analysis, the problem of inaccurate evaluation of code adoption is solved, the rapid and accurate judgment of AI generated code is achieved, and the capabilities of AI programming tools are optimized.
Patent Information
- Application Number
- CN202511008993.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The prior art cannot accurately evaluate the adoption of AI-generated code by user-submitted code.
By extracting the semantic information of the code, generating preset code vectors, using the vector database for similarity search, and conducting in-depth semantic analysis with large language models to judge the adoption between codes.
It significantly improves the accuracy and reliability of identifying code adoption, and realizes the ability of AI-assisted programming tools to facilitate optimization of tool performance.
Smart Images

Figure CN120508837A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and more specifically, to a method, device, electronic device, computer-readable storage medium, and computer program product for determining code adoption status. Background Art
[0002] With the rapid development of artificial intelligence (AI), particularly breakthroughs in natural language processing and machine learning, large language models (LLMs) are increasingly being used in software development. LLMs, powered by artificial intelligence (AI), can generate new code snippets based on natural language descriptions or existing code, significantly improving development efficiency and innovation.
[0003] However, related technologies are unable to accurately evaluate users' adoption of AI-generated code. Summary of the Invention
[0004] The present application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for determining code adoption status, to at least address the problem in related technologies of being unable to accurately assess the adoption status of user-submitted code relative to AI-generated code.
[0005] The present application provides a method for determining code adoption status, comprising: upon obtaining a first code submitted by a target object, extracting semantic information from the first code, and generating a first code vector in a preset format corresponding to the first code based on the semantic information; searching a pre-constructed vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a similarity requirement with the first code vector, wherein the vector database includes code vectors corresponding to respective codes in a preset artificial intelligence code library; based on the at least one candidate code vector, extracting at least one second code pointed to by the at least one candidate code vector from the artificial intelligence code library; inputting the first code and the at least one second code into a large language model, and determining the adoption status of the first code for the at least one second code based on an output result of the large language model, wherein the large language model is used to perform semantic understanding of the code to determine the similarity of the code.
[0006] The present application also provides a device for determining code adoption status, including: a vector conversion module, which is used to extract semantic information from the first code when obtaining the first code submitted by the target object, and generate a first code vector in a preset format corresponding to the first code based on the semantic information; a retrieval module, which is used to search a pre-constructed vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets the similarity requirement with the first code vector, wherein the vector database includes code vectors corresponding to each code in a preset artificial intelligence code library; an extraction module, which is used to extract at least one second code pointed to by the at least one candidate code vector in the artificial intelligence code library based on the at least one candidate code vector; an adoption determination module, which is used to input the first code and the at least one second code into a large language model, and determine the adoption status of the first code for the at least one second code based on the output result of the large language model, wherein the large language model is used to perform semantic understanding of the code to determine the similarity of the code.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for determining code adoption situations when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned methods for determining code adoption status are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for determining code adoption when the computer program is executed by a processor.
[0010] This application significantly improves the accuracy and reliability of identifying adoption status by extracting semantic features of code through a deep learning model, combined with efficient vector retrieval and deep semantic comparison. Based on the candidate code vectors retrieved from the first code vector, the system can accurately locate the source code of each candidate code vector in the AI code library. Then, through vector retrieval, it can initially search for the actual second code, significantly narrowing the search scope. Based on the understanding capabilities of the large language model, the system then performs in-depth semantic analysis and context-aware verification on the first code and at least one second code, accurately determining the adoption status between the first code and at least one second code. This two-level code recognition allows for rapid and accurate determination of the adoption status between the first code and at least one second code. Based on the adoption status, the actual value and impact of AI-assisted programming tools can be accurately assessed. This achieves accurate quantification of the capabilities of AI-assisted programming tools, facilitating targeted optimization of their capabilities. Therefore, it can address the technical issue of being unable to accurately assess the adoption status of user-submitted code relative to AI-generated code, achieving the technical effect of quickly and accurately determining the adoption status between the first code and at least one second code. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is a hardware structure block diagram of a server device according to a method for determining code adoption status in an embodiment of the present application;
[0013] Figure 2 is a flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0014] Figure 3 This is a second flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0015] Figure 4 This is a flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0016] Figure 5 This is a fourth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0017] Figure 6 This is a fifth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0018] Figure 7 This is a sixth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0019] Figure 8 This is a seventh flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0020] Figure 9 This is a flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0021] Figure 10 This is a ninth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0022] Figure 11 This is a tenth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0023] Figure 12 This is a flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0024] Figure 13 This is a twelfth flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0025] Figure 14 Flowchart 13 of a method for determining code adoption status according to an embodiment of the present application;
[0026] Figure 15 Flowchart 14 of a method for determining code adoption status according to an embodiment of the present application;
[0027] Figure 16 FIG15 is a flowchart of a method for determining code adoption status according to an embodiment of the present application;
[0028] Figure 17 This is a structural block diagram of a device for determining code adoption status according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the method for determining code adoption status depends, the specific application environment architecture or specific hardware architecture is described herein.
[0033] The method for determining the code adoption status provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a server device for determining a code adoption situation according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The server device may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the code adoption status in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to a server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0035] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0036] The embodiment of the present application provides a method for determining the code adoption situation, which is applied to the above-mentioned server device, and the method is described in detail in conjunction with the execution process of the method for determining the code adoption situation. Figure 2 As shown, the method includes the following steps S200-230:
[0037] Step S200 : When a first code submitted by a target object is obtained, semantic information in the first code is extracted, and a first code vector in a preset format corresponding to the first code is generated based on the semantic information.
[0038] Specifically, when the system receives a code snippet (first code) submitted by a target, its primary task is to perform deep semantic extraction on the code, rather than simply performing superficial text matching. This step aims to capture the code's functionality, logical structure, and underlying meaning, providing richer information for subsequent similarity comparisons. The system then leverages a pre-trained code embedding model to convert the extracted semantic information into a fixed-dimensional vector representation, the first code vector. This vectorization process provides a more concise representation of the code's semantic features, facilitating subsequent processing and storage.
[0039] Step S210: searching a pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a certain similarity with the first code vector.
[0040] Among them, the vector database includes code vectors corresponding to each code in the preset artificial intelligence code library.
[0041] Specifically, the system uses the first code vector to perform a similarity search in a pre-built vector database. This database contains vector representations of various code snippets generated in the past by artificial intelligence code generators (such as large language models) (called second code vectors). The goal of the search is to find candidate code vectors that have a high degree of similarity with the first code vector at the semantic level to further identify whether code adoption behavior has occurred. Meeting the similarity standard means that the similarity between the candidate code vector and the first code vector exceeds a preset threshold, which is usually achieved by calculating metrics such as cosine similarity or Euclidean distance between the vectors.
[0042] Step S220: Based on the at least one candidate code vector, extract at least one second code pointed to by the at least one candidate code vector from the artificial intelligence code library.
[0043] Specifically, after determining the candidate code vectors, the system needs to locate the original code snippets (second code) corresponding to these vectors in the artificial intelligence code library.
[0044] Exemplarily, code location: using metadata in the candidate code vector (such as code identifier (ID), generation time, model version, etc.) to find and extract the corresponding second code in the artificial intelligence code library.
[0045] Step S230: Input the first code and the at least one second code into the large language model, and determine the adoption of the first code for the at least one second code based on the output result of the large language model.
[0046] Among them, the large language model is used to perform semantic understanding of the code to determine the similarity of the code.
[0047] Specifically, to verify the semantic similarity between the first and second codes—that is, to determine whether the first code substantially adopts or is significantly inspired by the second code—the system feeds the two sets of codes into a pre-trained large language model. Leveraging its deep understanding capabilities, the large language model performs a high-level semantic comparison of the code pairs, determining whether there is substantial similarity between them and whether this similarity meets the adoption criteria.
[0048] For example, a structured prompt is constructed for each code pair, including a code snippet, metadata (such as generation time and file path), and explicit inquiries about code similarity and adoption. A large language model is then invoked to parse the judgment results and reasons returned by the model to determine the adoption relationship. The model may return a similarity score and a classification label (such as adopted, partially adopted, inspired, or unrelated).
[0049] In this embodiment, the semantic features of the code are extracted through a deep learning model, and then combined with efficient vector retrieval and deep semantic comparison, which significantly improves the accuracy and reliability of identifying adoption status. Based on the candidate code vectors retrieved by the first code vector, the system can accurately locate the source code of each candidate code vector in the artificial intelligence code library, and then preliminarily search for the actual second code through vector retrieval, greatly narrowing the scope of the search. Then, based on the understanding ability of the large language model, the first code and at least one second code are deeply semantically analyzed and context-awarely verified to accurately determine the adoption status between the first code and at least one second code. Thus, through two-level code recognition, the adoption status between the first code and at least one second code can be quickly and accurately determined. Based on the adoption status, the actual value and impact of AI-assisted programming tools can be accurately evaluated. Accurate quantification of the capabilities of AI-assisted programming tools is achieved, facilitating targeted optimization of the capabilities of AI-assisted programming tools.
[0050] In one embodiment, Figure 3 As shown, step S210 is to search the pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets the similarity requirement with the first code vector. It includes steps S300-S320:
[0051] Step S300: construct a first vector index corresponding to the first code vector according to the content of the first code vector and metadata corresponding to the first code.
[0052] Specifically, after obtaining the first code vector, the system needs to create an index for the first code vector in the vector database, combining metadata related to the first code (such as the code generation time, the version of the generated AI model, and the file path). This index is designed to store and organize vector information and related metadata, enabling the system to efficiently retrieve code vectors similar to the first code vector.
[0053] For example, you can choose a database system that supports vector storage and retrieval, such as Milvus or a relational database with a vector extension (Postgre Structured Query Language + pgvector, PostgreSQL + pgvector). Configure the database to support efficient data storage and retrieval. Ensure that when the first code vector is stored in the database, its associated metadata is saved to facilitate subsequent retrieval and analysis. Build efficient indexes on the vector field, such as Hierarchical Navigable Small World (HNSW) or Inverted File with Asymmetric Distance Computation (IVFADC), to improve the speed and accuracy of similarity searches.
[0054] Step S310 : searching a vector database according to the first vector index to determine a plurality of code vectors that meet a matching degree with the first vector index.
[0055] The matching degree between the multiple code vectors and the first vector index is used to indicate the similarity between the multiple code vectors and the first code vector.
[0056] Specifically, based on the constructed first vector index, the system performs a similarity search in the vector database, aiming to find multiple code vectors whose semantic similarity with the first code vector reaches or exceeds a preset matching threshold.
[0057] For example, cosine similarity, Euclidean distance, or other metrics suitable for vector similarity assessment can be used to set a matching threshold (e.g., cosine similarity > 0.7). Using the constructed vector index, an approximate nearest neighbor (ANN) search is performed to quickly identify multiple code vectors that meet the matching threshold. The search results are sorted by matching, and code vectors with matching scores above the preset threshold are retained for further verification.
[0058] Step S320: A code vector among the multiple code vectors whose similarity to the first code vector is greater than a first similarity threshold is selected as a candidate code vector.
[0059] Specifically, among the multiple retrieved code vectors, those whose similarity to the first code vector exceeds a first similarity threshold are further screened as candidates for subsequent adoption verification. This step aims to refine the search results and ensure that the candidate code vectors have substantial similarity with the first code vector, thereby improving the efficiency and accuracy of adoption verification.
[0060] For example, a high similarity threshold (e.g., cosine similarity > 0.85) is set to ensure that the candidate code vectors are semantically significantly similar to the first code vector. From the search results, all code vectors with similarities exceeding the first similarity threshold are selected to construct a candidate code vector set for subsequent deep semantic verification.
[0061] In this embodiment, by constructing a first vector index and an efficient retrieval mechanism, the system can quickly find code snippets similar to the target code in a large-scale code base, significantly improving the speed of code retrieval. At the same time, by setting a matching threshold, the semantic relevance of the retrieval results is ensured and the accuracy is improved. The setting of the first similarity threshold can effectively screen out code vectors with large semantic differences from the target code, avoid unnecessary deep verification of low-similarity code, and save computing resources and time costs. The determination of candidate code vectors lays the foundation for subsequent deep semantic verification, enabling the system to carefully analyze the code adoption situation, including direct adoption, partial adoption, and situations inspired by AI code, which enhances the accurate evaluation and traceability of code adoption.
[0062] In one embodiment, Figure 4 As shown, in step S200, when the first code submitted by the target object is obtained, semantic information in the first code is extracted, and a first code vector in a preset format corresponding to the first code is generated based on the semantic information. It includes: steps S400-S420:
[0063] Step S400 : performing segmentation processing on the first code to segment the first code into multiple code segments.
[0064] Specifically, the acquired original code (first code) is broken down into smaller, more manageable units—multiple code snippets. This segmentation process is based on the code's structure and logic, ensuring that each code snippet contains relatively independent semantic information, facilitating subsequent semantic parsing and vectorization.
[0065] Exemplarily, it can be based on the abstract syntax tree (AST) segmentation, using a syntax parser to build an abstract syntax tree (AST) of the first code, and segmenting the code into multiple logical units, such as functions, classes or important control flow blocks, according to the structure of the AST.
[0066] In step S410 , a preset semantic parsing model is used to perform semantic parsing on the plurality of code snippets respectively, so as to generate semantic vectors of preset dimensions corresponding to the plurality of code snippets respectively.
[0067] Among them, the semantic parsing model is used to capture the grammatical structure in the code snippet and parse the semantic information of the code snippet based on the grammatical structure.
[0068] Specifically, we perform deep semantic parsing on the code snippets obtained from the segmentation, and use a pre-trained semantic parsing model to generate fixed-length semantic vectors that match each code snippet. These semantic vectors not only reflect the grammatical structure of the code, but more importantly, express the deep semantics and functional characteristics of the code.
[0069] For example, a pre-trained semantic parsing model with code optimization is selected and loaded into memory. Each code snippet is converted to the model's input format, which may include tokenization, padding, or truncating sequences to a preset length. The pre-processed code snippet is then fed into the semantic parsing model, which outputs a fixed-dimensional semantic vector. Each code snippet will have its own semantic vector representation, reflecting its unique semantic characteristics.
[0070] Step S420 : Combining the semantic vectors of the preset dimensions corresponding to the plurality of code snippets to generate a first code vector.
[0071] Specifically, after obtaining semantic vectors for multiple code snippets, these vectors are combined into a single, integrated representation, known as the first code vector. This step integrates the semantic features of the individual code snippets through mathematical operations or algorithms, enabling the first code vector to fully reflect the semantic structure and functional characteristics of the entire code snippet.
[0072] For example, a vector integration strategy is selected, such as average pooling (averaging multiple semantic vectors), weighted summation (weighting different code snippets based on their importance), max pooling (selecting the highest value), or an attention mechanism (more granular consideration of the importance of different snippets). The selected integration strategy is then applied to the semantic vectors of multiple code snippets to generate a final first code vector. The generated first code vector can also be normalized to ensure comparability and stability in subsequent similarity matching.
[0073] In this embodiment, by structurally segmenting the first code and performing semantic parsing and vectorization on each code snippet, the system can gain a deeper understanding of the code's semantic structure, going beyond superficial textual matching. This significantly improves the depth and accuracy of code understanding and comparison. Each code snippet is converted into a fixed-dimensional semantic vector. This vectorization enables the system to represent the code's semantic features from multiple perspectives, such as grammatical structure, function calls, and variable usage, enhancing the richness of code representation. The generation of the first code vector provides a unified basis for subsequent similarity comparisons with vectors of historical AI-generated code. Inter-vector comparisons (such as cosine similarity) are more efficient than textual matching, reducing comparison time in large code bases and improving system performance. The generation of the first code vector makes it possible to accurately assess which portions of the developer-submitted code actually adopted the AI-generated code. By comparing the vectors with historical AI-generated code, adoption patterns can be more accurately identified, avoiding false positives and missed detections, and improving the accuracy and reliability of code adoption analysis.
[0074] In one embodiment, Figure 5 As shown, step S400 is to segment the first code to segment the first code into multiple code segments. It includes steps S500-S520:
[0075] Step S500: determining the programming language used by the first code according to feature information of the first code.
[0076] Specifically, before processing the first code, the system needs to accurately identify the programming language used in the code. This is achieved based on the programming language's unique syntax, keywords, and structure, laying the foundation for subsequent code segmentation and normalization. Accurate programming language identification ensures the applicability and efficiency of subsequent processing steps.
[0077] Step S510 : segmenting the first code using a corresponding segmentation strategy according to the programming language used by the first code to generate a plurality of code segments.
[0078] Specifically, based on the programming language determined in the previous step, the system uses a corresponding code segmentation strategy to decompose the first code into multiple code fragments, each of which contains relatively independent semantic information. This step facilitates subsequent code standardization and semantic vectorization, as smaller code fragments are easier to manage and analyze. The segmentation granularity can be set according to actual needs, such as function level or statement block level, to meet the needs of different scenarios.
[0079] Step S520 : performing normalization processing on the plurality of code snippets to eliminate non-semantic information in the plurality of code snippets and converting the plurality of code snippets into a preset standard format.
[0080] Specifically, through normalization processing, unnecessary information that does not affect the semantics of the code snippet, such as comments and whitespace, is eliminated, and the code snippet is converted into a unified standard format to facilitate subsequent vectorization and similarity comparison.
[0081] In this embodiment, the accurate identification of the programming language ensures the correct application of the subsequent code segmentation strategy, avoids inconsistent processing due to language recognition errors, and improves the accuracy and efficiency of code processing. Code segmentation helps to decompose the first code into logically independent and more manageable code fragments, each of which carries specific semantic information, which is convenient for subsequent normalization processing and semantic analysis. Through normalization processing, non-essential factors that do not reflect semantics in code differences are eliminated, the accuracy of code text comparison is improved, and mismatches caused by different formats are reduced. Converting code fragments into a preset standard format unifies the coding style, simplifies the subsequent vectorization, storage and similarity comparison processes, and improves the processing speed of the overall system and the convenience of data management. It can significantly optimize the processing flow of the code submitted by developers, improve the accuracy and efficiency of code understanding and adoption analysis, and reduce the consumption of computing resources.
[0082] In one embodiment, Figure 6 As shown, step S510, according to the programming language used by the first code, the first code is segmented using a corresponding segmentation strategy to generate multiple code fragments. It includes: steps S600-S640:
[0083] Step S600: Loading a corresponding parser according to the programming language used by the first code.
[0084] Specifically, before processing the code, the system first needs to identify the programming language used in the first code. This step involves automatic language detection based on code features (such as keywords and grammatical structure). Once the programming language is identified, the system will load a parser that matches the language. A parser is a tool that understands the grammatical structure of a specific programming language, providing the foundation for subsequent code analysis and segmentation.
[0085] Step S610: parse the first code using a parser to generate an abstract syntax tree corresponding to the first code.
[0086] Specifically, an Abstract Syntax Tree (AST) is a tree-like data structure used to represent the structure of source code. By generating an AST for the first code segment, the system can understand the logical structure and grammatical features of the code, which serves as the basis for subsequent segmentation. The AST provides a structured representation of the grammatical structure of the code, enabling the system to more accurately identify logical units such as functions, classes, and methods within the code.
[0087] Step S620 : determining multiple semantic logical scopes of the first code based on the logical structure of the first code indicated by the abstract syntax tree.
[0088] Specifically, based on the AST, the system further analyzes the logical structure of the first code and identifies scopes within the code that have independent semantic logic. These scopes can be functions, classes, methods, or control flow blocks, which constitute the logical units of the code and are the basic units for subsequent code fragment segmentation.
[0089] For example, we traverse the AST, identify logical units (such as functions, class definitions, and loops), and record their locations in the code (such as the starting and ending lines). We define a scope for each logical unit, which encompasses its complete code implementation to ensure logical integrity.
[0090] Step S630 : segmenting the first code based on the semantic logical scope to divide the first code into multiple code segments.
[0091] Each code snippet includes a complete semantic logic range.
[0092] Specifically, based on the determined semantic logic scope, the system divides the first code into multiple code fragments, each of which contains a complete semantic logic unit. This helps accurately analyze the semantic information of the code fragments and also facilitates subsequent normalization and vectorization.
[0093] For example, based on the logical unit range defined in the AST, corresponding code snippets are cut out from the first code, ensuring the integrity and independence of each snippet. Each code snippet is stored, and metadata of the snippet is recorded, such as the starting position, ending position, and logical unit type of the snippet.
[0094] Step S640 : establishing an association relationship between each code snippet and the first code based on the metadata of each code snippet.
[0095] The association relationship includes location information of the code snippet in the first code.
[0096] Specifically, to track and understand the relationship between code snippets and the original code, the system creates an association for each code snippet, which contains the exact location of the code snippet within the original code. This step ensures that even if the code is split into multiple snippets, the original location of each snippet can be traced, facilitating subsequent code adoption analysis and auditing.
[0097] For example, metadata is recorded for each code snippet, including the starting and ending lines of the code snippet in the first code, the file path to which it belongs, and possibly the function or class name. When storing each code snippet, an associated database or data structure is established to record the association between the snippet and the first code, facilitating subsequent retrieval and analysis.
[0098] In this embodiment, by loading a parser specific to the programming language and generating an AST, the system can more deeply understand the logical structure of the code, laying a solid foundation for the subsequent segmentation of code snippets, thereby improving the depth of code understanding and analysis. Based on the logical structure indicated by the AST, the system determines the semantic logical scope in the code, ensuring that each code snippet contains complete semantic information rather than a random code snippet, which improves the accuracy and effectiveness of subsequent code processing. By establishing an association relationship with the first code for each code snippet, the system can track the original location of the code snippet, which is crucial for code adoption analysis, ensuring accurate identification and complete tracing of adoption situations, and avoiding misjudgments and omissions in code adoption analysis. This improves the depth and efficiency of code understanding and processing, and significantly enhances the accuracy and tracing capabilities of code adoption analysis.
[0099] In one embodiment, Figure 7 As shown, after the first code is parsed by a parser in step S610 to generate an abstract syntax tree corresponding to the first code, the method further includes steps S700-S720:
[0100] Step S700: When it is determined that the parser fails to parse the first code, the position of the blank line in the first code is determined.
[0101] Specifically, when a standard code parser fails to correctly parse the first code, it indicates that the code may contain syntax errors or abnormal formatting. To avoid completely abandoning the analysis of this part of the code, the system instead searches for and uses the positions of blank lines in the code as split points. The reason for this is that in most programming languages, blank lines are usually used to separate different logical parts of the code. Using them as split points can help the system restore the logical structure of the code as much as possible, even in the case of syntax errors or confusing formatting.
[0102] For example, a regular expression or string manipulation function can be used to detect the position of blank lines in the code. All blank line positions found are then recorded as reference points for subsequent code segmentation.
[0103] Step S710: Using blank line positions in the first code as segmentation positions to segment the first code into multiple code segments.
[0104] Specifically, once the locations of blank lines in the first code are determined, the system will split the first code into multiple smaller code fragments based on these locations. Each fragment is a logically independent code block. Although it may contain grammatical errors or disordered formatting, it still retains some semantic information for subsequent processing.
[0105] For example, based on the recorded blank line positions, the string slicing function is used to generate multiple code snippets. These code snippets are stored, and the start and end positions of each snippet, as well as possible context information, are recorded to prepare for subsequent analysis.
[0106] Alternatively, in step S720, when it is determined that the parser fails to parse the first code, the first code is segmented in a sliding window manner according to a preset number of lines to segment the first code into multiple code segments, and the number of lines of each code segment is less than or equal to the preset number of lines.
[0107] Specifically, as an alternative strategy, when the parser fails to parse the first code and cannot reliably use blank lines for segmentation, the system will use a fixed number of lines as a window for sliding segmentation of the code. This method ensures that even in complex code structures or with a large amount of continuous logic, suitable code fragments can be generated. Although this may lead to logical discontinuities between fragments, it is an effective means of ensuring data availability when more precise segmentation is not possible.
[0108] For example, a reasonable line number threshold, such as 50 lines, is pre-set as the size of the sliding window. Starting from the first line of the first code, code snippets are generated according to the preset line number threshold, and then the window is moved to continue generating subsequent snippets until the entire code is covered.
[0109] In this embodiment, by identifying and utilizing blank lines as segmentation points, the system can still decompose the code and generate relatively independent code fragments when encountering a first code with a grammatical error or complex format, thereby increasing the stability and robustness of the code processing flow. The sliding window segmentation strategy provides a flexible way to generate processable code fragments even when the code has no obvious logical separation, thereby ensuring the data processing capability of the system under various circumstances. After the parser fails to parse the first code, the above two alternative strategies are adopted to minimize the impact of the parsing failure on the overall code processing flow and ensure the maximum utilization of the data. Therefore, when the standard code parsing method encounters an obstacle, the system can still effectively process the first code through the alternative code segmentation strategy and generate code fragments that can be used for subsequent analysis, which not only improves the robustness and flexibility of the code processing flow, but also ensures the comprehensiveness and accuracy of the code adoption analysis, and realizes more effective tracking and evaluation of the code adoption situation in the software development process.
[0110] In one embodiment, step S520, normalizing the plurality of code snippets, includes:
[0111] Removed comments from multiple code snippets.
[0112] Specifically, while comments help understand the intent of the code, they often don't directly reflect the code's functionality or structure in code similarity comparisons and semantic analysis. Therefore, the system removes comments from code snippets to eliminate non-semantic differences and focus on the code's logical and structural features.
[0113] And / or, trim leading and trailing whitespace from each line of code in multiple code snippets.
[0114] Specifically, whitespace (including spaces and tabs) is used in code for formatting and readability but does not contribute to the code's functionality. Removing leading and trailing whitespace from each line of code standardizes the code format and avoids misjudgment of code similarity due to formatting differences.
[0115] And / or, replace consecutive spaces in multiple code snippets with a single space.
[0116] Specifically, in code, multiple consecutive spaces are often used for alignment or formatting, but these are unnecessary for semantic analysis. Replacing consecutive spaces with single spaces can further standardize code formatting and reduce the impact of formatting differences on similarity comparisons.
[0117] And / or, replace local variable names in multiple code snippets with preformatted identifiers.
[0118] Specifically, the choice of local variable names may be influenced by the developer's personal style and have little relevance to the semantic information of the code. Replacing these variable names with uniform, pre-formatted identifiers can focus on the code structure and algorithmic logic, improving the accuracy of code comparison, especially when code snippets are highly similar but the variable names differ.
[0119] And / or, make the capitalization of letters consistent across multiple code snippets.
[0120] Specifically, while the case of keywords and function names in some programming languages can affect the functionality of the code, in most cases, case differences do not change the semantics of the code. Standardizing the case of letters can reduce unnecessary code differences and enhance the accuracy and efficiency of code comparison.
[0121] After the code snippets are normalized, an association relationship is established between the code snippets before normalization and the code snippets after normalization.
[0122] Specifically, to ensure that subsequent analysis can be traced back to the original code, the system needs to establish a relationship between the code snippet before normalization and the normalized code snippet after normalization. This relationship records the changes in the code snippet during the normalization process, making it easier to restore the original state of the code during adoption analysis, thereby accurately evaluating adoption status.
[0123] For example, metadata for each code snippet can be recorded, including the original code text, normalized code text, variable name replacement records, case conversion records, etc. In the code snippet database, an associated entry is created for each normalized code snippet, containing information about the original code snippet and the specific changes made during the normalization process, ensuring that subsequent analysis can accurately trace back to the original code state.
[0124] In this embodiment, the removal of comments, standardization of whitespace, and unification of variable names and capitalization reduce format differences in code text, enhance similarity comparisons between code snippets, and improve the accuracy of adoption analysis. Normalized code snippets are easier to process and vectorize because they eliminate unnecessary format differences, simplify the input of the vectorization model, and improve vectorization efficiency and quality. Even in the case of inconsistent code formats, the system can ensure data consistency through normalization, enhancing its ability to handle complex or poorly formatted code snippets. By recording code snippets and their changes before and after normalization, the system can accurately trace back to the original state of the code, which is crucial for detailed analysis and auditing of code adoption.
[0125] In one embodiment, Figure 8 As shown, step S220, based on at least one candidate code vector, extracts at least one second code pointed to by at least one candidate code vector from the artificial intelligence code library. It includes: steps S800-S820:
[0126] Step S800: parse at least one candidate code vector to determine metadata and code content corresponding to the at least one candidate code vector.
[0127] Specifically, after the similarity search, the system needs to decode the candidate code vectors and recover the code snippets represented by these vectors and their accompanying metadata. This is to further analyze whether the code snippets are truly associated with the target code and provide detailed information for subsequent semantic similarity verification.
[0128] For example, candidate code vectors can be read from a vector database. Using the key stored in the database, metadata associated with the vector (such as programming language, file path, code generation time, user identifier (ID), etc.) and code text can be retrieved. The retrieved metadata and code text are reassembled to restore the original state and context of each candidate code snippet, ensuring the accuracy of subsequent analysis.
[0129] Step S810: Locate a code snippet pointed to by at least one candidate code vector in the artificial intelligence code library based on the metadata and code content.
[0130] Specifically, using the metadata of the candidate code vector (e.g., file path, code ID), the system can accurately locate the corresponding code snippet in the AI code library. This step is the cornerstone for ensuring that subsequent analysis can compare with the correct code snippet.
[0131] Step S820: Construct at least one second code pointed to by at least one candidate code vector according to the code segment pointed to by at least one candidate code vector.
[0132] Specifically, the system needs to reorganize the located code snippets into a second code for the next step of semantic similarity verification. The second code should contain the complete content of all code snippets parsed from the candidate code vectors, so as to perform a comprehensive comparative analysis with the target code.
[0133] For example, based on the location information (e.g., line number, paragraph identifier) and code content of the located code snippets, they are reassembled according to the original code structure to generate secondary code. This ensures that the reassembled secondary code retains its original context in the source code repository, including file name, code generation time, and possible comments, to maintain code integrity and traceability.
[0134] In this embodiment, by parsing candidate code vectors and locating specific code snippets, the system can accurately compare code based on metadata and code content, avoiding the false positives and false negatives that can occur with simple matching methods based on text similarity. Leveraging code metadata, the system can trace the origins of each candidate code snippet, clarifying its relevance to the target code and its adoption path, enabling precise code tracing. Locating and constructing a secondary code prepares for subsequent in-depth semantic analysis.
[0135] In one embodiment, Figure 9 As shown, in step S230, before inputting the first code and at least one second code into the large language model and determining the adoption of the first code for the at least one second code based on the output result of the large language model, the method further includes steps S900-S920:
[0136] Step S900: classify at least one second code according to a preset similarity threshold.
[0137] Among them, the second code whose similarity with the first code is greater than the first threshold belongs to the first category, the second code whose similarity with the first code is less than or equal to the first threshold and greater than or equal to the second threshold belongs to the second category, and the second code whose similarity with the first code is less than the second threshold belongs to the third category.
[0138] Specifically, the system uses a pre-set similarity threshold to categorize the second code into three categories, based on the similarity calculated between the first and second codes. This classification allows the system to identify highly similar, moderately similar, and dissimilar second codes, providing a basis for subsequent matching decisions.
[0139] Exemplarily, a similarity score between a first code and a second code is calculated using cosine similarity or a code-specific similarity metric, such as a metric based on edit distance. Two thresholds are defined, with the first threshold being higher than the second threshold. For example, the first threshold is set to 0.8, and the second threshold is set to 0.3, representing the boundary between high and moderate similarity. A second code with a similarity greater than the first threshold is classified as the first category (highly similar); a second code with a similarity less than or equal to the first threshold and greater than or equal to the second threshold is classified as the second category (moderately similar); and a second code with a similarity less than the second threshold is classified as the third category (dissimilar).
[0140] Step S910: When at least one second code belongs to the first category, determining that the first code adopts the at least one second code.
[0141] Specifically, when the similarity score of the second code exceeds a first threshold, the system automatically considers these code snippets to be adopted by the first code, meaning that the first code clearly incorporates the logic or structure of the second code. This step is a preliminary judgment of the adoption decision, providing a candidate set for subsequent in-depth verification.
[0142] For example, if the similarity score of the second code exceeds a preset first threshold, the system automatically marks the first code as having adopted the second code, records the adoption relationship, and prepares for subsequent in-depth verification. The code adoption relationship database is updated to record the adoption relationship between the second code belonging to the first category and the first code in detail, including the adopted code snippet, similarity score, and so on.
[0143] Step S920: When the at least one second code belongs to the third category, it is determined that the first code does not adopt the at least one second code.
[0144] Specifically, for second codes belonging to the third category, i.e., code snippets with a similarity score below the second threshold, the system determines that the first code does not adopt these second codes. This helps filter out code snippets that are irrelevant to the logical structure of the first code, avoiding unnecessary in-depth verification and analysis later.
[0145] For example, if the similarity score of the second code is lower than the second threshold, the system regards this part of the code fragment as not adopted and does not perform further in-depth verification.
[0146] In this embodiment, by classifying the second code according to a preset similarity threshold, the system can quickly eliminate code fragments that are irrelevant to the first code and focus on highly similar and moderately similar codes, thereby improving the efficiency and accuracy of adoption decisions. The classification mechanism effectively filters out a large number of dissimilar second codes, reduces the candidate set for deep verification, avoids unnecessary consumption of computing resources, and improves the overall operating efficiency of the system. Dividing the similarity into three categories can not only quickly mark obvious adoption situations, but also retain moderately similar second codes for further deep verification, support refined analysis of the degree of code adoption, and increase the depth and breadth of adoption analysis. Through preliminary classification, the system can allocate computing resources more reasonably and avoid wasting resources on irrelevant code fragments. This is especially important when processing large-scale code bases and helps improve the overall performance and resource utilization of the system.
[0147] In one embodiment, Figure 10 As shown, in step S230, the first code and at least one second code are input into the large language model, and the adoption of the first code for the at least one second code is determined based on the output result of the large language model. It includes: steps S1000-S1010:
[0148] Step S1000: When at least one second code belongs to the second category, the second code belonging to the second category is input into the large language model, and structured prompt words are used to instruct the large language model to perform semantic understanding of the first code and the second code to obtain the output result of the large language model.
[0149] Specifically, for secondary codes that have a moderate level of similarity to the primary code, the system leverages a large language model for in-depth semantic understanding and comparison. This is because traditional similarity calculations may not fully capture the logical and functional relationships between codes, especially after the secondary code has been modified, reorganized, or variables renamed. By providing structured prompts to the large language model, the system guides the model to understand the potential connections between the two code segments, resulting in more accurate adoption decisions.
[0150] For example, a specific prompt template can be designed for the large language model. The template should include the content of the first and second code, as well as possible metadata (such as the code generation time and programming language). The model is required to evaluate the semantic similarity and adoption likelihood of the two code segments. The large language model is called and the structured prompt is passed to the model; the model will return the evaluation results of the similarity and adoption of the two code segments, including the probability score of adoption or the explicit adoption judgment, as well as the logical analysis supporting the judgment.
[0151] Step S1010 , when the output result of the large language model is that the first code adopts at least one second code, determining that the first code adopts at least one second code.
[0152] Specifically, based on the output of the large language model, if the model determines that the first code adopts the second code, the system will formally confirm this adoption relationship. This step is the decisive part of the adoption analysis, relying on the deep understanding and advanced reasoning capabilities of the large model to accurately determine whether the modified code actually adopts the AI-generated code logic.
[0153] For example, the system analyzes the output of a large language model. If the model indicates an adoption, the system records the adoption relationship as part of the final analysis. The adoption relationship database is updated to record the adoption event, including the adopted second code snippet, evidence of adoption (such as the analysis text output by the model), and the similarity score between the two code snippets.
[0154] In this embodiment, the introduction of a large language model significantly improves the system's ability to understand and judge code adoption, especially for code snippets that have been modified, reorganized, or renamed by developers, enabling a more accurate assessment of their adoption relationship with the original code. Through the deep semantic understanding of the large model, the system can not only identify surface similarities in code, but also conduct in-depth analysis of adoption at the logical and functional levels, increasing the depth and breadth of adoption analysis. In the adoption judgment of the second category of code snippets, the output of the large model becomes the basis for scientific decision-making, avoiding the misjudgment or omission that may result from relying solely on similarity calculations, and improving the rationality and credibility of adoption decisions. Although calling the large model may increase computational overhead, since the second category of second code has already been screened through similarity classification, the second code has already been screened once and is smaller in number. The system can effectively control the scale of in-depth analysis and avoid excessive resource consumption, thereby ensuring analysis accuracy while not excessively sacrificing system operational efficiency. In-depth adoption analysis is achieved for code snippets with medium similarity belonging to the second category. Therefore, the large language model is used for secondary judgment of code snippets with intermediate similarity and less certainty, balancing computing resources and judgment effectiveness.
[0155] In one embodiment, Figure 11 As shown, the method further includes: steps S1100-S1140:
[0156] Step S1100 : collecting interaction data between the target object and the large language model.
[0157] Among them, the interaction data includes artificial intelligence code generated by the large language model and the contextual content related to the artificial intelligence code.
[0158] Specifically, the system proactively captures all interactions between the target entity (e.g., a software developer) and the large language model. This includes, but is not limited to, code snippets generated by the large model based on the target entity's requests, the model's response text, and metadata associated with each interaction, such as timestamps, context, and question content. The purpose of collecting this information is to build a comprehensive and traceable AI codebase, ensuring that future analysis of code adoption is based on detailed data.
[0159] For example, by monitoring the target object's interaction with the large model (e.g., API calls and IDE plugin communication logs), code snippets and natural language descriptions generated by each interaction are collected in real time or periodically. During this collection process, each request (question text) and corresponding response (including the generated code snippet and additional model explanation information) is tied to the interaction metadata (e.g., timestamp, user ID, and conversation turn) for later analysis.
[0160] Step S1110: parse the interaction data to separate the artificial intelligence code and natural language description in the interaction data.
[0161] Specifically, the system further processes the collected interaction data to separate specific code snippets and their accompanying natural language descriptions, laying the foundation for subsequent code analysis and metadata association. This separation process requires identifying the boundaries between code and text descriptions to ensure the completeness and accuracy of the information.
[0162] For example, regular expressions or specialized code parsing libraries are used to identify code snippets and natural language descriptions in the response text. The identified code snippets and natural language descriptions are separated and stored separately, while their relative positions in the original interaction text are preserved so that context information can be restored later when needed.
[0163] Step S1120: Establish an association relationship between the artificial intelligence code and the metadata corresponding to the interaction data.
[0164] The metadata corresponding to the interaction data includes at least the identification information of the target object, timestamp, conversation turn, and question text.
[0165] Specifically, for each piece of isolated AI code, the system records and associates metadata about its generation context, including but not limited to the code generator (the identity of the target object), the time the code was generated, and the specific text of the question. This allows subsequent code adoption analysis to track the generation context of each piece of code, ensuring in-depth and accurate analysis.
[0166] For example, a unique identifier is generated for each piece of AI code, and relevant metadata (including question text, user ID, timestamp, conversation turn, etc.) is associated with the identifier. A metadata dictionary or database record is constructed, in which the identifier of each code snippet corresponds to its entire contextual information, facilitating subsequent query and analysis.
[0167] Step S1130: Combining the artificial intelligence code and metadata corresponding to the interaction data into preset structured data.
[0168] Specifically, to facilitate management and analysis, the system integrates AI code and its metadata into a structured data format, such as JavaScript Object Notation Object (JSON) or specific database records. This structured data format can be easily read and parsed by subsequent processing programs, ensuring data consistency and ease of use.
[0169] For example, AI code and corresponding metadata are encapsulated into a structured data structure, such as a JSON object, according to pre-set rules. Each metadata item is assigned a key and value to facilitate subsequent program access. This structured data is persistently stored, for example, in a database, to ensure data security and long-term availability.
[0170] Step S1140: Build an artificial intelligence code library based on the structured data.
[0171] Specifically, the system builds an AI-powered codebase based on the processed and integrated structured data. This codebase contains not only code snippets but also detailed metadata, providing a comprehensive dataset for subsequent code adoption metrics and in-depth analysis.
[0172] For example, a database system suitable for storing structured data, such as a relational database, is designed or selected. This structured data is then batch-inserted into the code repository to ensure that each code snippet and its metadata are correctly stored. Furthermore, indexes or other data structures are created to optimize query performance, ensuring rapid retrieval of the required code snippets and metadata.
[0173] In this embodiment, by recording the context of each code generation in detail, including question text, timestamps, and conversation turns, the system can conduct in-depth analysis of code adoption based on accurate background information, avoiding misjudgments or omissions that may result from relying solely on code text similarity calculations, greatly improving the accuracy and reliability of the analysis. The constructed artificial intelligence code library not only contains code snippets, but also integrates detailed metadata, establishing a clear record of the generation context for each piece of code. This provides a valuable traceability path for code adoption analysis and helps understand the evolution and adoption of code snippets. By integrating the code and metadata generated by the large language model into a structured data format, the system greatly simplifies the subsequent code adoption analysis process, making code search, comparison, and related queries more efficient and intuitive, and reducing the complexity of data processing. Storing code and metadata in a structured data format not only facilitates data management and query, but also supports compatibility with multiple data processing and analysis tools, enhancing the system's scalability and future adaptability. The constructed artificial intelligence code library provides a data foundation for the system to evaluate the accuracy and performance of the adoption analysis algorithm. By regularly analyzing the adoption rate trends and algorithm performance, the system can continuously tune parameters and strategies to form a closed-loop optimization mechanism, continuously improving the efficiency and quality of adoption analysis.
[0174] In one embodiment, Figure 12 As shown, step S1110 parses the interaction data to separate the artificial intelligence code and natural language description in the interaction data. It includes steps S1200-S1220:
[0175] Step S1200: parsing the interaction data to determine multiple rounds of dialogue between the target object and the large language model.
[0176] Specifically, the system identifies and isolates multi-round conversations between the target user and the large language model from the collected interaction data. This is because interactions between developers and AI models are often not a single question and response, but rather involve multiple rounds of communication, including questions, AI responses, and requests for further clarification or modification. By analyzing these multi-round conversations, the system can more fully understand the context and dynamic evolution of each code generation task, which is crucial for subsequent code adoption metrics.
[0177] Exemplarily, conversation tracking algorithms or time series analysis are used to identify conversational records belonging to different turns of the same conversation. For example, if two consecutive questions and responses occur within a short period of time and involve the same or related questions, the system will mark them as part of the same conversation. Based on the conversation ID, each group of conversational records is separated, and each conversation contains all relevant questions and answers, including the artificial intelligence code generated by the large language model.
[0178] Step S1210: extract information from the text content in the multiple rounds of conversations to determine metadata in the interaction data.
[0179] Specifically, metadata refers to data about data. For each interaction in a multi-turn conversation, the system extracts relevant metadata, including but not limited to question time, answer time, questioner information, session ID, and conversation turn number. This metadata plays a key role in tracing the context of code generation and understanding the generation background and evolution of code snippets. It serves as an important basis for subsequent code adoption analysis.
[0180] For example, regular expressions or other text analysis tools are used to extract key information from conversation logs. For example, the sending time of each message can be determined by matching strings in timestamp format. Metadata tags such as question time, answer time, questioner ID, conversation ID, and conversation turn number are added to each conversation log to ensure that this information can be clearly recorded and tracked.
[0181] Step S1220: Code recognition is performed on the text content in multiple rounds of dialogue to extract the artificial intelligence code generated by the large language model in the interaction data.
[0182] Specifically, within the text content of multi-round conversations, the system must identify and extract code snippets generated by a large language model. This is the core output of the AI-based code generation function and is crucial for subsequent code adoption metrics, code library construction, and code similarity analysis. Code recognition must accurately distinguish between code and natural language descriptions to facilitate independent processing and analysis of code.
[0183] For example, a code block matching algorithm or regular expressions is used to identify code blocks from the conversation transcript. These blocks are often used as code examples or suggestions in responses. The identified code snippets are extracted from the text and formatted, such as removing extraneous whitespace and comments, to ensure cleanliness and consistency for subsequent semantic comparison and vectorization.
[0184] In this embodiment, through the identification of multiple rounds of conversations and the extraction of metadata, the system can more carefully analyze the dynamic changes in the code adoption process, understand the context before and after the generation of code snippets, and thus more accurately evaluate the actual adoption of AI code. Accurately identifying and extracting artificial intelligence code from interactive data can ensure that the construction of the code base is both comprehensive and efficient, and each code snippet can be accurately indexed to facilitate subsequent retrieval and analysis. The parsing of multiple rounds of conversations and detailed metadata records provide a solid foundation for code tracing, help trace the generation history of code snippets, understand the modifications and optimizations during the adoption process, and increase the transparency and credibility of code adoption metrics. The separation of code identification and metadata extraction ensures the purity of code snippets and the richness of metadata, provides a clear data structure for subsequent code adoption metrics and similarity comparisons, avoids interference from non-code elements, and improves the accuracy and efficiency of analysis.
[0185] In one embodiment, Figure 13 As shown, the method further includes: steps S1300-S1340:
[0186] Step S1300: Configure access credentials and access path for the target code repository.
[0187] The target code repository stores all code managed by the target server's version control system. A code repository is a centralized location for project source code, documentation, scripts, and version history, typically hosted by a version control system (Global Information Tracker, Git). This repository can include source code files, configuration files (such as requirements.txt, pom.xml, and Docker files), test cases and scripts, documentation (such as READMEs, API documentation, and design specifications), and version history (commit records, branches, and tags).
[0188] Specifically, system administrators or automated tools need to configure access information for the target repository, including access credentials (such as username, password, API key, etc.) and access paths, so that they can subsequently access and retrieve all code data stored in the repository. This step is the basis for subsequent code submission access records and new code identification, ensuring that the system can legally and smoothly obtain information from the target repository.
[0189] For example, enter the target repository's login credentials, including but not limited to a username and password, or an API key and authentication token, into the system's configuration file to ensure the system can authenticate and access the repository. Set the target repository's access path, which could be the repository's Uniform Resource Locator (URL) address or the network path of a repository in an on-premises deployment environment, to ensure the system can accurately locate the target repository.
[0190] Step S1310 , in response to the analysis request of the target object, access the target code repository along the access path using the access credentials, and call the code submission records of the version control system within the target time range.
[0191] The analysis request carries a target time range.
[0192] Specifically, when the target entity initiates an analysis request, the system uses the previously configured access credentials and path to connect to the target code repository and retrieve commit records from the version control system (Git) within the specified timeframe. This operation ensures that the system can obtain historical information on all relevant code commits within the target entity within the target timeframe, which is crucial for subsequently identifying new code and building incremental code repositories.
[0193] For example, the system receives and parses an analysis request for a target object, identifying the target time range contained in the request. Using access credentials, the system logs into the code repository and calls the version control system API to retrieve all commit records within the target time range. This may include metadata such as committer information, commit time, commit description, and a list of modified files, as well as specific code change information.
[0194] Step S1320 : determining a plurality of codes submitted by the target object within the target time range according to the code submission records of the version control system within the target time range.
[0195] Specifically, based on the acquired commit records, the system further analyzes the specific code commits of the target object within the target time range. This includes identifying all code files and code changes submitted by the target object, preparing data for further code similarity analysis.
[0196] For example, filter the commits initiated by the target (or its team members) from the commit history to ensure that only code changes related to the target are focused on. Use the functions provided by the version control system, such as Git's git show or git diff commands, to parse the commit history to identify the specific code changes in each commit, including added, deleted, and modified lines of code.
[0197] Step S1330 : determining a new code among the multiple codes whose similarity to the code stored in the target code repository is lower than a preset level.
[0198] Specifically, the system needs to identify and filter out "new code" submitted within the target timeframe that has a low similarity to previously stored code in the repository. This allows for more precise analysis of which code may be AI-generated or significantly influenced by AI, as only new code has a higher probability of matching AI-generated code snippets.
[0199] For example, a code similarity algorithm (such as Levenshtein-based edit distance calculation or vector-based semantic similarity calculation) is used to compare code submitted within the target timeframe with existing code in the repository to determine similarity. A threshold is set. For example, if a code segment's similarity score falls below a preset value (e.g., 70%), it is considered "new code" and meets the criteria for subsequent analysis.
[0200] Step S1340: construct an incremental code library based on the new code and metadata corresponding to the new code.
[0201] The first code is obtained from the incremental code library.
[0202] Specifically, the system constructs an incremental code repository containing the selected new code and its corresponding metadata (such as submitter, submission time, and code change description). This repository serves as the primary data source for subsequent code adoption analysis. The first code, the target code for analysis, is selected from this incremental code repository to assess its adoption relationship with the AI-generated code.
[0203] For example, design the data structure and storage format for the incremental code repository to ensure efficient storage and retrieval of code snippets and their metadata. This could be a table database, a document database, or a customized data storage system. Store the new code and its metadata in the incremental code repository according to the designed data structure, ensuring that each code snippet has its complete context.
[0204] In this embodiment, by configuring access credentials and paths, the system can quickly connect to the target code repository and retrieve commit records within a specified time range. This avoids the tedious process of manually searching and filtering code, significantly improving data acquisition speed and laying a foundation for rapid and accurate subsequent analysis. The filtered new code (i.e., code with a similarity to existing code below a preset value) is more likely to be a potential candidate for adoption by AI-generated code. This narrows the scope of subsequent analysis, improves the targetedness and accuracy of the analysis, and reduces the possibility of misjudgments or omissions. The constructed incremental code repository not only contains new code but also integrates detailed metadata, providing rich contextual information for subsequent code adoption analysis, helping to trace the evolution and adoption process of the code and increasing the transparency and credibility of code adoption metrics. The construction of the incremental code repository enables centralized management and efficient retrieval of new code, avoids repeated processing of all code in the repository, reduces data processing time and computing resource consumption, and makes the code adoption analysis process smoother and more efficient. By configuring access credentials and paths, automatically obtaining submission records within the target time range, screening new code and building an incremental code base, the system not only improves the efficiency and accuracy of code adoption analysis, but also enhances code traceability and process optimization, ensuring data security and compliance.
[0205] In one embodiment, Figure 14 As shown, step S1330 is to determine a new code among the multiple codes whose similarity with the code stored in the target code repository is lower than a preset level. It includes steps S1400-S1430:
[0206] Step S1400 : Analyze the plurality of codes to determine newly added code lines in the plurality of codes compared to the codes stored in the target code repository.
[0207] Specifically, the system analyzes all code submitted by the target entity and identifies those lines of code that are completely new compared to the previous version in the target repository. This is because these new lines of code are more likely to be adopted parts of the AI-generated code and are the focus of subsequent in-depth analysis. These new lines of code are marked as "new code" and prepared for further processing.
[0208] For example, a code difference analysis tool, such as Git's git diff command, is used to compare the newly submitted code of the target object with the latest version of the same file in the target code repository to identify newly added lines. All newly added lines are filtered from the difference analysis results. These lines, which do not appear in the previous code version, are considered "new code."
[0209] Step S1410: The newly added code line is treated as the new code.
[0210] Step S1420 : Analyze the multiple codes to determine the modification degree of each of the multiple codes compared to the code stored in the target code repository.
[0211] Step S1430 : The code whose modification degree is greater than a preset degree among the multiple codes is used as a new code.
[0212] Specifically, in addition to identifying newly added lines of code, the system also needs to assess the extent of code modifications, especially those code snippets that have undergone significant changes. This is because, even if the code lines already exist in the repository, if the modifications are significant enough to significantly alter the code's logic and functionality, these code snippets may also be affected by the AI-generated code. By setting a threshold for the degree of modification, the system can filter out code snippets with a high degree of modification and treat them as "new code" for subsequent adoption measurement.
[0213] Exemplarily, an algorithm is developed or adopted to quantify the degree of code modification. For example, the edit distance, Jaccard similarity index, or a more complex syntax tree difference metric (such as the rate of change based on the abstract syntax tree AST) can be used. A preset degree threshold is set. For example, if the degree of modification of a code snippet exceeds 30%, it is considered that the degree of modification is large and may involve the introduction of new concepts or functions. For those code snippets whose degree of modification exceeds the threshold, the system regards them as "new code", even if some parts of these snippets have been recorded in the repository. Such a strategy can ensure that even if the code is significantly modified on the original basis, it can be included in the analysis scope of the adopted metric.
[0214] In this embodiment, by clearly distinguishing between newly added lines of code and significantly modified code snippets, the system can more finely categorize the code adoption analysis, avoiding treating all modified code as "new code." This allows the system to focus on truly adopted code while avoiding wasted resources and improving the relevance and efficiency of adoption metrics. Quantitative assessment of modification severity helps the system identify code snippets that appear to have undergone minor changes but may actually introduce new logic or functionality, ensuring the comprehensiveness and depth of subsequent adoption metric analysis and improving the accuracy of the analysis. By setting a preset threshold for modification severity, the system can flexibly adjust the analysis strategy based on different project requirements or code nature. For example, in a project with frequent code refactoring, the modification severity threshold can be appropriately relaxed to include more "new code," thereby supporting diverse adoption metric requirements. The sections ultimately marked as "new code," because they are either completely new or significantly modified, are more likely to represent the actual adoption of AI-generated code and can serve as more reliable and practical metrics to help evaluate the effectiveness of AI-assisted programming tools.
[0215] In one embodiment, Figure 15 As shown, the method further includes: steps S1500-S1530:
[0216] Step S1500: When it is determined that the first code has an adoption relationship with the at least one second code, a third code having the highest similarity to the first code is determined in the at least one second code.
[0217] Specifically, once the system has confirmed an adoption relationship between a first code and at least one second code (i.e., an AI-generated code), the next task is to find the segment of these second codes that is most similar to the first code and mark it as the third code. This operation aims to accurately locate the specific content of the adoption, facilitating subsequent adoption calculations and adoption degree analysis.
[0218] For example, all second codes confirmed to have an adoption relationship can be sorted based on the similarity scores calculated. The second code ranked first after sorting is selected as the third code, indicating that the second code has the highest semantic similarity with the first code. From the sorted results, the second code segment with the highest similarity to the first code is extracted as the third code for subsequent adoption calculation.
[0219] Step S1510: Determine the adoption rate of the first code for the third code generated by artificial intelligence based on the similarity between the first code and the third code.
[0220] Specifically, after determining the third code, the system calculates the amount of the third code adopted by the first code based on the specific similarity between the first and third codes. The adoption amount here may refer to the number of lines of code, the number of characters, or the number of logical units adopted, depending on the granularity and needs of the analysis.
[0221] For example, based on the similarity between the third code and the first code, the adoption rate can be calculated in various ways. For example, if the similarity is based on the number of lines of code, the adoption rate is the number of lines in the first code that match the third code. A more advanced calculation method may be based on logical unit matching based on the abstract syntax tree (AST), calculating which function, class, or method fragments have been adopted. This calculation method can better reflect the adoption rate at the code logic level.
[0222] Step S1520 , counting the total amount of new code in the incremental code library within a preset analysis period and the total amount of new code adopted from the AI-generated code.
[0223] Specifically, the system needs to count all new codes within the entire preset analysis cycle (i.e., new codes in the incremental code base or codes with significant modifications), and calculate the total adoption of these new codes for AI-generated codes, which provides the necessary data basis for subsequent adoption rate calculations.
[0224] For example, a preset analysis period is set, which can be one day, one week, one month, or any custom time period. The incremental code base is traversed, and the total amount of new code within the analysis period and the total number of lines or logical units of AI code adopted in these new codes are counted.
[0225] Step S1530: Determine the adoption rate of the new code for the AI-generated code based on the total amount of the new code and the total amount of adoption of the new code for the AI-generated code.
[0226] Specifically, based on the data collected in the first three steps, the adoption rate is calculated. This is a key metric that reflects the proportion of new code that is directly adopted or significantly influenced by AI code.
[0227] For example, the basic formula for calculating the adoption rate is (total number of adoptions / total number of new codes) * 100%. Substituting the total number of adoptions and the total number of new codes counted in step S1520 into the formula, the adoption rate can be calculated.
[0228] Adoption rate output: After the system calculates the adoption rate, it outputs the result, which may be displayed in the form of a graphical report or made available to other systems through an API, allowing the management team to intuitively understand the contribution of AI-generated code in software development.
[0229] In this embodiment, the system accurately locates the specific content of adoption by identifying the highest-similarity code snippets in the adoption relationship. Combining the adoption volume and total volume statistics, the system calculates the adoption rate. This method, based on deep comparison and quantitative statistics, significantly improves the accuracy of the analysis and the scientific nature of the results. The system can not only calculate the total number of lines or characters adopted, but also calculate the adoption volume based on logical units (such as functions and classes). This flexibility adapts to different analysis needs and provides a more detailed perspective on adoption measurement. Through automated statistics and adoption rate calculation, the system avoids the tedious manual analysis, reduces errors in the analysis process, and makes the assessment of code adoption more efficient and standardized. As a key metric, the adoption rate can provide software development teams and management with an intuitive and quantifiable AI code contribution, which can assist in decision-making, such as evaluating the return on investment of AI-assisted programming tools or optimizing developer workflows.
[0230] In one embodiment, Figure 16As shown, the method further includes: steps S1600-S1620:
[0231] Step S1600: When it is determined that the first code has an adoption relationship with the at least one second code according to the adoption status of the first code for the at least one second code, the first code is determined to be an adopted code.
[0232] Specifically, after analyzing and comparing the first code (the developer's submitted code snippet) with the second code (the AI-generated code snippet), if the system determines that the first code clearly adopts part or all of the second code, it marks the first code as "adopted code." This signifies that an adoption relationship has been confirmed, meaning that a portion of the first code originated from or was significantly influenced by the second code.
[0233] For example, based on the code similarity calculation results and semantic verification of the large language model, if the similarity between the first code and the second code exceeds a preset threshold, and the model determines that an adoption relationship exists, the first code is marked as adopted. The identification information of the adopted code, including but not limited to the file name, code snippet location, version control commit ID, etc., is recorded in the adoption relationship library for subsequent query and analysis.
[0234] Step S1610: query the metadata of the first code as the adopted code in the incremental code library to obtain the traceability information of the first code as the adopted code.
[0235] The metadata of the first code includes the submission time when the first code was generated and identification information of the submitter.
[0236] Specifically, for the first code marked as adopted, the system queries its metadata from the incremental code repository to obtain the developer's submission time, submitter identification information, and other information as traceability information for the adoption behavior. This helps understand and track the code adoption process, confirming the time point of adoption and the responsible party.
[0237] For example, through the incremental code repository's database query function, using the adopted code's identification information (such as file path and code snippet location) as a query key, metadata such as the corresponding submission time and submitter ID is obtained. The queried metadata is integrated into the adopted code's traceability information record, ensuring that each adoption record is accompanied by complete contextual information.
[0238] Step S1620: query the metadata of the second code adopted by the adopted code in the artificial intelligence code library to obtain the traceability information of the first code as the adopted code.
[0239] The metadata of the second code includes the conversation text and generation time of generating the second code.
[0240] Specifically, the system needs to query the metadata of the second code (the adopted AI-generated code) from the AI code library, including the text content of the generated dialogue, the generation time, etc., to further enrich the traceability information of the adopted code and ensure that the original source of the AI-generated code can be traced.
[0241] For example, the identification information of the second code (such as the query ID or model response ID corresponding to the code generation) is used to query the AI code library to obtain metadata such as the conversation text and generation time of the second code. The metadata of the second code is associated with the traceability information of the first code to form a complete adoption chain, including the time when the developer submitted the code, the submitter information, and the AI code generation conversation and generation time.
[0242] In this embodiment, by explicitly marking adopted code and extracting relevant metadata, the system can provide detailed traceability information for each adoption behavior, including the generation time and background of the code snippet, the developer submission time, and submitter information. This greatly improves the transparency of adoption analysis and facilitates audits or reviews by team members and management. The system not only relies on code similarity calculations but also combines deep semantic verification with a large language model to more accurately identify adoption behaviors, avoiding the false positives or omissions that may exist in simple comparisons, and ensuring the accuracy of adoption analysis. By directly querying and integrating metadata after determining the adoption relationship, the system simplifies the adoption analysis process, avoids redundant queries and data processing, and improves analysis efficiency. The integrated traceability information not only includes the source of the code, but also covers the conversation context of the code generation and the developer's submission details, which helps to deeply understand the generation process and adoption motivation of the adopted code. In summary, the method of this embodiment improves the transparency and accuracy of code adoption analysis, optimizes the analysis process, and strengthens the understanding and tracking capabilities of adopted code.
[0243] In one embodiment, to ensure the continued effectiveness and accuracy of the system, a set of evaluation and optimization methods can also be provided. Specifically, they include:
[0244] Through manual sampling and annotation, a gold standard verification dataset containing the matching relationship (positive example) and non-matching relationship (negative example) of "first code (Git submission code)-second code (AI generated code)" was constructed.
[0245] The dataset is regularly updated and expanded to cover more scenarios and edge cases.
[0246] The method of this embodiment described above was used on the validation dataset to determine adoption status and calculate key performance indicators: Precision: The proportion of code that the system determined to be adopted that was actually adopted. Recall: The proportion of all AI code that was actually adopted that was recognized by the system. F1-Score: The harmonic mean of precision and recall.
[0247] The system can also periodically or randomly select a subset of matching results (especially those where the LLM results are inconsistent with the vector search results (for example, the search result indicates that the first and second codes have an adoption relationship, but the LLM model output indicates that the first and second codes do not have an adoption relationship), or where the similarity is in the fuzzy range) and submit them to human reviewers for review. The system then collects reviewer feedback and analyzes the causes of false positives and false negatives.
[0248] Based on the evaluation results and manual feedback, key parameters in the system are adjusted, including: code segmentation strategy and granularity, and vector retrieval similarity threshold.
[0249] Then, you can adjust the LLM prompts and assess whether the code embedding model and LLM model need to be replaced or fine-tuned. Accumulate difficult-to-judge cases or error cases to improve the model or rules, forming a closed loop of continuous optimization.
[0250] In this embodiment, a system evaluation and model iterative optimization solution is provided, which can evaluate the system performance (accuracy, efficiency, etc.) based on the system analysis results, and optimize the system based on the evaluation results to improve the accuracy of the adoption judgment.
[0251] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0252] The embodiment of the present application also provides a device for determining code adoption status. Figure 17 : is a structural block diagram of a device for determining code adoption status according to an embodiment of the present application, the device comprising:
[0253] The vector conversion module 1701 is configured to extract semantic information from the first code submitted by the target object upon obtaining the first code, and generate a first code vector in a preset format corresponding to the first code based on the semantic information.
[0254] Retrieval module 1702 is configured to search a pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a requirement for similarity to the first code vector, wherein the vector database includes code vectors corresponding to respective codes in a preset artificial intelligence code library.
[0255] The extraction module 1703 is used to extract at least one second code pointed to by the at least one candidate code vector from the artificial intelligence code library based on the at least one candidate code vector.
[0256] Adoption determination module 1704 is used to input the first code and at least one second code into a large language model, and determine the adoption of the first code with respect to the at least one second code based on the output of the large language model, wherein the large language model is used to perform semantic understanding of the codes to determine the similarity of the codes.
[0257] In an exemplary embodiment, retrieval module 1702 is further configured to construct a first vector index corresponding to the first code vector based on the content of the first code vector and metadata corresponding to the first code. A search is performed in a vector database based on the first vector index to identify multiple code vectors that meet a matching requirement with the first vector index. The matching degree between the multiple code vectors and the first vector index indicates the similarity between the multiple code vectors and the first code vector. Code vectors among the multiple code vectors whose similarity to the first code vector exceeds a first similarity threshold are selected as candidate code vectors.
[0258] In an exemplary embodiment, vector conversion module 1701 is further configured to segment the first code into multiple code snippets. Semantic parsing of the multiple code snippets is performed using a preset semantic parsing model to generate semantic vectors of preset dimensions corresponding to the multiple code snippets. The semantic parsing model is configured to capture the grammatical structure in the code snippets and parse the semantic information of the code snippets based on the grammatical structure. Semantic vectors of preset dimensions corresponding to the multiple code snippets are combined to generate a first code vector.
[0259] In an exemplary embodiment, vector conversion module 1701 is further configured to determine a programming language used by the first code based on feature information of the first code. The first code is segmented using a segmentation strategy corresponding to the programming language used by the first code to generate multiple code snippets. The multiple code snippets are normalized to eliminate non-semantic information in the multiple code snippets and to convert the multiple code snippets into a preset standard format.
[0260] In an exemplary embodiment, the vector conversion module 1701 is also used to load a corresponding parser according to the programming language used by the first code. The parser is used to parse the first code to generate an abstract syntax tree corresponding to the first code. Based on the logical structure of the first code indicated by the abstract syntax tree, multiple semantic logical scopes of the first code are determined. The first code is segmented based on the semantic logical scope to segment the first code into multiple code fragments, wherein each code fragment includes a complete semantic logical scope. Based on the metadata of each code fragment, an association relationship between each code fragment and the first code is established, wherein the association relationship includes the position information of the code fragment in the first code.
[0261] In an exemplary embodiment, the apparatus further comprises:
[0262] The blank line determination module is used to determine the position of the blank line in the first code when it is determined that the parser fails to parse the first code.
[0263] The first segmentation module is configured to use blank line positions in the first code as segmentation positions to segment the first code into multiple code segments.
[0264] The second segmentation module is used to perform sliding window segmentation on the first code according to a preset number of lines when it is determined that the parser fails to parse the first code, so as to segment the first code into multiple code fragments, and the number of lines of each code fragment is less than or equal to the preset number of lines.
[0265] In an exemplary embodiment, vector conversion module 1701 is further configured to remove comments from multiple code snippets. Furthermore, it may be configured to remove leading and trailing whitespace from each line of code within the multiple code snippets. Furthermore, it may be configured to replace multiple consecutive whitespace characters within the multiple code snippets with a single whitespace character. Furthermore, it may be configured to replace local variable names within the multiple code snippets with identifiers in a preset format. Furthermore, it may be configured to unify the capitalization of letters within the multiple code snippets. Furthermore, after normalizing the code snippets, an association may be established between the code snippets before normalization and the code snippets after normalization.
[0266] In an exemplary embodiment, extraction module 1703 is further configured to parse at least one candidate code vector to determine metadata and code content corresponding to the at least one candidate code vector. Based on the metadata and code content, the code fragment pointed to by the at least one candidate code vector is located in the artificial intelligence code library. Based on the code fragment pointed to by the at least one candidate code vector, at least one second code pointed to by the at least one candidate code vector is constructed.
[0267] In an exemplary embodiment, the apparatus further comprises:
[0268] A classification module is configured to classify at least one second code according to a preset similarity threshold, wherein a second code whose similarity to the first code is greater than the first threshold belongs to the first category, a second code whose similarity to the first code is less than or equal to the first threshold and greater than or equal to the second threshold belongs to the second category, and a second code whose similarity to the first code is less than the second threshold belongs to the third category.
[0269] The adoption determination module is configured to determine that the first code adopts the at least one second code when the at least one second code belongs to the first category.
[0270] The non-adoption determination module is configured to determine that the first code does not adopt the at least one second code when the at least one second code belongs to the third category.
[0271] In an exemplary embodiment, adoption determination module 1704 is further configured to, when at least one second code belongs to the second category, input the second code belonging to the second category into the large language model, use structured prompt words to instruct the large language model to perform semantic understanding of the first code and the second code, and obtain an output result of the large language model. If the output result of the large language model is that the first code adopts the at least one second code, it is determined that the first code adopts the at least one second code.
[0272] In an exemplary embodiment, the apparatus further comprises:
[0273] The data collection module is used to collect interaction data between the target object and the large language model, wherein the interaction data includes the artificial intelligence code generated by the large language model and the contextual content related to the artificial intelligence code.
[0274] The data parsing module is used to parse the interaction data to separate the artificial intelligence code and natural language description in the interaction data.
[0275] The relationship establishment module is used to establish an association relationship between the artificial intelligence code and the metadata corresponding to the interaction data, wherein the metadata corresponding to the interaction data includes at least the identification information of the target object, timestamp, dialogue turn, and question text.
[0276] The combination module is used to combine the metadata corresponding to the artificial intelligence code and the interaction data into preset structured data.
[0277] The first building block is used to build an artificial intelligence code base based on structured data.
[0278] In one exemplary embodiment, the data parsing module is further configured to parse the interaction data to identify multiple rounds of conversation between the target subject and the large language model. Information extraction is performed on the text content of the multiple rounds of conversation to determine metadata within the interaction data. Code recognition is performed on the text content of the multiple rounds of conversation to extract the artificial intelligence code generated by the large language model within the interaction data.
[0279] In an exemplary embodiment, the apparatus further comprises:
[0280] The configuration module is used to configure access credentials and access paths of the target code repository, wherein the target code repository stores all codes managed by the version control system of the target server.
[0281] The access module is used to respond to an analysis request of a target object, access a target code repository along an access path using access credentials, and call the code submission records of the version control system within a target time range, wherein the analysis request carries the target time range.
[0282] The range search module is used to determine multiple codes submitted by the target object within the target time range based on the code submission records of the version control system within the target time range.
[0283] The code determination module is used to determine a new code among the multiple codes whose similarity with the code stored in the target code repository is lower than a preset level.
[0284] The second construction module is used to construct an incremental code library based on the new code and metadata corresponding to the new code, wherein the first code is obtained from the incremental code library.
[0285] In an exemplary embodiment, the code determination module is further configured to analyze the plurality of codes to determine newly added lines of code in the plurality of codes compared to the code already stored in the target code repository. The newly added lines of code are then treated as new code. The plurality of codes are analyzed to determine the degree of modification of each of the plurality of codes compared to the code already stored in the target code repository. The code in the plurality of codes with a modification degree greater than a preset degree is treated as new code.
[0286] In an exemplary embodiment, the apparatus further comprises:
[0287] The similar code determination module is configured to determine a third code in the at least one second code that has the highest similarity to the first code when it is determined that the first code has an adoption relationship with the at least one second code.
[0288] The adoption amount determination module is used to determine the adoption amount of the first code for the third code generated by artificial intelligence based on the similarity between the first code and the third code.
[0289] The total adoption amount determination module is used to count the total amount of new code in the incremental code base within a preset analysis cycle and the total amount of new code adopted for the code generated by artificial intelligence.
[0290] The adoption rate determination module is used to determine the adoption rate of the new code to the code generated by artificial intelligence based on the total amount of new code and the total amount of adoption of the new code to the code generated by artificial intelligence.
[0291] In an exemplary embodiment, the apparatus further comprises:
[0292] The adopted code determination module is configured to determine that the first code is an adopted code when it is determined that the first code has an adopted relationship with the at least one second code based on the adoption status of the first code for the at least one second code.
[0293] The first traceability module is used to query the metadata of the first code as the adopted code in the incremental code library to obtain traceability information of the first code as the adopted code, wherein the metadata of the first code includes the submission time of generating the first code and the identification information of the submitter.
[0294] The first tracing module is used to query the metadata of the second code adopted by the adopted code in the artificial intelligence code library to obtain the tracing information of the first code as the adopted code, wherein the metadata of the second code includes the dialogue text and generation time of generating the second code.
[0295] For the description of the features in the embodiment corresponding to the device for determining the code adoption situation, please refer to the relevant description of the embodiment corresponding to the method for determining the code adoption situation, and no further details will be given here.
[0296] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for determining code adoption status.
[0297] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned code adoption determination method embodiments when running.
[0298] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0299] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned methods for determining code adoption status are implemented.
[0300] An embodiment of the present application further provides another computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned methods for determining code adoption are implemented.
[0301] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0302] The above is a detailed introduction to the method, device, electronic device, computer-readable storage medium, and computer program product for determining the adoption of a code provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for determining code adoption, characterized in that: The method comprises: When a first code submitted by a target object is obtained, extracting semantic information from the first code, and generating a first code vector in a preset format corresponding to the first code based on the semantic information; Searching a pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a requirement for similarity to the first code vector, wherein the vector database includes code vectors corresponding to respective codes in a preset artificial intelligence code library; Based on the at least one candidate code vector, extracting at least one second code pointed to by the at least one candidate code vector in the artificial intelligence code library; The first code and at least one second code are input into a large language model, and the adoption of the first code for the at least one second code is determined based on an output result of the large language model, wherein the large language model is used to perform semantic understanding of the code to determine the similarity of the code.
2. The method for determining code adoption according to claim 1, wherein: The searching a pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a standard similarity with the first code vector includes: constructing a first vector index corresponding to the first code vector according to the content of the first code vector and metadata corresponding to the first code; Searching the vector database according to the first vector index to determine a plurality of code vectors that meet a matching degree requirement with the first vector index, wherein the matching degrees of the plurality of code vectors with the first vector index indicate similarities between the plurality of code vectors and the first code vector; A code vector among the multiple code vectors whose similarity to the first code vector is greater than a first similarity threshold is used as the candidate code vector.
3. The method for determining code adoption according to claim 1, wherein: The method of extracting semantic information from the first code submitted by the target object and generating a first code vector in a preset format corresponding to the first code based on the semantic information includes: performing segmentation processing on the first code to segment the first code into multiple code segments; Performing semantic parsing on the multiple code snippets respectively using a preset semantic parsing model to generate semantic vectors of preset dimensions corresponding to the multiple code snippets, wherein the semantic parsing model is used to capture the grammatical structure in the code snippets and parse the semantic information of the code snippets based on the grammatical structure; The first code vector is generated by combining semantic vectors of preset dimensions corresponding to the multiple code fragments.
4. The method for determining code adoption according to claim 3, wherein: The segmenting of the first code into a plurality of code segments includes: determining a programming language used by the first code according to feature information of the first code; Splitting the first code using a corresponding segmentation strategy according to the programming language used by the first code to generate multiple code fragments; Normalization is performed on the multiple code snippets to eliminate non-semantic information in the multiple code snippets and convert the multiple code snippets into a preset standard format.
5. The method for determining code adoption according to claim 4, wherein: The step of segmenting the first code using a corresponding segmentation strategy according to the programming language used by the first code to generate a plurality of code fragments includes: Loading a corresponding parser according to the programming language used by the first code; Parsing the first code using the parser to generate an abstract syntax tree corresponding to the first code; Determining multiple segments of semantic logic scope of the first code based on the logical structure of the first code indicated by the abstract syntax tree; Segmenting the first code based on the semantic logical scope to divide the first code into multiple code segments, wherein each of the code segments includes a complete semantic logical scope; Based on the metadata of each code snippet, an association relationship between each code snippet and the first code is established, wherein the association relationship includes location information of the code snippet in the first code.
6. The method for determining code adoption according to claim 5, wherein: After parsing the first code using the parser to generate an abstract syntax tree corresponding to the first code, the method further includes: If it is determined that the parser fails to parse the first code, determining a position of a blank line in the first code; Using blank line positions in the first code as segmentation positions to segment the first code into the multiple code segments; Alternatively, when it is determined that the parser fails to parse the first code, the first code is segmented in a sliding window manner according to a preset number of lines to segment the first code into multiple code fragments, and the number of lines of each code fragment is less than or equal to the preset number of lines.
7. The method for determining code adoption according to claim 3, wherein: The normalizing of the plurality of code snippets includes: Removing comment information from the plurality of code snippets; and / or, removing leading and trailing whitespace characters from each line of code in the plurality of code snippets; and / or, replacing multiple consecutive space characters in the multiple code snippets with a single space character; and / or, replacing local variable names in the plurality of code snippets with identifiers in a preset format; and / or, unifying the uppercase and lowercase letters in the plurality of code snippets; After the code snippet is normalized, an association relationship is established between the code snippet before normalization and the code snippet after normalization.
8. The method for determining code adoption according to any one of claims 1 to 7, characterized in that: The extracting, from the artificial intelligence code library, at least one second code pointed to by the at least one candidate code vector based on the at least one candidate code vector includes: parsing the at least one candidate code vector to determine metadata and code content corresponding to the at least one candidate code vector; Locating, in the artificial intelligence code library, a code snippet pointed to by the at least one candidate code vector based on the metadata and the code content; At least one second code pointed to by the at least one candidate code vector is constructed according to the code segment pointed to by the at least one candidate code vector.
9. The method for determining code adoption according to any one of claims 1 to 7, characterized in that: Before inputting the first code and the at least one second code into the large language model and determining, based on an output result of the large language model, whether the first code is adopted for the at least one second code, the method further includes: Classifying the at least one second code according to a preset similarity threshold, wherein a second code having a similarity with the first code greater than a first threshold belongs to a first category, a second code having a similarity with the first code less than or equal to the first threshold and greater than or equal to a second threshold belongs to a second category, and a second code having a similarity with the first code less than the second threshold belongs to a third category; In a case where the at least one second code belongs to the first category, determining that the first code adopts the at least one second code; In a case where the at least one second code belongs to the third category, it is determined that the first code does not adopt the at least one second code.
10. The method for determining code adoption according to claim 9, wherein: Inputting the first code and the at least one second code into a large language model, and determining, based on an output result of the large language model, an adoption status of the first code for the at least one second code, includes: When the at least one second code belongs to the second category, inputting the second code belonging to the second category into the large language model, using structured prompt words to instruct the large language model to perform semantic understanding on the first code and the second code, and obtaining an output result of the large language model; When the output result of the large language model is that the first code adopts the at least one second code, it is determined that the first code adopts the at least one second code.
11. The method for determining code adoption according to any one of claims 1 to 7, characterized in that: The method further comprises: Collecting interaction data between the target object and the large language model, wherein the interaction data includes artificial intelligence code generated by the large language model and contextual content related to the artificial intelligence code; Parsing the interaction data to separate artificial intelligence code and natural language description in the interaction data; Establishing an association relationship between the artificial intelligence code and metadata corresponding to the interaction data, wherein the metadata corresponding to the interaction data includes at least identification information of the target object, a timestamp, a conversation turn, and a question text; Combining the artificial intelligence code and metadata corresponding to the interaction data into preset structured data; The artificial intelligence code library is constructed based on the structured data.
12. The method for determining code adoption according to claim 11, wherein: The parsing of the interaction data to separate the artificial intelligence code and the natural language description in the interaction data includes: Parsing the interaction data to determine multiple rounds of dialogue between the target object and the large language model; Extracting information from text content in the multiple rounds of conversations to determine metadata in the interaction data; Code recognition is performed on the text content in the multiple rounds of conversations to extract the artificial intelligence code generated by the large language model in the interaction data.
13. The method for determining code adoption according to any one of claims 1 to 7, characterized in that: The method further comprises: Configure access credentials and access paths for a target code repository, where the target code repository stores all code managed by the target server's version control system; In response to an analysis request for the target object, access the target code repository along the access path using the access credential, and call commit records of code in the version control system within a target time range, wherein the analysis request carries the target time range; Determining, based on the code submission records of the version control system within the target time range, multiple codes submitted by the target object within the target time range; Determining a new code among the plurality of codes whose similarity to the codes stored in the target code repository is lower than a preset level; An incremental code library is constructed based on the new code and metadata corresponding to the new code, wherein the first code is obtained from the incremental code library.
14. The method for determining code adoption according to claim 13, wherein: The determining of a new code among the plurality of codes whose similarity to the code stored in the target code repository is lower than a preset level includes: Analyzing the plurality of codes to determine new code lines in the plurality of codes compared to codes already stored in the target code repository; Using the newly added code line as the new code; Analyzing the plurality of codes to determine a modification degree of each of the plurality of codes compared to a code stored in the target code repository; The code whose modification degree is greater than a preset degree among the multiple codes is used as the new code.
15. The method for determining code adoption according to claim 13, wherein: The method further comprises: If it is determined that the first code has an adoption relationship with the at least one second code, determining a third code in the at least one second code that has the highest similarity to the first code; determining, based on a similarity between the first code and the third code, an adoption rate of the first code for the third code generated by artificial intelligence; Counting the total amount of new code in the incremental code base within a preset analysis period and the total amount of adoption of the new code for the AI-generated code; An adoption rate of the new code for the artificial intelligence-generated code is determined based on the total amount of the new code and the total amount of adoption of the new code for the artificial intelligence-generated code.
16. The method for determining code adoption according to claim 13, wherein: The method further comprises: If it is determined that the first code has an adoption relationship with the at least one second code according to the adoption status of the first code for the at least one second code, determining the first code as an adopted code; Querying metadata of the first code as the adopted code in the incremental code repository to obtain traceability information of the first code as the adopted code, wherein the metadata of the first code includes a submission time when the first code was generated and identification information of a submitter; The metadata of the second code adopted by the adopted code is queried in the artificial intelligence code library to obtain traceability information of the first code as the adopted code, wherein the metadata of the second code includes the dialogue text and generation time of generating the second code.
17. A device for determining code adoption, characterized in that: include: a vector conversion module, configured to, upon obtaining a first code submitted by a target object, extract semantic information from the first code and generate a first code vector in a preset format corresponding to the first code based on the semantic information; a retrieval module, configured to search a pre-built vector database based on the first code vector to determine at least one candidate code vector in the vector database that meets a requirement for similarity to the first code vector, wherein the vector database includes code vectors corresponding to respective codes in a preset artificial intelligence code library; an extraction module, configured to extract, from the artificial intelligence code library, at least one second code pointed to by the at least one candidate code vector based on the at least one candidate code vector; An adoption determination module is used to input the first code and at least one second code into a large language model, and determine the adoption of the first code for the at least one second code based on the output result of the large language model, wherein the large language model is used to perform semantic understanding of the code to determine the similarity of the code.
18. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 16 when executing the computer program.
19. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
A semantic similarity-based Java application program interface use mode recommendation method
CN109670022A
Large language model application service method and device
CN117370523A
Method, system and equipment for automatically generating user interface code and medium
CN119883252A
Code generation method and related equipment
CN119960823A
Document plagiarism judgment method and system based on semantic vector library and large language model
CN120234420A
Cited By
Code adoption rate determination method and device, medium, electronic equipment and product
CN121050713A