Method and system for discriminating AI code generation based on Levenshtein distance
Through the discriminating method based on Levenshtein distance, by masking and completing the code segment, generating perturbation samples, and calculating the average Levenshtein distance, the problem of low code detection accuracy in the existing technology is solved, and more efficient and accurate AI-generated code detection is achieved.
Patent Information
- Application Number
- CN202510021927.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to accurately identify whether code is generated by AI, especially in diverse programming languages, with low detection accuracy and unstable performance in different scenarios.
The AI-based code generation method based on Levenshtein distance is used to determine whether the code is generated by fixed masking and code completion of the original code segment, multiple perturbation samples are generated, and the average Levenshtein distance between the original code segment and the perturbation sample is calculated, and the threshold value is determined whether the code is AI-generated.
It improves the effectiveness and accuracy of code detection, reduces computing costs, is suitable for real-time or resource-limited scenarios, and can effectively detect AI-generated code in the absence of data.
Smart Images

Figure CN120123205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of code generation and detection of Artificial Intelligence Generated Content (AIGC), and mainly relates to code detection in AIGC to determine whether a piece of code is generated by a large language model. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, especially in the fields of natural language processing and code generation, AI large models (such as GPT series of OpenAI and BERT of Google, etc.) have made remarkable progress in programming assistance. The popularization of some tools such as GitHub Copilot has made AI-generated code a norm, which has also triggered in-depth discussions on the quality and source of the generated code. Identifying whether a piece of code is generated by AI or written by humans not only involves the correctness of the syntax and logic of the code, but also relates to its maintainability and security. AI-generated code performs excellently in specific situations, but in complex or specific scenarios, it may hide security vulnerabilities or performance issues, thus affecting the overall quality of the software system.
[0003] With the wide application of AI in the field of programming, intellectual property and ethical issues have gradually become the focus of attention. How to ensure the originality and traceability of code has become an important issue in software development and academic research. Moreover, in the field of programming education, the discussion of this issue is equally important. With the wide application of AI tools, students and beginners may rely more on these tools when learning programming, resulting in deficiencies in their code understanding and logical thinking abilities. Therefore, the ability to identify the source of code not only concerns the quality review of code, but also directly affects the learning effect and professional quality of students. However, there is currently a lack of effective methods and tools to accurately distinguish between human-written and AI-generated code. Therefore, carrying out relevant research not only has important academic value, but also provides a practical demand orientation for software development practice. In-depth research on this topic will provide a new perspective and solution for improving the effectiveness of code review and ensuring the security of software systems.
[0004] With the popularization of generative AI in the fields of programming and education, the need to detect AI-generated code is becoming more and more urgent. Current code detection methods can be roughly divided into three categories: training-based detection, watermark embedding technology, and zero-shot detection.
[0005] First, there is detection based on training, such as GPTZero. Such methods were originally designed to identify AI-generated natural language content (such as articles, comments, etc.). However, since these models are mainly optimized for natural language, they often have difficulty effectively identifying the syntax and structural features of code when dealing with code, resulting in low detection accuracy. Similarly, although OpenAI's AI text classifier can distinguish whether text is generated by AI, when faced with code such as Python, JavaScript, C++, etc., the detection effect will also be affected due to the lack of sample support for programming languages. These limitations indicate that existing text detection tools need to be improved in adapting to diverse programming languages.
[0006] Watermark embedding technology is another type of method that has been widely applied in the fields of image and text detection. However, watermark embedding in code may affect the quality and readability of the code. Although the SWEET method designed for the low-entropy characteristics of code (compared to natural language) attempts to embed watermarks only on high-entropy tokens to avoid affecting the quality of low-entropy tokens, this method requires access to the original model for generating watermarks or its simplified version, thus encountering certain difficulties in practical applications, especially when the source of the code is unclear, and the reliability is limited.
[0007] Finally, there are zero-shot detection methods, such as DetectGPT4Code. Such methods are based on the improvement of the text detection tool DetectGPT and identify AI-generated code by estimating the generation probability of the end token of a code snippet. Its greatest advantage is that it does not require a large number of training samples, so it can effectively detect in the case of data scarcity. However, DetectGPT4Code relies on a proxy model for probability estimation, which may lead to the detection accuracy being inferior to directly using the original model, and its performance may be unstable in different code types or scenarios.
[0008] Generally speaking, the detection of AI-generated code still faces many challenges, including the poor adaptability of existing tools to programming language features, the limitations of various detection methods in code detection, and the differences in the robustness of detection effects in different scenarios. This indicates that optimizing existing detection models and developing detection methods more suitable for code has important research significance. Through systematic analysis and innovative design, improving the detection accuracy and adaptability of AI-generated code will provide stronger technical support for programming education and software development. Summary of the Invention
[0009] Object of the Invention: Aiming at the problems existing in the above-mentioned background technology, the object of the present invention is to propose a method and system for discriminating AI-generated code based on the Levenshtein distance, improving the effectiveness of detection, and reducing the time and hardware resources required for detection.
[0010] Technical solution: To achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] In the first aspect, the present invention provides a method for discriminating AI-generated code based on the Levenshtein distance, including the following steps:
[0012] Perform fixed masking on the original code segment line by line, and then use at least one code completion model to complete the masked part, and each code completion model obtains multiple perturbed code segment samples;
[0013] Calculate the Levenshtein distance between the original code segment and each perturbed sample respectively and take the average to obtain the average Levenshtein distance;
[0014] Compare the obtained Levenshtein distance with a set threshold. If it is greater than the threshold, it is considered that the original code segment belongs to machine-generated code; otherwise, it is considered that the original code segment belongs to manually written code.
[0015] Preferably, first calculate the total number of lines of the original code segment, and then completely mask the last X% part of the original code segment, where X ∈ [50, 90], and the specific value is the optimal result tested in the experimental part. After masking, perform completion.
[0016] Preferably, for the masked part in the original code segment, use at least one code completion model to complete the masked part. The code completion model will obtain a masked code segment and a prompt "Please complete the maskpart according to the existing part of the code." When completing, only require filling it completely, without considering the correctness of the code. Require the code completion model to output Y samples after completing the masked part according to the same input, where Y ∈ [2, 20], and the specific value is the optimal result tested in the experimental part.
[0017] Preferably, the preset threshold is a value determined through experiments, and this value can be adjusted according to different code sample types and code completion models.
[0018] Preferably, completely mask the last 70% part of the original code segment, and obtain 13 perturbed samples after completing the masked part through one code completion model.
[0019] Furthermore, before using the Levenshtein distance as a similarity metric, it further includes:
[0020] Calculate the statistical difference in the Levenshtein distance between positive and negative class samples;
[0021] The Shapiro-Wilk test is used to perform a normality test on the Levenshtein distances of positive and negative class samples to determine whether the data conforms to a normal distribution;
[0022] If the result of the normality test shows that the data conforms to a normal distribution, an independent samples t-test is used to perform a significance test on the difference in Levenshtein distances between the two groups of samples;
[0023] If the result of the normality test shows that the data does not conform to a normal distribution, the Mann-Whitney U test is used for the significance test;
[0024] Calculate Cohen's d value, which is used to quantify the effect size of the Levenshtein distance between the two groups of samples, so as to measure the degree of difference between the machine-generated code and the manually written code.
[0025] In a second aspect, the present invention provides a system for discriminating AI-generated code based on the Levenshtein distance, including:
[0026] A perturbation module, which is used to perform fixed masking on the original code segment in units of lines, and then use at least one code completion model to complete the masked part, and each code completion model obtains multiple perturbed code segment samples;
[0027] A calculation module, which is used to calculate the Levenshtein distance between the original code segment and each perturbed sample respectively and take the average to obtain the average Levenshtein distance;
[0028] A discrimination module, which is used to compare the obtained Levenshtein distance with a set threshold. If it is greater than the threshold, it is considered that the original code segment belongs to machine-generated code, otherwise, it is considered that the original code segment belongs to manually written code.
[0029] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program / instructions stored on the memory and executable on the processor. When the computer program / instructions are executed by the processor, the steps of the method for discriminating AI-generated code based on the Levenshtein distance are implemented.
[0030] In a fourth aspect, the present invention provides a computer program product, including computer program / instructions. When the computer program / instructions are executed by the processor, the steps of the method for discriminating AI-generated code based on the Levenshtein distance are implemented.
[0031] Beneficial effects: The method for discriminating AI-generated code based on the Levenshtein distance provided by the present invention is a new zero-shot code detection scheme, which can effectively detect even in the case of lack of data. The present invention obtains multiple perturbed samples of the code to be detected through masking and complementing operations, sets a threshold based on the average Levenshtein distance, and realizes the discrimination between machine-generated code and manually written code based on the comparison result of the threshold. Since the Levenshtein distance is based on a direct string comparison algorithm and does not require external models to assist in the calculation, its computational cost is significantly reduced, making it more suitable for real-time or resource-limited scenarios. The present invention can be applied to fields such as automated review of programming code, code copyright identification, and identification of AI-generated code. Brief Description of the Drawings
[0032] Figure 1 It is a flowchart of the method for discriminating AI-generated code based on the Levenshtein distance provided by an embodiment of the present invention. Detailed Embodiments
[0033] The present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] As Figure 1 shown, an embodiment of the present invention provides a method for discriminating AI-generated code based on the Levenshtein distance, which mainly includes a perturbation step, a calculation of difference step, and a discrimination step. The specific implementation of each step is as follows:
[0035] S1. Perturbation step: Fixedly mask the original code segment in units of lines, and then use at least one code completion model to complete the masked part, and each code completion model obtains multiple perturbed code segment samples.
[0036] The purpose of the perturbation step is to interfere with the code segment to be determined, so as to generate multiple perturbed samples for subsequent Levenshtein distance calculation. A fixed masking operation is performed on the code segment to be determined. Specifically, the last X percent of the code segment is masked, where X is a variable parameter, and the optimal value can be determined through experiments. Since the masking ratio X value has a greater impact on the final discrimination result, the optimal value of X should be determined through experiments. In the experiment, by changing the value of X, the changes in discrimination accuracy and Levenshtein distance can be observed to find the masking ratio most suitable for this task. Generally speaking, the X value needs to balance the intensity of perturbation and the consumption of computing resources, and select an X value that can generate effective perturbations without overly affecting the detection result. In the specific experiment process, the optimal X value can be selected within the range of [50, 90]. In this embodiment, the last 70% of the original code segment is completely masked. The masked code part does not need to be a complete function or statement and can be a part of any structure to increase the diversity of perturbed samples. The masked code segment will be fed into a code completion model (such as ChatGPT, Codeparrot, etc.), and the model will complete the masked part to generate multiple perturbed samples. The code completion model will automatically predict and complete the missing part based on the masked code segment. The completed code segment is not required to be able to compile or pass static analysis because the purpose of the present invention is to evaluate the difference between AI-generated code and manually written code through the change of Levenshtein distance, rather than judging whether the function of the code is correct. Therefore, the completed content mainly focuses on syntactic rationality, structural consistency, and code context fluency, but it is not required that the code after completion must be executable. The task of the completion model is to try to fill in the missing part to make the generated code segment as "reasonable" as possible syntactically, although these codes may not meet the actual compilation requirements.
[0037] Exemplarily, the specific implementation method of this step is as follows:
[0038] 1. Select the code segment: First, select a code segment a from the code library to be determined and process it. Assume the length of the original code segment a is m, and the code segment is split into several lines by line, denoted as a 1 …a m 。
[0039] 2. Fix the last X percent for masking: Fix the last X percent of the code segment for masking. Specifically, select the last X% lines of the code segment for masking. The specific masking operation can be achieved by using a specific symbol or placeholder (such as " <mask>”) is achieved by replacing the masked part.
[0040] 3. Code Completion: Input the masked code segment into a code completion model (such as ChatGPT, etc.), and use the model to automatically generate the completed code for the masked part. Specifically, the code completion model will receive a masked code segment and a prompt "Please complete the mask part according to the existing part of the code.” When completing, only require filling it completely, without considering the correctness of the code. Require the code completion model to output Y samples after completing the masked part according to the same input. In the specific experimental process, the value range of Y can be [2, 20], and the specific value is the optimal result tested in the experimental part. The completion process does not require the completed part to be successfully compiled, as long as the completed code conforms to the context in terms of syntax and structure.
[0041] 4. Generate Perturbed Samples: Generate multiple completed versions through one or more completion models to obtain several perturbed samples b…b k , and each sample represents a different code completion result.
[0042] S2. Calculation of Difference Step: Calculate the Levenshtein distance between the original code segment and each perturbed sample respectively, and take the average to obtain the average Levenshtein distance.
[0043] The purpose of the calculation of difference step is to measure the difference between the original code segment a and multiple perturbed samples. After obtaining several perturbed samples, calculate the Levenshtein distance between the original code segment and each perturbed sample respectively, and calculate the average of these distances. In the present invention, the core of the calculation step is to use the Levenshtein distance to measure the difference between the original code segment and each perturbed sample. The Levenshtein distance is used to calculate the minimum number of operations required to convert one string to another, and these operations include insertion, deletion, and replacement. By calculating the Levenshtein distance between the original code segment and multiple perturbed samples, the structural and syntactic differences between them can be quantified.
[0044] The calculation of the Levenshtein distance adopts the method of dynamic programming. Suppose there are two strings, a = a 1 a 2 …a i …a m and b = b 1 b 2 …b i …b m , where m and n are the lengths of strings a and b respectively. Define the matrix lev a,b (i, j) to represent the minimum Levenshtein distance from string a 1 …a i to string b 1 …b j .
[0045] Its recurrence formula is as follows:
[0046]
[0047] is the indicator function. When the character a i is different from the character b j , the value is 1, representing the cost of the substitution operation; the costs of insertion and deletion operations are both 1.
[0048] Exemplarily, the specific implementation method of this step is as follows;
[0049] 1. Initialize the Levenshtein distance matrix: For each pair of the original code segment a and the perturbed sample b k (k = 1, 2, … K), use the Levenshtein distance algorithm to calculate the edit distance between them. The Levenshtein distance is a metric for measuring the difference between two strings, which converts one string into another through insertion, deletion, and substitution operations.
[0050] 2. Calculation of the Levenshtein distance: The calculation process is carried out through a dynamic programming algorithm. The specific steps are as follows:
[0051] Create a two-dimensional matrix D with size (m + 1) × (n + 1), where m is the length of the original code segment a and n is the length of the perturbed sample b k . Then initialize the first row and the first column of the matrix to represent the number of operations required to convert an empty string into other strings. Finally, fill the matrix according to the recurrence formula, and finally obtain the minimum Levenshtein distance D(a, b k ) between the original code segment and the perturbed sample.
[0052] 3. Calculate the average Levenshtein distance: For all perturbed samples, calculate the Levenshtein distance between the original code segment a and each perturbed sample bbb, and obtain the distance of each perturbed sample. Finally, calculate the average Levenshtein distance between all perturbed samples and the original code segment:
[0053]
[0054] Among them, K is the number of perturbed samples.
[0055] The Levenshtein distance provides a very precise way to quantify the differences between code segments. By calculating the differences between the original code and the perturbed samples, this method reduces the dependence on large-scale hardware resources and greatly improves the detection efficiency.
[0056] S3. Discrimination step: Compare the obtained Levenshtein distance with the set threshold. If it is greater than the threshold, it is considered that the original code segment belongs to machine-generated code; otherwise, it is considered that the original code segment belongs to manually written code.
[0057] The purpose of the discrimination step is to determine whether the original code segment is generated by AI or manually written based on the calculated Levenshtein distance. After obtaining the average Levenshtein distance between the original code segment and multiple perturbed samples, next, the determination step decides whether the original code segment is generated by AI or manually written by comparing this average Levenshtein distance with a preset threshold. The threshold is obtained through a large number of experiments and plays a key role in distinguishing between AI-generated code and manually written code. Through the experimental process, we observed that under specific perturbation conditions and masking ratios, the differences in the Levenshtein distance between manually written code and machine-generated code show a certain regularity.
[0058] Exemplarily, the specific implementation method of this step is as follows:
[0059] 1. Threshold setting: Through a large number of experiments, determine an optimal threshold T, which is used to distinguish between AI-generated code and manually written code. The experiment adjusts the threshold, analyzes the differences between the Levenshtein distance and manually written code and machine-generated code, and finally selects the threshold that can optimize the discrimination accuracy.
[0060] 2. Determine the type of the original code segment: Compare the calculated average Levenshtein distance with the threshold T:
[0061] If the average Levenshtein distance is greater than the threshold T: Determine that the code segment is AI-generated code. At this time, the difference between the original code segment and the perturbed samples is large, indicating that the original code segment is more likely to be generated by AI.
[0062] If the average Levenshtein distance is less than or equal to the threshold T: Determine that the code segment is manually written code. At this time, the difference between the original code segment and the perturbed samples is small, indicating that the original code segment is more likely to be manually written.
[0063] 3. Discrimination result output: Finally, output the discrimination result, indicating whether the original code segment is written manually or generated by AI.
[0064] Before using the Levenshtein distance as the similarity metric in the embodiments of the present invention, it further includes:
[0065] Step 1: Calculate the statistical difference in the Levenshtein distance between positive and negative samples (i.e., machine-generated code and manually-written code).
[0066] Step 2: Use the Shapiro-Wilk test to perform a normality test on the Levenshtein distance of positive and negative samples to determine whether the data conforms to a normal distribution.
[0067] Step 3: If the normality test result shows that the data conforms to a normal distribution, use an independent samples t-test to perform a significance test on the difference in the Levenshtein distance between the two groups of samples.
[0068] Step 4: If the normality test result shows that the data does not conform to a normal distribution, use the Mann-Whitney U test for significance testing.
[0069] Step 5: Calculate Cohen's d value, which is used to quantify the effect size of the Levenshtein distance between the two groups of samples, thereby measuring the degree of difference between machine-generated code and manually-written code.
[0070] The Levenshtein distance selected by the present invention has the following advantages:
[0071] The Levenshtein distance can effectively handle strings of different lengths and calculate the differences through insertion, deletion, and replacement operations. In contrast, other similarity metrics such as the Hamming distance can only handle strings of equal length, and the cosine similarity and Jaccard distance are less sensitive to the order of characters and cannot accurately describe the code editing behavior.
[0072] In the scenario of code comparison, the order of characters or syntax units is usually closely related to semantics. The Levenshtein distance can reflect the differences caused by order changes, while the cosine similarity and Jaccard distance are insensitive to order changes.
[0073] The Levenshtein distance allows adjusting the costs of insertion, deletion, and replacement operations according to specific needs, thus maintaining a high degree of adaptability in different application scenarios. In contrast, the calculation methods of other metrics are usually more fixed and difficult to adapt to diverse requirements.
[0074] Based on the same inventive concept, an embodiment of the present invention also discloses a system for discriminating AI-generated code based on the Levenshtein distance, including: a perturbation module, configured to perform fixed masking on the original code segment in units of lines, and then use at least one code completion model to complete the masked part, and each code completion model obtains a plurality of perturbed code segment samples; a calculation module, configured to calculate the Levenshtein distance between the original code segment and each perturbed sample respectively and take the average to obtain the average Levenshtein distance; a discrimination module, configured to compare the obtained Levenshtein distance with a set threshold, and if it is greater than the threshold, it is considered that the original code segment belongs to machine-generated code, otherwise, it is considered that the original code segment belongs to manually written code.
[0075] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program / instructions stored on the memory and executable on the processor. When the computer program / instructions are executed by the processor, the steps of the method for discriminating AI-generated code based on the Levenshtein distance are implemented.
[0076] An embodiment of the present invention also discloses a computer program product, including computer program / instructions. When the computer program / instructions are executed by the processor, the steps of the method for discriminating AI-generated code based on the Levenshtein distance are implemented.
[0077] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.< / mask>
Claims
1. A method for discriminating AI code generation based on Levenshtein distance, characterized in that: The steps include: The original code segment is fixedly masked in units of lines, and then the masked part is completed with the help of at least one code completion model, and each code completion model obtains multiple perturbed code segment samples; Calculate the Levenshtein distance between the original code segment and each perturbation sample and take the average Levenshtein distance; The obtained Levenshtein distance is compared with the set threshold. If it is greater than the threshold, the original code segment is considered to be machine-generated code. Otherwise, the original code segment is considered to be manually written code.
2. The method for generating codes by using AI based on Levenshtein distance according to claim 1, characterized in that: First, calculate the total number of lines in the original code segment, and then completely cover the last X% of the original code segment, X∈[50,90], where the specific value is the optimal result tested in the experimental part, and then complete it after covering it.
3. The method for generating codes by discriminating AI based on Levenshtein distance according to claim 1, characterized in that: For the masked part in the original code segment, use at least one code completion model to complete the masked part. The code completion model will obtain a masked code segment and a prompt word "Please complete the mask part according to the existing part of the code." When completing, it is only required to complete the code without considering the correctness of the code. The code completion model is required to output Y samples after the mask part is completed according to the same input, Y∈[2,20], and the specific value is the optimal result tested in the experimental part.
4. The method for generating codes by discriminating AI based on Levenshtein distance according to claim 1, characterized in that: The preset threshold is a value determined through experiments, and the value can be adjusted according to different code sample types and code completion models.
5. The method for generating codes by discriminating AI based on Levenshtein distance according to claim 1, characterized in that: The Levenshtein distance formula is: Among them, lev a,b (i,j) represents the Levenshtein distance between the original string and the target string at the i-th character and the j-th character. is an indicator function. When the character a i With the character b j If they are different, the value is 1, indicating the cost of the replace operation; the cost of the insert and delete operations is 1.
6. The method for generating codes by discriminating AI based on Levenshtein distance according to claim 1, characterized in that: The last 70% of the original code segment is completely masked, and 13 perturbation samples after the masked parts are completed are obtained through a code completion model.
7. The method for generating codes by discriminating AI based on Levenshtein distance according to claim 1, characterized in that: Before using Levenshtein distance as a similarity metric, further include: Calculate the statistical difference between positive and negative samples in Levenshtein distance; The Shapiro-Wilk test was used to test the normality of the Levenshtein distance of positive and negative samples to determine whether the data conformed to the normal distribution; If the normality test results show that the data conform to the normal distribution, the independent sample t test is used to perform a significance test on the difference in Levenshtein distance between the two groups of samples; If the normality test results showed that the data did not conform to the normal distribution, the Mann-Whitney U test was used for significance test; Cohen's d value was calculated to quantify the effect size of the Levenshtein distance between the two groups of samples, thus measuring the degree of difference between the machine-generated code and the human-written code.
8. A system for discriminating AI generated codes based on Levenshtein distance, characterized in that: include: A perturbation module is used to perform fixed masking of the original code segment in units of lines, and then complete the masked part with the help of at least one code completion model, and each code completion model obtains multiple perturbed code segment samples; A calculation module, used for respectively calculating the Levenshtein distance between the original code segment and each perturbation sample and taking the average to obtain the average Levenshtein distance; The discrimination module is used to compare the obtained Levenshtein distance with a set threshold. If the distance is greater than the threshold, the original code segment is considered to be a machine-generated code. Otherwise, the original code segment is considered to be a manually written code.
9. A computer system comprising a memory, a processor, and a computer program / instruction stored in the memory and executable on the processor, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method for generating codes based on Levenshtein distance discriminative AI are implemented according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method for generating codes based on Levenshtein distance discriminative AI are implemented according to any one of claims 1 to 7.