Code cloning detection method and device, equipment, medium and product

By using detection methods based on large language models in code cloning detection, obtaining the hash value of code snippets and performing semantic analysis, the problem of insufficient detection accuracy of traditional methods in complex scenarios is solved, and higher code cloning detection accuracy and lower model development costs are achieved.

CN119938135AActive Publication Date: 2025-05-06BEIJING BEIDA SOFTWARE ENG DEV CO LTD

Patent Information

Application Number
CN202510442636.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-06
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Traditional code cloning detection methods perform well when processing structured code or simple cloning, but there are great limitations in the detection of semantic cloning, especially in the scenarios of improving code complexity and cross-language code, the detection accuracy is low.

Method used

By obtaining the hash value of each code snippet in the code base to be detected, a high-similar code snippet pair is determined, and inputting it into a pre-trained code cloning detection model based on a large language model for detection, the ability of the large language model is used for semantic analysis to achieve accurate code functional similarity judgment.

Benefits of technology

It improves the accuracy of code cloning detection and can more effectively identify code snippets with semantic similarity, solves the problem of insufficient detection accuracy of traditional methods in complex scenarios, and reduces the cost of model development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938135A_ABST
    Figure CN119938135A_ABST
Patent Text Reader

Abstract

The invention discloses a code clone detection method and device, equipment, a medium and a product, and relates to the technical field of deep learning, and the method comprises the steps: obtaining a to-be-detected code library; wherein the to-be-detected code library comprises a plurality of to-be-detected code snippets; determining a hash value of each to-be-detected code snippet; based on the hash value of each to-be-detected code snippet, determining a high-similarity code snippet pair set; wherein the high-similarity code snippet pair set comprises at least one group of high-similarity code snippet pairs; the high-similarity code snippet pair set is input into a pre-trained code clone detection model based on a large language model, and a code clone detection result output by the code clone detection model is obtained, the capacity of the large language model can be fully utilized, excessive dependence of the model on code grammar surface features is avoided, and the code clone detection efficiency is improved. And accurate judgment on the code function similarity is realized based on semantic analysis on the code snippets, so that the accuracy of code clone detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to a code clone detection method, device, equipment, medium and product. Background Art

[0002] Code clones refer to code fragments with the same or similar functions but different implementations in a software system. These code clones may be caused by code reuse, repeated development, or other human factors, but too many code clones will lead to code redundancy, increase software maintenance costs, and may hide potential vulnerabilities or defects. Therefore, fast and accurate detection of code clones is of great significance to improving software quality and development efficiency.

[0003] Traditional code clone detection methods usually rely on matching techniques based on text, syntax, or Abstract Syntax Tree (AST). These methods perform well when dealing with structured code or simple clones, but have significant limitations in detecting clones that use different code writing methods to achieve the same function (i.e., semantic clones). With the increase in code complexity and the increase in cross-language code scenarios, traditional methods that rely only on static rules and simple feature extraction have low detection accuracy and cannot meet actual needs. Summary of the invention

[0004] The purpose of this application is to provide a code clone detection method, device, equipment, medium and product, which can improve the accuracy of code clone detection.

[0005] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a code clone detection method, comprising: Obtain a code library to be detected; wherein the code library to be detected contains multiple code snippets to be detected; Determine the hash value of each code snippet to be detected; Based on the hash value of each code snippet to be detected, a set of high-similarity code snippet pairs is determined; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in a group of high-similarity code snippet pairs is greater than a preset similarity; The set of highly similar code snippet pairs is input into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein the code clone detection result includes a clone code pair, and the two clone code snippets contained in the clone code pair implement the same function.

[0006] Optionally, the training method of the code clone detection model based on the large language model is specifically as follows: Build an initial code clone detection model based on a large language model; Using the first training data set, the initial code clone detection model is fine-tuned for multiple summarization enhancements to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions; Using a second training data set to perform multiple mixed summary cloning fine-tuning on the first code clone detection model, to obtain a second code clone detection model capable of determining whether a code clone is a semantic clone; The second code clone detection model is directly optimized multiple times using the third training data set to obtain a trained code clone detection model.

[0007] Optionally, using the first training data set to perform summary enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions specifically includes: Acquire first training data from a first training data set; wherein the first training data includes a first training code segment and a second training code segment; Acquire first standard function description information of the first training code segment, second standard function description information of the second training code segment, and a first real code clone tag; wherein the first real code clone tag indicates whether the first standard function description information and the second standard function description information are actually the same; Inputting the first training code segment, the second training code segment and the detection instruction into the initial code clone detection model, obtaining first function description information of the first training code segment, second function description information of the second training code segment and a first code clone judgment result output by the initial code clone detection model; wherein the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same; Using a loss function, calculating the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first real code clone label to obtain a first loss value; If the first loss value is greater than or equal to a first preset loss threshold, fine-tuning the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein the adjusted initial code clone detection model obtained in any summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation; If the first loss value is less than the first preset loss threshold, the initial code clone detection model in this summary enhancement fine-tuning operation is determined as a first code clone detection model that can analyze code functions and determine whether it is a semantic clone based on the code functions.

[0008] Optionally, the code clone detection method further includes: Construct a second training data set; wherein the second training data set includes second initial training data of two data types, the two data types include pure code type and code-function type; the second initial training data of the pure code type includes a third training code segment and a fourth training code segment; the second initial training data of the code-function type includes a third training code segment, a fourth training code segment, third standard function description information of the third training code segment, and fourth standard function description information of the fourth training code segment.

[0009] Optionally, the using the second training data set to perform a mixed summary clone fine-tuning on the first code clone detection model to obtain a second code clone detection model capable of determining whether a semantic clone is present specifically includes: Determine a second initial training data randomly obtained from the second training data set as the second training data; Acquire a second real code clone tag corresponding to the third training code segment and the fourth training code segment in the second training data; wherein the second real code clone tag indicates whether the function information of the third training code segment and the fourth training code segment is actually the same; Inputting the second training data into the first code clone detection model to obtain a second code clone determination result output by the first code clone detection model; wherein the second code clone determination result indicates whether the function information of the third training code segment and the fourth training code segment is the same; Using the loss function, calculating the second code clone judgment result and the second real code clone label to obtain a second loss value; If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein the adjusted first code clone detection model obtained in any hybrid summary cloning fine-tuning operation is used as the first code clone detection model in the next summary enhancement fine-tuning operation; If the second loss value is less than the second preset loss threshold, the first code clone detection model in this hybrid summary cloning fine-tuning operation is determined as the second code clone detection model that can determine whether semantic cloning occurs.

[0010] Optionally, the using the third training data set to perform a direct preference optimization on the second code clone detection model to obtain a trained code clone detection model specifically includes: Acquire third training data from a third training data set; wherein the third training data includes a fifth training code segment and a sixth training code segment; Acquire a third real code clone tag corresponding to the fifth training code segment and the sixth training code segment; wherein the third real code clone tag indicates whether the function information of the fifth training code segment and the sixth training code segment is actually the same; Inputting the third training data into the second code clone detection model to obtain a third code clone determination result output by the second code clone detection model; wherein the third code clone determination result indicates whether the function information of the fifth training code segment and the sixth training code segment is the same; Comparing the third code clone determination result with the third real code clone label to obtain a comparison result; wherein the comparison result indicates whether the third code clone determination result is the same as the third real code clone label; Using a target loss function, calculating the third code clone judgment result, the third real code clone label, and the comparison result to obtain a third loss value; If the third loss value is greater than or equal to the third preset loss threshold, the second code clone detection model is directly preferred to be optimized to obtain an optimized second code clone detection model; wherein the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization; If the third loss value is less than the third preset loss threshold, the second code clone detection model in this direct preference optimization is determined as the trained code clone detection model.

[0011] In a second aspect, the present application provides a code clone detection device, comprising: An acquisition unit is used to acquire a code library to be detected; wherein the code library to be detected contains multiple code fragments to be detected; A first determining unit, used to determine a hash value of each code fragment to be detected; A second determination unit is used to determine a set of high-similarity code snippet pairs based on the hash value of each code snippet to be detected; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in one group of high-similarity code snippet pairs is greater than a preset similarity; An input unit is used to input the set of high-similarity code snippet pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein the code clone detection result includes a clone code pair, and the two clone code snippets contained in the clone code pair implement the same function.

[0012] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any one of the above-described code clone detection methods.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the above-described code clone detection methods.

[0014] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned code clone detection methods.

[0015] In a sixth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction, and when the processor executes the program or instruction, the steps of any one of the above-mentioned code clone detection methods are implemented.

[0016] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides a code clone detection method, apparatus, device, medium and product, which can calculate a hash value for each code snippet to be detected in a code library to be detected, and group the code snippets to be detected with high similarity into high-similarity code snippet pairs according to the hash value of each code snippet to be detected, and then use the trained code clone detection model based on a large language model to perform code clone detection on the high-similarity code snippet pairs, which can make full use of the capabilities of the large language model, avoid the model's excessive reliance on the surface features of the code syntax, and then accurately judge the similarity of code functions based on the semantic analysis of the code snippets, thereby improving the accuracy of code clone detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A schematic diagram of a code clone detection method in an embodiment of the present application; Figure 2 A schematic diagram of functional modules of a code cloning detection device provided in one embodiment of the present application.

[0019] Figure 3 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0022] In an exemplary embodiment, Figure 1 As shown, a code clone detection method is provided, which is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method includes the following steps 101 to 104. Among them: Step 101, obtaining a code library to be detected.

[0023] In the embodiment of the present application, the code library to be detected includes multiple code snippets to be detected. The code snippet to be detected refers to a relatively small, independent and functional code set, which is a part of the code as a whole. A code snippet to be detected can be a few lines of code to complete a specific task, or it can be a partial implementation of a function, a class or a module. The code snippet to be detected usually has a relatively independent function and can complete a small task alone, such as file reading, data conversion, etc. Even if it is separated from the overall code, it can run normally under a suitable environment.

[0024] Step 102: determine the hash value of each code snippet to be detected.

[0025] In an embodiment of the present application, the method for determining the hash value of the code snippet to be detected can be: converting the text content of the code snippet to be detected to obtain a byte stream (binary format) of the code snippet to be detected, thereby ensuring the consistency of different encodings (such as UTF-8); and then using a hash algorithm to generate a hash value of the byte stream. Among them, the hash algorithm used can be MD5 (128 bits, fast but less secure), SHA-1 (160 bits, better security than MD5), SHA-256 (256 bits, secure and widely used) or SHA-512 (512 bits, suitable for high-security scenarios), etc., and this embodiment of the present application does not limit this.

[0026] Step 103: Determine a set of highly similar code snippet pairs based on the hash value of each code snippet to be detected.

[0027] In an embodiment of the present application, the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in a group of high-similarity code snippet pairs is greater than a preset similarity.

[0028] Step 104: input the set of high-similar code snippet pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model.

[0029] In the embodiment of the present application, the code clone detection model is constructed based on a large language model (LLM). The code clone detection result includes a clone code pair, and the two clone code fragments included in the clone code pair implement the same function.

[0030] In the embodiment of the present application, the training method of the code clone detection model based on the large language model is specifically as follows: Build an initial code clone detection model based on a large language model; Using the first training data set, the initial code clone detection model is fine-tuned for multiple summarization enhancements to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions; Using a second training data set to perform multiple mixed summary cloning fine-tuning on the first code clone detection model, to obtain a second code clone detection model capable of determining whether a code clone is a semantic clone; The second code clone detection model is directly optimized multiple times using the third training data set to obtain a trained code clone detection model.

[0031] Among them, the implementation of this implementation method improves the code semantic understanding ability through summary enhancement and fine-tuning, and hybrid summary cloning training integrates multimodal features to enhance generalization. Finally, direct preference optimization is used to achieve precise adaptation of the model to actual application scenarios. It effectively solves the problems of insufficient generalization ability of traditional code cloning detection models, difficulty in multimodal feature fusion, and poor business adaptability. It significantly reduces the model development cost while ensuring detection accuracy.

[0032] As an optional implementation, a method of using the first training data set to perform summary enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is a semantic clone based on the code functions may include the following steps: Acquire first training data from a first training data set; wherein the first training data includes a first training code segment and a second training code segment; Acquire first standard function description information of the first training code segment, second standard function description information of the second training code segment, and a first real code clone tag; wherein the first real code clone tag indicates whether the first standard function description information and the second standard function description information are actually the same; Inputting the first training code segment, the second training code segment and the detection instruction into the initial code clone detection model, obtaining first function description information of the first training code segment, second function description information of the second training code segment and a first code clone judgment result output by the initial code clone detection model; wherein the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same; Using a loss function, calculating the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first real code clone label to obtain a first loss value; If the first loss value is greater than or equal to a first preset loss threshold, fine-tuning the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein the adjusted initial code clone detection model obtained in any summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation; If the first loss value is less than the first preset loss threshold, the initial code clone detection model in this summary enhancement fine-tuning operation is determined as a first code clone detection model that can analyze code functions and determine whether it is a semantic clone based on the code functions.

[0033] Among them, implementing this implementation method, by obtaining the first training data containing the first training code segment and the second training code segment from the first training data set, and obtaining the corresponding standard function description information and the real code clone label, provides a rich and accurate data basis for model training, which helps the model learn the real characteristics of the code segment function description and clone relationship. Secondly, multiple summary enhancement fine-tuning operations combined with the calculation of the loss function can continuously evaluate the difference between the model output and the actual situation. By continuously adjusting the model parameters, the model gradually approaches the optimal state, thereby improving the accuracy and reliability of the model for code clone detection. Fine-tuning is stopped when the first loss value of the model is less than the first preset loss threshold, ensuring that the model achieves better performance within a reasonable error range and avoiding overtraining or undertraining.

[0034] In an embodiment of the present application, the training data in the first training data set can be divided into two categories. One category is a data set based on programming competitions, such as GoogleCodeJam and OJClone, etc.; the other category is a data set based on real projects, such as BigCloneBench and GPTClone. In order to effectively train the model, two data sets are preferred: GoogleCodeJam and GPTClone. The reason for choosing them is that the code of GoogleCodeJam is more complex than that of OJClone, which helps to improve the effect of model training. At the same time, we maintain a uniform distribution among different types of codes, which has been proven to be very beneficial for model training and evaluation. In addition, the data quality of GPTClone is better than that of BigCloneBench. According to previous studies, there is some low-quality data in BigCloneBench, which will affect the training and evaluation of the model.

[0035] We introduced 3 new programming languages ​​(C++, Python, Php) by extending the GoogleCodeJam2 dataset. In addition, in order to effectively enhance the clone detection ability of the model, we used a large language model (LLM) to generate short code summaries for the dataset and incorporated them into the subsequent model training (see Section III-C1). Finally, our dataset contains 4 programming languages ​​(Java, C++, Python, Php), with a total of 37,364 code snippets. Since GPTCloneBench only contains real clone pairs, we use it as an evaluation dataset.

[0036] In the embodiment of the present application, the first training data may be randomly acquired from the first training data set.

[0037] In the embodiment of the present application, the loss function Can be: Among them, for each word in the sequence, the loss function calculates the probability of the next word appearing, L represents the length of the sequence (the sequence includes the first function description information, the second function description information and the first code clone judgment result), D represents the data set (including the first standard function description information of the first training data, the second standard function description information and the first real code clone label), represents the initial code clone detection model, Represents the i-th word in the sequence.

[0038] In an embodiment of the present application, a first code clone detection model that can simultaneously analyze code functions and detect semantic clones based on these functions is constructed through summary enhancement fine-tuning. However, it can be seen from the loss function that during the initial code clone detection model training process, the code clone judgment result only accounts for a small part of the loss calculation because the generated part includes the code function description. This results in limited improvement in the accuracy of code clone detection. In order to solve this problem and ensure that the model output conforms to the expected format, we designed a second-stage hybrid summary clone fine-tuning method.

[0039] Optionally, the method for constructing the second training data set may include the following steps: Construct a second training data set; wherein the second training data set includes second initial training data of two data types, the two data types include pure code type and code-function type; the second initial training data of the pure code type includes a third training code segment and a fourth training code segment; the second initial training data of the code-function type includes a third training code segment, a fourth training code segment, third standard function description information of the third training code segment, and fourth standard function description information of the fourth training code segment.

[0040] Among them, implementing this implementation method, by mixing pure code type and code-function type data to construct a second training data set, can dynamically balance the multimodal feature input of model training and effectively enhance the adaptability of the code clone detection model to complex scenarios. Specifically, pure code type data can enhance the model's ability to recognize code structure features, and code-function type data can enhance the model's understanding of semantic similarity by integrating functional description information. The random combination of the two avoids the limitations of a single data type and stimulates the model's feature generalization ability through data diversity, laying a solid foundation for multimodal feature fusion in the subsequent mixed summary cloning fine-tuning stage, and significantly improving the model's clone detection accuracy and robustness in real complex code environments.

[0041] In the embodiment of the present application, the number of second initial training data of pure code type contained in the second training data set is the same as the number of second initial training data of code-function type.

[0042] As an optional implementation, the method of fine-tuning the first code clone detection model by using the second training data set to perform a hybrid summary clone fine-tuning to obtain the second code clone detection model capable of determining whether a semantic clone is present may include the following steps: Determine a second initial training data randomly obtained from the second training data set as the second training data; Acquire a second real code clone tag corresponding to the third training code segment and the fourth training code segment in the second training data; wherein the second real code clone tag indicates whether the function information of the third training code segment and the fourth training code segment is actually the same; Inputting the second training data into the first code clone detection model to obtain a second code clone determination result output by the first code clone detection model; wherein the second code clone determination result indicates whether the function information of the third training code segment and the fourth training code segment is the same; Using the loss function, calculating the second code clone judgment result and the second real code clone label to obtain a second loss value; If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein the adjusted first code clone detection model obtained in any hybrid summary cloning fine-tuning operation is used as the first code clone detection model in the next summary enhancement fine-tuning operation; If the second loss value is less than the second preset loss threshold, the first code clone detection model in this hybrid summary cloning fine-tuning operation is determined as the second code clone detection model that can determine whether semantic cloning occurs.

[0043] Among them, the implementation of this implementation method effectively enhances the model's ability to understand the dual features of code semantics and structure by dynamically fusing pure code structure features with multimodal data of code-function description, so that the model can not only capture the clone relationship at the code syntax level, but also deeply explore the similarity of functional semantics during the iterative optimization process, significantly improving the accuracy and generalization ability of code clone detection in complex scenarios. The mechanism of randomly selecting training data avoids the limitations of a single data type, and combined with the supervised training of the second real code clone label, ensures that the model continuously calibrates the judgment logic of code function similarity during the optimization process, and finally stops training when the preset loss threshold is reached, which not only ensures model performance but also prevents overfitting, laying a solid foundation for the robust performance of subsequent models in actual engineering code detection.

[0044] Optionally, the third standard function description information of the third training code segment and the fourth standard function description information of the fourth training code segment included in the second initial training data of the code-function type may be generated by the first code clone detection model that has completed summary enhancement and fine-tuning.

[0045] In order to retain the instruction following ability of the model, the second initial training data of pure code type and code-function type are mixed in a 1:1 ratio in the mixed summary clone fine-tuning of this stage. In this stage, the second loss value of the second code clone judgment result output by the first code clone detection model is calculated by continuing to use the same loss function as the summary enhancement fine-tuning stage. Since the output of this stage only contains the code clone judgment result, the model can focus more on improving the accuracy of clone detection.

[0046] As an optional implementation, a method of performing a direct preference optimization on the second code clone detection model using the third training data set to obtain a trained code clone detection model may include the following steps: Acquire third training data from a third training data set; wherein the third training data includes a fifth training code segment and a sixth training code segment; Acquire a third real code clone tag corresponding to the fifth training code segment and the sixth training code segment; wherein the third real code clone tag indicates whether the function information of the fifth training code segment and the sixth training code segment is actually the same; Inputting the third training data into the second code clone detection model to obtain a third code clone determination result output by the second code clone detection model; wherein the third code clone determination result indicates whether the function information of the fifth training code segment and the sixth training code segment is the same; Comparing the third code clone determination result with the third real code clone label to obtain a comparison result; wherein the comparison result indicates whether the third code clone determination result is the same as the third real code clone label; Using a target loss function, calculating the third code clone judgment result, the third real code clone label, and the comparison result to obtain a third loss value; If the third loss value is greater than or equal to the third preset loss threshold, the second code clone detection model is directly preferred to be optimized to obtain an optimized second code clone detection model; wherein the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization; If the third loss value is less than the third preset loss threshold, the second code clone detection model in this direct preference optimization is determined as the trained code clone detection model.

[0047] Among them, by implementing this implementation method, by inputting the third training data into the second code clone detection model and obtaining the comparison results, it is possible to accurately capture the judgment deviation of the model in the actual application scenario and effectively calibrate the model's evaluation logic for the similarity of code functions. Specifically, the optimization process dynamically adjusts the model parameters so that the model can give priority to the code features whose real code clone labels are inconsistent with the predicted results, significantly improving the detection accuracy of the model in complex code scenarios. At the same time, the third training data selection mechanism based on the third training data set ensures the representativeness and consistency of the training data, avoiding the problem of model performance degradation caused by differences in data distribution. This optimization strategy that combines manual labeling preferences with model autonomous learning not only enhances the model's ability to adapt to actual business needs, but also realizes the continuous iteration of model performance through a real feedback loop mechanism, providing key technical guarantees for the stable and reliable detection performance of the finally trained code clone detection model in the actual engineering environment.

[0048] In the embodiment of the present application, the direct preference optimization of the second code clone detection model can be implemented by using the target loss function of the direct preference optimization (DPO) algorithm. The target loss function of the DPO algorithm is: in, and Represents the second code clone detection model given by x input or The cumulative probability of is the language model strategy, is the benchmark reference strategy, and β is the parameter that controls the deviation between the two. In this way, the DPO algorithm can use alternative parameterization to fit the implicit reward, and its optimal strategy is .

[0049] Generally speaking, the DPO algorithm works by increasing the log probability of preferred samples through the left part of the objective loss function, while reducing the log probability of non-preferred sample responses through the right part. The DPO algorithm treats human preference alignment as a classification problem, which makes it particularly suitable for the fine-tuning task in this paper. By treating preference alignment as a classification task, the DPO algorithm enables us to optimize the second code clone detection model to better align with the real clone judgment results. In this study, this process involves accurately detecting code clones.

[0050] Specifically, for the third training data input, the DPO algorithm is used to train the second code clone detection model. This goal is achieved without introducing significant fluctuations in the output, thereby avoiding the situation where the model fails to accurately determine whether two code snippets are code clones after correctly analyzing them.

[0051] By implementing the above steps 101 to 104, the capabilities of the large language model can be fully utilized to avoid the model's excessive reliance on the surface features of the code syntax, and then based on the semantic analysis of the code snippets, accurate judgment of the functional similarity of the code can be achieved, thereby improving the accuracy of code clone detection. In addition, the present application can also solve the problems of insufficient generalization ability of traditional code clone detection models, difficulty in multimodal feature fusion, and poor business adaptability, while ensuring detection accuracy. The cost of model development is significantly reduced. In addition, the present application can also improve the accuracy and reliability of the model for code clone detection. In addition, the present application can also improve the accuracy and robustness of clone detection of the model in a real complex code environment. In addition, the present application can also prevent overfitting while ensuring model performance. In addition, the present application can also enhance the model's ability to adapt to actual business needs, and also achieve continuous iteration of model performance through a real feedback loop mechanism.

[0052] Based on the same inventive concept, the embodiment of the present application also provides a code clone detection device for implementing the code clone detection method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more code clone detection device embodiments provided below can refer to the limitations of the code clone detection method above, and will not be repeated here.

[0053] In an exemplary embodiment, Figure 2 As shown, a code clone detection device is provided, comprising: The acquisition unit 201 is used to acquire a code library to be detected; wherein the code library to be detected contains multiple code fragments to be detected; A first determining unit 202 is used to determine a hash value of each code segment to be detected; A second determining unit 203 is used to determine a set of high-similarity code snippet pairs based on the hash value of each code snippet to be detected; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in a group of high-similarity code snippet pairs is greater than a preset similarity; The input unit 204 is used to input the set of high-similarity code snippet pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein the code clone detection result includes a clone code pair, and the two clone code snippets contained in the clone code pair implement the same function.

[0054] In the embodiment of the present application, the training method of the code clone detection model based on the large language model is specifically as follows: Build an initial code clone detection model based on a large language model; Using the first training data set, the initial code clone detection model is fine-tuned for multiple summarization enhancements to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions; Using a second training data set to perform multiple mixed summary cloning fine-tuning on the first code clone detection model, to obtain a second code clone detection model capable of determining whether a code clone is a semantic clone; The second code clone detection model is directly optimized multiple times using the third training data set to obtain a trained code clone detection model.

[0055] Among them, the implementation of this implementation method improves the code semantic understanding ability through summary enhancement and fine-tuning, and hybrid summary cloning training integrates multimodal features to enhance generalization. Finally, direct preference optimization is used to achieve precise adaptation of the model to actual application scenarios. It effectively solves the problems of insufficient generalization ability of traditional code cloning detection models, difficulty in multimodal feature fusion, and poor business adaptability. It significantly reduces the model development cost while ensuring detection accuracy.

[0056] As an optional implementation, a method of using the first training data set to perform summary enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is a semantic clone based on the code functions may include the following steps: Acquire first training data from a first training data set; wherein the first training data includes a first training code segment and a second training code segment; Acquire first standard function description information of the first training code segment, second standard function description information of the second training code segment, and a first real code clone tag; wherein the first real code clone tag indicates whether the first standard function description information and the second standard function description information are actually the same; Inputting the first training code segment, the second training code segment and the detection instruction into the initial code clone detection model, obtaining first function description information of the first training code segment, second function description information of the second training code segment and a first code clone judgment result output by the initial code clone detection model; wherein the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same; Using a loss function, calculating the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first real code clone label to obtain a first loss value; If the first loss value is greater than or equal to a first preset loss threshold, fine-tuning the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein the adjusted initial code clone detection model obtained in any summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation; If the first loss value is less than the first preset loss threshold, the initial code clone detection model in this summary enhancement fine-tuning operation is determined as a first code clone detection model that can analyze code functions and determine whether it is a semantic clone based on the code functions.

[0057] Among them, implementing this implementation method, by obtaining the first training data containing the first training code segment and the second training code segment from the first training data set, and obtaining the corresponding standard function description information and the real code clone label, provides a rich and accurate data basis for model training, which helps the model learn the real characteristics of the code segment function description and clone relationship. Secondly, multiple summary enhancement fine-tuning operations combined with the calculation of the loss function can continuously evaluate the difference between the model output and the actual situation. By continuously adjusting the model parameters, the model gradually approaches the optimal state, thereby improving the accuracy and reliability of the model for code clone detection. Fine-tuning is stopped when the first loss value of the model is less than the first preset loss threshold, ensuring that the model achieves better performance within a reasonable error range and avoiding overtraining or undertraining.

[0058] Optionally, the method for constructing the second training data set may include the following steps: Construct a second training data set; wherein the second training data set includes second initial training data of two data types, the two data types include pure code type and code-function type; the second initial training data of the pure code type includes a third training code segment and a fourth training code segment; the second initial training data of the code-function type includes a third training code segment, a fourth training code segment, third standard function description information of the third training code segment, and fourth standard function description information of the fourth training code segment.

[0059] Among them, implementing this implementation method, by mixing pure code type and code-function type data to construct a second training data set, can dynamically balance the multimodal feature input of model training and effectively enhance the adaptability of the code clone detection model to complex scenarios. Specifically, pure code type data can enhance the model's ability to recognize code structure features, and code-function type data can enhance the model's understanding of semantic similarity by integrating functional description information. The random combination of the two avoids the limitations of a single data type and stimulates the model's feature generalization ability through data diversity, laying a solid foundation for multimodal feature fusion in the subsequent mixed summary cloning fine-tuning stage, and significantly improving the model's clone detection accuracy and robustness in real complex code environments.

[0060] As an optional implementation, the method of fine-tuning the first code clone detection model by using the second training data set to perform a hybrid summary clone fine-tuning to obtain the second code clone detection model capable of determining whether a semantic clone is present may include the following steps: Determine a second initial training data randomly obtained from the second training data set as the second training data; Acquire a second real code clone tag corresponding to the third training code segment and the fourth training code segment in the second training data; wherein the second real code clone tag indicates whether the function information of the third training code segment and the fourth training code segment is actually the same; Inputting the second training data into the first code clone detection model to obtain a second code clone determination result output by the first code clone detection model; wherein the second code clone determination result indicates whether the function information of the third training code segment and the fourth training code segment is the same; Using the loss function, calculating the second code clone judgment result and the second real code clone label to obtain a second loss value; If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein the adjusted first code clone detection model obtained in any hybrid summary cloning fine-tuning operation is used as the first code clone detection model in the next summary enhancement fine-tuning operation; If the second loss value is less than the second preset loss threshold, the first code clone detection model in this hybrid summary cloning fine-tuning operation is determined as the second code clone detection model that can determine whether semantic cloning occurs.

[0061] Among them, the implementation of this implementation method effectively enhances the model's ability to understand the dual features of code semantics and structure by dynamically fusing pure code structure features with multimodal data of code-function description, so that the model can not only capture the clone relationship at the code syntax level, but also deeply explore the similarity of functional semantics during the iterative optimization process, significantly improving the accuracy and generalization ability of code clone detection in complex scenarios. The mechanism of randomly selecting training data avoids the limitations of a single data type, and combined with the supervised training of the second real code clone label, ensures that the model continuously calibrates the judgment logic of code function similarity during the optimization process, and finally stops training when the preset loss threshold is reached, which not only ensures model performance but also prevents overfitting, laying a solid foundation for the robust performance of subsequent models in actual engineering code detection.

[0062] As an optional implementation, a method of performing a direct preference optimization on the second code clone detection model using the third training data set to obtain a trained code clone detection model may include the following steps: Acquire third training data from a third training data set; wherein the third training data includes a fifth training code segment and a sixth training code segment; Acquire a third real code clone tag corresponding to the fifth training code segment and the sixth training code segment; wherein the third real code clone tag indicates whether the function information of the fifth training code segment and the sixth training code segment is actually the same; Inputting the third training data into the second code clone detection model to obtain a third code clone determination result output by the second code clone detection model; wherein the third code clone determination result indicates whether the function information of the fifth training code segment and the sixth training code segment is the same; Comparing the third code clone determination result with the third real code clone label to obtain a comparison result; wherein the comparison result indicates whether the third code clone determination result is the same as the third real code clone label; Using a target loss function, calculating the third code clone judgment result, the third real code clone label, and the comparison result to obtain a third loss value; If the third loss value is greater than or equal to the third preset loss threshold, the second code clone detection model is directly preferred to be optimized to obtain an optimized second code clone detection model; wherein the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization; If the third loss value is less than the third preset loss threshold, the second code clone detection model in this direct preference optimization is determined as the trained code clone detection model.

[0063] Among them, by implementing this implementation method, by inputting the third training data into the second code clone detection model and obtaining the comparison results, it is possible to accurately capture the judgment deviation of the model in the actual application scenario and effectively calibrate the model's evaluation logic for the similarity of code functions. Specifically, the optimization process dynamically adjusts the model parameters so that the model can give priority to the code features whose real code clone labels are inconsistent with the predicted results, significantly improving the detection accuracy of the model in complex code scenarios. At the same time, the third training data selection mechanism based on the third training data set ensures the representativeness and consistency of the training data, avoiding the problem of model performance degradation caused by differences in data distribution. This optimization strategy that combines manual labeling preferences with model autonomous learning not only enhances the model's ability to adapt to actual business needs, but also realizes the continuous iteration of model performance through a real feedback loop mechanism, providing key technical guarantees for the stable and reliable detection performance of the finally trained code clone detection model in the actual engineering environment.

[0064] By implementing the above-mentioned implementation methods, the capabilities of large language models can be fully utilized to avoid the model's excessive reliance on the surface features of code syntax, and then based on the semantic analysis of code snippets, accurate judgment of code functional similarity can be achieved, thereby improving the accuracy of code clone detection. In addition, the present application can also solve the problems of insufficient generalization ability of traditional code clone detection models, difficulty in multimodal feature fusion, and poor business adaptability, while ensuring detection accuracy and significantly reducing the model development cost. In addition, the present application can also improve the accuracy and reliability of the model for code clone detection. In addition, the present application can also improve the accuracy and robustness of clone detection of the model in a real complex code environment. In addition, the present application can also prevent overfitting while ensuring model performance. In addition, the present application can also enhance the model's adaptability to actual business needs, and also achieve continuous iteration of model performance through a real feedback loop mechanism.

[0065] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store code clone detection data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a code clone detection method is implemented.

[0066] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0067] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0068] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0069] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0070] In an exemplary embodiment, a chip is provided, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the above-mentioned method embodiments and achieve the same technical effects. To avoid repetition, they are not described here.

[0071] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0072] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0073] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0074] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0075] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0076] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A code clone detection method, characterized in that: The code clone detection method comprises: Obtain a code library to be detected; wherein the code library to be detected contains multiple code snippets to be detected; Determine the hash value of each code snippet to be detected; Based on the hash value of each code snippet to be detected, a set of high-similarity code snippet pairs is determined; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in a group of high-similarity code snippet pairs is greater than a preset similarity; The set of highly similar code snippet pairs is input into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein the code clone detection result includes a clone code pair, and the two clone code snippets contained in the clone code pair implement the same function.

2. The code clone detection method according to claim 1, characterized in that: The training method of the code clone detection model based on the large language model is specifically as follows: Build an initial code clone detection model based on a large language model; Using the first training data set, the initial code clone detection model is fine-tuned for multiple times by summary enhancement, so as to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions; Using a second training data set to perform multiple mixed summary cloning fine-tuning on the first code clone detection model, to obtain a second code clone detection model capable of determining whether a code clone is a semantic clone; The second code clone detection model is directly optimized multiple times using the third training data set to obtain a trained code clone detection model.

3. The code clone detection method according to claim 2, characterized in that: The method of using the first training data set to perform summary enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and determining whether a code clone is semantically cloned based on the code functions specifically includes: Acquire first training data from a first training data set; wherein the first training data includes a first training code segment and a second training code segment; Acquire first standard function description information of the first training code segment, second standard function description information of the second training code segment, and a first real code clone tag; wherein the first real code clone tag indicates whether the first standard function description information and the second standard function description information are actually the same; Inputting the first training code segment, the second training code segment and the detection instruction into the initial code clone detection model, obtaining first function description information of the first training code segment, second function description information of the second training code segment and a first code clone judgment result output by the initial code clone detection model; wherein the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same; Using a loss function, calculating the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first real code clone label to obtain a first loss value; If the first loss value is greater than or equal to a first preset loss threshold, fine-tuning the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein the adjusted initial code clone detection model obtained in any summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation; If the first loss value is less than the first preset loss threshold, the initial code clone detection model in this summary enhancement fine-tuning operation is determined as a first code clone detection model that can analyze code functions and determine whether it is a semantic clone based on the code functions.

4. The code clone detection method according to claim 3, characterized in that: The code clone detection method further includes: Construct a second training data set; wherein the second training data set includes second initial training data of two data types, the two data types include pure code type and code-function type; the second initial training data of the pure code type includes a third training code segment and a fourth training code segment; the second initial training data of the code-function type includes a third training code segment, a fourth training code segment, third standard function description information of the third training code segment, and fourth standard function description information of the fourth training code segment.

5. The code clone detection method according to claim 4, characterized in that: The method of fine-tuning the first code clone detection model by using the second training data set to perform a mixed summary clone fine-tuning to obtain a second code clone detection model capable of determining whether a semantic clone exists specifically includes: Determine a second initial training data randomly obtained from the second training data set as the second training data; Acquire a second real code clone tag corresponding to the third training code segment and the fourth training code segment in the second training data; wherein the second real code clone tag indicates whether the function information of the third training code segment and the fourth training code segment is actually the same; Inputting the second training data into the first code clone detection model to obtain a second code clone determination result output by the first code clone detection model; wherein the second code clone determination result indicates whether the function information of the third training code segment and the fourth training code segment is the same; Using the loss function, calculating the second code clone judgment result and the second real code clone label to obtain a second loss value; If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein the adjusted first code clone detection model obtained in any hybrid summary cloning fine-tuning operation is used as the first code clone detection model in the next summary enhancement fine-tuning operation; If the second loss value is less than the second preset loss threshold, the first code clone detection model in this hybrid summary cloning fine-tuning operation is determined as the second code clone detection model that can determine whether semantic cloning occurs.

6. The code clone detection method according to any one of claims 3 to 5, characterized in that: The step of performing a direct preference optimization on the second code clone detection model using the third training data set to obtain a trained code clone detection model specifically includes: Acquire third training data from a third training data set; wherein the third training data includes a fifth training code segment and a sixth training code segment; Acquire a third real code clone tag corresponding to the fifth training code segment and the sixth training code segment; wherein the third real code clone tag indicates whether the function information of the fifth training code segment and the sixth training code segment is actually the same; Inputting the third training data into the second code clone detection model to obtain a third code clone determination result output by the second code clone detection model; wherein the third code clone determination result indicates whether the function information of the fifth training code segment and the sixth training code segment is the same; Comparing the third code clone determination result with the third real code clone label to obtain a comparison result; wherein the comparison result indicates whether the third code clone determination result is the same as the third real code clone label; Using a target loss function, calculating the third code clone judgment result, the third real code clone label, and the comparison result to obtain a third loss value; If the third loss value is greater than or equal to the third preset loss threshold, the second code clone detection model is directly preferred to be optimized to obtain an optimized second code clone detection model; wherein the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization; If the third loss value is less than the third preset loss threshold, the second code clone detection model in this direct preference optimization is determined as the trained code clone detection model.

7. A code clone detection device, characterized in that: The code clone detection device comprises: An acquisition unit is used to acquire a code library to be detected; wherein the code library to be detected contains multiple code fragments to be detected; A first determining unit, used to determine a hash value of each code fragment to be detected; A second determination unit is used to determine a set of high-similarity code snippet pairs based on the hash value of each code snippet to be detected; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values ​​of the two target code snippets to be detected included in one group of high-similarity code snippet pairs is greater than a preset similarity; An input unit is used to input the set of high-similarity code snippet pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein the code clone detection result includes a clone code pair, and the two clone code snippets contained in the clone code pair implement the same function.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the code clone detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the code clone detection method described in any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the code clone detection method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Clone code detection method and device

    CN110990273A

  • Code cloning detection method and device and electronic equipment

    CN111124487A

  • Repeated code detection method and device, terminal equipment and storage medium

    CN117573519A

  • Code detection method and device and electronic equipment

    CN118152000A

  • Systems and methods for code understanding and generation

    US20220382527A1

Cited By

  • Code cloning detection method and device, electronic equipment and storage medium

    CN121742899A