A method, apparatus, device, medium and product for code clone detection
By building a code cloning detection model based on a large language model, using a multi-stage training data set to improve the semantic understanding of code functions and multimodal feature fusion ability of the model, the limitations of traditional methods in semantic cloning detection are solved, and high accuracy and robust code cloning detection are achieved.
Patent Information
- Application Number
- CN202510442636.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing code cloning detection methods perform well when processing structured code or simple cloning, but there are great limitations for semantic cloning detection that uses different code writing methods to implement the same function, resulting in low detection accuracy and inability to meet actual needs.
By obtaining the hash value of the code snippet in the code base to be detected, a code cloning detection model based on a large language model is constructed, and a multi-stage training data set is used to fine-tune the model, including abstract enhancement fine-tune, mixed abstract cloning fine-tune and direct preference optimization, improving the model's semantic understanding ability of code functions and multimodal feature fusion ability.
It significantly improves the accuracy and robustness of code cloning detection, solves the problems of insufficient generalization capabilities of traditional methods and difficulty in fusion of multimodal features in complex scenarios, reduces model development costs, and shows stable detection performance in actual engineering environments.
Smart Images

Figure CN119938135B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a code clone detection method, device, equipment, medium and product. Background Art
[0002] Code clones refer to code fragments with the same or similar functions but different implementations in a software system. These code clones may be caused by code reuse, repeated development, or other human factors, but too many code clones will lead to code redundancy, increase software maintenance costs, and may hide potential vulnerabilities or defects. Therefore, fast and accurate detection of code clones is of great significance to improving software quality and development efficiency.
[0003] Traditional code clone detection methods usually rely on matching techniques based on text, syntax, or Abstract Syntax Tree (AST). These methods perform well when dealing with structured code or simple clones, but have significant limitations in detecting clones that use different code writing methods to achieve the same function (i.e., semantic clones). With the increase in code complexity and the increase in cross-language code scenarios, traditional methods that rely only on static rules and simple feature extraction have low detection accuracy and cannot meet actual needs. Summary of the invention
[0004] The purpose of this application is to provide a code clone detection method, device, equipment, medium and product, which can improve the accuracy of code clone detection.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a code clone detection method, comprising:
[0007] Obtain a code library to be detected; wherein the code library to be detected contains multiple code snippets to be detected;
[0008] Determine the hash value of each code snippet to be detected;
[0009] Based on the hash value of each code snippet to be detected, a set of high-similarity code snippet pairs is determined; wherein the set of high-similarity code snippet pairs includes at least one group of high-similarity code snippet pairs, the high-similarity code snippet pairs include two target code snippets to be detected, and the similarity of the hash values of the two target code snippets to be detected included in a group of high-similarity code snippet pairs is greater than a preset similarity;
[0010] Input the set of high - similarity code snippet pairs into a pre - trained code clone detection model based on a large language model to obtain the code clone detection result output by the code clone detection model; wherein, the code clone detection result includes clone code pairs, and the two clone code snippets included in the clone code pair implement the same function.
[0011] Optionally, the training method of the code clone detection model based on a large language model is specifically as follows:
[0012] Construct an initial code clone detection model based on a large language model;
[0013] Use the first training dataset to perform multiple abstract augmentation fine - tuning on the initial code clone detection model to obtain a first code clone detection model that can analyze code functions and judge whether it is a semantic clone based on code functions;
[0014] Use the second training dataset to perform multiple hybrid abstract clone fine - tuning on the first code clone detection model to obtain a second code clone detection model that can judge whether it is a semantic clone;
[0015] Use the third training dataset to perform multiple direct preference optimizations on the second code clone detection model to obtain a trained code clone detection model.
[0016] Optionally, the step of using the first training dataset to perform one - time abstract augmentation fine - tuning on the initial code clone detection model to obtain a first code clone detection model that can analyze code functions and judge whether it is a semantic clone based on code functions specifically includes:
[0017] Obtain first training data from the first training dataset; wherein, the first training data includes a first training code segment and a second training code segment;
[0018] Obtain the first standard function description information of the first training code segment, the second standard function description information of the second training code segment, and a first true code clone label; wherein, the first true code clone label indicates whether the first standard function description information and the second standard function description information are actually the same;
[0019] Input the first training code segment, the second training code segment, and the detection instruction into the initial code clone detection model to obtain the first function description information of the first training code segment, the second function description information of the second training code segment, and the first code clone judgment result output by the initial code clone detection model; wherein, the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same;
[0020] Use a loss function to calculate the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first true code clone label to obtain a first loss value;
[0021] If the first loss value is greater than or equal to the first preset loss threshold, fine-tune the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein, the adjusted initial code clone detection model obtained in any one summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation;
[0022] If the first loss value is less than the first preset loss threshold, determine the initial code clone detection model in this summary enhancement fine-tuning operation as the first code clone detection model that can analyze the code function and judge whether it is a semantic clone based on the code function.
[0023] Optionally, the code clone detection method further includes:
[0024] Construct a second training data set; wherein, the second training data set contains second initial training data of two data types, and the two data types include pure code type and code-function type; the second initial training data of the pure code type contains a third training code segment and a fourth training code segment; the second initial training data of the code-function type contains a third training code segment, a fourth training code segment, the third standard function description information of the third training code segment, and the fourth standard function description information of the fourth training code segment.
[0025] Optionally, the step of using the second training data set to perform a mixed summary clone fine-tuning on the first code clone detection model to obtain a second code clone detection model that can judge whether it is a semantic clone specifically includes:
[0026] Determine a piece of second initial training data randomly obtained from the second training data set as the second training data;
[0027] Obtain a second true code clone label corresponding to the third training code segment and the fourth training code segment in the second training data; wherein, the second true code clone label indicates whether the functional information of the third training code segment and the fourth training code segment is actually the same;
[0028] Input the second training data into the first code clone detection model to obtain a second code clone judgment result output by the first code clone detection model; wherein, the second code clone judgment result indicates whether the functional information of the third training code segment and the fourth training code segment is the same;
[0029] Use the loss function to calculate the second code clone judgment result and the second true code clone label to obtain a second loss value;
[0030] If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein, the adjusted first code clone detection model obtained in any mixed abstract clone fine-tuning operation is used as the first code clone detection model in the next abstract enhancement fine-tuning operation;
[0031] If the second loss value is less than the second preset loss threshold, determine the first code clone detection model in the current mixed abstract clone fine-tuning operation as a second code clone detection model capable of judging semantic cloning.
[0032] Optionally, performing a direct preference optimization on the second code clone detection model using the third training dataset to obtain a trained code clone detection model, specifically including:
[0033] Obtain third training data from the third training dataset; wherein, the third training data includes a fifth training code segment and a sixth training code segment;
[0034] Obtain a third true code clone label corresponding to the fifth training code segment and the sixth training code segment; wherein, the third true code clone label indicates whether the functional information of the fifth training code segment and the sixth training code segment is actually the same;
[0035] Input the third training data into the second code clone detection model to obtain a third code clone judgment result output by the second code clone detection model; wherein, the third code clone judgment result indicates whether the functional information of the fifth training code segment and the sixth training code segment is the same;
[0036] Compare the third code clone judgment result with the third true code clone label to obtain a comparison result; wherein, the comparison result indicates whether the third code clone judgment result is the same as the third true code clone label;
[0037] Use the target loss function to calculate the third code clone judgment result, the third true code clone label, and the comparison result to obtain a third loss value;
[0038] If the third loss value is greater than or equal to the third preset loss threshold, directly optimize the second code clone detection model to obtain an optimized second code clone detection model; wherein, the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization;
[0039] If the third loss value is less than the third preset loss threshold, determine the second code clone detection model in the current direct preference optimization as the trained code clone detection model.
[0040] In a second aspect, the present application provides a code clone detection device, including:
[0041] An acquisition unit, configured to acquire a code library to be detected; wherein, the code library to be detected contains multiple code segments to be detected;
[0042] A first determination unit, configured to determine the hash value of each code segment to be detected;
[0043] A second determination unit, configured to determine a set of high-similarity code segment pairs based on the hash value of each code segment to be detected; wherein, the set of high-similarity code segment pairs contains at least one set of high-similarity code segment pairs, each high-similarity code segment pair contains two target code segments to be detected, and the similarity of the hash values of the two target code segments contained in a set of high-similarity code segment pairs is greater than a preset similarity;
[0044] An input unit, configured to input the set of high-similarity code segment pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein, the code clone detection result includes clone code pairs, and the two clone code segments included in the clone code pair implement the same function.
[0045] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the steps of the code clone detection method described in any one of the above.
[0046] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the code clone detection method described in any one of the above are implemented.
[0047] Fifthly, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the code clone detection method described in any one of the above are implemented.
[0048] Sixthly, the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions. When the processor executes the programs or instructions, the steps of the code clone detection method described in any one of the above are implemented.
[0049] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0050] The present application provides a code clone detection method, device, equipment, medium and product, which can calculate hash values for each code snippet to be detected in the code library to be detected, and form high-similarity code snippet pairs according to the hash values of each code snippet to be detected. Furthermore, the trained code clone detection model based on the large language model is used to detect code clones for the high-similarity code snippet pairs, which can make full use of the capabilities of the large language model, avoid the model's over-reliance on the surface features of code syntax, and then accurately judge the functional similarity of codes based on the semantic analysis of code snippets, thereby improving the accuracy of code clone detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a schematic flowchart of a code clone detection method in an embodiment of the present application;
[0053] Figure 2 It is a schematic diagram of the functional modules of a code clone detection device provided in an embodiment of the present application.
[0054] Figure 3 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0056] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments.
[0057] In an exemplary embodiment, as Figure 1 shown, a code clone detection method is provided. This method is executed by a computer device, and specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, this method includes the following steps 101 to 104. Among them:
[0058] Step 101, obtain the code library to be detected.
[0059] In the embodiments of the present application, the code library to be detected contains multiple code fragments to be detected. A code fragment to be detected refers to a relatively small, independent, and code collection with a specific function. It is a part of the overall code. A code fragment to be detected can be several lines of code that complete a specific task, or it may be a function, a class, or a partial implementation of a module. A code fragment to be detected usually has a relatively independent function and can complete a small task alone, such as file reading, data conversion, etc. Even if it is separated from the overall code, it can run normally in a suitable environment.
[0060] Step 102, determine the hash value of each code fragment to be detected.
[0061] In the embodiments of the present application, the method for determining the hash value of a code fragment to be detected can be: perform code conversion on the text content of the code fragment to be detected to obtain the byte stream (binary format) of the code fragment to be detected, so as to ensure the consistency of different encodings (such as UTF-8); then use a hash algorithm to generate the hash value of this byte stream. Among them, the hash algorithm used can be MD5 (128 bits, fast but with relatively low security), SHA-1 (160 bits, with better security than MD5), SHA-256 (256 bits, secure and widely used), or SHA-512 (512 bits, suitable for high-security scenarios), etc. For this, the embodiments of the present application do not make any limitations.
[0062] Step 103, based on the hash value of each code fragment to be detected, determine a set of code fragment pairs with high similarity.
[0063] In the embodiments of the present application, the set of high - similarity code fragment pairs contains at least one set of high - similarity code fragment pairs. Each high - similarity code fragment pair contains two target code fragments to be detected, and the similarity of the hash values of the two target code fragments contained in a set of high - similarity code fragment pairs is greater than a preset similarity.
[0064] Step 104: Input the set of high - similarity code fragment pairs into a pre - trained code clone detection model based on a large language model to obtain the code clone detection result output by the code clone detection model.
[0065] In the embodiments of the present application, the code clone detection model is constructed based on a large language model (LLM). The code clone detection result includes clone code pairs, and the two clone code fragments contained in the clone code pair implement the same function.
[0066] In the embodiments of the present application, the training method of the code clone detection model based on a large language model is specifically as follows:
[0067] Construct an initial code clone detection model based on a large language model;
[0068] Use the first training data set to perform multiple abstract - enhancement fine - tunings on the initial code clone detection model to obtain a first code clone detection model that can analyze the code function and determine whether it is a semantic clone based on the code function;
[0069] Use the second training data set to perform multiple hybrid - abstract clone fine - tunings on the first code clone detection model to obtain a second code clone detection model that can determine whether it is a semantic clone;
[0070] Use the third training data set to perform multiple direct - preference optimizations on the second code clone detection model to obtain a trained code clone detection model.
[0071] Among them, by implementing this implementation method, the code semantic understanding ability is improved through abstract - enhancement fine - tuning, the generalization ability is enhanced by fusing multi - modal features in hybrid - abstract clone training, and finally, the accurate adaptation of the model to the actual application scenario is achieved through direct - preference optimization, effectively solving problems such as insufficient generalization ability, difficulty in fusing multi - modal features, and poor business adaptability of traditional code clone detection models, and significantly reducing the model development cost while ensuring the detection accuracy.
[0072] As an optional implementation method, the method of using the first training data set to perform one abstract - enhancement fine - tuning on the initial code clone detection model to obtain a first code clone detection model that can analyze the code function and determine whether it is a semantic clone may include the following steps:
[0073] Obtain first training data from a first training dataset; wherein, the first training data includes a first training code segment and a second training code segment;
[0074] Obtain first standard function description information of the first training code segment, second standard function description information of the second training code segment, and a first true code clone label; wherein, the first true code clone label indicates whether the first standard function description information and the second standard function description information are actually the same;
[0075] Input the first training code segment, the second training code segment, and a detection instruction into the initial code clone detection model to obtain first function description information of the first training code segment, second function description information of the second training code segment, and a first code clone judgment result output by the initial code clone detection model; wherein, the detection instruction is used to enable the initial code clone detection model to perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same;
[0076] Use a loss function to calculate the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first true code clone label to obtain a first loss value;
[0077] If the first loss value is greater than or equal to a first preset loss threshold, fine-tune the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein, the adjusted initial code clone detection model obtained in any one summary enhancement fine-tuning operation is used as the initial code clone detection model in the next summary enhancement fine-tuning operation;
[0078] If the first loss value is less than the first preset loss threshold, determine the initial code clone detection model in this summary enhancement fine-tuning operation as a first code clone detection model that can analyze code functions and judge whether there is a semantic clone based on the code functions.
[0079] Among them, implementing this implementation method provides a rich and accurate data basis for model training by obtaining the first training data containing the first training code segment and the second training code segment from the first training dataset, and obtaining the corresponding standard function description information and true code clone labels, which helps the model learn the true features of the code segment function description and clone relationship. Secondly, multiple abstract augmentation fine-tuning operations combined with the calculation of the loss function can continuously evaluate the difference between the model output and the actual situation. By continuously adjusting the model parameters, the model gradually approaches the optimal state, improving the accuracy and reliability of the model for code clone detection. When the first loss value of the model is less than the first preset loss threshold, the fine-tuning is stopped, ensuring that the model achieves good performance within a reasonable error range and avoiding overtraining or under-training.
[0080] In the embodiments of this application, the training data in the first training dataset can be divided into two categories. One is the dataset based on programming competitions, such as GoogleCodeJam and OJClone, etc.; the other is the dataset based on real projects, such as BigCloneBench and GPTClone. To effectively train the model, two datasets are preferably selected: GoogleCodeJam and GPTClone. The reason for choosing them is that the code in GoogleCodeJam is more complex than that in OJClone, which helps to improve the effect of model training. At the same time, we maintain a uniform distribution in different types of codes, which has been proven to be very beneficial for both model training and evaluation. In addition, the data quality of GPTClone is better than that of BigCloneBench. According to previous research, there is some low-quality data in BigCloneBench, which will affect the training and evaluation of the model.
[0081] We introduced 3 new programming languages (C++, Python, Php) by expanding the GoogleCodeJam2 dataset. In addition, to effectively enhance the clone detection ability of the model, we used a large language model (LLM) to generate short code summaries for the dataset and incorporated them into subsequent model training (see Section III-C1). Finally, our dataset contains 4 programming languages (Java, C++, Python, Php), with a total of 37,364 code segments. Since GPTCloneBench only contains real clone pairs, we use it as the evaluation dataset.
[0082] In the embodiments of this application, the first training data can be randomly obtained from the first training dataset.
[0083] In the embodiments of this application, the loss function can be:
[0084]
[0085] Among them, for each word in the sequence, the loss function calculates the probability of the next word appearing. L represents the length of the sequence (the sequence includes the first functional description information, the second functional description information, and the first code clone judgment result), and D represents the data set (including the first standard functional description information, the second standard functional description information, and the first true code clone label of the first training data). represents the initial code clone detection model. represents the i-th word in the sequence.
[0086] In the embodiments of the present application, through abstract-enhanced fine-tuning, a first code clone detection model that can simultaneously analyze code functions and detect semantic clones based on these functions is constructed. However, it can be seen from the loss function that during the training process of the initial code clone detection model, the code clone judgment result only accounts for a small part of the loss calculation because the generation part includes code function descriptions. This results in limited improvement in code clone detection accuracy. To solve this problem and ensure that the model output conforms to the expected format, we designed a second-stage hybrid abstract clone fine-tuning method.
[0087] Optionally, the construction method of the second training data set may include the following steps:
[0088] Construct a second training data set; among them, the second initial training data of two data types are included in the second training data set, and the two data types include pure code type and code-function type; the second initial training data of the pure code type includes a third training code segment and a fourth training code segment; the second initial training data of the code-function type includes a third training code segment, a fourth training code segment, the third standard functional description information of the third training code segment, and the fourth standard functional description information of the fourth training code segment.
[0089] Among them, by implementing this implementation method, constructing the second training data set by mixing pure code type and code-function type data can dynamically balance the multi-modal feature input of model training, effectively enhance the adaptability of the code clone detection model to complex scenarios. Specifically, the pure code type data can strengthen the model's ability to recognize code structure features, and the code-function type data can improve the model's understanding depth of semantic similarity by integrating functional description information. The random combination of the two not only avoids the limitations of a single data type but also stimulates the model's feature generalization ability through data diversity, laying a solid multi-modal feature fusion foundation for the subsequent hybrid abstract clone fine-tuning stage, and significantly improving the clone detection accuracy and robustness of the model in a real complex code environment.
[0090] In the embodiment of the present application, the number of the second initial training data of the pure code type included in the second training data set is the same as the number of the second initial training data of the code - function type.
[0091] As an optional implementation manner, the method for obtaining the second code clone detection model capable of judging semantic cloning by performing a single hybrid summary clone fine - tuning on the first code clone detection model using the second training data set may include the following steps:
[0092] Determine a second initial training data randomly obtained from the second training data set as the second training data;
[0093] Obtain a second true code clone label corresponding to the third training code segment and the fourth training code segment in the second training data; wherein, the second true code clone label indicates whether the function information of the third training code segment and the fourth training code segment is actually the same;
[0094] Input the second training data into the first code clone detection model to obtain a second code clone judgment result output by the first code clone detection model; wherein, the second code clone judgment result indicates whether the function information of the third training code segment and the fourth training code segment is the same;
[0095] Calculate the second loss value by using the loss function for the second code clone judgment result and the second true code clone label;
[0096] If the second loss value is greater than or equal to the second preset loss threshold, fine - tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein, the adjusted first code clone detection model obtained in any single hybrid summary clone fine - tuning operation serves as the first code clone detection model in the next summary enhancement fine - tuning operation;
[0097] If the second loss value is less than the second preset loss threshold, determine the first code clone detection model in the current hybrid summary clone fine - tuning operation as the second code clone detection model capable of judging semantic cloning.
[0098] Among them, by implementing this implementation method, through dynamically integrating the pure code structure features and the multi-modal data of code-function descriptions, the understanding ability of the model for the dual features of code semantics and structure is effectively enhanced. During the iterative optimization process, the model can not only capture the cloning relationships at the code syntax level but also deeply explore the similarity of functional semantics, significantly improving the accuracy and generalization ability of code clone detection in complex scenarios. The mechanism of randomly selecting training data avoids the limitations of a single data type. Combined with the supervised training of the second true code clone labels, it ensures that the model continuously calibrates the judgment logic of code functional similarity during the optimization process and finally stops training when the preset loss threshold is reached, which not only guarantees the model performance but also prevents overfitting, laying a solid foundation for the robust performance of the subsequent model in actual engineering code detection.
[0099] Optionally, the third standard function description information of the third training code segment and the fourth standard function description information of the fourth training code segment included in the second initial training data of the code-function type can be generated by the first code clone detection model that has completed the abstract enhancement fine-tuning.
[0100] To retain the instruction-following ability of the model, in the hybrid abstract clone fine-tuning at this stage, the second initial training data of the pure code type and the code-function type are mixed in a 1:1 ratio. In this stage, the same loss function as in the abstract enhancement fine-tuning stage is continued to calculate the second loss value of the second code clone judgment result output by the first code clone detection model. Since the output of this stage only contains the code clone judgment result, the model can focus more on improving the accuracy of clone detection.
[0101] As an optional implementation method, the way of using the third training dataset to directly optimize the second code clone detection model once to obtain the trained code clone detection model may include the following steps:
[0102] Obtain the third training data from the third training dataset; wherein, the third training data includes the fifth training code segment and the sixth training code segment;
[0103] Obtain the third true code clone labels corresponding to the fifth training code segment and the sixth training code segment; wherein, the third true code clone labels indicate whether the functional information of the fifth training code segment and the sixth training code segment is actually the same;
[0104] Input the third training data into the second code clone detection model to obtain the third code clone judgment result output by the second code clone detection model; wherein, the third code clone judgment result indicates whether the functional information of the fifth training code segment and the sixth training code segment is the same;
[0105] Compare the third code clone judgment result with the third true code clone label to obtain a comparison result; wherein, the comparison result indicates whether the third code clone judgment result is the same as the third true code clone label;
[0106] Use the target loss function to calculate the third code clone judgment result, the third true code clone label, and the comparison result to obtain a third loss value;
[0107] If the third loss value is greater than or equal to the third preset loss threshold, directly optimize the second code clone detection model to obtain an optimized second code clone detection model; wherein, the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization;
[0108] If the third loss value is less than the third preset loss threshold, determine the second code clone detection model in this direct preference optimization as the trained code clone detection model.
[0109] Among them, by implementing this implementation method, by inputting the third training data into the second code clone detection model and obtaining the comparison result, the judgment deviation of the model in the actual application scenario can be accurately captured, and the evaluation logic of the model for code function similarity can be effectively calibrated. Specifically, this optimization process dynamically adjusts the model parameters, enabling the model to give priority to the code features where the true code clone label is inconsistent with the prediction result, significantly improving the detection accuracy of the model in complex code scenarios. At the same time, the third training data selection mechanism based on the third training data set ensures the representativeness and consistency of the training data, avoiding the problem of model performance degradation caused by data distribution differences. This optimization strategy that combines manual annotation preferences with model self-learning not only strengthens the adaptability of the model to actual business requirements but also realizes the continuous iteration of the model performance through the real feedback loop mechanism, providing key technical support for the finally trained code clone detection model to play a stable and reliable detection effect in the actual engineering environment.
[0110] In the embodiments of this application, the direct preference optimization of the second code clone detection model can be implemented using the target loss function of the Direct Preference Optimization (DPO) algorithm. The target loss function of the DPO algorithm is:
[0111]
[0112] Among them, And represents the output of the second code clone detection model under the input of x or cumulative probability of, is the language model policy, is the baseline reference policy, and β is a parameter that controls the deviation between the two. In this way, the DPO algorithm can use alternative parameterization to fit the implicit reward, and its optimal policy is .
[0113] Generally speaking, the working principle of the DPO algorithm is: by increasing the log probability of preferred samples through the left part of the objective loss function, and at the same time reducing the log probability of non-preferred sample responses through the right part. The DPO algorithm regards human preference alignment as a classification problem, which makes it particularly suitable for the fine-tuning task in this article. By regarding preference alignment as a classification task, the DPO algorithm enables us to optimize the second code clone detection model to better align with the true clone judgment results. In this study, this process involves accurately detecting code clones.
[0114] Specifically, for the input third training data, the DPO algorithm is used to train the second code clone detection model. This goal is achieved without introducing significant fluctuations in the output, thus avoiding the situation where the model fails to accurately judge whether two code segments are code clones after correctly analyzing them.
[0115] Implementing the above steps 101 to 104 can make full use of the capabilities of the large language model, avoid the model's over-reliance on the surface features of code syntax, and then accurately judge the functional similarity of code based on the semantic analysis of code segments, thereby improving the accuracy of code clone detection. In addition, this application can also solve problems such as insufficient generalization ability, difficulty in multi-modal feature fusion, and poor business adaptability of traditional code clone detection models, significantly reducing the model development cost while ensuring the detection accuracy. In addition, this application can also improve the accuracy and reliability of the model for code clone detection. In addition, this application can also improve the clone detection accuracy and robustness of the model in a real and complex code environment. In addition, this application can prevent overfitting while ensuring the model performance. In addition, this application can strengthen the model's adaptability to actual business needs and also achieve continuous iteration of the model performance through a real feedback loop mechanism.
[0116] Based on the same inventive concept, the embodiment of this application also provides a code clone detection device for implementing the above-mentioned code clone detection method. The solution provided by this device to solve problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the code clone detection device provided below can refer to the limitations on the code clone detection method in the above text and will not be repeated here.
[0117] In an exemplary embodiment, as Figure 2 shown, a code clone detection device is provided, including:
[0118] An acquisition unit 201, configured to acquire a code library to be detected; wherein, the code library to be detected contains a plurality of code segments to be detected;
[0119] A first determination unit 202, configured to determine the hash value of each code segment to be detected;
[0120] A second determination unit 203, configured to determine a set of high-similarity code segment pairs based on the hash value of each code segment to be detected; wherein, the set of high-similarity code segment pairs contains at least one set of high-similarity code segment pairs, each high-similarity code segment pair contains two target code segments to be detected, and the similarity of the hash values of the two target code segments contained in a set of high-similarity code segment pairs is greater than a preset similarity;
[0121] An input unit 204, configured to input the set of high-similarity code segment pairs into a pre-trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein, the code clone detection result includes clone code pairs, and the two clone code segments included in the clone code pairs implement the same function.
[0122] In the embodiment of the present application, the training method of the code clone detection model based on the large language model is specifically:
[0123] Construct an initial code clone detection model based on the large language model;
[0124] Use a first training data set to perform multiple abstract enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and judging whether it is semantic cloning based on code functions;
[0125] Use a second training data set to perform multiple mixed abstract clone fine-tuning on the first code clone detection model to obtain a second code clone detection model capable of judging whether it is semantic cloning;
[0126] Use a third training data set to perform multiple direct preference optimizations on the second code clone detection model to obtain a trained code clone detection model.
[0127] Among them, implementing this implementation method enhances the code semantic understanding ability through abstract augmentation fine-tuning, mixes abstract cloning training to fuse multi-modal features to enhance generalization, and finally achieves precise adaptation of the model to the actual application scenario through direct preference optimization, effectively solving the problems of insufficient generalization ability, difficult multi-modal feature fusion, and poor business adaptability of traditional code clone detection models, while ensuring the detection accuracy and significantly reducing the model development cost.
[0128] As an optional implementation method, the method of using the first training data set to perform one-time abstract augmentation fine-tuning on the initial code clone detection model to obtain the first code clone detection model capable of analyzing the code function and judging whether it is a semantic clone may include the following steps:
[0129] Obtain the first training data from the first training data set; wherein, the first training data includes the first training code segment and the second training code segment;
[0130] Obtain the first standard function description information of the first training code segment, the second standard function description information of the second training code segment, and the first true code clone label; wherein, the first true code clone label indicates whether the first standard function description information and the second standard function description information are actually the same;
[0131] Input the first training code segment, the second training code segment, and the detection instruction into the initial code clone detection model to obtain the first function description information of the first training code segment, the second function description information of the second training code segment, and the first code clone judgment result output by the initial code clone detection model; wherein, the detection instruction is used to make the initial code clone detection model perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same;
[0132] Use a loss function to calculate the first function description information, the second function description information, the first code clone judgment result, the first standard function description information, the second standard function description information, and the first true code clone label to obtain a first loss value;
[0133] If the first loss value is greater than or equal to the first preset loss threshold, fine-tune the parameters of the initial code clone detection model to obtain the adjusted initial code clone detection model; wherein, the adjusted initial code clone detection model obtained in any one-time abstract augmentation fine-tuning operation is used as the initial code clone detection model in the next abstract augmentation fine-tuning operation;
[0134] If the first loss value is less than the first preset loss threshold, the initial code cloning detection model in the current abstract enhancement fine-tuning operation is determined as the first code cloning detection model that can analyze the code function and determine whether it is a semantic clone based on the code function.
[0135] Among them, implementing this implementation method, by obtaining the first training data containing the first training code segment and the second training code segment from the first training dataset, and obtaining the corresponding standard function description information and true code clone label, it provides a rich and accurate data basis for model training, which helps the model learn the true features of the code segment function description and clone relationship. Secondly, multiple abstract enhancement fine-tuning operations combined with the calculation of the loss function can continuously evaluate the difference between the model output and the real situation. By continuously adjusting the model parameters, the model gradually approaches the optimal state, improving the accuracy and reliability of the model for code clone detection. When the first loss value of the model is less than the first preset loss threshold, the fine-tuning is stopped, which ensures that the model achieves better performance within a reasonable error range and avoids overtraining or under-training.
[0136] Optionally, the construction method of the second training dataset may include the following steps:
[0137] Construct a second training dataset; wherein, the second training dataset contains second initial training data of two data types, and the two data types include pure code type and code-function type; the second initial training data of the pure code type contains a third training code segment and a fourth training code segment; the second initial training data of the code-function type contains a third training code segment, a fourth training code segment, the third standard function description information of the third training code segment, and the fourth standard function description information of the fourth training code segment.
[0138] Among them, implementing this implementation method, by constructing the second training dataset by mixing pure code type and code-function type data, it can dynamically balance the multi-modal feature input of model training and effectively enhance the adaptability of the code clone detection model to complex scenarios. Specifically, the pure code type data can strengthen the model's ability to recognize code structure features, and the code-function type data can improve the model's understanding depth of semantic similarity by integrating function description information. The random combination of the two not only avoids the limitations of a single data type, but also stimulates the model's feature generalization ability through data diversity, laying a solid multi-modal feature fusion foundation for the subsequent mixed abstract clone fine-tuning stage, and significantly improving the clone detection accuracy and robustness of the model in a real complex code environment.
[0139] As an alternative implementation, the method of using the second training dataset to perform a single hybrid abstract clone fine-tuning on the first code clone detection model to obtain a second code clone detection model capable of judging semantic cloning may include the following steps:
[0140] Determine a second initial training data randomly obtained from the second training dataset as the second training data;
[0141] Obtain a second true code clone label corresponding to the third training code segment and the fourth training code segment in the second training data; wherein, the second true code clone label indicates whether the function information of the third training code segment and the fourth training code segment is actually the same;
[0142] Input the second training data into the first code clone detection model to obtain a second code clone judgment result output by the first code clone detection model; wherein, the second code clone judgment result indicates whether the function information of the third training code segment and the fourth training code segment is the same;
[0143] Use the loss function to calculate the second code clone judgment result and the second true code clone label to obtain a second loss value;
[0144] If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein, the adjusted first code clone detection model obtained in any single hybrid abstract clone fine-tuning operation is used as the first code clone detection model in the next abstract enhancement fine-tuning operation;
[0145] If the second loss value is less than the second preset loss threshold, determine the first code clone detection model in this hybrid abstract clone fine-tuning operation as the second code clone detection model capable of judging semantic cloning.
[0146] Among them, implementing this implementation mode effectively enhances the model's understanding ability of the dual characteristics of code semantics and structure by dynamically fusing pure code structure features and multi-modal data of code-function descriptions, enabling the model to capture clone relationships at the code syntax level and deeply explore the similarity of functional semantics during the iterative optimization process, significantly improving the accuracy and generalization ability of code clone detection in complex scenarios. The mechanism of randomly selecting training data avoids the limitations of a single data type, and combined with the supervised training of the second true code clone label, ensures that the model continuously calibrates the judgment logic of code function similarity during the optimization process, and finally stops training when the preset loss threshold is reached, which not only guarantees the model performance but also prevents overfitting, laying a solid foundation for the robust performance of the subsequent model in actual engineering code detection.
[0147] As an alternative implementation, the method for directly optimizing the second code clone detection model once using the third training dataset to obtain the trained code clone detection model may include the following steps:
[0148] Obtain third training data from the third training dataset; wherein, the third training data includes a fifth training code segment and a sixth training code segment;
[0149] Obtain third true code clone labels corresponding to the fifth training code segment and the sixth training code segment; wherein, the third true code clone labels indicate whether the functional information of the fifth training code segment and the sixth training code segment is actually the same;
[0150] Input the third training data into the second code clone detection model to obtain a third code clone judgment result output by the second code clone detection model; wherein, the third code clone judgment result indicates whether the functional information of the fifth training code segment and the sixth training code segment is the same;
[0151] Compare the third code clone judgment result with the third true code clone label to obtain a comparison result; wherein, the comparison result indicates whether the third code clone judgment result is the same as the third true code clone label;
[0152] Calculate a third loss value using a target loss function for the third code clone judgment result, the third true code clone label, and the comparison result;
[0153] If the third loss value is greater than or equal to a third preset loss threshold, directly optimize the second code clone detection model to obtain an optimized second code clone detection model; wherein, the optimized second code clone detection model obtained in any direct optimization is used as the second code clone detection model in the next direct optimization;
[0154] If the third loss value is less than the third preset loss threshold, determine the second code clone detection model in this direct optimization as the trained code clone detection model.
[0155] Among them, by implementing this implementation method, by inputting the third training data into the second code clone detection model and obtaining the comparison result, it is possible to accurately capture the judgment deviation of the model in the actual application scenario, and effectively calibrate the evaluation logic of the model for code function similarity. Specifically, this optimization process dynamically adjusts the model parameters, enabling the model to prioritize the code features where the true code clone label is inconsistent with the prediction result, significantly improving the detection accuracy of the model in complex code scenarios. At the same time, the third training data selection mechanism based on the third training data set ensures the representativeness and consistency of the training data, avoiding the problem of model performance degradation caused by data distribution differences. This optimization strategy that combines manual annotation preferences with model self-learning not only strengthens the model's adaptability to actual business needs but also realizes the continuous iteration of model performance through the real feedback loop mechanism, providing key technical support for the finally trained code clone detection model to play a stable and reliable detection efficiency in the actual engineering environment.
[0156] Implementing the above implementation method can make full use of the capabilities of the large language model, avoid the model's over-reliance on the surface features of code syntax, and then accurately judge the code function similarity based on the semantic analysis of code fragments, thereby improving the accuracy of code clone detection. In addition, this application can also solve problems such as the insufficient generalization ability, difficulty in multi-modal feature fusion, and poor business adaptability of traditional code clone detection models, significantly reducing the model development cost while ensuring the detection accuracy. In addition, this application can also improve the accuracy and reliability of the model for code clone detection. In addition, this application can also improve the clone detection accuracy and robustness of the model in a real and complex code environment. In addition, this application can also prevent overfitting while ensuring the model performance. In addition, this application can also strengthen the model's adaptability to actual business needs and realize the continuous iteration of model performance through the real feedback loop mechanism.
[0157] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal, and its internal structure diagram can be as Figure 3As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store code clone detection data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a code clone detection method.
[0158] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0159] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.
[0160] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0161] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0162] In an exemplary embodiment, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the steps in the above method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0163] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-chip, system chip, chip system, or system-on-a-chip, etc.
[0164] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0165] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0166] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0167] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered to be within the scope described in this specification.
[0168] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A code clone detection method, characterized in that, The described code clone detection method includes: Obtain the code library to be detected; wherein, the code library to be detected contains multiple code snippets to be detected; Determine the hash value of each code snippet to be detected; Based on the hash values of each code snippet to be detected, determine a set of highly similar code snippet pairs; wherein, the set of highly similar code snippet pairs contains at least one set of highly similar code snippet pairs, and each highly similar code snippet pair contains two target code snippets to be detected, and the similarity of the hash values of the two target code snippets contained in a set of highly similar code snippet pairs is greater than a preset similarity; Input the set of highly similar code snippet pairs into a pre-trained code clone detection model based on a large language model to obtain the code clone detection result output by the code clone detection model; wherein, the code clone detection result includes clone code pairs, and the two clone code snippets contained in the clone code pair implement the same function; Wherein, the training method of the code clone detection model based on the large language model is specifically: Construct an initial code clone detection model based on the large language model; Use the first training data set to perform multiple abstract enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and judging whether it is a semantic clone based on the code functions; Use the second training data set to perform multiple hybrid abstract clone fine-tuning on the first code clone detection model to obtain a second code clone detection model capable of judging whether it is a semantic clone; Use the third training data set to perform multiple direct preference optimizations on the second code clone detection model to obtain a trained code clone detection model.
2. The code clone detection method according to claim 1, characterized in that The step of using the first training data set to perform one abstract enhancement fine-tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and judging whether it is a semantic clone based on the code functions specifically includes: Obtain the first training data from the first training data set; wherein, the first training data includes a first training code segment and a second training code segment; Obtain the first standard function description information of the first training code segment, the second standard function description information of the second training code segment, and the first true code clone label; wherein, the first true code clone label indicates whether the first standard function description information and the second standard function description information are actually the same; Input the first training code segment, the second training code segment, and the detection instruction into the initial code clone detection model to obtain the first function description information of the first training code segment, the second function description information of the second training code segment, and the first code clone judgment result output by the initial code clone detection model; wherein, the detection instruction is used to make the initial code clone detection model perform code clone detection on the first training code segment and the second training code segment, and the first code clone judgment result indicates whether the first function description information and the second function description information are the same; Calculate a first loss value by using a loss function for the first functional description information, the second functional description information, the first code clone judgment result, the first standard functional description information, the second standard functional description information, and the first true code clone label; If the first loss value is greater than or equal to a first preset loss threshold, fine-tune the parameters of the initial code clone detection model to obtain an adjusted initial code clone detection model; wherein, the adjusted initial code clone detection model obtained in any abstract enhancement fine-tuning operation is used as the initial code clone detection model in the next abstract enhancement fine-tuning operation; If the first loss value is less than the first preset loss threshold, determine the initial code clone detection model in the current abstract enhancement fine-tuning operation as a first code clone detection model capable of analyzing code functions and judging semantic cloning based on code functions.
3. The code clone detection method according to claim 2, wherein The code clone detection method further includes: Construct a second training data set; wherein, the second training data set contains second initial training data of two data types, and the two data types include pure code type and code-function type; the second initial training data of the pure code type contains a third training code segment and a fourth training code segment; the second initial training data of the code-function type contains a third training code segment, a fourth training code segment, the third standard functional description information of the third training code segment, and the fourth standard functional description information of the fourth training code segment.
4. The code clone detection method according to claim 3, characterized in that The use of the second training data set to perform a mixed abstract clone fine-tuning on the first code clone detection model to obtain a second code clone detection model capable of judging semantic cloning specifically includes: Determine a second initial training data randomly obtained from the second training data set as the second training data; Obtain a second true code clone label corresponding to the third training code segment and the fourth training code segment in the second training data; wherein, the second true code clone label indicates whether the functional information of the third training code segment and the fourth training code segment is actually the same; Input the second training data into the first code clone detection model to obtain a second code clone judgment result output by the first code clone detection model; wherein, the second code clone judgment result indicates whether the functional information of the third training code segment and the fourth training code segment is the same; Calculate a second loss value by using the loss function for the second code clone judgment result and the second true code clone label; If the second loss value is greater than or equal to a second preset loss threshold, fine-tune the parameters of the first code clone detection model to obtain an adjusted first code clone detection model; wherein, the adjusted first code clone detection model obtained in any mixed abstract clone fine-tuning operation is used as the first code clone detection model in the next abstract enhancement fine-tuning operation; If the second loss value is less than the second preset loss threshold, then the first code clone detection model in the current hybrid digest cloning fine-tuning operation is determined as the second code clone detection model capable of judging semantic cloning.
5. The code clone detection method according to any one of claims 2 to 4, characterized in that, The step of performing one direct preference optimization on the second code clone detection model using the third training dataset to obtain a trained code clone detection model specifically includes: Obtain third training data from the third training dataset; wherein, the third training data includes a fifth training code segment and a sixth training code segment; Obtain third true code clone labels corresponding to the fifth training code segment and the sixth training code segment; wherein, the third true code clone labels indicate whether the functional information of the fifth training code segment and the sixth training code segment is actually the same; Input the third training data into the second code clone detection model to obtain a third code clone judgment result output by the second code clone detection model; wherein, the third code clone judgment result indicates whether the functional information of the fifth training code segment and the sixth training code segment is the same; Compare the third code clone judgment result with the third true code clone label to obtain a comparison result; wherein, the comparison result indicates whether the third code clone judgment result is the same as the third true code clone label; Use a target loss function to calculate the third code clone judgment result, the third true code clone label, and the comparison result to obtain a third loss value; If the third loss value is greater than or equal to the third preset loss threshold, then perform direct preference optimization on the second code clone detection model to obtain an optimized second code clone detection model; wherein, the optimized second code clone detection model obtained in any direct preference optimization is used as the second code clone detection model in the next direct preference optimization; If the third loss value is less than the third preset loss threshold, then determine the second code clone detection model in the current direct preference optimization as the trained code clone detection model.
6. A code clone detection device, characterized in that, The code clone detection device includes: An acquisition unit, configured to acquire a code library to be detected; wherein, the code library to be detected contains multiple code segments to be detected; A first determination unit, configured to determine the hash value of each code segment to be detected; A second determination unit, configured to determine a set of high similarity code segment pairs based on the hash values of each code segment to be detected; wherein, the set of high similarity code segment pairs contains at least one set of high similarity code segment pairs, each set of high similarity code segment pairs contains two target code segments to be detected, and the similarity of the hash values of the two target code segments to be detected included in one set of high similarity code segment pairs is greater than a preset similarity; An input unit for inputting the set of high - similarity code snippet pairs into a pre - trained code clone detection model based on a large language model to obtain a code clone detection result output by the code clone detection model; wherein, the code clone detection result includes clone code pairs, and the two clone code snippets included in the clone code pair implement the same function. Wherein, the training method of the code clone detection model based on the large language model is specifically as follows: Construct an initial code clone detection model based on the large language model; Use a first training data set to perform multiple abstract enhancement fine - tuning on the initial code clone detection model to obtain a first code clone detection model capable of analyzing code functions and judging whether it is a semantic clone based on code functions; Use a second training data set to perform multiple hybrid abstract clone fine - tuning on the first code clone detection model to obtain a second code clone detection model capable of judging whether it is a semantic clone; Use a third training data set to perform multiple direct preference optimizations on the second code clone detection model to obtain a trained code clone detection model.
7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the code clone detection method according to any one of claims 1 - 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the code clone detection method according to any one of claims 1 - 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the code clone detection method according to any one of claims 1 - 5.
Citation Information
Patent Citations
Repeated code detection method and device, terminal equipment and storage medium
CN117573519A