Code clone detection method, system and device, medium and product

By training the cloned code detection model of multi-task learning, using contrast learning and translation enhancement loss function, it solves the problem that traditional methods are difficult to recognize code cloning with different syntax but the same functions, and achieves higher code cloning detection accuracy.

CN119960825AActive Publication Date: 2025-05-09BEIJING BEIDA SOFTWARE ENG DEV CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510449628.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Traditional code cloning detection methods are difficult to effectively identify Type-4 code clones with different syntax but the same functions, resulting in low detection accuracy.

Method used

By training a multi-task learning cloned code detection model, using comparative learning loss function and translation enhancement learning loss function, the semantic features of the code segment and capture fine-grained semantic relationships, thereby improving the accuracy of the code representation vector.

Benefits of technology

Improve the accuracy of code cloning detection, ensure that the model can accurately translate the code segment to be detected into a code representation vector that is equivalent to its semantics, and effectively identify various code cloning types including Type-4.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960825A_ABST
    Figure CN119960825A_ABST
Patent Text Reader

Abstract

The invention discloses a code clone detection method, system and device, a medium and a product, and relates to the field of software engineering.The method comprises the steps that to-be-detected code segments are obtained from a to-be-detected code warehouse, and all the to-be-detected code segments are input into a trained clone code detection model; outputting a first code representation vector corresponding to the code segment to be detected; and randomly selecting two first code representation vectors as code pairs, and marking the code pairs of which the semantic distance is smaller than a preset threshold value as clone codes. The code clone detection method and device can improve the accuracy of code clone detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of software engineering, and in particular to a code clone detection method, system, device, medium and product. Background Art

[0002] To speed up the development process, developers usually rely on frameworks to provide infrastructure and common functions to avoid duplication of work. However, this practice leads to the phenomenon of code cloning, that is, the appearance of identical or similar code segments in the software system.

[0003] Studies have shown that a large number of code clones can have a negative impact on software systems. For example, reusing code segments containing undetected defects may cause defects to propagate, thereby reducing the stability of information systems. In addition, if code clones in a system are not properly managed, the code base may expand unnecessarily, resulting in code redundancy and increased maintenance costs. Therefore, identifying code clones in software systems is an important research topic in the field of software engineering.

[0004] Code clones can be divided into four types according to the degree of similarity between code snippets: Type-1: Code pairs that differ only in comment and whitespace characters; Type-2: In addition to the differences in Type-1, there are also code pairs with changed variable names and constant values; Type-3: In addition to the differences between Type-1 and Type-2, code pairs also include added or deleted statements; Type-4: Code pairs with different syntax but identical functionality.

[0005] Traditional code clone detection techniques based on text similarity perform well in dealing with Type-1 to Type-3 code clones, as these types of clones have high similarity at the text level. However, when it comes to Type-4 clones, due to their grammatical differences, traditional code clone detection methods have difficulty effectively identifying these code segments that essentially have the same functions but different forms. Summary of the invention

[0006] The purpose of this application is to provide a code clone detection method, system, device, medium and product, which can improve the accuracy of code clone detection.

[0007] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a code clone detection method, comprising: Acquire the code segments to be detected from the code repository to be detected, input all the code segments to be detected into the trained clone code detection model, and output the first code representation vector corresponding to the code segments to be detected; arbitrarily selecting two of the first code representation vectors as a code pair, and marking the code pair whose semantic distance is less than a preset threshold as a clone code; The clone code detection model is trained in the following way: Randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, each pair of the source code segment and the first target code segment that have the unique correspondence relationship is a clone code of each other, and the unique correspondence relationship refers to semantic equivalence or grammatical differences but the same function; All the source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, wherein the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

[0008] In a second aspect, the present application provides a code clone detection system, comprising: An acquisition module, configured to acquire code segments to be detected from a code repository to be detected, input all the code segments to be detected into a trained clone code detection model, and output a first code representation vector corresponding to the code segments to be detected; a marking module, configured to arbitrarily select two of the first code representation vectors as a code pair, and mark the code pair whose semantic distance is less than a preset threshold as a clone code; The marking module is also used to train the clone code detection model in the following manner: Randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, each pair of the source code segment and the first target code segment that have the unique correspondence relationship is a clone code of each other, and the unique correspondence relationship refers to semantic equivalence or grammatical differences but the same function; All the source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, wherein the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

[0009] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any one of the above-described code clone detection methods.

[0010] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the above-described code clone detection methods.

[0011] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned code clone detection methods.

[0012] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides a code clone detection method, system, device, medium and product. The method obtains a code segment to be detected from a code repository to be detected, inputs all the code segments to be detected into a trained clone code detection model, and outputs a first code representation vector corresponding to the code segment to be detected; arbitrarily selects two first code representation vectors as code pairs, and marks the code pairs whose semantic distance is less than a preset threshold as clone codes; wherein the clone code detection model is trained in the following manner: randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments with a unique corresponding relationship, each pair of source code segments with a unique corresponding relationship and the first target code segment are clone codes of each other, and the unique corresponding relationship refers to semantic equivalence or different syntax but the same function; all source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, and the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function. This application combines the contrastive learning loss function with the translation enhancement learning loss function. The contrastive learning loss function is used to learn the semantic features of the source code segment, so that the cloned code is closer in the embedding space. The translation enhancement learning loss function is used to learn to translate each source code segment into its corresponding cloned code, capturing the fine-grained semantic relationship between the code segments. By combining the contrastive learning loss function with the translation enhancement learning loss function, the accuracy of the first code representation vector is improved, thereby making the code clone detection result more accurate. Ensure that the clone code detection model can accurately translate the code segment to be detected into the first code representation vector that is semantically equivalent to it. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0014] Figure 1 This is an application environment diagram of a code clone detection method in an embodiment of the present application; Figure 2 A flowchart of a code clone detection method provided in one embodiment of the present application; Figure 3 A flowchart of a code clone detection method provided in another embodiment of the present application; Figure 4 A schematic diagram of functional modules of a code clone detection system provided in one embodiment of the present application; Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0016] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0017] The code clone detection method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the source code segment to be processed and the first target code segment to the server 104. After the server 104 receives the source code segment to be processed and the first target code segment, for the source code segment to be processed and the first target code segment, the server 104 inputs all the source code segments and the first target code segment into the initial clone code detection model, and trains the initial clone code detection model through the multi-task learning loss function to obtain the clone code detection model. The server 104 obtains the code segment to be detected from the code repository to be detected, and inputs all the code segments to be detected into the clone code detection model, and outputs the first code representation vector corresponding to the code segment to be detected; arbitrarily select two first code representation vectors as code pairs, and mark the code pairs whose semantic distance is less than a preset threshold as clone code. The server 104 can feedback the obtained clone code to the terminal 102. In addition, in some embodiments, the code clone detection method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the source code segment to be processed, or the server 104 can obtain the source code segment to be processed from the data storage system and process the source code segment to be processed.

[0018] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.

[0019] In an exemplary embodiment, Figure 2 As shown, a code clone detection method is provided, which is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in FIG. 1 is taken as an example to illustrate the method, which includes the following steps S201 to S204. Steps S201 to S204 correspond to Figure 3 The model fine-tuning process and model inference process in . Among them: In step S201, code segments to be detected are obtained from a code repository to be detected, and all the code segments to be detected are input into a trained clone code detection model, and a first code representation vector corresponding to the code segments to be detected is output.

[0020] In step S202, two of the first code representation vectors are arbitrarily selected as a code pair, and the code pair whose semantic distance is less than a preset threshold is marked as a clone code.

[0021] The clone code detection model is trained in the following way: In step S203, random sampling is performed from the code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, and each pair of source code segments and first target code segments that have a unique correspondence relationship are clone codes of each other.

[0022] Specifically, N pairs of clone codes are randomly found from the code clone detection dataset, one code segment in each pair of clone codes is used as the source code segment, and the other code segment is used as the first target code segment. The present disclosure selects the open source code clone detection dataset GoogleCodeJam. The GoogleCodeJam dataset contains codes in multiple programming languages ​​(such as Python, Java, C++, etc.), which can cover different programming styles and techniques. The code segments in the dataset are derived from real programming competition scenarios, reflecting the code cloning problems that programmers may encounter in actual development. The code segments in the dataset have been annotated as clone or non-clone relationships, which provides a reliable benchmark for subsequent model training and verification. In the GoogleCodeJam dataset, different answers to each question are regarded as Type-4 Clones. Type-4 code clones refer to code segments that are semantically similar but may have completely different grammatical structures. This type of clone detection is more difficult because it requires the model to have a deep understanding of the code semantics, not just grammatical matching.

[0023] In order to train the model, multiple sets of data need to be randomly sampled from the Google Code Jam dataset. The purpose of sampling is to build a dataset containing cloned and non-cloned code segments for contrastive learning of the model. Contrastive learning is a method of training a model through positive samples (clone code pairs) and negative samples (non-clone code pairs), which can effectively improve the model's ability to identify code clones.

[0024] The source code segments and target code segments have been annotated with clone labels or non-clone labels, which will be used for comparative learning of the model to help the model distinguish between clone code pairs and non-clone code pairs.

[0025] In step S204, all source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model. The multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

[0026] In one embodiment, step S204 corresponds to Figure 3 The multi-task learning process in step S204 of "training the initial clone code detection model through multi-task learning to obtain a clone code detection model" includes the following sub-steps S2041-S2045: S2041. The source code segment is encoded by an encoder, and a second code representation vector corresponding to the source code segment is output. Each source code segment and the corresponding second code representation vector are attached with the same label, and the label is used to characterize the function of the source code segment.

[0027] Specifically, the initial clone code detection model includes an encoder and a decoder, and the initial clone code detection model is built based on the pre-trained CodeT5p model. CodeT5p is an open code large-scale language model that uses an encoder-decoder architecture and can run in different modes, including encoder-only, decoder-only, and encoder-decoder modes. This architecture enables CodeT5p to support a wide range of code understanding and generation tasks. In the clone code detection task, the model built based on CodeT5p can use its powerful semantic representation capabilities to extract the semantic features of the code segment, thereby completing the code clone detection task more accurately.

[0028] The labels of source code segments are predefined and are used to characterize the functions or semantic categories of the code segments. For example, if the function of a source code segment is to implement a sorting algorithm, its label may be "sort"; if the function of a code segment is to implement data search, its label may be "search". These labels are annotated based on the functional characteristics of the code. After being processed by the encoder, source code segments 1-source code segment N will be converted into second code representation vectors 1-second code representation vectors N. The second code representation vector is a numerical representation of the semantic and structural features of the source code segment, which can capture the core information of the code. The role of the encoder is to convert the text form of the code into a vector form that the model can understand and process. The source code segment and its corresponding second code representation vector are given the same label to ensure that the two are semantically consistent. The label of the source code segment reflects its function, while the second code representation vector is a numerical representation of the semantic features of the source code segment. By giving the same label, the model can learn the direct correspondence between the second code representation vector and the function of the source code segment. This consistency enables the model to better understand the semantics of the code.

[0029] S2042: Pair the source code segments in pairs to obtain multiple code pairs, and determine whether the code pairs are clone code pairs based on the matching of the labels of the source code segments. If the labels match, mark the code pairs as clone code pairs; otherwise, mark the code pairs as non-clone code pairs.

[0030] Specifically, in a batch, for N source code segments, each source code segment Ci is processed by the encoder to obtain a second code representation vector Ei. This Ei is obtained by linear transformation from the [CLS] tag output of the encoder. Each source code segment also has a label. The extracted source code segments are paired to form multiple code pairs. For the case of n code segments, For each code pair, the label information of the two source code segments is compared. If the labels of the two source code segments completely match, the code pair is marked as a clone code pair, otherwise, it is marked as a non-clone code pair.

[0031] S2043. Calculate the contrastive learning loss of the clone code pair according to the contrastive learning loss function.

[0032] The contrastive learning loss function is expressed by the following formula: ; in, represents contrastive learning loss; N represents the number of source code segments in a training batch; and A second code representation vector representing the correspondence between any two source code segments; express The positive sample of source code segments forming clone code pairs; express and The cosine similarity between ; Used for calculation and its positive sample The similarity index of Used for calculation The sum of similarity indices with all other source code segments; represents a hyperparameter used to control the distribution, and its empirical value is set to 0.11; the goal of the contrastive learning loss function is to make each source code segment With its positive sample The similarity of has a higher relative importance among all code segments.

[0033] S2044. All second code representation vectors are decoded by a decoder, a second target code segment corresponding to the second code representation vector is output, and a translation enhancement loss of the second target code segment relative to the first target code segment is calculated by a translation enhancement learning loss function.

[0034] Specifically, the code translation task is processed using an encoder-decoder architecture. The source code segment Ci is input into the encoder, and the encoder outputs a hidden state, i.e., a second code representation vector, through encoding. The second code representation vector contains the semantic information and structural information of the input source code segment. The translation instruction is spliced ​​into the input of the decoder. This instruction can be "translate the code into Java", "translate the code into Python", etc., depending on the language of the second target code segment. The purpose of this instruction is to guide the decoder on how to decode the encoder's hidden state into the second target code segment.

[0035] The first target code segment is the original code segment, the second target code segment is the translated code segment, the first target code segment is the original code that the model attempts to translate into the second target code segment through the translation task, and the first target code segment is the correct answer or target summarized by the model learning process. The difference between the second target code segment and the first target code segment is minimized by the translation reinforcement learning loss function. The translation reinforcement learning loss function is expressed by the following formula: ; Wherein, N represents the number of first target code segments in a training batch; T represents the length of the first target code segment, that is, the total number of tokens in the first code segment; represents the token that the model tries to generate at time step t; indicates that a tag sequence has been generated in the first target code segment before time step t, represents the second code representation vector generated by the encoder, Indicates that given a generated sequence of tokens and the second code represents the vector Under the condition of probability.

[0036] The translation enhancement loss function is based on the standard causal language model design. The translation enhancement loss function calculates the token that the model attempts to generate at time step t (the predicted token). The actual tth token in the first target code segment The difference between. is not the actual token in the first target code segment. By minimizing the translation enhancement loss function, the model learns how to better represent the vector according to the input second target code. and the generated sequence , to predict the next token During the training process, the parameters of the model are adjusted through optimization algorithms, such as gradient descent, to minimize the loss function , the model parameters are updated according to the gradient to reduce the difference between the predicted second target code segment and the actual first target code segment. Through multiple iterative operations, the model gradually learns how to better replace the second target code segment with a code segment that is similar in semantics and structure to the first target code segment.

[0037] Since the loss function quantifies the gap between the model prediction and the actual result, the input data is propagated through the model once to calculate the model output. The gap between the predicted output and the actual output is calculated by the loss function to obtain the loss value. The gradient of the loss function with respect to the model parameters is calculated, that is, the partial derivative of the loss function with respect to each model parameter. These gradients are used to indicate the direction in which the loss function rises fastest in the parameter space. Therefore, in order to reduce the loss, the parameters should be updated in the opposite direction of the gradient. According to the calculated gradient, the model parameters are updated using an optimization algorithm, such as the gradient descent method. The optimization algorithm determines the step size and direction of the parameter update based on hyperparameters such as the gradient and the learning rate. Repeat the steps of forward propagation, loss calculation, backward propagation, and parameter update until the loss value converges to a lower level or reaches a preset number of iterations. In the above process of adjusting the model parameters through the loss value, the model parameters are continuously updated, the loss value gradually decreases, and the model performance gradually improves, and finally a model that can accurately predict is obtained.

[0038] Through the translation enhancement task, the model is forced to learn how to convert code snippets from the form of representation vectors to the form of target code, which helps to enhance the semantic and structural representation of code snippets. By splicing translation instructions into the input of the decoder, the model can learn to translate the same code representation into different or the same language, thereby supporting cross-language code clone detection.

[0039] It should be noted that the encoder converts the input source code segment into a fixed-length second code representation vector, which captures the semantic and structural information of the code and is independent of the programming language. In other words, no matter what language the input code is in, the encoder will extract the core meaning of the code. The decoder uses this vector to generate a second target code segment, which can be an equivalent code in a different language or the same language. For example, if the input code is Java, the target code can be Python, or still Java, but with a different style.

[0040] The translation instruction (such as "translate the code to Python") is spliced ​​into the input of the decoder to instruct the decoder how to decode the encoder's hidden state into the target code segment. This instruction tells the decoder which language of code to generate, but does not participate in the loss calculation and the encoding process of the encoder. The loss function only calculates the difference between the second target code segment generated by the decoder and the first target code segment. The loss function only focuses on whether the generated second target code segment is semantically equivalent to the first target code segment, regardless of which language is generated. The vector representation generated by the encoder is independent of the programming language and only captures the semantic and structural information of the code. This means that the same code logic will produce the same vector representation in different languages. The decoder generates code in different languages ​​based on the translation instruction and the same vector representation. Although these codes are in different languages, they are functionally equivalent. What the model learns is how to generate equivalent code based on the semantic and structural information of the input code segment, rather than relying on a specific programming language. Therefore, the model is able to translate the same code representation into different or the same language without affecting the representation results. This method is suitable for dealing with the problem of code cloning between different programming languages ​​because it allows the model to translate code between different languages ​​while maintaining the semantic consistency of the code. For example, it can detect functionally equivalent code snippets in different languages, or translate code from one language to another while keeping the code semantically consistent.

[0041] In one embodiment, the second target code segment is input into the encoder to generate a representation vector, and then the representation vector is input into the decoder to generate a third target code segment, and the difference between the third target code segment and the second target code segment is calculated by the translation reinforcement learning loss function.

[0042] Specifically, for the source code segment Ci and the first target code segment Cj, they are functionally equivalent code clone pairs, which means that they perform the same task but may use different programming languages ​​or code styles. Step S2024 performs a forward translation task: the training model translates the source code segment Ci into the second target code segment Cj. , which involves inputting the source code segment Ci into the encoder, generating a second code representation vector, and then inputting the second code representation vector into the decoder to generate a second target code segment , Cj is calculated by translation enhancement learning loss function and The reverse translation task is to convert the second target code segment Input into the encoder to generate a representation vector, which is then input into the decoder to generate the third target code , the third target code is calculated by the translation reinforcement learning loss function and the second target code segment The difference between.

[0043] The translation reinforcement learning loss functions of the forward translation task and the back translation task are combined and the model is trained simultaneously. This ensures that the model learns the conversion between code segments from two directions, which helps to improve the generalization ability and robustness of the model.

[0044] S2045. A multi-task learning loss is obtained by contrasting learning loss and translation enhancement learning loss. Model parameters of the initial clone code detection model are continuously iterated by the multi-task learning loss, so that the model continuously learns how to capture the semantics of the source code segment, thereby obtaining the clone code detection model.

[0045] The model parameters of the initial clone code detection model are continuously updated and iterated through multi-task learning loss to minimize the multi-task learning loss function. Each iterative learning includes the encoding process of the encoder, the decoding process of the decoder, the calculation of the contrastive learning loss, the calculation of the translation enhancement loss, and the optimization of this iteration. The contrastive learning loss function and the translation enhancement learning loss function are combined to obtain the multi-task learning loss function, which is expressed by the following formula: ; in, represents a hyperparameter that controls the balance between contrastive learning loss and translation enhancement loss. =1 means that the contrastive learning loss accounts for 0.1 of the final loss.

[0046] In the code clone detection task, each set of data (code pair) is trained by combining contrastive learning loss and translation enhancement loss to update the parameters of the model. The core of contrastive learning loss is to distinguish the similarity between positive sample pairs and negative sample pairs. It helps the model learn the semantic features of source code segments by shortening the distance between cloned code pairs in the feature space and pushing away the distance between non-cloned code pairs. Translation loss focuses on optimizing the generation ability of the model to ensure that the model can accurately translate the source code segment into a target code segment that is semantically equivalent to it. This combination not only improves the model's ability to identify code clones, but also strengthens the model's deep understanding of code semantics. Through this training method, the model can more accurately capture the semantic information of the code in a complex code environment, thereby effectively identifying cloned code pairs. The present disclosure combines the contrastive learning loss function with the translation enhancement learning loss function, learns the semantic features of the source code segment through the contrastive learning loss function, makes the cloned code closer in the embedding space, and learns to translate each source code segment into its corresponding cloned code through the translation enhancement learning loss function, captures the fine-grained semantic relationship between the code segments, and improves the accuracy of the first code representation vector by combining the contrastive learning loss function with the translation enhancement learning loss function, thereby making the code clone detection result more accurate. It is ensured that the clone code detection model can accurately translate the code segment to be detected into the first code representation vector that is semantically equivalent to it.

[0047] In one embodiment, the step S201 of "obtaining the code segment to be detected from the code repository to be detected" includes the following sub-steps S2011-S2013: S2011, reading the code file to be detected from the code repository to be detected; S2012, passing the read content of the code file to be detected to a pre-specified AST parsing tool, and converting the code file to be detected into a structured tree representation through the AST parsing tool; S2013. Find the function definition nodes of the code file to be detected by traversing the tree representation, and extract the code segment corresponding to each function definition node as the function-level code segment to be detected.

[0048] Specifically, a code repository is a place where code is stored, which can be a local file system, a version control system (such as Git), or a remote code repository (such as GitHub, Bitbucket, etc.). The code files to be detected are files that need to be analyzed to find cloned code. They can be source code files, header files, or any other files containing executable code. An AST parsing tool is a software tool that parses code files and generates an abstract syntax tree. This tree structure represents the grammatical structure of the code, including all statements, expressions, and the relationships between them. Before starting the analysis, you need to select an AST parsing tool for a suitable code language. For example, for Python code, you can use the ast module; for Java code, you can use JavaParser.

[0049] The code file is read through the AST parsing tool, and the code is parsed according to the grammatical rules to generate a tree structure. In this tree structure, each node represents a structure of the code (such as function definition, loop, conditional statement, etc.). This structured representation makes code analysis easier because the code can be decomposed into its basic components and the relationship between them can be examined. Once the AST of the code is available, the tree structure can be traversed to find a specific node. In code clone detection, the function definition node represents the function or method definition in the code. Each function definition node contains the function name, parameters, return type, and function body.

[0050] For each function definition node found, extract the corresponding code segment. This code segment includes the entire definition of the function, starting from the function declaration and ending with the function body. The code segment to be detected at the function level is the basic unit for clone detection. By comparing the code segments of different function definition nodes, potential cloned code can be identified. Through steps S2011-S2013, the code file to be detected is converted into a structured tree representation, and then the code segment at the function level is extracted to prepare for subsequent clone detection. This structured representation makes code analysis more systematic and automated, which helps to improve the accuracy and efficiency of code clone detection.

[0051] In step S202, two first code representation vectors are arbitrarily selected as code pairs, and code pairs whose semantic distance is less than a preset threshold are marked as clone codes.

[0052] In one embodiment, step S202 includes the following sub-steps S2021-S2023: S2021. Calculate the Euclidean distance between any two first code representation vectors, and measure the similarity between any two first code representation vectors using the Euclidean distance calculation result.

[0053] Euclidean distance is used to measure the straight-line distance between two points in a plane or space. The calculation formula of Euclidean distance is as follows: ; in, and They represent the coordinate values ​​of point 1 and point 2 in the i-th dimension respectively. In the process of code clone detection, the representation vector of each code segment can be regarded as a point. Calculating the Euclidean distance between two points can measure the similarity of the two code segments.

[0054] S2022. Classify similar first code representation vectors into the same category according to the Euclidean distance calculation result.

[0055] Specifically, the similarity between any two first code representation vectors can be measured by calculating the Euclidean distance. The smaller the distance, the more similar the two code segments are. This is because a smaller distance means that the two vectors are closer in position in the multidimensional space, so their semantics and structures are more likely to be similar. Based on the calculated Euclidean distance, similar code representation vectors can be classified into the same category. This usually involves setting a threshold, and when the Euclidean distance of two code segments is less than this threshold, they are considered similar and are classified into the same category.

[0056] S2023. Perform pairwise distance comparison on the first code representation vectors within each category, and mark the code pairs whose semantic distance between the first code representation vectors within each category is less than a preset threshold as the clone code pairs.

[0057] Specifically, within each category, the first code representation vectors are compared pairwise. For each code pair within each category, the semantic distance between them is calculated. If the semantic distance between the code pairs is less than a preset threshold, the pair of codes is marked as a clone code pair. The semantic distance can be Euclidean distance or other types of distances, such as cosine similarity or Jaccard similarity.

[0058] The first code representation vector is clustered by calculating the Euclidean distance, and then the distance comparison is performed between each category. This method considers both the structure of the code and the similarity at the text level, making the result more accurate.

[0059] After training the model in a multi-task framework, the decoder is usually discarded and only the encoder is retained. This is because the encoder has learned how to convert the input code snippet into a fixed-length vector representation that captures the semantic and structural information of the code snippet. For N samples, the encoder is used to calculate the vector representation of each sample. Then, the cosine distance between these vectors is calculated to determine whether the two code snippets are cloned code snippets. Cosine similarity is used to measure the angle between two vectors. The closer the cosine similarity value is to 1, the more similar the two vectors are. The formula for calculating cosine similarity is as follows: ; Among them, A and B are two vectors to be compared, and the value of cosine similarity ranges from -1 to 1. The closer to 1, the more similar the vectors are. The preset threshold is an empirical value used to determine whether the similarity between two code representation vectors is high enough that they can be marked as clone code pairs. If the cosine similarity of two code segments is greater than this preset threshold, then the two code segments are considered semantically similar and can therefore be marked as clone code pairs. This threshold is usually determined based on experiments and may need to be adjusted based on specific application scenarios and data sets. In actual operation, a cosine similarity greater than 0.5 or 0.6 is generally considered to be similar, but this value can be adjusted based on specific circumstances.

[0060] Based on the same inventive concept, the embodiment of the present application also provides a system for implementing the above-mentioned code clone detection system. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in one or more code clone detection system embodiments provided below can refer to the above-mentioned limitations on the code clone detection system, and will not be repeated here.

[0061] In an exemplary embodiment, Figure 4 As shown, a code clone detection system is provided, including: An acquisition module 410 is used to acquire code segments to be detected from a code repository to be detected, input all the code segments to be detected into a trained clone code detection model, and output a first code representation vector corresponding to the code segment to be detected; A marking module 420, configured to arbitrarily select two of the first code representation vectors as a code pair, and mark the code pair whose semantic distance is less than a preset threshold as a clone code; The marking module is also used to train the clone code detection model in the following manner: Randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, each pair of the source code segment and the first target code segment that have the unique correspondence relationship is a clone code of each other, and the unique correspondence relationship refers to semantic equivalence or grammatical differences but the same function; All the source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, wherein the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

[0062] As an optional implementation, the initial clone code detection model includes an encoder and a decoder.

[0063] As an optional implementation, the code clone detection system is also used to: The source code segment is encoded by an encoder, and a second code representation vector corresponding to the source code segment is output, wherein each source code segment and the corresponding second code representation vector are attached with the same label, and the label is used to represent the function of the source code segment; Pairing all the source code segments in pairs to obtain a plurality of code pairs, and judging whether the code pairs are clone code pairs according to the matching of the labels of the source code segments; if the labels match, marking the code pairs as clone code pairs, otherwise, marking the code pairs as non-clone code pairs; Calculating the contrastive learning loss of the clone code pair according to the contrastive learning loss function; Decoding all the second code representation vectors through a decoder, outputting a second target code segment corresponding to the second code representation vector, and calculating a translation enhancement loss of the second target code segment relative to the first target code segment by using a translation enhancement learning loss function; A multi-task learning loss is obtained through the contrastive learning loss and the translation enhancement learning loss. The model parameters of the initial clone code detection model are continuously iterated through the multi-task learning loss, so that the model continuously learns how to capture the semantics of the source code segment, thereby obtaining the clone code detection model.

[0064] As an optional implementation, in the aspect of obtaining the code segment to be detected from the code repository to be detected, the acquisition module 410 is specifically used to: Read the code file to be tested from the code repository to be tested; Passing the read content of the code file to be detected to a pre-specified AST parsing tool, and converting the code file to be detected into a structured tree representation through the AST parsing tool; By traversing the tree representation, the function definition nodes of the code file to be detected are found, and the code segment corresponding to each function definition node is extracted as the code segment to be detected at the function level.

[0065] As an optional implementation, the marking module 420 is specifically configured to: Calculate the Euclidean distance between any two of the first code representation vectors, and measure the similarity between any two of the first code representation vectors by using the Euclidean distance calculation result; Classifying similar first code representation vectors into the same category according to the Euclidean distance calculation result; The first code representation vectors within each of the categories are compared with each other in distance, and the code pairs whose semantic distance between the first code representation vectors within each of the categories is less than a preset threshold are marked as clone code pairs.

[0066] As an optional implementation, the initial clone code detection model is constructed based on a pre-trained CodeT5p model.

[0067] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store source code segments and first target code segments. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a code cloning detection method is implemented.

[0068] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0069] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0070] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0071] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0072] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0073] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0074] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0075] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0076] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A code clone detection method, characterized in that: The code clone detection method comprises: Acquire the code segments to be detected from the code repository to be detected, input all the code segments to be detected into the trained clone code detection model, and output the first code representation vector corresponding to the code segments to be detected; arbitrarily selecting two of the first code representation vectors as a code pair, and marking the code pair whose semantic distance is less than a preset threshold as a clone code; The clone code detection model is trained in the following way: Randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, each pair of the source code segment and the first target code segment that have the unique correspondence relationship is a clone code of each other, and the unique correspondence relationship refers to semantic equivalence or grammatical differences but the same function; All the source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, wherein the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

2. The code clone detection method according to claim 1, characterized in that: The initial clone code detection model includes an encoder and a decoder.

3. The code clone detection method according to claim 2, characterized in that: The initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, including: The source code segment is encoded by an encoder, and a second code representation vector corresponding to the source code segment is output, wherein each source code segment and the corresponding second code representation vector are attached with the same label, and the label is used to represent the function of the source code segment; Pairing all the source code segments in pairs to obtain a plurality of code pairs, and judging whether the code pairs are clone code pairs according to the matching of the labels of the source code segments; if the labels match, marking the code pairs as clone code pairs, otherwise, marking the code pairs as non-clone code pairs; Calculating the contrastive learning loss of the clone code pair according to the contrastive learning loss function; Decoding all the second code representation vectors through a decoder, outputting a second target code segment corresponding to the second code representation vector, and calculating a translation enhancement loss of the second target code segment relative to the first target code segment by using a translation enhancement learning loss function; A multi-task learning loss is obtained through the contrastive learning loss and the translation enhancement learning loss. The model parameters of the initial clone code detection model are continuously iterated through the multi-task learning loss, so that the model continuously learns how to capture the semantics of the source code segment, thereby obtaining the clone code detection model.

4. The code clone detection method according to claim 1, characterized in that: The obtaining of the code segment to be detected from the code repository to be detected includes: Reading a code file to be detected from the code repository to be detected; Passing the read content of the code file to be detected to a pre-specified AST parsing tool, and converting the code file to be detected into a structured tree representation through the AST parsing tool; By traversing the tree representation, the function definition nodes of the code file to be detected are found, and the code segment corresponding to each function definition node is extracted as the code segment to be detected at the function level.

5. The code clone detection method according to claim 1, characterized in that: The arbitrarily selecting two of the first code representation vectors as a code pair, and marking the code pair whose semantic distance is less than a preset threshold as a clone code, comprises: Calculating the Euclidean distance between any two of the first code representation vectors, and measuring the similarity between any two of the first code representation vectors by using the Euclidean distance calculation result; Classifying similar first code representation vectors into the same category according to the Euclidean distance calculation result; The first code representation vectors within each of the categories are compared with each other in distance, and the code pairs whose semantic distance between the first code representation vectors within each of the categories is less than the preset threshold are marked as clone code pairs.

6. The code clone detection method according to claim 1, characterized in that: The initial clone code detection model is built based on the pre-trained CodeT5p model.

7. A code clone detection system, characterized in that: The code clone detection system comprises: An acquisition module is used to acquire code segments to be detected from a code repository to be detected, input all the code segments to be detected into a trained clone code detection model, and output a first code representation vector corresponding to the code segment to be detected; a marking module, configured to arbitrarily select two of the first code representation vectors as a code pair, and mark the code pair whose semantic distance is less than a preset threshold as a clone code; The marking module is also used to train the clone code detection model in the following manner: Randomly sampling from a code clone detection data set to obtain the same number of source code segments and first target code segments that have a unique correspondence relationship, each pair of the source code segment and the first target code segment that have the unique correspondence relationship is a clone code of each other, and the unique correspondence relationship refers to semantic equivalence or grammatical differences but the same function; All the source code segments and the first target code segment are input into an initial clone code detection model, and the initial clone code detection model is trained by a multi-task learning loss function to obtain a clone code detection model, wherein the multi-task learning loss function includes a contrastive learning loss function and a translation enhancement learning loss function.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the code clone detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the code clone detection method described in any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the code clone detection method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Code clone detection method and system based on byte code and neural network

    CN114064117A

  • Software clone detection method based on plagiarism-detector confrontation

    CN116578336A

  • Code similarity detection-oriented cross-programming language migration method and system

    CN117608651A

  • Code detection method and device and electronic equipment

    CN118152000A

  • Code clone detection method and device based on hierarchical analysis, equipment and medium

    CN118916077A