Code searching method and device, electronic equipment and storage medium

Through a two-stage architecture combining dual encoder and cross encoder, the problem of taking into account both code search efficiency and accuracy is solved, efficient and accurate code search results are achieved, and computing resource consumption is optimized through knowledge distillation and code simplification technology.

CN120011546APending Publication Date: 2025-05-16SUN YAT SEN UNIV

Patent Information

Application Number
CN202510243488.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art cannot take into account the search efficiency and accuracy of code search. Dual encoders are difficult to capture the complex interaction between query and code, while cross-encoders are less efficient when dealing with large code bases.

Method used

Using a two-stage architecture, the natural language query is first encoded through a preset dual encoder to obtain the query encoding vector, match the candidate code fragments in the dense vector space, and simplify the program; then, the knowledge distillation of the preset cross code encoder is obtained to obtain the distilled cross code encoder, calculate the similarity between the query encoding vector and the simplified code fragment, and sort it.

Benefits of technology

The balance of efficiency and accuracy is achieved, and the accuracy and efficiency of code search is significantly improved through efficient search by dual encoders and precise matching of cross-encoders, and the consumption of computing resources is reduced through knowledge distillation and code simplification techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011546A_ABST
    Figure CN120011546A_ABST
Patent Text Reader

Abstract

The invention discloses a code search method and device, electronic equipment and a storage medium. The technical problem that in the prior art, the search efficiency and precision of code search cannot be considered at the same time is solved. The method comprises the following steps: receiving a natural language query input by a user; encoding the natural language query through preset double encoders to obtain a query encoding vector; matching the query code vector in a preset dense vector space to obtain a plurality of candidate code snippets; performing program simplification on the candidate code snippets to obtain simplified code snippets; performing knowledge distillation on a preset cross encoder to obtain a distillation cross encoder; calculating the similarity between the query code vector and the simplified code snippet through the distillation cross encoder; and sorting the simplified code snippets according to the similarity to obtain a code search result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of code search technology, and in particular to a code search method, device, electronic device and storage medium. Background Art

[0002] Code search is a key development tool that helps developers quickly find code snippets that match natural language queries in large code bases. Traditional code search methods rely on keyword matching or simple statistical models. However, as the size of code bases increases and the code structure becomes more complex, these methods have difficulty returning high-quality code snippets accurately and quickly. Therefore, researchers have begun to explore code search methods based on deep learning, which can better understand the semantic information of the code and improve the accuracy and practicality of code search.

[0003] Code search methods based on deep learning usually adopt an encoder model to encode natural language queries and code snippets into feature vectors, and measure the relevance of code snippets by the similarity of feature vectors. With the advancement of technology, existing code search methods have gradually evolved into two main architectures: dual encoder and cross encoder. These two architectures have their own advantages and disadvantages, so there is a trade-off between efficiency and accuracy:

[0004] Dual encoder: The dual encoder architecture contains two independent encoders, which are used to encode natural language queries and code snippets respectively. The advantage of this method is that the feature vectors of all code snippets can be pre-calculated in the offline stage. When querying, you only need to encode the query and match the similarity with the code feature vector, thereby improving the retrieval efficiency. However, in the dual encoder architecture, the query and code are encoded separately, which makes it difficult for the model to capture the complex interactive relationship between the query and the code, reducing the accuracy of code search.

[0005] Cross Encoder: The cross encoder uses a single encoder to jointly encode queries and code snippets, and captures the deep semantic relationship between queries and code snippets through a self-attention mechanism. Compared with the dual encoder, the cross encoder can more accurately match the relevance of queries and codes, so it has an advantage in search accuracy. However, the cross encoder needs to encode all candidate code snippets in real time at each query, which has a large computational overhead and cannot pre-compute the vector representation of the code snippet, so it is less efficient when processing large code bases. Summary of the invention

[0006] The present invention provides a code search method, device, electronic device and storage medium, which are used to solve the technical problem that the prior art cannot take into account both the search efficiency and accuracy of code search.

[0007] The present invention provides a code search method, comprising:

[0008] receiving a natural language query input by a user;

[0009] Encode the natural language query by using a preset dual encoder to obtain a query encoding vector;

[0010] Matching the query encoding vector in a preset dense vector space to obtain a plurality of candidate code fragments;

[0011] Simplifying the candidate code snippet to obtain a simplified code snippet;

[0012] Perform knowledge distillation on the preset cross encoder to obtain a distilled cross encoder;

[0013] Calculating the similarity between the query encoding vector and the simplified code snippet by the distilled cross encoder;

[0014] The simplified code snippets are sorted according to the similarities to obtain code search results.

[0015] Optionally, before the step of receiving a natural language query input by a user, the step further includes:

[0016] Obtaining original code, and splitting the original code into multiple code fragments;

[0017] Encoding the code segment by using the preset dual encoder to obtain a code vector;

[0018] The code vectors are clustered to obtain a plurality of clusters, and the clusters are mapped into the dense vector space.

[0019] Optionally, the step of matching the query encoding vector in a preset dense vector space to obtain a plurality of candidate code snippets includes:

[0020] Selecting a target cluster corresponding to the query encoding vector in a preset dense vector space;

[0021] Calculating the distance between the query encoding vector and the code vectors in the target cluster;

[0022] A plurality of candidate code snippets are screened from the target cluster according to the distance.

[0023] Optionally, the candidate code snippet includes a plurality of tokens, and the step of simplifying the candidate code snippet to obtain a simplified code snippet includes:

[0024] Calculate the self-attention weights of various tokens in the candidate code snippet;

[0025] Filtering a target sentence from the candidate code snippets according to the self-attention weight;

[0026] Filter the target token according to the self-attention weight of each token in the target sentence;

[0027] The target token is used to form a simplified code snippet.

[0028] The present invention also provides a code search device, comprising:

[0029] A natural language query receiving module, used to receive a natural language query input by a user;

[0030] A query encoding vector generation module, used for encoding the natural language query by using a preset dual encoder to obtain a query encoding vector;

[0031] A candidate code snippet matching module, used to match the query encoding vector in a preset dense vector space to obtain a plurality of candidate code snippets;

[0032] A program simplification module, used to simplify the candidate code snippet to obtain a simplified code snippet;

[0033] A knowledge distillation module is used to perform knowledge distillation on a preset cross encoder to obtain a distilled cross encoder;

[0034] A similarity calculation module, used for calculating the similarity between the query encoding vector and the simplified code snippet through the distillation cross encoder;

[0035] The reordering module is used to sort the simplified code snippets according to the similarity to obtain code search results.

[0036] Optionally, it also includes:

[0037] A code snippet generation module, used for obtaining original code and splitting the original code into multiple code snippets;

[0038] A code vector generation module, used for encoding the code fragment by using the preset dual encoder to obtain a code vector;

[0039] A clustering module is used to cluster the code vectors to obtain a plurality of clusters, and map the clusters into the dense vector space.

[0040] Optionally, the candidate code snippet matching module includes:

[0041] A target clustering screening submodule, used to screen the target clustering corresponding to the query encoding vector in a preset dense vector space;

[0042] A distance calculation submodule, used to calculate the distance between the query encoding vector and the code vector in the target cluster;

[0043] The candidate code fragment screening submodule is used to screen a number of candidate code fragments from the target cluster according to the distance.

[0044] Optionally, the candidate code snippet includes multiple tokens, and the program simplification module includes:

[0045] A self-attention weight calculation submodule is used to calculate the self-attention weights of various tokens in the candidate code snippet;

[0046] A target sentence screening submodule, used for screening target sentences from the candidate code snippets according to the self-attention weights;

[0047] A target token screening submodule, used to screen the target token according to the self-attention weight of each token in the target sentence;

[0048] The simplified code snippet forming submodule is used to form a simplified code snippet using the target token.

[0049] The present invention also provides an electronic device, the device comprising a processor and a memory:

[0050] The memory is used to store program code and transmit the program code to the processor;

[0051] The processor is used to execute any of the above code searching methods according to the instructions in the program code.

[0052] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the code search method as described in any one of the above items.

[0053] It can be seen from the above technical scheme that the present invention has the following advantages: The present invention provides a code search method, and specifically discloses: receiving a natural language query input by a user; encoding the natural language query by a preset dual encoder to obtain a query encoding vector; matching the query encoding vector in a preset dense vector space to obtain a number of candidate code snippets; simplifying the candidate code snippets to obtain a simplified code snippet; performing knowledge distillation on a preset cross encoder to obtain a distilled cross encoder; calculating the similarity between the query encoding vector and the simplified code snippet by a distilled cross encoder; sorting the simplified code snippets according to the similarity to obtain code search results. The present invention achieves a balance between efficiency and accuracy by designing a two-stage architecture, combining the efficient retrieval of the dual encoder with the precise matching of the cross encoder in the reordering stage. At the same time, the introduction of knowledge distillation and code simplification technology effectively reduces the consumption of computing resources in the reordering stage and alleviates the demand for computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0055] Figure 1 A flowchart of a code search method provided by an embodiment of the present invention;

[0056] Figure 2 A flowchart of a code search method provided by another embodiment of the present invention;

[0057] Figure 3 A logic flow chart of a code search method provided by an embodiment of the present invention;

[0058] Figure 4 A structural block diagram of a code search device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The embodiments of the present invention provide a code search method, device, electronic device and storage medium, which are used to solve the technical problem that the prior art cannot take into account both the search efficiency and accuracy of code search.

[0060] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0061] See also Figure 1 , Figure 1 A flowchart of a code search method provided by an embodiment of the present invention.

[0062] The present invention provides a code search method, which may specifically include the following steps:

[0063] Step 101, receiving a natural language query input by a user;

[0064] Natural language query is a query based on natural language expression, which is one of the contents of natural language processing. In artificial intelligence, the process of users inputting natural language expressions so that computers can accept and process natural language, process natural language information, and understand natural language, so as to provide feedback on information, is called natural language query.

[0065] Step 102, encoding the natural language query by using a preset dual encoder to obtain a query encoding vector;

[0066] Dual encoders generally refer to structures used in some machine learning and deep learning models, where two encoders process different types of information or encode information from different perspectives. This architecture is common in natural language processing (NLP), computer vision, and other fields that need to process complex data patterns.

[0067] In an embodiment of the present invention, a dual encoder may be used as a CodeBERT model to encode a natural language query, thereby obtaining a query encoding vector.

[0068] Step 103, matching the query encoding vector in a preset dense vector space to obtain a number of candidate code fragments;

[0069] A dense vector space is a vector space in which each dimension of a vector has a value (usually a real number) and these values ​​are not sparsely distributed, that is, the values ​​in most dimensions are not zero. In contrast, a sparse vector space is a space in which the values ​​in most dimensions are zero and only a few dimensions have non-zero values.

[0070] In the embodiment of the present invention, the dense vector space can store the vector representation of the code in advance. When the user searches for the code through natural language query, it can be converted into a vector representation through the query encoder, and then the query encoding vector is matched in the dense vector space through the vector retrieval system (Faiss), thereby obtaining several candidate code fragments. Faiss greatly reduces the amount of retrieval calculation by constructing an effective index structure and data compression technology.

[0071] Step 104, simplifying the candidate code snippet to obtain a simplified code snippet;

[0072] In the specific implementation, program simplification can reduce the input length by cutting out low-importance tokens in the candidate code snippets, thereby reducing the computational complexity.

[0073] Step 105, performing knowledge distillation on the preset cross encoder to obtain a distilled cross encoder;

[0074] Knowledge distillation is a model compression technique that aims to transfer the knowledge of a large and complex model (usually called a teacher model) to a smaller and less computationally expensive model (student model). In this way, the student model can inherit the capabilities of the teacher model without significant loss of performance, making it suitable for resource-constrained environments such as mobile devices or embedded systems.

[0075] In an embodiment of the present invention, by performing knowledge distillation on the cross encoder, the number of model parameters and computational complexity can be significantly reduced, thereby improving the operating efficiency of the reordering stage while ensuring an accuracy close to that of the teacher model.

[0076] Step 106, calculating the similarity between the query encoding vector and the simplified code snippet through a distilled cross encoder;

[0077] Step 107, sorting the simplified code snippets according to similarity to obtain code search results.

[0078] After completing the knowledge distillation of the cross encoder, the simplified code snippet and the query encoding vector can be input into the distilled cross encoder to calculate the similarity between the two, so as to sort the simplified code snippets according to the similarity and obtain the code search result.

[0079] The present invention achieves a balance between efficiency and accuracy by designing a two-stage architecture, combining the efficient retrieval of the dual encoders with the precise matching of the cross encoders in the reordering stage. At the same time, the introduction of knowledge distillation and code simplification technology effectively reduces the consumption of computing resources in the reordering stage and alleviates the demand for computing resources.

[0080] See also Figure 2 , Figure 2 A flowchart of a code search method provided by another embodiment of the present invention. Specifically, the following steps may be included:

[0081] Step 201, obtaining the original code, and splitting the original code into multiple code fragments;

[0082] Step 202, encoding the code fragment by using a preset dual encoder to obtain a code vector;

[0083] Step 203, clustering the code vectors to obtain a number of clusters, and mapping the clusters into a dense vector space;

[0084] In an embodiment of the present invention, the vector representation of the original code can be calculated and stored in advance in the offline stage, so that when querying, it is only necessary to encode the natural language query and then calculate the similarity with the pre-stored code vector, thereby greatly reducing the overhead of real-time calculation.

[0085] The specific implementation process is as follows: first, decompose the original code into several code fragments, then use the dual-encoder version of the CodeBERT model to independently encode the code fragments, and map the encoded code vectors to a dense vector space (shared embedding space) to form a vector representation of fixed dimension.

[0086] Furthermore, in order to further improve the retrieval efficiency, the embodiment of the present invention adopts reverse indexing and clustering algorithms to accelerate the retrieval process. Specifically, Faiss clusters the code vectors, and the clustering process is as follows: first, randomly initialize k cluster centers, and then assign each data point to the cluster where the nearest cluster center is located, and then update the cluster center according to the mean of all data points in each cluster. This process is repeated until the change in the cluster center is lower than the specified threshold or the maximum number of iterations is reached. The clustering process makes it possible to search for the nearest code snippet in the relevant clusters, because only part of the clusters need to be searched, and the distance between the vector in the cluster and the query vector is measured to find the data point closest to the query vector. This method greatly reduces the number of similarity calculations, thereby improving the retrieval speed.

[0087] Step 204, receiving a natural language query input by a user;

[0088] Step 205, encoding the natural language query by using a preset dual encoder to obtain a query encoding vector;

[0089] In the embodiment of the present invention, steps 204-205 are the same as steps 101-102. For details, please refer to the description of steps 101-102, which will not be repeated here.

[0090] Step 206, matching the query encoding vector in a preset dense vector space to obtain a number of candidate code fragments;

[0091] In this embodiment of the present invention, step 206 may include the following sub-steps:

[0092] S61, screening the target cluster corresponding to the query encoding vector in a preset dense vector space;

[0093] S62, calculating the distance between the query encoding vector and the code vector in the target cluster;

[0094] S63, screening a number of candidate code snippets from the target cluster according to the distance.

[0095] In a specific implementation, when a natural language query from a user is received, a target cluster of the query encoding vector can be screened in the dense vector space, and the distance between the query encoding vector and each code vector in the target cluster can be calculated in turn, thereby screening several candidate code snippets based on the distance.

[0096] In an example, the distance between the query code vector and each code vector in the target cluster may be a Euclidean distance, and the distance threshold for screening the candidate code segments may be set according to actual needs, which is not specifically limited in the embodiment of the present invention.

[0097] Step 207, simplifying the candidate code snippet to obtain a simplified code snippet;

[0098] In the specific implementation, program simplification can reduce the input length by cutting out low-importance tokens in the candidate code snippets, thereby reducing the computational complexity.

[0099] In one example, the candidate code snippet includes multiple tokens, and the steps of simplifying the candidate code snippet to obtain the simplified code snippet include:

[0100] S71, calculate the self-attention weights of various tokens in the candidate code snippet;

[0101] S72, select the target sentence from the candidate code snippets according to the self-attention weight;

[0102] S73, filtering the target token according to the self-attention weights of each token in the target sentence;

[0103] S74, using the target token to form a simplified code snippet.

[0104] In the embodiment of the present invention, token refers to a code unit after word segmentation, which is a basic building block of the code. Each token represents a logical unit, such as a keyword, an identifier, an operator, a literal, etc.

[0105] In the implementation of the present invention, the candidate code snippets obtained by the above process may still contain some codes that do not completely match the query, so it is necessary to further improve the accuracy of the results by re-ranking.

[0106] Before reordering, the candidate codes are first screened using the program simplification method DietCode to reduce the amount of calculation.

[0107] In an example, the simplified implementation process of the program is as follows:

[0108] 1. Sentence selection: Sentence selection is performed based on the self-attention weight of each type of token in the cross encoder, and sentences with higher weights are retained as target sentences to ensure that the main information is not lost during the simplification process.

[0109] 2. Token pruning: In the target statement, further prune low-weight tokens and retain high-weight tokens, and finally obtain a simplified code snippet.

[0110] By simplifying the program, the input length of the cross encoder can be significantly reduced, the computational complexity of the model can be reduced, and the core information of the code snippet can be retained, thereby improving the execution efficiency of the reordering stage.

[0111] Step 208, performing knowledge distillation on the preset cross encoder to obtain a distilled cross encoder;

[0112] In the embodiment of the present invention, the process of performing knowledge distillation on the preset cross encoder is specifically as follows:

[0113] 1. Teacher model adjustment: First, the cross encoder as the teacher model is fine-tuned so that it can accurately generate soft labels for the target task.

[0114] 2. Student model learning: The student model not only imitates the final prediction output of the teacher model, but also learns the CLS token representation of each layer by extracting features from the intermediate layers of the teacher model.

[0115] Objective function: There are two main loss terms in the knowledge distillation process: classification loss and distillation loss. The final objective function combines these two loss terms and achieves a balance between accuracy and efficiency by adjusting the weight parameters.

[0116] Step 209, calculating the similarity between the query encoding vector and the simplified code snippet through a distillation cross encoder;

[0117] Step 210, sorting the simplified code snippets according to similarity to obtain code search results.

[0118] After completing the knowledge distillation of the cross encoder, the simplified code snippet and the query encoding vector can be input into the distilled cross encoder to calculate the similarity between the two, so as to sort the simplified code snippets according to the similarity and obtain the code search result.

[0119] The present invention achieves a balance between efficiency and accuracy by designing a two-stage architecture, combining the efficient retrieval of the dual encoders with the precise matching of the cross encoders in the reordering stage. At the same time, the introduction of knowledge distillation and code simplification technology effectively reduces the consumption of computing resources in the reordering stage and alleviates the demand for computing resources.

[0120] See also Figure 3 , Figure 3 A logic flow chart of a code search method provided by an embodiment of the present invention.

[0121] like Figure 3 As shown, the code search method provided by the embodiment of the present invention includes two stages: recall (stage one) and rearrangement (stage two).

[0122] Among them, the main task of the recall phase is to quickly retrieve the candidate code snippets that are most relevant to the query from a large-scale code base. In the specific implementation, the original code is first preprocessed to form code snippets, and then the dual-encoder version of the CodeBERT model is used to generate vector representations for all code snippets in the code base in advance and store them in the vector retrieval library to reduce the need for real-time calculations. After the user enters a natural language query, the query encoder is used to convert the query into a vector representation, and then the vector retrieval system Faiss is used to quickly find the top k high-similarity candidate code snippets in the pre-stored vector set.

[0123] The re-ranking stage uses a cross-encoder-based model to finely sort the candidate code snippets to ensure the accuracy of the returned results. The specific process is as follows: First, the program simplification method is used to simplify the candidate code snippets recalled in the first stage, and unimportant tokens are deleted to speed up the subsequent similarity calculation process. In addition, in order to further reduce the computational complexity of the cross-encoder model, the original large model is compressed into a small model through the knowledge distillation method, which greatly speeds up the similarity calculation process. Finally, based on the similarity calculation results of the cross-encoder after knowledge distillation, the candidate codes are re-ranked according to the similarity between the query vector and the query vector, and the re-ranked results are returned.

[0124] See also Figure 4 , Figure 4A structural block diagram of a code search device provided by an embodiment of the present invention.

[0125] An embodiment of the present invention provides a code search device, comprising:

[0126] A natural language query receiving module 401 is used to receive a natural language query input by a user;

[0127] A query encoding vector generating module 402, configured to encode a natural language query by using a preset dual encoder to obtain a query encoding vector;

[0128] A candidate code snippet matching module 403 is used to match the query encoding vector in a preset dense vector space to obtain a number of candidate code snippets;

[0129] A program simplification module 404 is used to simplify the candidate code snippet to obtain a simplified code snippet;

[0130] A knowledge distillation module 405 is used to perform knowledge distillation on a preset cross encoder to obtain a distilled cross encoder;

[0131] A similarity calculation module 406, used to calculate the similarity between the query encoding vector and the simplified code snippet through a distilled cross encoder;

[0132] The reordering module 407 is used to order the simplified code snippets according to similarity to obtain code search results.

[0133] In an embodiment of the present invention, it also includes:

[0134] A code snippet generation module is used to obtain the original code and split the original code into multiple code snippets;

[0135] A code vector generation module, used for encoding the code fragment by using a preset dual encoder to obtain a code vector;

[0136] The clustering module is used to cluster the code vectors to obtain several clusters and map the clusters into a dense vector space.

[0137] In the embodiment of the present invention, the candidate code fragment matching module 403 includes:

[0138] A target clustering screening submodule is used to screen the target clustering corresponding to the query encoding vector in a preset dense vector space;

[0139] A distance calculation submodule, used to calculate the distance between the query encoding vector and the code vector in the target cluster;

[0140] The candidate code snippet screening submodule is used to screen several candidate code snippets from the target cluster according to the distance.

[0141] In the embodiment of the present invention, the candidate code snippet includes multiple tokens, and the program simplification module 404 includes:

[0142] The self-attention weight calculation submodule is used to calculate the self-attention weights of various tokens in the candidate code snippets;

[0143] The target sentence screening submodule is used to screen the target sentence from the candidate code snippets according to the self-attention weight;

[0144] The target token filtering submodule is used to filter the target token according to the self-attention weights of each token in the target sentence;

[0145] The simplified code snippet formation submodule is used to form a simplified code snippet using the target token.

[0146] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:

[0147] The memory is used to store the program code and transmit the program code to the processor;

[0148] The processor is used to execute the code search method of the embodiment of the present invention according to the instructions in the program code.

[0149] The embodiment of the present invention further provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the code search method of the embodiment of the present invention.

[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0151] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0152] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0153] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0154] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0156] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0157] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0158] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A code search method, characterized in that: include: receiving a natural language query input by a user; Encode the natural language query by using a preset dual encoder to obtain a query encoding vector; Matching the query encoding vector in a preset dense vector space to obtain a plurality of candidate code fragments; Simplifying the candidate code snippet to obtain a simplified code snippet; Perform knowledge distillation on the preset cross encoder to obtain a distilled cross encoder; Calculating the similarity between the query encoding vector and the simplified code snippet by the distilled cross encoder; The simplified code snippets are sorted according to the similarities to obtain code search results.

2. The method according to claim 1, characterized in that Before the step of receiving a natural language query input by a user, the method further includes: Obtaining original code, and splitting the original code into multiple code fragments; Encoding the code segment by using the preset dual encoder to obtain a code vector; The code vectors are clustered to obtain a plurality of clusters, and the clusters are mapped into the dense vector space.

3. The method according to claim 1, characterized in that The step of matching the query encoding vector in a preset dense vector space to obtain a plurality of candidate code fragments includes: Selecting a target cluster corresponding to the query encoding vector in a preset dense vector space; Calculating the distance between the query encoding vector and the code vectors in the target cluster; A plurality of candidate code snippets are screened from the target cluster according to the distance.

4. The method according to claim 1, characterized in that: The candidate code snippet includes a plurality of tokens, and the step of simplifying the candidate code snippet to obtain a simplified code snippet includes: Calculate the self-attention weights of various tokens in the candidate code snippet; Filtering a target sentence from the candidate code snippets according to the self-attention weight; Filter the target token according to the self-attention weight of each token in the target sentence; The target token is used to form a simplified code snippet.

5. A code search device, characterized in that: include: A natural language query receiving module, used to receive a natural language query input by a user; A query encoding vector generation module, used for encoding the natural language query by using a preset dual encoder to obtain a query encoding vector; A candidate code snippet matching module, used to match the query encoding vector in a preset dense vector space to obtain a plurality of candidate code snippets; A program simplification module, used to simplify the candidate code snippet to obtain a simplified code snippet; A knowledge distillation module is used to perform knowledge distillation on a preset cross encoder to obtain a distilled cross encoder; A similarity calculation module, used for calculating the similarity between the query encoding vector and the simplified code snippet through the distillation cross encoder; The reordering module is used to sort the simplified code snippets according to the similarity to obtain code search results.

6. The device according to claim 5, characterized in that Also includes: A code snippet generation module, used for obtaining original code and splitting the original code into multiple code snippets; A code vector generation module, used for encoding the code fragment by using the preset dual encoder to obtain a code vector; A clustering module is used to cluster the code vectors to obtain a plurality of clusters, and map the clusters into the dense vector space.

7. The device according to claim 5, characterized in that The candidate code fragment matching module includes: A target clustering screening submodule, used to screen the target clustering corresponding to the query encoding vector in a preset dense vector space; A distance calculation submodule, used to calculate the distance between the query encoding vector and the code vector in the target cluster; The candidate code fragment screening submodule is used to screen a number of candidate code fragments from the target cluster according to the distance.

8. The device according to claim 5, characterized in that The candidate code snippet includes a plurality of tokens, and the program simplification module includes: A self-attention weight calculation submodule is used to calculate the self-attention weights of various tokens in the candidate code snippet; A target sentence screening submodule, used for screening target sentences from the candidate code snippets according to the self-attention weights; A target token screening submodule, used to screen the target token according to the self-attention weight of each token in the target sentence; The simplified code snippet forming submodule is used to form a simplified code snippet using the target token.

9. An electronic device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the code search method according to any one of claims 1 to 4 according to the instructions in the program code.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the code search method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Semantic matching method and device

    CN114579704A

  • Code search method based on multi-modal representation

    CN117390130A

Cited By

  • Information retrieval method and device and electronic equipment

    CN121144486A