Code retrieval method and device, electronic equipment and storage medium
By using the MdCoQA method to represent the relationship between code snippets and natural language queries through multi-dimensional spatial feature attributes, the problems of low computational efficiency and insufficient semantic understanding in existing technologies are solved, achieving efficient and accurate code search, which is suitable for real-time search of large-scale code libraries.
Patent Information
- Application Number
- CN202510819453.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-14
AI Technical Summary
Existing code retrieval methods suffer from low computational efficiency, low memory efficiency, and insufficient semantic understanding when dealing with large-scale codebases, making it difficult to perform well in scenarios with high real-time requirements.
We employ multi-dimensional spatial feature attributes to represent the relationship between code snippets and natural language queries. We transform code retrieval questions into question-answering matching questions using the MdCoQA method. We then utilize multi-dimensional spatial embedding technology and a simplified architecture design to calculate similarity scores and output matching results.
It achieves efficient and accurate code search, reduces computational and memory overhead, improves semantic accuracy and adaptability, and is suitable for real-time search scenarios of large-scale code bases.
Smart Images

Figure CN120950704A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of code retrieval technology, and in particular to a code retrieval device, electronic device, and storage medium. Background Technology
[0002] In the field of modern software development, code search is an indispensable part of developers' daily work.
[0003] With the rapid growth of code repositories on open-source platforms such as GitHub and Stack Overflow, developers need to quickly find code that matches their natural language requirements from massive amounts of code snippets.
[0004] Traditional code retrieval methods often face serious computational and memory efficiency issues when dealing with large-scale codebases. Furthermore, the methods in related technologies typically rely on complex matching mechanisms and attention models, which can lead to long retrieval times and high resource consumption, especially in real-time applications where performance is poor. Summary of the Invention
[0005] This application provides a code retrieval device, electronic device, and storage medium that utilizes multi-dimensional spatial feature attributes to represent the relationship between code fragments and natural language queries, thereby achieving code retrieval.
[0006] The embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, embodiments of this application provide a code retrieval method, wherein the retrieval method includes:
[0008] In response to a code retrieval request, generate a combination of question and answer pairs;
[0009] Calculate the similarity score between the pairs of questions and answers; and
[0010] The final matching result is output based on the similarity score;
[0011] In this process, the combination of questions and answers is generated using a code retrieval model to generate multi-dimensional feature attributes that characterize the relationship between the code snippet to be queried and the natural language query.
[0012] In some embodiments, generating a question-and-answer pair in response to a code retrieval request includes:
[0013] Based on the code retrieval request, a combination of question and answer pairs is generated using the code retrieval model.
[0014] The code retrieval model is used to learn representations of question-and-answer pairs in a multi-dimensional space, so that correct question-and-answer pairs are closer together in the multi-dimensional space, while incorrect question-and-answer pairs are farther apart.
[0015] In some embodiments, the code retrieval model is further used for:
[0016] The question-and-answer pairs are embedded into the multi-dimensional space, and a multi-dimensional distance function is used to measure the relationship between the natural language query and the query code.
[0017] In some embodiments, the code retrieval model is further used for:
[0018] Similarity calculation based on multi-dimensional distance calculates distance in the multi-dimensional space and transforms the multi-dimensional distance into a similarity score through a linear layer.
[0019] In some embodiments, the response prior to the code retrieval request further includes:
[0020] The pre-trained BERT model is used to transform natural language queries and code snippets into initial vector representations.
[0021] In some embodiments, the pre-trained BERT model employs a static BERT model.
[0022] In some embodiments, the initial vector representation is converted into a word representation suitable for code retrieval tasks via a projection layer.
[0023] In some embodiments, the projection layer employs a single-layer neural network that uses a non-linear activation function to process each word in the natural language query and the code snippet to be queried, and the parameters of the projection layer are shared between the query and the code.
[0024] In some embodiments, a holistic representation of a natural language query and a code snippet to be queried is generated by aggregating all word representations in the projection layer sequence.
[0025] In some embodiments, the method further includes:
[0026] The code retrieval model is trained using a pairwise hinge loss function.
[0027] In some embodiments, the method further includes:
[0028] Riemannian optimization techniques were used to update the code retrieval model parameters.
[0029] Secondly, embodiments of this application also provide a code retrieval device, wherein the retrieval device includes:
[0030] The generation module is used to generate question-and-answer pairs in response to code retrieval requests;
[0031] A similarity module is used to calculate a similarity score between pairs of questions and answers; and
[0032] The matching module is used to output the final matching result based on the similarity score;
[0033] In this process, the combination of questions and answers is generated using a code retrieval model to generate multi-dimensional feature attributes that characterize the relationship between the code snippet to be queried and the natural language query.
[0034] Thirdly, embodiments of this application also provide an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, which, when executed, cause the processor to perform the above-described method.
[0035] Fourthly, embodiments of this application also provide a computer-readable storage medium that stores one or more programs, which, when executed by an electronic device including multiple applications, cause the electronic device to perform the above-described method.
[0036] The at least one technical solution adopted in this application embodiment can achieve the following beneficial effects: In response to a code retrieval request, a combination of questions and answers is generated. Then, a similarity score between the question and answer combination is calculated, and the final matching result is output based on the similarity score. When generating the question and answer combination, a multi-dimensional feature attribute is generated using a code retrieval model to characterize the relationship between the queried code fragment and the natural language query. The above method achieves efficient and accurate code search by using multi-dimensional feature attributes to represent the relationship between code fragments and natural language queries. Attached Figure Description
[0037] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0038] Figure 1 This is a schematic diagram illustrating the implementation process of the code retrieval method in the embodiments of this application;
[0039] Figure 2 This is a flowchart illustrating the code retrieval method in an embodiment of this application;
[0040] Figure 3 This is a schematic diagram illustrating the implementation principle of the code retrieval method in the embodiments of this application;
[0041] Figure 4 This is a schematic diagram of the code retrieval device in an embodiment of this application;
[0042] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] During their research, the inventors discovered that code retrieval methods primarily rely on traditional lexical information retrieval techniques, such as keyword matching or API call-based searches. While these techniques are simple and easy to use, they have limitations in understanding the semantic meaning of code and struggle to handle complex query requirements.
[0045] With the rise of deep learning technology, the field of code retrieval has seen significant progress. Researchers have begun to leverage neural network models to improve code search performance. For example, CodeBERT, a pre-trained model based on the Transformer architecture, is trained by combining bimodal data from natural and programming languages to generate contextual representations of code and text. CodeBERT uses hybrid objective functions such as substitution marker detection for optimization, which improves retrieval accuracy to some extent. Another representative method is CodeRetriever, which introduces contrastive learning techniques to construct code pairs in an unsupervised manner and generate code-text pairs using documents and annotations, further enhancing the model's semantic understanding capabilities. Furthermore, CoCoSoDa proposed a multimodal momentum contrastive learning and soft data augmentation approach, combining multiple data sources and augmentation techniques to improve code retrieval performance.
[0046] Furthermore, CodeBERT relies on large-scale pre-training and attention mechanisms, making it suitable for scenarios requiring high-precision semantic matching; CodeRetriever optimizes code and text representations through contrastive learning, reducing its dependence on labeled data; and CoCoSoDa further enriches the model's input information through multimodal learning. These approaches have driven the development of code retrieval technology, gradually shifting it from simple keyword matching to deeper semantic understanding. However, these methods still face some common problems in practical applications, especially in terms of computational efficiency and memory usage, requiring further improvement.
[0047] Code retrieval methods in related technologies typically rely on complex matching mechanisms and attention models, resulting in long retrieval times and high resource consumption, especially in real-time application scenarios where performance is poor.
[0048] Furthermore, code retrieval methods in related technologies are insufficient in understanding the deep semantic relationship between code and queries, making it difficult to accurately capture the developer's intent.
[0049] To address the aforementioned shortcomings, the code retrieval method in this application embodiment is mainly used to solve the problems of high computational complexity, low memory efficiency, and insufficient semantic understanding ability of existing code retrieval technologies. It adopts an efficient, accurate code search solution applicable to large-scale code libraries, providing developers with more convenient tools and improving software development efficiency.
[0050] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0051] like Figure 1 As shown, from bottom to top, the process includes multi-dimensional indexing, vectorization processing, a vector database, similarity score calculation, and the final matching result. The multi-dimensional index includes, but is not limited to, Code, Description, and AST. Here, Code represents the code dimension, Description represents the description dimension, and AST represents the abstract syntax tree dimension. After vectorization processing based on these multi-dimensional dimensions, each element is stored in the vector database. When a user's question request is received, a similarity score is calculated for each element, and the latest result is output as the matching result.
[0052] This application provides a code retrieval method, such as... Figure 2 The diagram shows a flowchart of a code retrieval method in an embodiment of this application. The method includes at least the following steps S210 to S230:
[0053] Step S210: In response to the code retrieval request, generate a combination of question and answer pairs.
[0054] like Figure 1 As shown, the code retrieval method in this embodiment generates a combination of question and answer based on the user's question, i.e., the code retrieval request. It can be understood that the combination of question and answer can be retrieved from an existing vector database.
[0055] It is important to note that, such as Figure 1 As shown, when generating question-and-answer pairs, multiple vector databases may be called, or only one vector database may be called. The specific implementation of this application is not limited in this way.
[0056] Preferably, in order to speed up code retrieval results, it is usually possible to automatically detect which files have been modified, so that when generating question-and-answer pairs, only the modified parts are indexed and the results are found in the corresponding vector database.
[0057] Step S220: Calculate the similarity score between the pairs of questions and answers.
[0058] A similarity calculation method is used to calculate similarity scores between multiple question-answer pairs. It can be understood that the similarity score is used to evaluate the degree of matching between question-answer pairs, thereby providing a basis for subsequent ranking and retrieval.
[0059] Step S230: Output the final matching result based on the similarity score; wherein, when generating the question and answer combination pair, a multi-dimensional feature attribute is generated using a code retrieval model to characterize the relationship between the code fragment to be queried and the natural language query.
[0060] The final matching result is output based on the similarity score of the best match.
[0061] Specifically, when generating question-and-answer pairs, a multi-dimensional feature attribute is generated using a code retrieval model to characterize the relationship between the code snippet to be queried and the natural language query.
[0062] Specifically, this application proposes a code retrieval method called "Multi-dimensional Code QAMatching" (MdCoQA). The core of MdCoQA lies in utilizing the unique attributes of a multi-dimensional space to represent the relationship between code snippets and natural language queries, thereby achieving efficient and accurate code search. For example... Figure 1 As shown, the overall architecture of MdCoQA includes a complete architecture of modules such as BERT embedding layer, projection layer of task-specific word representation, generation of QA representation, embedding interaction in multi-dimensional space, and similarity calculation based on multi-dimensional distance.
[0063] In this application's embodiment, MdCoQA redefines the traditional code retrieval problem as a question-and-answer (QA) matching framework. In this framework, natural language queries are considered "questions (Q)," while code snippets are considered "answers (A)." By learning representations of these QA pairs in a multi-dimensional space, the code retrieval model can bring correct QA pairs closer together and incorrect QA pairs further apart. This approach avoids the complex interaction layers and attention mechanisms found in traditional neural network models, significantly reducing computational and memory overhead while improving the semantic accuracy of retrieval. The redefined QA matching framework transforms the code retrieval problem into a question-and-answer matching problem by learning representations of QA pairs in a multi-dimensional space and using multi-dimensional distance to measure similarity.
[0064] By combining multi-dimensional spatial embedding with a QA pair matching framework, the limitations of traditional code retrieval methods are overcome. Furthermore, MdCoQA's simplified architecture and efficient computation make it particularly suitable for real-time search scenarios in large-scale codebases, while maintaining high retrieval accuracy.
[0065] By employing the above methods and leveraging the application of multi-dimensional spatial embedding, MdCoQA has for the first time introduced multi-dimensional spatial embedding technology into the field of code retrieval, utilizing its hierarchical representation capabilities to capture the deep semantic relationship between queries and code.
[0066] Using the above method, the redefinition of the QA pair matching framework in MdCoQA transforms the code retrieval problem into a question-answering matching problem by learning the distance representation of QA pairs in a multi-dimensional space and using multi-dimensional distance to measure similarity.
[0067] Unlike related technologies, which generally have high computational complexity (e.g., CodeBERT and CoCoSoDa rely on complex attention mechanisms and multi-layered neural networks, requiring significant computational resources when processing large-scale codebases, resulting in slow retrieval speeds and difficulty meeting real-time requirements), the method in this application avoids storing large amounts of intermediate representations and parameters, making MdCoQA more economical in terms of memory usage. Compared to models like CodeBERT and CodeRetriever, which require massive memory support, MdCoQA is more suitable for deployment in resource-constrained environments.
[0068] Unlike related technologies, low memory efficiency is a prominent issue. These models typically require storing large amounts of intermediate representations and parameters. For example, CodeRetriever needs to maintain embedding representations of multiple data pairs during contrastive learning, which places high demands on memory resources, making deployment difficult, especially in resource-constrained environments. The method in this application, through its simplified architecture and multi-dimensional distance calculation, avoids high computational overhead. This allows the model to maintain efficient retrieval speed even when processing large-scale codebases.
[0069] Unlike related technologies, MdCoQA still has limitations in deep semantic understanding. Although deep learning models improve semantic matching capabilities, they are still insufficient in capturing the hierarchical structure and complex relationships between code and queries, resulting in less than ideal accuracy of retrieval results. The method in this application, by utilizing the hierarchical representation capabilities of a multi-dimensional space, can more accurately capture the semantic relationships between code and queries. Specific experimental results show that its average MRR on the CodeSearchNet dataset is 3.5% to 4% higher than existing state-of-the-art methods (such as CoCoSoDa), and it performs exceptionally well across six programming languages.
[0070] Furthermore, the MdCoQA provided in this embodiment outperforms the baseline model across multiple programming languages, demonstrating its good adaptability to different languages and scenarios. This generalization ability gives it broader application prospects in practical software development.
[0071] In one embodiment of this application, generating question-and-answer pairs in response to a code retrieval request includes: generating question-and-answer pairs based on the code retrieval model according to the code retrieval request, wherein the code retrieval model is configured to learn representations of the question-and-answer pairs in a multi-dimensional space so that correct question-and-answer pairs are closer together in the multi-dimensional space, while incorrect question-and-answer pairs are farther apart.
[0072] Specifically, when generating question-and-answer pairs based on the code retrieval model, correct question-and-answer pairs are closer in the multi-dimensional space, while incorrect pairs are farther apart. This distance-based approach avoids the complex interaction layers and attention mechanisms of traditional methods, significantly reducing computational and memory overhead while improving the semantic accuracy of the retrieval.
[0073] In one embodiment of this application, the code retrieval model is further configured to: embed the question-and-answer pair representation into the multi-dimensional space, and use a multi-dimensional distance function to measure the relationship between the natural language query and the query code.
[0074] MdCoQA in this application provides embedded interaction methods in a multi-dimensional space, specifically including:
[0075] This approach embeds vector representations of QA pairs into a multi-dimensional space and uses a multi-dimensional distance function to measure the relationship between queries and codes. Specifically, the multi-dimensional space is a non-Euclidean space whose distance function naturally captures hierarchical structures and semantic relationships. This method enables the model to efficiently represent the complex semantic connections between codes and queries. By redefining the QA pair matching framework, the code retrieval problem is transformed into a question-answering matching problem. This is achieved by learning representations of QA pairs in a multi-dimensional space and using multi-dimensional distance to measure similarity.
[0076] Preferably, the multi-dimensional distance calculation is based on the Poincaré sphere model, which evaluates the matching degree of two vectors by measuring their distance in a multi-dimensional space. By introducing multi-dimensional space embedding technology into the field of code retrieval, its hierarchical representation capabilities are utilized to capture the deep semantic relationship between queries and code.
[0077] In one embodiment of this application, the code retrieval model is further configured to: calculate distance in the multi-dimensional space based on multi-dimensional distance similarity calculation, and convert the multi-dimensional distance into a similarity score through a linear layer.
[0078] The MdCoQA in this application provides a similarity calculation scheme based on multi-dimensional distance. Specifically, after calculating the distance in a multi-dimensional space, MdCoQA transforms the multi-dimensional distance into a similarity score through a linear layer. The similarity score is used to evaluate the matching degree of QA pairs, providing a basis for subsequent ranking and retrieval. The parameters of the linear layer are trained to map spatial relationships into quantifiable similarity values.
[0079] Preferably, by introducing hyperbolic space embedding technology, which is naturally suited to representing hierarchical structures, deep semantic relationships between code and queries can be captured with lower computational complexity. Unlike existing methods that rely on complex interaction layers, this simplified architecture reduces the demand for computational and memory resources while improving retrieval accuracy and efficiency. This not only overcomes the shortcomings of existing technologies but also provides a more practical technical path for large-scale code search.
[0080] In one embodiment of this application, the response to the code retrieval request further includes: using a pre-trained BERT model to convert the natural language query and the code fragment to be queried into an initial vector representation.
[0081] like Figure 2 As shown in the vectorization process, the BERT embedding layer first uses a pre-trained BERT model to transform queries and code snippets into initial vector representations.
[0082] It is important to note that BERT is a Transformer-based model that can capture contextual information of text, providing high-quality vector embeddings for subsequent processing.
[0083] In one embodiment of this application, the pre-trained BERT model is a static BERT model.
[0084] like Figure 2 As shown in the vectorization process, to improve efficiency, static BERT is preferred, i.e., its parameters are frozen, and embeddings are generated only using its pre-training capabilities without additional fine-tuning. This design reduces computational costs while maintaining representation quality.
[0085] In one embodiment of this application, the initial vector representation is converted into a word representation suitable for code retrieval tasks through a projection layer.
[0086] After vectorizing the task-specific word representations, MdCoQA transforms these embeddings into word representations suitable for code retrieval tasks by designing a projection layer after obtaining the initial embeddings.
[0087] Alternatively, the common ReLU can be used as the non-linear activation function.
[0088] In one embodiment of this application, the projection layer employs a single-layer neural network, using a non-linear activation function to process each word in the natural language query and the code snippet to be queried, and the parameters of the projection layer are shared between the query and the code.
[0089] Specifically, this projection layer is a single-layer neural network that uses a non-linear activation function to process each word in the query and code snippet. Furthermore, the parameters of the projection layer are shared between the query and the code, ensuring a consistent representation transformation.
[0090] In one embodiment of this application, a holistic representation of a natural language query and a code fragment to be queried is generated by aggregating all word representations in the projection layer sequence.
[0091] As a method for generating QA pair representations, the overall representation of the query and code snippet is generated by aggregating the embeddings of all words in the sequence. Specifically, the generation of QA pair representations involves the code retrieval model generating the overall representation of the query and code snippet by aggregating the embeddings of all words in the sequence. In other words, the representation is subsequently normalized to the unit sphere, ensuring that the vector norm is less than or equal to 1, in order to adapt to subsequent processing in multi-dimensional spaces.
[0092] It can be understood that the above-mentioned aggregation of all word representations in the projection layer sequence is actually a semantic representation of the code, i.e., a mapping relationship between computer language and natural language. Based on this mapping relationship, the corresponding QA pairs at the same position can be obtained.
[0093] In one embodiment of this application, the method further includes: training a code retrieval model using a pairwise hinge loss function.
[0094] Optionally, to train the model, the MdCoQA in this embodiment employs a pairwise hinge loss function to ensure that correct QA pairs score higher than incorrect QA pairs. OA pairs can be used as a representation of task feature words.
[0095] It is understandable that Hinge loss is a surrogate function of 0-1 loss function, and the typical classifier that uses Hinge loss is the SVM algorithm.
[0096] The above methods are used to optimize the similarity score.
[0097] In one embodiment of this application, the method further includes: using Riemannian optimization techniques to update the code retrieval model parameters.
[0098] Optionally, due to the non-Euclidean nature of multidimensional space, the model in this embodiment uses Riemannian optimization to update parameters. This optimization method combines the characteristics of multidimensional geometry, ensuring the accuracy and convergence of gradient updates.
[0099] As can be understood, Riemannian optimization techniques are a mathematical method for performing optimization tasks on Riemannian manifolds, often used to handle problems with complex constraints.
[0100] The above methods are used to optimize the similarity score.
[0101] The method in this embodiment offers lower computational complexity, higher memory efficiency, and stronger generalization ability. It not only improves code retrieval performance but also enhances its practicality, making MdCoQA a powerful tool for developers navigating massive codebases.
[0102] like Figure 3 As shown, the architecture of MdCoQA in this embodiment consists of the following key components, each working together to achieve efficient code retrieval:
[0103] In the BERT embedding layer, MdCoQA uses a pre-trained BERT model to transform queries and code snippets into initial vector representations. BERT is a Transformer-based model that captures contextual information from the text, providing high-quality embeddings for subsequent processing.
[0104] Task-Specific Word Representations: After obtaining the initial embeddings, MdCoQA transforms these embeddings into word representations suitable for the code retrieval task through a projection layer. This projection layer is a single-layer neural network that processes each word in the query and code snippet using a non-linear activation function.
[0105] The QA representation generation and code retrieval model generate a holistic representation of the query and code snippets by aggregating the embeddings of all words in the sequence. These representations are then normalized to a unit sphere to accommodate subsequent processing in a multi-dimensional space.
[0106] In MdCoQA, QA embeds QA representations into a multi-dimensional space and uses a multi-dimensional distance function to measure the relationship between queries and code. The multi-dimensional space is a non-Euclidean space, and its distance function naturally captures hierarchical structures and semantic relationships.
[0107] Based on multi-dimensional distance similarity calculation, MdCoQA transforms the distance into a similarity score through a linear layer after calculating it in a multi-dimensional space. This score is used to evaluate the matching degree of QA pairs, providing a basis for subsequent ranking and retrieval. The parameters of the linear layer are trained to map spatial relationships into quantifiable similarity values.
[0108] To optimize the MdCoQA model, a pairwise hinge loss function is used to ensure that correct QA pairs score higher than incorrect ones. Simultaneously, due to the non-Euclidean nature of the multi-dimensional space, the model employs Riemannian optimization to update parameters. This optimization method combines the characteristics of multi-dimensional geometry, guaranteeing the accuracy and convergence of gradient updates.
[0109] This application embodiment also provides a code retrieval device 400, such as Figure 4 As shown, a schematic diagram of the structure of the code retrieval device in this application embodiment is provided. The code retrieval device 400 includes at least: a generation module 410, a similarity module 420, and a matching module 430, wherein:
[0110] In one embodiment of this application, the generation module 410 is specifically used to: generate a combination of questions and answers in response to a code retrieval request.
[0111] like Figure 1 As shown, the code retrieval method in this embodiment generates a combination of question and answer based on the user's question, i.e., the code retrieval request. It can be understood that the combination of question and answer can be retrieved from an existing vector database.
[0112] It is important to note that, such as Figure 1 As shown, when generating question-and-answer pairs, multiple vector databases may be called, or only one vector database may be called. The specific implementation of this application is not limited in this way.
[0113] Preferably, in order to speed up code retrieval results, it is usually possible to automatically detect which files have been modified, so that when generating question-and-answer pairs, only the modified parts are indexed and the results are found in the corresponding vector database.
[0114] In one embodiment of this application, the similarity module 420 is specifically used to: calculate the similarity score between the combined pairs of questions and answers.
[0115] A similarity calculation method is used to calculate similarity scores between multiple question-answer pairs. It can be understood that the similarity score is used to evaluate the degree of matching between question-answer pairs, thereby providing a basis for subsequent ranking and retrieval.
[0116] In one embodiment of this application, the matching module 430 is specifically used to: output the final matching result based on the similarity score; wherein, when generating the combination of question and answer, a multi-dimensional feature attribute is generated using a code retrieval model to characterize the relationship between the code fragment to be queried and the natural language query.
[0117] The final matching result is output based on the similarity score of the best match.
[0118] Specifically, when generating question-and-answer pairs, a multi-dimensional feature attribute is generated using a code retrieval model to characterize the relationship between the code snippet to be queried and the natural language query.
[0119] Specifically, this application proposes a code retrieval method called "Multi-dimensional Code QAMatching" (MdCoQA). The core of MdCoQA lies in utilizing the unique attributes of a multi-dimensional space to represent the relationship between code snippets and natural language queries, thereby achieving efficient and accurate code search. For example... Figure 1As shown, the overall architecture of MdCoQA includes a complete architecture of modules such as BERT embedding layer, projection layer of task-specific word representation, generation of QA representation, embedding interaction in multi-dimensional space, and similarity calculation based on multi-dimensional distance.
[0120] In one embodiment of this application, generating a question-and-answer pair in response to a code retrieval request includes:
[0121] Based on the code retrieval request, a combination of question and answer pairs is generated using the code retrieval model.
[0122] The code retrieval model is used to learn representations of question-and-answer pairs in a multi-dimensional space, so that correct question-and-answer pairs are closer together in the multi-dimensional space, while incorrect question-and-answer pairs are farther apart.
[0123] In one embodiment of this application, the code retrieval model is further used for:
[0124] The question-and-answer pairs are embedded into the multi-dimensional space, and a multi-dimensional distance function is used to measure the relationship between the natural language query and the query code.
[0125] In one embodiment of this application, the code retrieval model is further used for:
[0126] Similarity calculation based on multi-dimensional distance calculates distance in the multi-dimensional space and transforms the multi-dimensional distance into a similarity score through a linear layer.
[0127] In one embodiment of this application, the response to the code retrieval request further includes:
[0128] The pre-trained BERT model is used to transform natural language queries and code snippets into initial vector representations.
[0129] In one embodiment of this application, the pre-trained BERT model is a static BERT model.
[0130] In one embodiment of this application, the initial vector representation is converted into a word representation suitable for code retrieval tasks through a projection layer.
[0131] In one embodiment of this application, the projection layer employs a single-layer neural network, using a non-linear activation function to process each word in the natural language query and the code snippet to be queried, and the parameters of the projection layer are shared between the query and the code.
[0132] In one embodiment of this application, a holistic representation of a natural language query and a code fragment to be queried is generated by aggregating all word representations in the projection layer sequence.
[0133] In one embodiment of this application, the method further includes:
[0134] The code retrieval model is trained using a pairwise hinge loss function.
[0135] In one embodiment of this application, the method further includes:
[0136] Riemannian optimization techniques were used to update the code retrieval model parameters.
[0137] It is understood that the above-described code retrieval device can implement each step of the code retrieval method provided in the foregoing embodiments. The relevant explanations of the code retrieval method are applicable to the code retrieval device and will not be repeated here.
[0138] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 5 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0139] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0140] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0141] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a code retrieval mechanism at the logical level. The processor executes the program stored in memory and specifically performs the following operations:
[0142] In response to a code retrieval request, generate a combination of question and answer pairs;
[0143] Calculate the similarity score between the pairs of questions and answers; and
[0144] The final matching result is output based on the similarity score;
[0145] In this process, the combination of questions and answers is generated using a code retrieval model to generate multi-dimensional feature attributes that characterize the relationship between the code snippet to be queried and the natural language query.
[0146] The above is as stated in this application. Figure 2 The method executed by the code retrieval device disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0147] The electronic device can also perform Figure 2 The method executed by the code retrieval device, and the implementation of the code retrieval device in... Figure 2 The functions of the embodiments shown are not described in detail here.
[0148] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform... Figure 2 The method executed by the code retrieval device in the illustrated embodiment is specifically used to perform:
[0149] In response to a code retrieval request, generate a combination of question and answer pairs;
[0150] Calculate the similarity score between the pairs of questions and answers; and
[0151] The final matching result is output based on the similarity score;
[0152] In this process, the combination of questions and answers is generated using a code retrieval model to generate multi-dimensional feature attributes that characterize the relationship between the code snippet to be queried and the natural language query.
[0153] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0155] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0156] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0157] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0158] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0159] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0160] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0161] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0162] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A code retrieval method, wherein, The retrieval method includes: In response to a code retrieval request, generate a combination of question and answer pairs; Calculate the similarity score between the pairs of questions and answers; and The final matching result is output based on the similarity score; In this process, the combination of questions and answers is generated using a code retrieval model to generate feature attributes in a multi-dimensional space, representing the relationship between the code fragment to be queried and the natural language query.
2. The method as described in claim 1, wherein, The process of generating a question-and-answer pair in response to a code retrieval request includes: Based on the code retrieval request, a combination of question and answer pairs is generated using the code retrieval model. The code retrieval model is used to learn representations of question-and-answer pairs in a multi-dimensional space, so that correct question-and-answer pairs are closer together in the multi-dimensional space, while incorrect question-and-answer pairs are farther apart.
3. The method as described in claim 2, wherein, The code retrieval model is also used for: The question-and-answer pairs are embedded into the multi-dimensional space, and a multi-dimensional distance function is used to measure the relationship between the natural language query and the query code.
4. The method as described in claim 2, wherein, The code retrieval model is also used for: Similarity calculation based on multi-dimensional distance calculates distance in the multi-dimensional space and transforms the multi-dimensional distance into a similarity score through a linear layer.
5. The method as described in claim 1, wherein, The response prior to the code retrieval request also includes: The pre-trained BERT model is used to transform natural language queries and code snippets into initial vector representations.
6. The method of claim 5, wherein, The pre-trained BERT model is a static BERT model.
7. The method of claim 5, wherein, The initial vector representation is transformed into a word representation suitable for code retrieval tasks through a projection layer.
8. The method of claim 7, wherein, The projection layer employs a single-layer neural network, using a non-linear activation function to process each word in the natural language query and the code snippet to be queried, and the parameters of the projection layer are shared between the query and the code.
9. The method of claim 7, wherein, By aggregating all word representations in the projection layer sequence, an overall representation of the natural language query and the code fragment to be queried is generated.
10. The method of claim 1, wherein, The method further includes: The code retrieval model is trained using a pairwise hinge loss function.
11. The method of claim 1, wherein, The method further includes: Riemannian optimization techniques were used to update the code retrieval model parameters.
12. A code retrieval device, wherein, The retrieval device includes: The generation module is used to generate question-and-answer pairs in response to code retrieval requests; A similarity module is used to calculate a similarity score between pairs of questions and answers; and The matching module is used to output the final matching result based on the similarity score; In this process, the combination of questions and answers is generated using a code retrieval model to generate feature attributes in a multi-dimensional space, representing the relationship between the code fragment to be queried and the natural language query.
13. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the method of any one of claims 1 to 11.
14. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the method of any one of claims 1 to 11.