Code search method, system and equipment based on large language model classification and prototype integration

By applying large language models to classify and train expert models in code search, combining multimodal loss and integration modules, the semantic ambiguity and performance bottleneck problems caused by ambiguous queries in the existing technology are solved, and more efficient and accurate code search is achieved.

CN119961381AActive Publication Date: 2025-05-09GUANGDONG UNIV OF TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510438207.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Existing code search technologies have degraded performance when handling ambiguous manual queries, making it difficult to accurately reflect the core intentions of the code, resulting in semantic ambiguity and performance bottlenecks.

Method used

The code search method based on the classification and prototype integration of large language models is adopted, query-code pairs are classified through large language models, expert models are trained for different categories, multimodal hard negative sample loss is used for model training, and coarse-grained classification screening and fine-grained multimodal integration module are combined to optimize the search results.

Benefits of technology

Effectively narrow the semantic gap between query and code, improve the performance and accuracy of code search, enhance the robustness and generalization capabilities of the model, and be suitable for flexible adaptation of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961381A_ABST
    Figure CN119961381A_ABST
Patent Text Reader

Abstract

The invention discloses a code search method, system and device based on large language model classification and prototype integration, and relates to the technical field of code semantic analysis, and the method comprises the following steps: S01, cleaning a to-be-processed corpus segment, and extracting a query source Token and a code source Token to obtain a cleaned "query-code" source Token pair; s02, using a large language model to classify the data; s03, inputting the query-code source Token pairs of different categories into a pre-training model for model training to obtain expert models of different categories; s04, performing code search by utilizing the expert model to obtain a preliminary search result; s05, screening a preliminary search result; and S06, integrating the screened code search results to obtain a final search result. By adopting the method, the system and the equipment, the semantic difference between the query and the code can be effectively reduced, and the problem of semantic fuzziness possibly caused by the ambiguous query is solved, so that the code search performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code semantic analysis, and in particular to a code search method, system and device based on large language model classification and prototype integration. Background Art

[0002] To help developers retrieve relevant code snippets from open source platforms such as GitHub, Stack Overflow, etc., researchers have proposed many methods, known as code search, which requires developers to input natural language from websites to retrieve semantically matching code snippets during the development process.

[0003] Existing code search techniques can be roughly divided into three categories: information retrieval (IR)-based methods, deep learning methods without fusion, and deep learning methods with fusion strategies. The first category of research focuses on traditional information retrieval techniques to extract relevant code from software libraries. These methods rely heavily on regular expression matching and find it difficult to capture the deeper semantic relationship between queries and code. With the emergence of deep learning technology, more and more research has turned to using deep learning models to learn features from code. These methods belong to the second category and usually use a single model. Although deep learning has effectively improved the performance of code search, these methods still face challenges due to the semantic gap between query language and code language. This gap leads to performance bottlenecks, which was partially alleviated after the introduction of pre-trained models. In addition, some methods based on pre-trained models use techniques such as contrastive learning to improve code semantic understanding and search performance. For example, CoCoSoDa, a momentum contrastive learning algorithm, is introduced in the pre-trained model UniXcoder to learn better feature representations through more negative samples, which shows convincing performance on the large-scale benchmark dataset CodeSearchNet. However, these methods still struggle when dealing with semantically ambiguous human-written queries. The third category of methods is multimodal code search with fusion strategies, which combine multiple representations or results from pre-trained models to improve performance. But not all techniques can achieve good performance alone.

[0004] Current approaches mainly leverage pre-trained models and feature representation techniques such as contrastive learning to narrow the semantic distance between matching query-code pairs, and improve performance through fusion strategies. However, in real-world search scenarios, manually written queries are often ambiguous and may not accurately reflect the core intent of the code. Existing approaches often suffer from performance degradation in the face of such ambiguous queries. Recently, large language model techniques have been widely used in generation tasks such as code summarization and test generation, and have shown superior performance to traditional pre-trained models. We hypothesize that large language models, with their large parameter size, abundant training data, advanced architecture, and multi-task learning capabilities, can also outperform existing models in understanding ambiguous code queries. However, large language models have not yet been widely used in code search, and previous studies using retrieval-augmented generation techniques have shown limited success in applying large language models to code search. Summary of the invention

[0005] The purpose of the present invention is to provide a code search method, system and device based on large language model classification and prototype integration, which can effectively narrow the semantic gap between query and code, solve the problem of semantic ambiguity that may be caused by ambiguous queries, and thus improve the performance of code search.

[0006] To achieve the above object, the present invention provides a code search method based on large language model classification and prototype integration, the steps are as follows: S01. After performing data cleaning on the to-be-processed corpus segment containing multiple "query-code" pairs, the query source Token and the code source Token are extracted to obtain the cleaned "query-code" source Token pairs; S02. Classify the “query-code” source Token pairs using a large language model to obtain “query-code” source Token pairs of different categories; S03, inputting the "query-code" source Token pairs of different categories into multiple pre-trained models, and performing model training using a multimodal hard negative sample loss based on category characteristics to obtain trained expert models of different categories; S04. Use the trained expert models of different categories to perform code searches respectively to obtain preliminary search results; S05, using a coarse-grained large language model classification and screening module to screen the preliminary search results to obtain screened code search results; S06. Using a fine-grained multimodal integration module, the screened code search results are integrated to obtain a final search result.

[0007] Preferably, in S01, the data cleaning includes lowercase conversion, underscore splitting, camel case splitting and word form restoration.

[0008] Preferably, in S03, a multimodal hard negative sample loss based on category characteristics is used to perform model training to obtain trained expert models of different categories, including: 1) Use the multimodal hard negative sample loss based on category characteristics to calculate the loss value of the model's output result; 2) Based on the loss value, update the parameters in the improved model by back-propagating the output result; 3) Repeat 1) and 2) until the parameters converge to obtain a trained improved TranS0former model.

[0009] Preferably, the multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; the expression of the multimodal hard negative sample loss based on category characteristics is: , In the formula, represents the focal loss, represents the triple multimodal loss, represents the weight of the focal loss, represents the weight of triplet loss; The focal loss improves the convergence speed by focusing on the hard negative samples in the training data; the expression of the focal loss is: , In the formula, Indicates the prediction The probability that a query matches a code, Indicates that the The labels of the training samples are converted to one-hot encoding. is the amount of data to be predicted, is the adjustment coefficient for the number of positive and negative samples, The adjustment coefficient for classifying difficult and easy samples; The triple multimodal loss includes inter-modal loss and intra-modal loss; the inter-modal loss represents the loss between code and query, and the intra-modal loss represents the loss between code and code, and query and query; the expression of the triple multimodal loss is: , , , , , In the formula, Represents two features and The cosine distance between Representation and Code Anchors Positive samples of codes with the same category, Representation and Code Anchors There are negative samples of codes of different categories, Representation and Query Anchors Positive query samples with the same category, Representation and Query Anchors have represents the specific training category of the model, and are the minimum distances that need to be maintained between positive and negative pairs, and Greater than m.

[0010] Preferably, in S03, during the model training process, after the expert model training is completed, for the difficult samples for which the expert model of each category performs poorly, the enhanced expert model is retrained, and the expression is: , In the formula, and Indicates The more difficult samples in the data where the expert model performs poorly, and It represents the feature representation that the model is continuously optimized through the loss function during the training process. Represents an enhanced expert model.

[0011] Preferably, in S05, the coarse-grained large language model classification and screening module uses the large language model to predict a category for the query and the code respectively, and assigns a higher confidence to the search results predicted to be of the same category. The codes and queries of the same category should have a relatively high similarity score, and the codes and queries of different categories should have a lower similarity score. The expression of the similarity score is: , In the formula, Indicates A sample of queries, Indicates Code samples, Represents the confidence coefficient added when the predictions are exactly the same category, Indicates the confidence coefficient added when there is an intersection between the predicted categories. and They represent the first The query sample and Classification labels for code samples.

[0012] Preferably, in S06, the fine-grained multimodal integration module integrates the preliminary search results of expert models of different categories using a prototype-based integration method, which is expressed as: , In the formula, Indicates A sample of queries, Indicates Code samples, Indicates The code search results after screening by category experts, represents an integrated method; Ensemble methods The purpose is to accurately select the best expert based on the characteristics of the input data, generate a probability distribution of the expert's prediction accuracy based on the input query, and the final output is the weighted sum of all expert outputs, which is expressed as: , In the formula, Indicates The query features, Indicates The prototype of the class query data, represents the cosine similarity, Represents the first The classification labels of query samples, Indicates the predicted category With category When matching, add the confidence coefficient; Among them, the prototype It represents a representative feature vector that can capture the essence of a class of data. Assuming that query samples in the same category will show similar features after training, these features are defined as prototypes. Features of different categories have different prototypes. The prototype of a category is represented by taking the average of all features of the category trained: , In the formula, represents the first The query features, Indicates the training category The total number of query samples.

[0013] A code search system based on large language model classification and prototype integration is applied to any of the above methods, and the system includes: A data processing module is used to perform data cleaning on the corpus to be processed containing multiple "query-code" pairs and then extract the query source token and the code source token; A data classification module, used to classify the “query-code” source Token pairs using a large language model; Code search module for: Inputting the "query-code" source Token pairs of different categories into multiple pre-trained models, and using a multimodal hard negative sample loss based on category characteristics to train the models, to obtain trained expert models of different categories; the multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; The code search is performed respectively using trained expert models of different categories to obtain preliminary search results; the preliminary search results are screened using a coarse-grained large language model classification and screening module to obtain code search results after screening the preliminary search results of experts of different categories; the coarse-grained large language model classification and screening module uses a large language model to predict a category for the query and the code respectively, and gives a higher confidence to the search results predicted to be of the same category; the filtered code search results are integrated using a fine-grained multimodal integration module to obtain the final search results; the fine-grained multimodal integration module uses a prototype-based integration method to integrate the preliminary search results of expert models of different categories.

[0014] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned methods when executing the computer program.

[0015] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects: (1) Improving search accuracy and relevance Through the large language model, the query-code pairs are finely classified, and expert models for different categories are trained so that each model focuses on the features of a specific field and reduces cross-domain noise interference. Combining focal loss (handling sample imbalance) and triple multimodal loss (optimizing cross-modal alignment) enhances the model's ability to distinguish difficult samples and improves the accuracy of query-code matching. Through prototype weighted integration of the outputs of different expert models, multi-category features are integrated to avoid single model bias and further optimize the final result.

[0016] (2) Enhance model robustness and generalization The code and query representations are unified through operations such as lowercase and camel case splitting, reducing the impact of grammatical differences on the model and improving the adaptability to diverse inputs. The enhanced model is retrained for samples where the expert model performs poorly to avoid overfitting and improve the coverage of edge cases. The triple loss simultaneously constrains the similarity between modalities (code-query) and within modalities (code-code, query-query), enhancing the model's capture of semantic consistency.

[0017] (3) Accelerate training convergence efficiency By adjusting the weights of difficult and easy samples (parameter γ) and the category balance coefficient (α), the model can accelerate the learning of key samples and shorten the convergence time. The coarse-grained classification and screening module quickly filters low-confidence results to reduce subsequent calculation overhead; the fine-grained integration module only performs weighted fusion on highly relevant candidates to improve overall efficiency.

[0018] (4) Support flexible adaptation of complex scenarios The prototype representation dynamically integrates text and code features, which is suitable for complex code search needs across languages ​​and frameworks. The expert model can be expanded as needed (such as adding new categories), and the modular design of the system facilitates function iteration and maintenance.

[0019] (5) Practical application value Based on pre-trained models and Transformer architecture, it can be adapted to GPU acceleration to meet the real-time search requirements of large-scale code bases. It provides interpretability of search results through classification labels to help developers understand the matching logic.

[0020] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0022] Figure 1 A flowchart of a code search method based on large language model classification and prototype integration according to an embodiment of the subject title of the present invention; Figure 2 A schematic diagram of model training and actual prediction after training according to an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a code search system according to an embodiment of the present invention; Figure 4 The figure is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.

[0023] Reference numerals 100. Electronic device; 110. Memory; 120. Processor. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. Example

[0026] like Figure 1 As shown in the figure, the code search method based on large language model classification and prototype integration has the following steps: S01. After performing data cleaning on a corpus segment to be processed containing multiple "query-code" pairs, query source Tokens and code source Tokens are extracted to obtain cleaned "query-code" source Token pairs.

[0027] In the application, Python toolkits and specific cleaning methods can be used to clean the data of the corpus to be processed and extract its tokens. Examples of data cleaning include lowercase, underscore splitting, camel case splitting, and word form restoration.

[0028] Data cleaning can remove noise data such as comments, blank lines, non-ASCII characters, etc. in the code to ensure the purity of the source token. The cleaned code is segmented to generate a source token sequence. At the same time, the target token sequence in the corresponding target language (such as natural language) can be prepared as a summary reference.

[0029] During the training process, the corpus segments to be processed are the sub-datasets of six code languages ​​in the CodeSearchNet dataset, which can include Ruby dataset, JavaScript dataset, Java dataset, Go dataset, Php dataset and Python dataset; the Ruby dataset contains a total of 27,588 data instances, 24,927 instances are used for training, and 1,400 instances are used for verification; the Ruby dataset contains a total of 27,588 data instances, 24,927 instances are used for training, and 1,400 instances are used for verification; the JavaScript dataset contains a total of 65,201 data instances, 5 8,025 instances are used for training and 3,885 instances are used for validation; the Java dataset contains 181,061 data instances, of which 164,923 instances are used for training and 5,183 instances are used for validation; the Go dataset contains a total of 182,735 data instances, of which 167,288 instances are used for training and 7,325 instances are used for validation; the Php dataset contains a total of 268,237 data instances, of which 241,241 instances are used for training and 12,982 instances are used for validation; the Python dataset contains a total of 280,652 data instances, of which 251,820 instances are used for training and 13,914 instances are used for validation.

[0030] The following operations can be performed when cleaning the data of the corpus to be processed: Tokenize the code snippets in the corpus to be processed and convert each token to lowercase; Use the NLTK library to split each identifier based on camel case and underscores; for example, the identifier "IndexOutOfBoundError" will be broken down into "index", "out", "of", "bound", and "error"; Each word was lemmatized using spaCy2, standardizing the form to improve uniformity.

[0031] S02. Use a large language model to classify the "query-code" source Token pairs to obtain "query-code" source Token pairs of different categories.

[0032] S03, inputting the "query-code" source token pairs of different categories into multiple pre-trained models, and using multimodal hard negative sample loss based on category characteristics to train the models, to obtain trained expert models of different categories. Including: 1) Use the multimodal hard negative sample loss based on category characteristics to calculate the loss value of the model's output result; 2) Based on the loss value, update the parameters in the improved model by back-propagating the output result; 3) Repeat 1) and 2) until the parameters converge to obtain a trained improved TranSformer model.

[0033] like Figure 2 As shown in the upper part of the figure, the classified source tokens are input into multiple models for training. On the right is the multimodal hard negative sample loss based on category characteristics used in the training process. The model trained in this way can be regarded as a code search expert that only focuses on a certain category.

[0034] The multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; the expression of the multimodal hard negative sample loss based on category characteristics is: , In the formula, represents the focal loss, represents the triple multimodal loss, represents the weight of the focal loss, represents the weight of triplet loss; The focal loss improves the convergence speed by focusing on the hard negative samples in the training data; the expression of the focal loss is: , In the formula, Indicates the prediction The probability that a query matches a code, Indicates that the The labels of the training samples are converted to one-hot encoding. is the amount of data to be predicted, is the adjustment coefficient for the number of positive and negative samples, The adjustment coefficient for classifying difficult and easy samples; The triple multimodal loss includes inter-modal loss and intra-modal loss; the inter-modal loss represents the loss between code and query, and the intra-modal loss represents the loss between code and code, and query and query; the expression of the triple multimodal loss is: , , , , , In the formula, Represents two features and The cosine distance between Representation and Code Anchors Positive samples of codes with the same category, Representation and Code Anchors There are negative samples of codes of different categories, Representation and Query Anchors Positive query samples with the same category, Representation and Query Anchors have represents the specific training category of the model, and are the minimum distances that need to be maintained between positive and negative pairs, and Greater than m.

[0035] After completing the training of the expert model, for the more difficult samples where the expert model of each category performs poorly, the enhanced expert model is retrained, and the expression is: ,

[0036] In the formula, and Indicates The more difficult samples in the data where the expert model performs poorly, and It represents the feature representation that the model is continuously optimized through the loss function during the training process. Represents an enhanced expert model.

[0037] S04. Use the trained expert models of different categories to perform code searches and obtain preliminary search results.

[0038] S05, using the coarse-grained large language model classification and screening module, screening the preliminary search results to obtain screened code search results; Figure 2 The initial search results are coarsely screened by a coarse-grained large language model classification and screening module CCR. This module uses a large language model to predict a category for the query and code respectively, and gives higher confidence to search results predicted to be of the same category. Codes and queries of the same category should have relatively high similarity scores, while codes and queries of different categories should have lower similarity scores. The expression of the similarity score is: , In the formula, Indicates A sample of queries, Indicates Code samples, Represents the confidence coefficient added when the predictions are exactly the same category, indicates the confidence coefficient added when there is an intersection in the predicted categories, and They represent the first The query sample and Classification labels for code samples.

[0039] S06. Using a fine-grained multimodal integration module, the filtered code search results are integrated to obtain a final search result. The filtered code search results are passed through a fine-grained multimodal integration module FMPI to obtain the final result. This module uses a prototype-based integration method to integrate the preliminary search results of expert models of different categories, and its expression is: , In the formula, Indicates A sample of queries, Indicates Code samples, Indicates The code search results after screening by category experts, Represents an ensemble method.

[0040] Ensemble methods The purpose is to accurately select the best expert based on the characteristics of the input data, generate a probability distribution of the expert's prediction accuracy based on the input query, and the final output is the weighted sum of all expert outputs, which is expressed as: , In the formula, Indicates The query features, Indicates The prototype of the class query data, represents the cosine similarity, Represents the first The classification labels of query samples, Indicates the predicted category With category When matching, add a confidence factor.

[0041] Among them, the prototype It represents a representative feature vector that can capture the essence of a class of data. Assuming that query samples in the same category will show similar features after training, these features are defined as prototypes. Features of different categories have different prototypes. The prototype of a category is represented by taking the average of all features of the category trained: , In the formula, represents the first The query features, Indicates the training category k The total number of query samples.

[0042] Through the above process, accurate code search can be achieved. Corresponding to the above application function implementation method embodiment, the present invention also provides a code search system, device and corresponding embodiment based on large language model classification and prototype integration. Figure 3 shown.

[0043] The code search system includes: The data processing module is used to extract the query source token and code source token after performing data cleaning on the corpus to be processed (including multiple "query-code" pairs); A data classification module is used to classify the “query-code” source Token pairs using a large language model: Code search module for: Inputting the "query-code" source Token pairs of different categories into multiple pre-trained models, and using a multimodal hard negative sample loss based on category characteristics to train the models, to obtain trained expert models of different categories; the multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; The code search is performed respectively using trained expert models of different categories to obtain preliminary search results; the preliminary search results are screened using a coarse-grained large language model classification and screening module to obtain code search results after screening the preliminary search results of experts of different categories; the coarse-grained large language model classification and screening module uses a large language model to predict a category for the query and the code respectively, and gives a higher confidence to the search results predicted to be of the same category; the filtered code search results are integrated using a fine-grained multimodal integration module to obtain the final search results; the fine-grained multimodal integration module uses a prototype-based integration method to integrate the preliminary search results of expert models of different categories.

[0044] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.

[0045] An embodiment of the present invention further provides an electronic device, such as Figure 4 As shown, the electronic device 100 includes a memory 110 and a processor 120 .

[0046] The processor 120 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0047] The memory 110 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, ROM can store static data or instructions required by the processor 120 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even after the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as a permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as a dynamic random access memory. The system memory may store some or all instructions and data required by the processor at runtime. In addition, the memory 110 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 110 may include a readable and / or writable removable storage device, such as a laser disc (CD), a read-only digital versatile disc (such as a DVD-ROM, a double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (such as an SD card, a mini SD card, a Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.

[0048] The memory 110 stores executable codes, and when the executable codes are processed by the processor 120 , the processor 120 can execute part or all of the above-mentioned methods.

[0049] The remaining technical features in the above embodiments can be flexibly selected by those skilled in the art according to actual conditions to meet different specific practical needs. However, it is obvious to those skilled in the art that it is not necessary to adopt these specific details to implement the present invention. In other examples, in order to avoid confusing the present invention, the well-known components, structures or parts are not specifically described, which are all within the technical protection scope defined by the technical solution claimed for protection in the claims of the present invention.

[0050] Modifications and changes made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the scope of protection of the claims attached to the present invention. In the above description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details are not necessary to practice the present invention. In other examples, in order to avoid confusing the present invention, well-known technologies, such as specific construction details, operating conditions and other technical conditions, are not specifically described.

[0051] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A code search method based on large language model classification and prototype integration, characterized in that: Here are the steps: S01. After performing data cleaning on the corpus segment to be processed containing multiple "query-code" pairs, the query source Token and the code source Token are extracted to obtain the cleaned "query-code" source Token pair; S02. Classify the “query-code” source Token pairs using a large language model to obtain different categories of “query-code” source Token pairs; S03, inputting the "query-code" source Token pairs of different categories into multiple pre-trained models, and performing model training using a multimodal hard negative sample loss based on category characteristics to obtain trained expert models of different categories; S04. Use the trained expert models of different categories to perform code searches respectively to obtain preliminary search results; S05, using a coarse-grained large language model classification and screening module to screen the preliminary search results to obtain screened code search results; S06. Using a fine-grained multimodal integration module, the screened code search results are integrated to obtain a final search result.

2. The code search method based on large language model classification and prototype integration according to claim 1 is characterized in that: In S01, the data cleaning includes lowercase conversion, underscore splitting, camel case splitting and word form restoration.

3. The code search method based on large language model classification and prototype integration according to claim 1, characterized in that: In S03, the model is trained using a multimodal hard negative sample loss based on category characteristics to obtain trained expert models of different categories, including: 1) Use the multimodal hard negative sample loss based on category characteristics to calculate the loss value of the model's output result; 2) Based on the loss value, update the parameters in the improved model by back-propagating the output result; 3) Repeat 1) and 2) until the parameters converge to obtain a trained improved TranSformer model.

4. The code search method based on large language model classification and prototype integration according to claim 3 is characterized in that: The multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; the expression of the multimodal hard negative sample loss based on category characteristics is: , In the formula, represents the focal loss, represents the triple multimodal loss, represents the weight of the focal loss, represents the weight of triplet loss; The focal loss improves the convergence speed by focusing on the hard negative samples in the training data; the expression of the focal loss is: , In the formula, Indicates the prediction The probability that a query matches a code, Indicates that the The labels of the training samples are converted to one-hot encoding. is the amount of data to be predicted, is the adjustment coefficient for the number of positive and negative samples, The adjustment coefficient for classifying difficult and easy samples; The triple multimodal loss includes inter-modal loss and intra-modal loss; the inter-modal loss represents the loss between code and query, and the intra-modal loss represents the loss between code and code, and query and query; the expression of the triple multimodal loss is: , , , , In the formula, Represents two features and The cosine distance between Representation and Code Anchors Positive samples of codes with the same category, Representation and Code Anchors There are negative samples of codes of different categories, Representation and Query Anchors Positive query samples with the same category, Representation and Query Anchors have represents the specific training category of the model, and are the minimum distances that need to be maintained between positive and negative pairs, and Greater than m.

5. The code search method based on large language model classification and prototype integration according to claim 4 is characterized in that: In S03, during the model training process, after the expert model training is completed, for the difficult samples where the expert model of each category performs poorly, the enhanced expert model is retrained, and the expression is: , In the formula, and Indicates The more difficult samples in the data where the expert model performs poorly, and It represents the feature representation that the model is continuously optimized through the loss function during the training process. Represents an enhanced expert model.

6. The code search method based on large language model classification and prototype integration according to claim 1, characterized in that: In S06, the fine-grained multimodal integration module integrates the preliminary search results of expert models of different categories using a prototype-based integration method, which is expressed as: , In the formula, Indicates A sample of queries, Indicates Code samples, Represents the confidence coefficient added when the predictions are exactly the same category, Indicates the confidence coefficient added when there is an intersection between the predicted categories. and They represent the first The query sample and Classification labels for code samples.

7. The code search method based on large language model classification and prototype integration according to claim 1, characterized in that: In S06, the fine-grained multimodal integration module integrates the preliminary search results of expert models of different categories using a prototype-based integration method, which is expressed as: , In the formula, Indicates A sample of queries, Indicates Code samples, Indicates The code search results after screening by category experts, represents an integrated method; Ensemble methods The purpose is to accurately select the best expert based on the characteristics of the input data, generate a probability distribution of the expert's prediction accuracy based on the input query, and the final output is the weighted sum of all expert outputs, which is expressed as: , In the formula, Indicates The query features, Indicates The prototype of the class query data, represents the cosine similarity, Represents the first The classification labels of query samples, Indicates the predicted category With category When matching, add the confidence coefficient; Among them, the prototype It represents a representative feature vector that can capture the essence of a class of data. Assuming that query samples in the same category will show similar features after training, these features are defined as prototypes. Features of different categories have different prototypes. The prototype of a category is represented by taking the average of all features of the category trained: , In the formula, represents the first The query features, Indicates the training category The total number of query samples.

8. A code search system based on large language model classification and prototype integration, characterized in that: The method applied to any one of claims 1 to 7, wherein the system comprises: The data processing module is used to perform data cleaning on the corpus to be processed containing multiple "query-code" pairs and then extract the query source token and the code source token; A data classification module, used to classify the "query-code" source Token pairs using a large language model; Code search module for: Inputting the "query-code" source Token pairs of different categories into multiple pre-trained models, and using the multimodal hard negative sample loss based on category characteristics to train the models, to obtain trained expert models of different categories; the multimodal hard negative sample loss based on category characteristics includes focal loss and triple multimodal loss; The code search is performed respectively using trained expert models of different categories to obtain preliminary search results; the preliminary search results are screened using a coarse-grained large language model classification and screening module to obtain code search results after screening the preliminary search results of experts of different categories; the coarse-grained large language model classification and screening module uses a large language model to predict a category for the query and the code respectively, and gives a higher confidence to the search results predicted to be of the same category; the filtered code search results are integrated using a fine-grained multimodal integration module to obtain the final search results; the fine-grained multimodal integration module uses a prototype-based integration method to integrate the preliminary search results of expert models of different categories.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Code searching method based on annotation semantic information

    CN112507065A

  • Code search method based on post-interaction mechanism

    CN117112851A

  • Efficient code searching method based on product quantization

    CN117951251A

  • Code search model training method, code search method and equipment

    CN117992643A

  • Code searching method and device, electronic equipment and storage medium

    CN118035424A