Code retrieval method and device, equipment, medium and product

By building a multi-dimensional index library and semantic understanding, the problems of inaccurate results and low efficiency in traditional code retrieval methods are solved, and efficient and accurate code retrieval is achieved.

CN120632170APending Publication Date: 2025-09-12SAIC MOTOR OVERSEAS INTELLIGENT MOBILITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510725051.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional code retrieval methods rely on simple keyword matching, resulting in inaccurate retrieval results, low retrieval efficiency, and difficulty in meeting the needs of developers.

Method used

Build multiple index libraries, including indexes of code structure, important variables, code comments, and original code dimensions, use large language models for semantic understanding, convert query instructions into semantic vectors and keyword phrases, merge search results and calculate similarity weights, and adjust scores to retain the highest-scoring results.

Benefits of technology

It improves the accuracy and efficiency of code retrieval, can fully capture the structure and relationships in the code, and provide high-quality, easy-to-understand retrieval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632170A_ABST
    Figure CN120632170A_ABST
Patent Text Reader

Abstract

The invention relates to a code retrieval method and device, equipment, a medium and a product. The code retrieval method comprises the steps that multiple indexes are constructed, multiple index databases of the multiple indexes are correspondingly generated, and each index comprises a keyword index and / or a vector index; a query instruction input by a user is converted into a semantic vector and a keyword group, and based on the semantic vector and the keyword group, the index databases are retrieved to obtain retrieval results; the retrieval results are combined, repeated code snippets are removed, and the retrieval results after combination and duplicate removal are divided into class retrieval results and method retrieval results; respectively calculating the weighted sum of the similarity of each class retrieval result and the method retrieval result so as to determine the score of each retrieval result; the scores of the class retrieval results are adjusted, and n class retrieval results with the highest scores are reserved; the scores of the method retrieval results are adjusted, and m method retrieval results with the highest scores are reserved; and returning n class retrieval results and m method retrieval results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer data processing, and in particular to a code retrieval method, apparatus, device, medium, and product. Background Art

[0002] With the continuous expansion of software development and the dramatic increase in code volume, fast and efficient code retrieval has become increasingly important. Traditional code retrieval methods often rely on simple keyword matching, failing to fully understand the structure and semantics of the code. This leads to inaccurate search results and requires constant optimization of search keywords and filtering of search results based on these keywords. This results in low retrieval efficiency and fails to meet the needs of developers. Summary of the Invention

[0003] In order to solve the problem that the current code retrieval results in inaccurate clothing replacement, thereby resulting in low retrieval efficiency, the present application proposes a code retrieval method.

[0004] The code retrieval method includes: constructing multiple indexes and generating multiple index libraries corresponding to the multiple indexes; converting the query instructions input by the user into semantic vectors and keyword phrases, and based on the semantic vectors and the keyword phrases, searching the multiple index libraries respectively to obtain retrieval results; merging the retrieval results, removing duplicate code fragments, and dividing the merged and deduplicated retrieval results into class retrieval results and method retrieval results; calculating the weighted sum of the similarities of each class retrieval result and the method retrieval result respectively to determine the score of each retrieval result; adjusting the score of the class retrieval result based on the method retrieval result, and retaining the n class retrieval results with the highest scores; and adjusting the score of the method retrieval result based on the class retrieval result, and retaining the m method retrieval results with the highest scores; returning the n class retrieval results and the m method retrieval results.

[0005] Optionally, the multiple indexes are constructed based on multiple information dimensions, wherein each index includes a keyword index and / or a vector index.

[0006] Optionally, the information dimension includes: code structure dimension, important variable dimension, code annotation dimension, code interpretation dimension and original code dimension.

[0007] Optionally, constructing multiple indexes for the code based on the information dimension of the code includes:

[0008] For the code structure dimension, a code structure index is constructed based on the path information, class name and function name of the code;

[0009] For the important variable dimension, a large language model is used to identify important variables in the code, and an important variable index is constructed based on the name and type of the important variables;

[0010] For the code comment dimension, construct a code comment index based on the document strings or comments of the code;

[0011] For the code interpretation dimension, a large language model is used to generate a semantic description of the code, and a code interpretation index is constructed based on the semantic description, wherein the semantic description includes an explanation of the function of the method or class of the code, the role of the parameters, the meaning of the return value, and the dependency relationship between the codes;

[0012] For the original code dimension, an original code index is constructed based on the source code of the code.

[0013] Optionally, the query instruction is in natural language form.

[0014] Optionally, the query instruction is parsed based on a natural language form of the query instruction, and keywords in the query instruction are extracted to convert the query instruction into a keyword group.

[0015] Optionally, the query instruction is semantically encoded using a large language model to convert the query instruction into a semantic vector.

[0016] Optionally, adjusting the score of the class search result based on the method search result includes:

[0017] When the method search results corresponding to one or more of the class search results have high scores, the scores of the one or more class search results are increased.

[0018] Optionally, adjusting the score of the method retrieval result based on the class retrieval result includes:

[0019] When the class search results corresponding to one or more of the method search results have high scores, the scores of the one or more method search results are increased.

[0020] Another aspect of the present application further provides a code retrieval device, comprising:

[0021] An index building unit, configured to build a plurality of indexes and generate a plurality of index libraries corresponding to the plurality of indexes;

[0022] An instruction retrieval unit, configured to convert a query instruction input by a user into a semantic vector and a keyword group, and to search the plurality of index libraries based on the semantic vector and the keyword group to obtain a retrieval result;

[0023] A merging and deduplication unit is used to merge the search results, remove duplicate code snippets, and divide the merged and deduplicated search results into class search results and method search results;

[0024] a score calculation unit, configured to calculate a weighted sum of similarities between each of the class search results and the method search results, to determine a score for each of the search results;

[0025] a score adjustment unit, configured to adjust the scores of the class retrieval results based on the method retrieval results and retain the n class retrieval results with the highest scores; and to adjust the scores of the method retrieval results based on the class retrieval results and retain the m method retrieval results with the highest scores;

[0026] The result output unit is used to return the n class search results and the m method search results.

[0027] Another aspect of the present application also provides an electronic device, which includes a memory storing computer-executable instructions and a processor; when the instructions are executed by the processor, the device implements any of the methods described above.

[0028] Another aspect of the present application further provides a computer-readable medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement any of the methods described above.

[0029] Another aspect of the present application further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method as described above is implemented.

[0030] This application constructs indexes of multiple different information dimensions based on multiple information dimensions, and the index of each information dimension can include keyword index and vector index, so as to capture classes, methods, variables, comments, etc. in the code and their relationships with each other, rather than just searching through simple keywords, which greatly improves the accuracy of the search results and thus improves the efficiency of the search. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A flow chart of a code retrieval method according to an embodiment of the present application is shown.

[0032] Figure 2 A schematic diagram of an index library of different information dimensions according to an embodiment of the present application is shown.

[0033] Figure 3 A schematic diagram of a code retrieval device according to an embodiment of the present application is shown.

[0034] Figure 4 A schematic diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0035] In order to make the above-mentioned purposes, features and advantages of this application more obvious and easy to understand, the early warning method for ocean transportation of new energy vehicles proposed in this application is further explained in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] Figure 1 It is a flow chart of the code retrieval method proposed in this application.

[0037] In step 101, multiple indexes are constructed, and multiple index libraries corresponding to the multiple indexes are generated. The multiple indexes can be constructed based on different information dimensions. For example, different indexes can be constructed based on code structure, important variables, code annotations, code interpretation, and original code. Different methods can be used to construct corresponding indexes for different information dimensions.

[0038] Next, combine Figure 2 An example of constructing an index library based on different information dimensions is described below, where each index may include a keyword index and / or a vector index.

[0039] Code structure is typically represented in the format file.class.method, where file is the path to the code file, class is the class name to which the code belongs, and method is the function name. This structure reflects the code's hierarchical structure and ensures uniqueness of the index. For the code structure dimension, keywords can be extracted from the path information, function names, and class names to construct a keyword index for the code structure dimension. By semantically encoding the path information, function names, and class names, a vectorized model is used to construct a vector index for the code structure dimension.

[0040] Important variables are parameters or member variables that significantly affect method and class behavior, such as function input parameters, key return values, or attribute variables. For the important variable dimension, a large language model can be used to globally understand functions and classes in conjunction with the code context. This allows the identification of important variables within these functions and classes, as well as information such as their names and types, and their functions, to be constructed into a keyword index for this dimension. Typically, variable names reflect their function and purpose, and constructing a keyword index captures the matching relationship between user queries and code variables. The global understanding provided by the large language model ensures that the selected important variables are accurate and representative.

[0041] Code comments contain textual explanations of the developer's explanations of the code's functionality and logic. They are an important dimension for understanding the purpose of the code and are typically expressed as doc strings or comments. For the code comment dimension, keywords can be extracted from the content of the code comments to construct a keyword index for the code comment dimension. Alternatively, the content of the code comments can be vectorized and converted into computable vectors to construct a vector index for the code comment dimension, thereby supporting complex semantic retrieval. Furthermore, because processing code comments requires parsing natural language content, natural language processing (NLP) technology can be used to perform code comment analysis, keyword extraction, and entity recognition on the code comments. These can then be associated with the code context when building the index to ensure that the correlation between code comments and code is fully considered during retrieval.

[0042] Code interpretation includes the functionality of a function or class, the purpose of its parameters, the meaning of its return value, and dependencies between code. For the code interpretation dimension, a large language model can be used to interpret the functionality of a code class or function, generating explanations in the form of natural language semantic descriptions. Keywords can then be extracted from the explanations to construct a keyword index for the code interpretation dimension. Alternatively, the explanations can be vectorized to generate vector representations to construct a vector index for semantic retrieval in the code interpretation dimension. To provide rich code information during retrieval, the large language model must consider the code's context and combine method calls, comments, and code logic to generate comprehensive and accurate explanations in the form of language descriptions.

[0043] Original code refers to the original form of the code, that is, the code form itself. For the original code dimension, keywords, operators, class names, method names, etc. in the code can be extracted to build a keyword index for the original code dimension; you can also use a vectorization model suitable for the code to convert the code into a semantic vector to build a vector index for the original code dimension. Building an index based on the code itself ensures that the literal information contained in the code can be directly matched and retrieved. The vectorization model can use any pre-trained model suitable for code, such as CodeBERT, GraphCodeBERT, or CodeT5. When you need to directly query the details of the code implementation, the index built using the original code dimension can achieve rapid retrieval and recall of relevant content.

[0044] It is understandable that different information dimensions can be considered based on different needs and standards. For example, in order to achieve visualization, the code structure, variables, comments and the relationship between them can be visually represented by introducing the knowledge graph dimension, thereby building a graph database to achieve more efficient retrieval. It is also possible to analyze the context dependencies of the code and establish a context dependency graph for understanding the logical flow and dependencies of the code, thereby building a context dependency graph index library. In the index library of different information dimensions, there may be different forms of indexes. In addition, as mentioned above, the code structure index, code comment index, code interpretation index and original code index can all include keyword indexes and vector indexes of their respective information dimensions, and the important variable index can only include keyword indexes. However, it can be decided to construct different forms of indexes for each information dimension according to needs, and in addition to keyword indexes and vector indexes, more forms of indexes can also be included, such as building a visual data index in the aforementioned graph database.

[0045] Through the above method, code information can be constructed into multiple different indexes based on multiple information dimensions. Each index based on information dimensions can also include keyword indexes and / or vector indexes to meet different retrieval requirements.

[0046] In step 102, the query instruction input by the user is converted into a semantic vector and a keyword phrase, and based on the semantic vector and the keyword phrase, multiple index libraries are searched separately to obtain search results. Since the index construction includes a keyword index and a vector index, the query instruction in the natural language form input by the user needs to be converted into a semantic vector and a keyword phrase first so that the search can be performed more quickly and accurately. The query text of the query instruction in the natural language form can be parsed to decompose it into a keyword phrase containing multiple keywords. Among them, the keywords can match the variable names, function names, class names and comment contents in the code, etc., which are expressed in the form of words or phrases. At the same time, a large language model can be used to semantically encode the query instruction to convert the query instruction into a semantic vector. The semantic vector can capture the semantic information in the query instruction so that code fragments with similar meanings can be searched in the index library, thereby supporting fuzzy matching and more complex semantic retrieval.

[0047] After searching different index libraries according to the query instructions, multiple search results with different similarities will be returned. In order to improve the data processing efficiency of subsequent steps and the relevance and simplicity of the search results, a similarity threshold can be set manually to eliminate search results with similarities lower than the similarity threshold.

[0048] In step 103, the search results are merged, duplicate code snippets are removed, and the merged and deduplicated search results are divided into class search results and method search results. After retrieving multiple search results from the index library of each information dimension, the multiple search results are merged, and the duplicate code snippets are deduplicated to improve the simplicity of the code snippets. In addition, although it is necessary to dedupe the duplicate code snippets, if the duplicate code snippets are retrieved from the index library based on different information dimensions, the similarity of the code snippets retrieved in different dimensions can be recorded for subsequent operations. After the search results are merged and deduplicated, the search results need to be divided into class search results and method search results based on the type of code corresponding to the search results, and a class result list and a method result list can be formed respectively. Among them, the class search results in the class result list are classes related to the user query, and the method search results in the method result list are functions or methods related to the user query.

[0049] In step 104, the weighted sum of the similarities of each class search result and method search result is calculated to determine the score of each search result. For example, for a class search result A, its similarities in different dimensions are a1, a2, a3, a4, and a5 respectively. The similarities in the five dimensions correspond to five weights w1, w2, w3, w4, and w5 respectively. Then, the score G of the class search result A can be determined by the following formula:

[0050]

[0051] Similarly, the score of each search result can be determined. Different weights can be assigned to the similarities corresponding to different information dimensions based on actual needs, particularly considering the intent of the user's query, the characteristics of the code, and the importance of each dimension. The optimal weight configuration can be determined through dynamic adjustment or model learning. For example, if the code comment dimension is more important for a particular query, a higher weight can be assigned to the similarity of the code comment dimension, thereby changing the score ranking of the search results.

[0052] In step 105, based on the method retrieval results, the scores of the class retrieval results are adjusted, and the n class retrieval results with the highest scores are retained; and based on the class retrieval results, the scores of the method retrieval results are adjusted, and the m method retrieval results with the highest scores are retained.

[0053] Because classes and methods in code are correlated, there must also be a correlation between class and method search results. In other words, when a class's methods have a high score, the class should also have a high score; and when a method's corresponding class has a high score, the method should also have a high score. Therefore, the scores of method and class search results can be adjusted accordingly. When using method search results to adjust the scores of class search results, if a method in a class is highly relevant to the user's query, that is, if the method search result corresponding to that method has a high score, the score of the class search result can be increased accordingly to reflect the overall relevance of the class to the user's query. Conversely, when using class search results to adjust the scores of method search results, if the class corresponding to a method is highly relevant to the user's query, that is, if the class search result corresponding to that method has a high score, the score of the method search result can be increased accordingly to reflect the contribution of the class to the user's query. Finally, in order to further improve the accuracy of the retrieval results, the n class retrieval results with the highest scores (i.e., the first n after sorting the scores from large to small) can be retained according to actual needs to ensure that they better match the user's query intention, and the m method retrieval results with the highest scores (i.e., the first m after sorting the scores from large to small) can be retained to ensure that they can reflect the semantic structure, contextual association and other important information dimensions of the code.

[0054] In step 106, n class search results and m method search results are returned. Finally, the retained n class search results and m method search results are returned to the user, thus completing the user's query.

[0055] This method also builds multiple index libraries based on multiple information dimensions, which can comprehensively capture the structure and relationships in the code, associate the classes, methods, variables and comments in the code snippets, and effectively improve the accuracy of retrieval.

[0056] In addition, using a large language model to perform semantic understanding of the code can capture the function, logic, and contextual information of the code, converting the code into an easy-to-understand semantic form that is suitable for various query instructions raised by users, greatly improving retrieval performance.

[0057] Finally, the contextual dependencies and quality indicators of the code are fully considered during the retrieval process, so that the retrieval results can not only match the user query, but also filter out high-quality, easy-to-understand and reusable code snippets, providing users with more layered and practical retrieval results.

[0058] Another aspect of the present application provides a code retrieval device 300, such as Figure 3 The code retrieval device 300 includes the following units:

[0059] An index building unit 301 is configured to build multiple indexes and generate multiple index libraries corresponding to the multiple indexes, wherein each index includes a keyword index and / or a vector index;

[0060] The instruction retrieval unit 302 is used to convert the query instruction input by the user into a semantic vector and a keyword group, and search the multiple index libraries based on the semantic vector and the keyword group to obtain a search result;

[0061] A merging and deduplication unit 303 is used to merge the search results, remove duplicate code snippets, and divide the merged and deduplicated search results into class search results and method search results;

[0062] A score calculation unit 304 is used to calculate the weighted sum of the similarities between each of the class search results and the method search results to determine the score of each of the search results;

[0063] A score adjustment unit 305 is configured to adjust the scores of the class search results based on the method search results and retain the n class search results with the highest scores; and to adjust the scores of the method search results based on the class search results and retain the m method search results with the highest scores;

[0064] The result output unit 306 is used to return the n class search results and the m method search results.

[0065] The code retrieval device can be implemented on various computers and servers to efficiently retrieve and accurately output code snippets that meet the user's search intention and expectations in response to the user's query instructions.

[0066] Now refer to Figure 4 , which is a block diagram of an electronic device 400 according to one embodiment of the present application. The electronic device 400 may include one or more processors 402, a system control logic 408 connected to at least one of the processors 402, a system memory 404 connected to the system control logic 408, a non-volatile memory (NVM) 406 connected to the system control logic 408, and a network interface 410 connected to the system control logic 408.

[0067] Processor 402 may include one or more single-core or multi-core processors. Processor 402 may include any combination of general-purpose processors and specialized processors (e.g., graphics processors, application processors, baseband processors, etc.). In the embodiments herein, processor 402 may be configured to execute one or more embodiments of providing early warning for abnormal conditions of new energy vehicles in ocean transportation as proposed in this application.

[0068] In some embodiments, system control logic 408 may include any suitable interface controller to provide any suitable interface to at least one of processors 402 and / or any suitable device or component in communication with system control logic 408 .

[0069] In some embodiments, the system control logic 408 may include one or more memory controllers to provide an interface to the system memory 404. The system memory 404 may be used to load and store data and / or instructions. In some embodiments, the system memory 404 of the electronic device 400 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).

[0070] The non-volatile memory 406 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 406 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of an HDD (Hard Disk Drive), a CD (Compact Disc) drive, and a DVD (Digital Versatile Disc) drive.

[0071] The non-volatile memory 406 may include a portion of storage resources installed on the device of the electronic device 400, or it may be accessible to the device but not necessarily a part of the device. For example, the non-volatile memory 406 may be accessed over a network via the network interface 410.

[0072] In particular, system memory 404 and non-volatile memory 406 may each include a temporary copy and a permanent copy of instructions 420. Instructions 420 may include instructions that, when executed by at least one of processors 402, cause electronic device 400 to implement the methods provided herein. In some embodiments, instructions 420, hardware, firmware, and / or software components thereof may additionally or alternatively be located in system control logic 408, network interface 410, and / or processor 402.

[0073] In some embodiments, the network interface 410 may be integrated with other components of the electronic device 400. For example, the network interface 410 may be integrated with at least one of the processor 402, the system memory 404, the non-volatile memory 406, and a firmware device (not shown) having instructions, and when at least one of the processors 402 executes the instructions, the electronic device 400 implements one or more of the various embodiments described herein. The network interface 410 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface.

[0074] In one embodiment, at least one of the processors 402 may be packaged together with logic for one or more controllers of the system control logic 408 to form a system-in-package (SiP). In one embodiment, at least one of the processors 402 may be integrated on the same die with logic for one or more controllers of the system control logic 408 to form a system-on-chip (SoC).

[0075] The electronic device 400 may further include an input / output (I / O) device 412. The input / output (I / O) device 412 may include a user interface to enable a user to interact with the electronic device 400; and a peripheral component interface may be designed to enable peripheral components to interact with the electronic device 400.

[0076] In some embodiments, the user interface may include, but is not limited to, a display (e.g., an LCD display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.

[0077] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0078] It should be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 400. In other embodiments of the present application, the electronic device 400 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0079] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0080] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described herein is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0081] One or more aspects of at least one embodiment may be implemented as representative instructions stored on a computer-readable storage medium, which represent various logic within a processor and, when read by a machine, causes the machine to fabricate logic for performing the techniques described herein. These representations, known as "IP cores," may be stored on a tangible, computer-readable storage medium and supplied to various customers or manufacturing facilities to load into fabrication machines that actually manufacture the logic or processor.

[0082] An embodiment of the present application discloses a computer-readable medium storing one or more programs executable by one or more processors to implement the early warning method for ocean transportation of new energy vehicles of the present application.

[0083] One embodiment of the present application discloses a computer program product, including a computer program, which, when executed by a processor, implements the early warning method for ocean transportation of new energy vehicles of the present application.

[0084] The above is an explanation of the embodiments of the present application by specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Although the description of the present application will be introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this embodiment. In addition, in order to avoid confusion or blurring the focus of the present application, some specific details will be omitted in the description. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other unless there is a conflict.

[0085] Furthermore, various operations will be described as multiple discrete operations in a manner that is most helpful in understanding the illustrative embodiments; however, the order of description should not be construed as implying that these operations are necessarily order dependent. In particular, these operations do not need to be performed in the order presented.

[0086] Unless the context dictates otherwise, the terms "comprising," "having," and "including" are synonymous. The phrase "A / B" means "A or B." The phrase "A and / or B" means "(A and B) or (A or B)."

[0087] As used herein, the term "module" or "unit" may refer to, be or include: an application specific integrated circuit (ASIC), an electronic circuit, a (shared, dedicated or group) processor and / or memory that executes one or more software or firmware programs, a combinational logic circuit and / or other suitable components that provide the described functionality.

[0088] In the accompanying drawings, some structural or method features are shown in a specific arrangement and / or order. However, it should be understood that such specific arrangement and / or order may not be required. In some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0089] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements or data, these elements or data should not be limited by these terms. These terms are used only to distinguish one feature from another. For example, a first feature may be referred to as a second feature, and similarly, a second feature may be referred to as a first feature without departing from the scope of the exemplary embodiments.

[0090] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0091] While the present invention has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the invention.

Claims

1. A code retrieval method, characterized in that: include: Constructing multiple indexes and generating multiple index libraries corresponding to the multiple indexes; Converting the query command input by the user into a semantic vector and a keyword group, and searching the plurality of index libraries based on the semantic vector and the keyword group to obtain search results; Merging the search results, removing duplicate code snippets, and dividing the merged and deduplicated search results into class search results and method search results; Calculating the weighted sum of the similarities between each of the class search results and the method search results to determine a score for each of the search results; Based on the retrieval results of the method, adjusting the scores of the class retrieval results, and retaining the n class retrieval results with the highest scores; and adjusting the scores of the method search results based on the class search results, and retaining the m method search results with the highest scores; Return the n class search results and the m method search results.

2. The method according to claim 1, characterized in that The multiple indexes are constructed based on multiple information dimensions, wherein each index includes a keyword index and / or a vector index.

3. The method according to claim 2, characterized in that The information dimensions include: code structure dimension, important variable dimension, code annotation dimension, code explanation dimension and original code dimension.

4. The method according to claim 3, characterized in that The constructing of multiple indexes for the code based on the information dimension of the code includes: For the code structure dimension, a code structure index is constructed based on the path information, class name and function name of the code; For the important variable dimension, a large language model is used to identify important variables in the code, and an important variable index is constructed based on the name and type of the important variables; For the code comment dimension, construct a code comment index based on the document strings or comments of the code; For the code interpretation dimension, a large language model is used to generate a semantic description of the code, and a code interpretation index is constructed based on the semantic description, wherein the semantic description includes an explanation of the function of the method or class of the code, the role of the parameters, the meaning of the return value, and the dependency relationship between the codes; For the original code dimension, an original code index is constructed based on the source code of the code.

5. The method according to claim 1, wherein The query instruction is in natural language form.

6. The method according to claim 5, characterized in that The query instruction is parsed based on the natural language form of the query instruction, and keywords in the query instruction are extracted to convert the query instruction into a keyword group.

7. The method according to claim 6, characterized in that The query instruction is semantically encoded using a large language model to convert the query instruction into a semantic vector.

8. The method according to claim 1, characterized in that The adjusting the score of the class search result based on the search result of the method includes: When the method search results corresponding to one or more of the class search results have high scores, the scores of the one or more class search results are increased.

9. The method according to claim 1, characterized in that The step of adjusting the score of the method search result based on the class search result includes: When the class search results corresponding to one or more of the method search results have high scores, the scores of the one or more method search results are increased.

10. A code retrieval device, characterized in that: include: An index building unit, configured to build a plurality of indexes and generate a plurality of index libraries corresponding to the plurality of indexes; An instruction retrieval unit, configured to convert a query instruction input by a user into a semantic vector and a keyword group, and to search the plurality of index libraries based on the semantic vector and the keyword group to obtain a retrieval result; A merging and deduplication unit is used to merge the search results, remove duplicate code snippets, and divide the merged and deduplicated search results into class search results and method search results; a score calculation unit, configured to calculate a weighted sum of similarities between each of the class search results and the method search results, to determine a score for each of the search results; A score adjustment unit, configured to adjust the scores of the class search results based on the search results of the method, and retain the n class search results with the highest scores; and adjusting the scores of the method search results based on the class search results, and retaining the m method search results with the highest scores; The result output unit is used to return the n class search results and the m method search results.

11. An electronic device, characterized in that: The device comprises a memory storing computer-executable instructions and a processor; when the instructions are executed by the processor, the device implements the method according to any one of claims 1 to 9.

12. A computer-readable medium, characterized in that The computer-readable medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.