Code recommendation method and device, electronic equipment and storage medium
By generating compressed instructions to crop code files and retaining key information, the problem of insufficient acquisition of context information in long code files is solved, and the accuracy and efficiency of code recommendations are improved.
Patent Information
- Application Number
- CN202510337901.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
When processing long code recommendation methods, existing code recommendation methods cannot obtain comprehensive and accurate context information, resulting in inaccurate recommendation results and cannot provide the most suitable code recommendation results for the requirements.
By generating compression instructions, the code file is cropped to meet the model input length limit, key information is retained, including the content before and after the cursor, similar functions and code structure information, and input the pre-trained code to recommend the model to generate recommended results.
Improve the accuracy of code recommendations, ensure that the model can understand the code context and functional requirements, and generate code recommendation results that are more in line with actual requirements.
Smart Images

Figure CN120276715A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer information processing technologies, and in particular to the field of intelligent code generation and optimization based on large models, which can be used in application scenarios such as code recommendation, code completion, and code generation. Specifically, the present disclosure relates to a code recommendation method, apparatus, electronic device, and storage medium. Background Art
[0002] In the code recommendation scenario, to ensure the recommendation speed, there is a limit on the length of the context information received by the model. When the code file is too long, the entire file cannot be input into the model. Traditional methods usually select a certain number of lines of content before and after the cursor as the context to input into the model. However, this context extraction method has obvious deficiencies: on the one hand, the concentration of the extracted context information is low, containing a large amount of irrelevant information; on the other hand, many useful information is not included in the context range because it is far from the cursor position, resulting in the model being unable to obtain comprehensive and accurate context information, thereby affecting its recommendation accuracy and being unable to provide the most suitable code recommendation results. Summary of the Invention
[0003] The present disclosure provides a code recommendation method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of the present disclosure, there is provided a code recommendation method, including: generating a compression instruction according to a code file and a model input length limit; executing the compression instruction to obtain a target content, the length of which conforms to the model input length limit; inputting the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model; wherein, executing the compression instruction to obtain the target content includes: determining a current function according to the code file and the cursor position; determining the content before the cursor and the content after the cursor of the current function according to the current function and the cursor position; determining similar functions of the current function according to the code file; determining code structure information according to the code file; compressing the content before the cursor, the content after the cursor, the similar functions, and the code structure information to generate the target content.
[0005] According to a second aspect of the present disclosure, there is provided a code recommendation device, including: an instruction generation module, configured to generate a compression instruction according to a code file and a model input length limit; a content compression module, configured to execute the compression instruction to obtain a target content, where the length of the target content meets the model input length limit; a model recommendation module, configured to input the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model; wherein, the content compression module includes: a function determination sub-module, configured to determine a current function according to the code file and a cursor position; a content determination sub-module, configured to determine the content before the cursor and the content after the cursor of the current function according to the current function and the cursor position; a similar function matching sub-module, configured to determine a similar function of the current function according to the code file; a code structure parsing sub-module, configured to determine code structure information according to the code file; a content generation sub-module, configured to compress the content before the cursor, the content after the cursor, the similar function, and the code structure information to generate the target content.
[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor, implements any method in the embodiments of the present disclosure.
[0009] Adopting the solution of the present disclosure can improve the accuracy of code recommendation.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 is a flowchart of a code recommendation method according to an embodiment of the present disclosure;
[0013] Figure 2It is a schematic flowchart of generating code recommendation results according to an embodiment of the present disclosure;
[0014] Figure 3 It is a schematic flowchart of generating target content according to an embodiment of the present disclosure;
[0015] Figure 4 It is a schematic structural diagram of a code recommendation device according to an embodiment of the present disclosure;
[0016] Figure 5 It is a schematic scenario diagram of a code recommendation method according to an embodiment of the present disclosure;
[0017] Figure 6 It is a structural diagram of an electronic device for implementing the code recommendation method according to an embodiment of the present disclosure. Detailed implementation manners
[0018] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0019] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in this article means any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C. The terms "first" and "second" in this article represent referring to multiple similar technical terms and distinguishing them, and do not mean limiting the order, or limiting to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.
[0020] In addition, for better explaining the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0021] Before introducing the technical solutions of the embodiments of the present disclosure, further explanations are made on the technical terms that may be used in the present disclosure:
[0022] Code recommendation: It refers to the recommendation of relevant code snippets or functions to users based on code context information and user input through technologies such as machine learning, supporting the improvement of code writing efficiency. Specifically, it analyzes the syntax and semantics of the code, combines the user's operation context, generates suggestions to help users complete complex programming tasks, while reducing errors and improving code quality.
[0023] In related technologies, the current code recommendation architecture usually intercepts a certain number of lines or a certain number of token contents before and after the cursor position in the code context as context information and sends it to the model according to the context length that the model can receive. This technical solution has the following deficiencies: When the code file is relatively long, if the function where the cursor is located is relatively long, the traditional interception scheme only takes a certain number of lines of content before and after the cursor position in this function as the context and passes it to the model. However, usually, the density of these context information is relatively low, and a lot of useful information is ignored because it is far from the cursor position.
[0024] To at least partially solve one or more of the above problems and other potential problems, the present disclosure proposes a code recommendation method, which can improve the accuracy of code recommendation.
[0025] The embodiments of the present disclosure provide a code recommendation method. Figure 1 It is a schematic flowchart of the code recommendation method according to the embodiments of the present disclosure. This code recommendation method can be applied to a code recommendation device. The code recommendation device is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and the servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to code recommendation devices, code completion devices, or code generation devices, etc. The code recommendation device, code completion device, or code generation device can be mobile phones, tablets, etc. In some possible implementation manners, this code recommendation method can also be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, this code recommendation method includes:
[0026] S101. Generate a compression instruction according to the code file and the model input length limit.
[0027] S102. Execute the compression instruction to obtain target content, and the length of the target content meets the model input length limit.
[0028] S103. Input the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model.
[0029] Here, the code file is a file written and stored by a developer for program code, usually existing in the file format of a specific programming language. In the embodiments of the present disclosure, the code file may include function definitions, class definitions, variable declarations, logical implementations, comments, debugging information, and other contents.
[0030] Here, the model input length limit refers to the maximum limit on the length of the input content when the pre-trained code recommendation model accepts input data. In the embodiments of the present disclosure, the model input length limit requires cropping or compressing an overly long code file to ensure that the input conforms to its model input length limit. Otherwise, the model cannot process it effectively, affecting the recommendation result.
[0031] Here, the compression instruction is an operation rule or method for processing and cropping the content of the code file, aiming to retain the key information of the code file while conforming to the model input length limit. In the embodiments of the present disclosure, the execution of the compression instruction ensures that the code content can be reduced to conform to the model input length limit while retaining the key information to support the operation of the code recommendation model.
[0032] In the embodiments of the present disclosure, the code area to be processed can be first identified from the code file, and then the compression instruction can be designed according to the model input length limit. Exemplarily, the compression instruction can include operations such as code snippet selection, deletion of duplicate and irrelevant code, and extraction of key logic. In particular, the compression instruction can be automatically generated by an algorithm or predefined by the developer.
[0033] Here, the target content is the code data obtained after executing the compression instruction, with a length conforming to the model input limit and containing the key information in the code file. In the embodiments of the present disclosure, the target content can be used as the actual input of the code recommendation model, ensuring the integrity of the code meaning and meeting the model's input length limit.
[0034] In the embodiments of the present disclosure, the compression instruction can be used to process the content of the code file, extract and compress the code. Exemplarily, the compression process can be achieved by removing comments, deleting redundant code blocks, simplifying expressions, etc. In particular, the compression process needs to ensure that the target content still retains the core information of the code after compression and the length conforms to the model input limit.
[0035] Here, the code recommendation model can be a machine learning-based model specifically used for analyzing the input code content and generating code completion, code optimization, or code suggestions. In the embodiments of the present disclosure, the code recommendation model can understand the code context, logical structure, and functional requirements from the target content, and then generate code recommendation results related to the developer's task. It should be noted that the present disclosure does not limit how to train the code recommendation model.
[0036] Here, the code recommendation model can be a large model. After training, the large model should at least have the following capabilities: understanding the context, logical structure, and functional requirements of the code from the target content, and then generating code recommendation results related to the developer's task. It should be noted that this disclosure does not limit how to train the large model to have the code recommendation function.
[0037] Here, the code recommendation result is the output generated by the code recommendation model according to the target content, which can be a recommended code snippet, a code optimization suggestion, or a functional implementation code. In the embodiments of this disclosure, the code recommendation result can help developers accelerate the code writing process, solve technical problems, or improve code quality, thereby improving development efficiency.
[0038] In the embodiments of this disclosure, the generated target content can be first formatted as input data and then input into a pre-trained code recommendation model. Subsequently, the code recommendation model generates a recommended code snippet or suggestion based on the target content and the programming language knowledge learned during its training process. Exemplarily, the output result may include code implementation suggestions, code completion, or optimization solutions.
[0039] Specifically, the process of executing the compression instruction to obtain the target content in step S102 includes: determining the current function according to the code file and the cursor position; determining the content before the cursor and the content after the cursor of the current function according to the current function and the cursor position; determining the similar functions of the current function according to the code file; determining the code structure information according to the code file; compressing the content before the cursor, the content after the cursor, the similar functions, and the code structure information to generate the target content.
[0040] Here, the cursor position is the specific position of the developer's current operation in the code editor, usually a character index or line number. It represents the specific position where the developer is viewing or editing the code. In the embodiments of this disclosure, the developer's current working focus can be determined through the cursor position, so as to extract the code snippet closely related to the cursor position.
[0041] In the embodiments of this disclosure, the code structure can be first parsed from the code file to identify the scope of function definitions. Exemplarily, a programming language parsing library can be used to read the code file and construct an abstract syntax tree of the code, so as to identify the scope of each function definition in the code file. Subsequently, the function where the cursor is located can be found according to the cursor position. Exemplarily, the current function where the cursor is located can be determined based on the matching of the cursor position and the function definition scope.
[0042] In the embodiments of this disclosure, after determining the current function, the function content can be divided into the part before the cursor and the part after the cursor. Exemplarily, according to the text content of the code file and the cursor position, the content before and after the cursor can be directly extracted by string splitting.
[0043] Here, a similar function refers to other functions in the code file that have a certain similarity with the current function in terms of function, structure, semantics, or code logic. In the embodiments of the present disclosure, by analyzing similar functions, reference information can be provided for code completion or optimization of the current function.
[0044] Here, the code structure information refers to the structured content and its relationships in the code file, which may include key code elements such as function declaration information, class variable definition information, and referenced package information. In the embodiments of the present disclosure, the code structure information mainly includes the hierarchical logical structure of the code, which helps the code recommendation model understand the organization form and context of the code file.
[0045] In the embodiments of the present disclosure, the structured information of the code can be generated by analyzing the overall structure of the entire code file. Exemplarily, a syntax parsing tool can be used to extract the code structure information.
[0046] In the embodiments of the present disclosure, a code analysis algorithm can be used to find functions similar to the current function. Exemplarily, the similarity can be calculated by comparing the names, parameters, and function implementations of the current function with other functions, or a syntax or semantic similarity algorithm can be used to analyze the similarity degree between functions.
[0047] In the embodiments of the present disclosure, the content before the cursor, the content after the cursor, similar functions, and code structure information extracted can be compressed to meet the model input length limit. Exemplarily, the compression methods can include information screening, deleting redundant information, simplifying expressions, etc. Among them, the information screening method can be achieved through the process of retaining the most critical information and filtering out redundant or irrelevant content; the method of deleting redundant information can be achieved through processes such as removing comments and duplicate code; the method of simplifying expressions can be achieved through the process of condensing complex code logic into a concise expression form.
[0048] The technical solution of the embodiments of the present disclosure can ensure that the input code content meets the model input length limit through the generation of compression instructions and target content. Even if the code file is large, it can be reasonably trimmed to avoid problems such as the model being unable to process or performance degradation caused by too long input. By compressing the code content, the length of the input data is reduced, the computational amount during model processing is reduced, the processing efficiency is improved, and resource occupancy is reduced. During the compression process, the content before the cursor, the content after the cursor, similar functions, and code structure information are retained, ensuring that the model input covers the key context and global background of the code. By extracting similar functions and code structure information, the model can better understand the running logic and potential functional requirements of the current code, thereby generating more accurate code recommendation results that meet the actual requirements and improving the accuracy of code recommendation.
[0049] In some embodiments, compression instructions are generated according to the code file and the model input length limit, including: obtaining the code file and the model input length limit; generating compression instructions according to the number of lines of the code file and the model input length limit.
[0050] In the embodiments of the present disclosure, the code file can be obtained from a development environment, a file system, memory, a file operation library, a plugin interface, etc. Exemplarily, the content of the code file can be read through a programming language or a tool library and stored in a string or text format for subsequent processing. Further, the input length limit can be obtained by reading the configuration file of the model or directly calling the model interface. Exemplarily, the model input length limit can be obtained through configuration file parameters or direct hard coding in the code.
[0051] In the embodiments of the present disclosure, the code file can be analyzed first to determine the number of lines of the code file. Exemplarily, the content of the code file can be split by line to count the number of lines; or the string operation method provided by the programming language can be used to split the text and count. Further, the relationship between the number of lines of the code file and the model input length limit can be judged. In particular, if the number of lines is less than or equal to the model input length limit, no compression is required. Further still, if the number of lines of the code file exceeds the model input length limit, compression instructions are generated to shorten the content of the code file. In particular, the compression instructions can include compression strategies such as deleting irrelevant content, extracting key content, and simplifying code logic.
[0052] In this way, through the generation of compression instructions, it is ensured that the code file can be shortened to meet the model input length limit, so that the code recommendation system can work properly, avoiding the problem of model processing failure caused by an overly long code file, improving the processing success rate of overly long code files, and also improving the quality of code recommendation results.
[0053] Figure 2 The flowchart showing the generation of code recommendation results is as Figure 2 shown, and this process includes the following steps.
[0054] S201. Trigger the recommendation process.
[0055] S202. Determine whether the number of lines of the code file exceeds the model input length limit; if so, execute S203; if not, execute S204.
[0056] S203. Perform context compression, and then execute S204.
[0057] When the number of lines of the code file exceeds the model input length limit, compression instructions are generated and the compression instructions are executed.
[0058] S204. Generate code recommendation results using the code recommendation model.
[0059] When the compression instruction is executed or the number of lines in the code file does not exceed the model input length limit, the code recommendation model is used to obtain the code recommendation result.
[0060] In some embodiments, according to the current function and the cursor position, the content before the cursor and the content after the cursor of the current function are determined, including: splitting the current function with the cursor position as the boundary; taking the content before the cursor position as the content before the cursor; and taking the content after the cursor position as the content after the cursor.
[0061] In the embodiments of the present disclosure, the current function range can be first obtained through a code structure analysis tool, and then the start position and end position of the function can be determined. Subsequently, the content of the current function can be divided into two parts according to the cursor position.
[0062] Here, the content before the cursor position refers to all the text / code segments on the left side or the front end of the cursor position, that is, the continuous character sequence from the start position of the current function to the cursor insertion point, also known as the prefix content. In the embodiments of the present disclosure, the code from the start of the function to the cursor position can be used as the content before the cursor. In particular, the content before the cursor may include the function name, definition, and / or the code that has been written in the current function. Exemplarily, the method of string extraction or line number extraction can be used to obtain the content before the cursor.
[0063] Here, the content after the cursor position refers to all the text / code segments on the right side or the back end of the cursor position, that is, the continuous character sequence from the cursor insertion point to the end position of the current function, also known as the suffix content.
[0064] In the embodiments of the present disclosure, the code from the cursor position to the end of the function can be used as the content after the cursor. In particular, the content after the cursor may include the function return statement, the tail code, or the code that has not been completed in the current function. Exemplarily, the method of string extraction or line number extraction can be used to obtain the content after the cursor.
[0065] In this way, by splitting the current function, the model input can focus on the content of the current function without being interfered by other codes. At the same time, the division of the content before and after the cursor helps the model understand the code block that the developer is currently focusing on, generates more accurate recommendation results that meet the actual needs, and thus improves the accuracy of code recommendation.
[0066] In some embodiments, determining the similar functions of the current function according to the code file includes: calculating the similarity between the current function and other functions in the code file respectively; and determining the similar functions according to the similarity.
[0067] Here, the similarity is used to measure the similarity between other functions and the current function in terms of function, structure, semantics, etc. In the embodiments of the present disclosure, the similarity is used to represent the degree of similarity between two functions. In particular, the similarity can be represented by a numerical value, and the higher the similarity, the stronger the similarity between the functions.
[0068] In the embodiments of the present disclosure, the code file can be first parsed to extract the definitions of all functions. Exemplarily, a code analysis tool can be used to complete the extraction of functions. Further, the similarity can be calculated for each function and the current function. Exemplarily, the process of calculating the similarity can be implemented by means of semantic analysis, embedding comparison, etc. Among them, the semantic analysis method can judge the syntactic similarity between functions by comparing function names, parameter lists, return value types, code structures, etc.; the embedding comparison method can convert the function content into feature vectors and calculate the similarity between functions through cosine similarity or distance.
[0069] In the embodiments of the present disclosure, screening can be performed to select functions with high similarity as similar functions. Exemplarily, a threshold can be set or several functions with the highest similarity can be selected as similar functions.
[0070] In the embodiments of the present disclosure, a syntax analysis tool can be used to parse the code file to extract the overall structure information of the code. Exemplarily, a language parsing library can be used to construct the code structure information.
[0071] In this way, by calculating the similarity, the system can accurately find the similar functions of the current function and use their logic as the recommendation basis, and the similar functions can provide highly relevant references for code completion or optimization, reducing the interference of irrelevant information. In addition, similar functions usually share certain context information, providing context support for the recommendation logic of the current function.
[0072] In some embodiments, calculating the similarity between other functions in the code file and the current function respectively includes: performing word segmentation processing on the current function to generate a current function word list; performing word segmentation processing on other functions in the code file to generate an auxiliary function word list; determining the similarity according to the current function word list and the auxiliary function word list.
[0073] Here, the current function is the target function being analyzed, usually the focus function that needs to find similar functions or perform code refactoring.
[0074] Here, the auxiliary function is all other functions in the code file except the current function, and is the object that needs to be compared with the current function for similarity one by one.
[0075] Here, the tokenization process is to split the code content of a function into individual words for subsequent calculation and comparison. In the embodiments of the present disclosure, the result of tokenization is a series of words or tokens that can represent the components in the function.
[0076] Here, the function vocabulary is a set of all the words or tokens extracted from the result of the tokenization process, used to represent the content of the function. The function vocabulary can be used to analyze the characteristics of the function and compare it with other functions. In the embodiments of the present disclosure, the function vocabulary can be formed in the form of a set or a frequency distribution.
[0077] In the embodiments of the present disclosure, the code content of the current function can be analyzed first, and then a tokenization tool or a code parsing tool can be used to convert the code content into words or tokens. Exemplarily, regular expressions or the syntax analysis library of a programming language can be used for tokenization, or the abstract syntax tree method can be used for parsing. Finally, the unique words can be extracted using a set to generate the current function vocabulary, or the function vocabulary can be generated by counting the word frequencies. In particular, common words in the vocabulary and keywords in the code language can be filtered out from the vocabulary, and the processed function vocabulary is used as the current function vocabulary.
[0078] In the embodiments of the present disclosure, the code content of other functions in the code file can be analyzed first, and then a tokenization tool or a code parsing tool can be used to convert the code content into words or tokens. Exemplarily, regular expressions or the syntax analysis library of a programming language can be used for tokenization, or the abstract syntax tree method can be used for parsing. Finally, the unique words can be extracted using a set to generate the auxiliary function vocabulary, or the function vocabulary can be generated by counting the word frequencies. In particular, common words in the vocabulary and keywords in the code language can be filtered out from the vocabulary, and the processed function vocabulary is used as the auxiliary function vocabulary. It should be noted that the tokenization processing methods for each of the auxiliary function vocabularies need to be consistent with those of the current function vocabulary for subsequent comparison.
[0079] In the embodiments of the present disclosure, the current function vocabulary and the auxiliary function vocabulary can be compared to calculate the similarity. Exemplarily, the Jaccard similarity can be used for calculation, which can be used for vocabularies in the form of a set. The ratio of the intersection and union of two sets is calculated, and the calculation process can be represented by the following formula:
[0080]
[0081] where S represents the Jaccard similarity, A represents the current function vocabulary, and B represents the auxiliary function vocabulary.
[0082] In particular, similarity can also be calculated using methods such as cosine similarity or machine learning embeddings. The above is only an exemplary illustration and does not limit all possible cases of calculating similarity. Here, exhaustive enumeration is not done.
[0083] In this way, through tokenization processing, the function code content is transformed into words or tokens, accurately expressing the logic and semantics inside the function. Using a vocabulary and similarity calculation can distinguish functions that are functionally similar, providing a highly relevant reference for code recommendation. The logical association between the current function and the similar functions can help the model generate more semantically consistent code completion suggestions, improving the quality of the recommendation results. Through the calculation of similar functions, the model can understand the current context and avoid recommending irrelevant code. Through similarity calculation, similar functions can be automatically located, providing support for code completion in complex scenarios and improving the quality and efficiency of code completion.
[0084] In some embodiments, determining similar functions according to similarity includes: sorting other functions in the code file according to similarity to generate a similarity sorted list; extracting at least one function with the maximum similarity from the similarity sorted list according to a preset rule as the similar function.
[0085] In the embodiments of the present disclosure, a sorting function can be used to sort functions according to similarity values to ensure that functions with high similarity are ranked in the front.
[0086] In the embodiments of the present disclosure, similar functions can be extracted from the sorted list according to a preset rule. Exemplarily, the preset rule can be to only select the function with the highest similarity as the similar function, or to select multiple functions ranked in the front of the similarity, or to select all functions with similarity higher than a certain threshold as the similar function.
[0087] In this way, by sorting all functions in the code file except the current function, the system can preferentially select the most relevant functions and avoid recommending irrelevant or low-quality code. Extracting functions according to a preset rule can flexibly adjust the quantity and quality of the recommended content and improve the accuracy of the recommendation.
[0088] In some embodiments, determining code structure information according to the code file includes: extracting function declaration information of each function in the code file; extracting class variable definition information of each class in the code file; extracting information of each referenced package in the code file; and using the function declaration information, class variable definition information, and referenced package information as code structure information.
[0089] Here, a function declaration refers to the statement that defines a function in the code, including function name, parameter list, return value information, etc. In the embodiments of the present disclosure, the function declaration can provide detailed definition information of the function.
[0090] Here, the class variable definition refers to the variables defined in a class. These variables generally belong to the members of the class and can be accessed or modified during the instantiation process of the class. In the embodiments of the present disclosure, the class variable definition provides context information for functions.
[0091] Here, the referenced package information refers to the import statement, which is the statement in the code file to import external modules, libraries, or packages, enabling the code to use externally defined functions, classes, or other resources. In the embodiments of the present disclosure, the referenced package information can provide the dependency environment of the code file for the code recommendation model.
[0092] In the embodiments of the present disclosure, the function declarations in the code can be extracted through a code parsing tool. Exemplarily, the function nodes can be directly located through an abstract syntax tree.
[0093] In the embodiments of the present disclosure, the class nodes can be extracted by using a syntax analysis tool, and the class variable definitions, including class variables and instance variables, can be obtained. Exemplarily, the class nodes can be extracted through a parsing tool or regular expressions.
[0094] In the embodiments of the present disclosure, a syntax analysis tool can be used to extract the import statements, thereby extracting the referenced package information. Exemplarily, regular expressions or language parsing tools can be used to extract the relevant statements.
[0095] In the embodiments of the present disclosure, the extracted function declaration information, class variable definition information, and referenced package information can be organized into structured data, and this structured data can be used as the code structure information. In particular, the code structure information can represent the components of the entire code file.
[0096] In this way, by extracting the function declaration information, class variable definition information, and referenced package information, the structured information of the code can be provided, facilitating the quick understanding of the content and logic of the code file. By extracting the code structure information, context support can be provided for code completion. Automatically extracting the code structure information replaces the process of manually analyzing the code file, significantly improving the efficiency of code analysis.
[0097] In some embodiments, the content before the cursor, the content after the cursor, the similar functions, and the code structure information are compressed to generate the target content, including: generating the following content according to the content after the cursor; generating the target content according to the following content, the content before the cursor, the similar functions, and the code structure information.
[0098] In the embodiments of the present disclosure, the content after the cursor can be analyzed to extract its key elements and generate the following content. Exemplarily, redundant information such as comments, spaces, and line breaks can be deleted first, then the unfinished code structure can be analyzed to extract the structural information, and finally, the content after the cursor can be supplemented based on the language rules to achieve syntax completion.
[0099] In the embodiments of the present disclosure, content before the cursor, subsequent content, similar functions, and code structure information can be compressed to remove duplicate or redundant data, and only the key parts related to code generation are retained. Subsequently, the compressed information of each part can be combined to generate a unified input for use by the code recommendation model. Finally, the code embedding model can be used to process the fused information to generate the target content.
[0100] In this way, by compressing the content before and after the cursor, similar functions, and code structure information into the target content, the quality and efficiency of code completion are significantly improved. By combining similar functions and code structure information, complex scenarios can be better handled in large code files or complex projects. Exemplarily, in the scenario of cross-module calls, traditional solutions rely on developers' memory and are prone to interface path errors, while the present solution can quickly infer legal call paths through code structure information. In the scenario of legacy code refactoring, traditional solutions are difficult to locate scattered similar logics, while the present solution can accurately locate duplicate code through similar function matching. In the scenario of design pattern implementation, the template code of traditional solutions is lengthy and prone to missing thread safety details, while the present solution can quickly generate best practice code by combining similar functions and structure information. The above is only an exemplary illustration and does not limit all possible complex scenarios, and only exhaustive listing is not done here. The present solution has the capabilities of accurate inference, pattern reuse, and architecture awareness in large complex projects, significantly improving development efficiency and code quality.
[0101] In some embodiments, to generate the target content according to subsequent content, content before the cursor, similar functions, and code structure information, it includes: placing the similar functions and code structure information before the content before the cursor, and placing the subsequent content after the content before the cursor to generate the target content.
[0102] In the embodiments of the present disclosure, the content of similar functions can be compressed or formatted according to actual needs to retain the key logic. At the same time, the code structure information can be compressed to reduce redundancy, and only the content related to the current function is retained. Finally, the similar functions and code structure information can be combined and inserted before the content before the cursor.
[0103] In the embodiments of the present disclosure, the content after the cursor can be syntactically analyzed to extract incomplete code fragments or structure information, and the extracted content is used as the subsequent content. Subsequently, the subsequent content can be formatted to make it concise and easy to integrate. Finally, the subsequent content is directly appended after the content before the cursor.
[0104] In the embodiments of the present disclosure, the similar functions, code structure information, content before the cursor, and subsequent content can be integrated into a complete context input, and the context input is used as the target content.
[0105] In this way, by integrating the content before the cursor, the content after the cursor, similar functions and code structure information, the context of the target content generation is more complete, which can improve the accuracy of code recommendation and enable the model to generate content that fits the current context. By automatically integrating multiple information and generating target content, the complexity of user operations can be significantly reduced, while improving the efficiency of code recommendation.
[0106] In some embodiments, target content is generated based on following content, content before a cursor, similar functions, and code structure information, including: generating preceding content based on content before a cursor; generating target content based on preceding content and following content; in response to the length of the target content exceeding a model input length limit, intercepting a portion that meets the model input length limit as the target content.
[0107] In the disclosed embodiment, the content before the cursor can be parsed to extract the implemented logic or code of the current function, and the extracted content can be used as the previous content. In particular, the previous content can be formatted to make it concise and easy to integrate.
[0108] In the disclosed embodiment, the previous content and the following content may be spliced to form a complete target content. In particular, the target content may be formatted to ensure that it complies with grammatical rules.
[0109] In the disclosed embodiment, when the length of the target content exceeds the model input length limit, the content can be compressed first. Exemplarily, the content can be compressed first to remove redundant information, such as comments, blank lines, etc., to ensure that more key information can be retained in the input. Further, the content can be intercepted. Exemplarily, the end of the content before the cursor can be intercepted in the previous content, and the beginning of the content after the cursor can be intercepted in the following content. In particular, the interception position can be dynamically adjusted according to the input length limit to make the content meet the model requirements.
[0110] In this way, the context of the current code can be easily understood by combining the content before and after the cursor. The logic of intercepting content ensures that the input length is within the model limit, avoiding errors caused by exceeding the length limit, and improving the success rate of processing overlong code files.
[0111] In some embodiments, target content is generated based on following content, content before a cursor, similar functions, and code structure information, including: generating preceding content based on content before a cursor and similar functions; generating target content based on preceding content and following content; in response to determining that the length of the target content exceeds a model input length limit, intercepting a portion that meets the model input length limit as the target content.
[0112] In the embodiments of the present disclosure, syntax analysis can be performed on the content before the cursor to extract the implemented logic or code of the current function. Further, the information of similar functions can be used as supplementary reference, and combined with the content before the cursor, to generate a more complete previous context. Exemplarily, similar functions can be added before the content before the cursor, and the integrated content can be used as the previous context. In particular, the previous context can be formatted to make it concise and easy to integrate.
[0113] In the embodiments of the present disclosure, the previous context and the following context can be concatenated to form a complete target content. In particular, the target content can be formatted to ensure that it conforms to the syntax rules.
[0114] In the embodiments of the present disclosure, when the length of the target content exceeds the model input length limit, content compression can be performed first. Exemplarily, content can be preferentially compressed to remove redundant information such as comments and blank lines, ensuring that more key information can be retained in the input. Further, content truncation can be performed. Exemplarily, similar functions can be truncated from the previous context, and the function name, parameter list, and core logic can be preferentially retained. In particular, the truncation position can be dynamically adjusted according to the input length limit to make the content meet the model requirements.
[0115] In this way, by combining similar functions, it is possible to better adapt to complex scenarios. And the similar functions provide reference implementations, making the generated target content more in line with the code style and logic specifications of the project.
[0116] In some embodiments, the target content is generated according to the following context, the content before the cursor, similar functions, and code structure information, including: generating the previous context according to the content before the cursor, similar functions, and function declaration information; generating the target content according to the previous context and the following context; and in response to determining that the length of the target content exceeds the model input length limit, truncating the part that meets the model input length limit as the target content.
[0117] In the embodiments of the present disclosure, syntax analysis can be performed on the content before the cursor to extract the implemented logic or code of the current function. Further, the information of similar functions can be used as supplementary reference. Furthermore, the declarations of all functions in the current code file can be extracted, and combined with the content before the cursor and similar functions, to generate a more complete previous context. Exemplarily, similar functions can be added before the similar functions, and the integrated content can be used as the previous context. In particular, the previous context can be formatted to make it concise and easy to integrate.
[0118] In the embodiments of the present disclosure, the previous context and the following context can be concatenated to form a complete target content. In particular, the target content can be formatted to ensure that it conforms to the syntax rules.
[0119] In the disclosed embodiment, when the length of the target content exceeds the model input length limit, the content can be compressed first. Exemplarily, the content can be compressed first to remove redundant information, such as comments, blank lines, etc., to ensure that more key information can be retained in the input. Further, content interception can be performed. Exemplarily, the function declaration information can be intercepted in the above content, and only the part related to the current operation is retained. In particular, the interception position can be dynamically adjusted according to the input length limit to make the content meet the model requirements.
[0120] In this way, by combining function declaration information and similar functions, the generated code style and logic are made more standardized, so that the target content generation is consistent with the current code environment, reducing unnecessary adjustments and improving the success rate of processing ultra-long code files.
[0121] In some embodiments, target content is generated based on following content, content before a cursor, similar functions, and code structure information, including: generating preceding content based on content before a cursor, similar functions, function declaration information, and class variable definition information; generating target content based on preceding content and following content; in response to determining that the length of the target content exceeds a model input length limit, intercepting a portion that meets the model input length limit as the target content.
[0122] In the disclosed embodiment, the disclosed embodiment can perform syntax analysis on the content before the cursor to extract the implemented logic or code of the current function. Further, the information of similar functions can be used as a supplementary reference. Still further, the declarations of all functions in the current code file can be extracted. Still further, the class variable definition information, such as the variable name, type and initial value, can be extracted, and combined with the content before the cursor, similar functions and function declaration information to generate a more complete above content. Exemplarily, the class variable definition information can be added before the function declaration information, and the integrated content can be used as the above content. In particular, the above content can be formatted to make it concise and easy to integrate.
[0123] In the disclosed embodiment, the previous content and the following content may be spliced to form a complete target content. In particular, the target content may be formatted to ensure that it complies with grammatical rules.
[0124] In the disclosed embodiment, when the length of the target content exceeds the model input length limit, the content can be compressed first. Exemplarily, the content can be compressed first to remove redundant information, such as comments, blank lines, etc., to ensure that more key information can be retained in the input. Further, content interception can be performed. Exemplarily, the information defined by the class variable can be intercepted in the above content, and only the key parts that are helpful for completion are retained. In particular, the interception position can be dynamically adjusted according to the input length limit to make the content meet the model requirements.
[0125] Thus, the above - mentioned content generated by combining the content before the cursor, similar functions, function declaration information, and class variable definition information provides a comprehensive context for code completion, reduces the deviation between the generated result and the actual requirements, and improves the success rate of processing ultra - long code files.
[0126] In some embodiments, target content is generated based on the following content, the content before the cursor, similar functions, and code structure information, including: generating the above - mentioned content based on the content before the cursor, similar functions, function declaration information, class variable definition information, and reference package information; generating the target content based on the above - mentioned content and the following content; and in response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
[0127] In the embodiments of the present disclosure, the content before the cursor can be syntactically analyzed to extract the implemented logic or code of the current function. Further, the information of similar functions can be used as supplementary reference. Furthermore, the declarations of all functions in the current code file can be extracted. Still further, class variable definition information such as variable name, type, and initial value can be extracted. In addition, all reference package information can be extracted, and combined with the content before the cursor, similar function information, function declaration information, and class variable definition information, to generate a more complete above - mentioned content. Exemplarily, the reference package information can be added before the class variable definition information, and the integrated content can be used as the above - mentioned content. Particularly, the above - mentioned content can be formatted to be concise and easy to integrate.
[0128] In the embodiments of the present disclosure, the above - mentioned content and the following content can be spliced to form a complete target content. Particularly, the target content can be formatted to ensure that it conforms to the syntax rules.
[0129] In the embodiments of the present disclosure, when the length of the target content exceeds the model input length limit, content compression can be performed first. Exemplarily, content can be preferentially compressed to remove redundant information such as comments and blank lines, ensuring that more key information can be retained in the input. Further, content interception can be performed. Exemplarily, the reference package information can be intercepted from the above - mentioned content, and only the compressed or filtered part relevant to the current completion task is retained. Particularly, the interception position can be dynamically adjusted according to the input length limit to make the content meet the model requirements.
[0130] Thus, the reference information for completion is increased through similar functions and function declaration information, making the generated code more accurate. At the same time, the global perspective is provided through class variable definition information and reference package information, which helps the result of code recommendation to be more in line with the project architecture and improves the success rate of processing ultra - long code files.
[0131] Figure 3 A flow diagram showing the generation of the target content is as followsFigure 3 As shown, the process includes the following steps.
[0132] S301. Determine whether the content length of the function where the cursor is located exceeds the model input length limit. If so, execute S302; if not, execute S303.
[0133] S302. Perform a truncation operation to make the length of the obtained target content meet the model input length limit, and then execute S316.
[0134] If the current function has already exceeded the model input length limit, then truncate the content before the cursor and the content after the cursor that meet the model input length limit to form the final target content.
[0135] S303. Combine the content before the cursor and the content after the cursor to form the current target content, and then execute S304.
[0136] If the current function does not exceed the model input length limit, then combine the content before the cursor and the content after the cursor to form the current target content.
[0137] S304. Determine whether the content length of the current function combined with similar functions exceeds the model input length limit. If so, execute S305; if not, execute S306.
[0138] S305. Perform a truncation operation to make the length of the obtained target content meet the model input length limit, and then execute S316.
[0139] If the content length of the current function combined with similar functions has already exceeded the model input length limit, then truncate the content before the cursor, the content after the cursor, and the similar functions that meet the model input length limit to form the final target content.
[0140] S306. Add the similar functions to the target content, and then execute S307.
[0141] S307. Determine whether the content length of the current function, similar functions combined with the function declaration exceeds the model input length limit. If so, execute S308; if not, execute S309.
[0142] S308. Perform a truncation operation to make the length of the obtained target content meet the model input length limit, and then execute S316.
[0143] If the content length of the current function, similar functions combined with the function declaration has already exceeded the model input length limit, then truncate the content before the cursor, the content after the cursor, the similar functions, and the function declaration that meet the model input length limit to form the final target content.
[0144] S309. Add the function declaration to the target content.
[0145] S310. Determine whether the combined content length of the current function, similar functions, function declarations, and class variable definitions exceeds the model input length limit. If so, execute S311; if not, execute S312.
[0146] S311. Perform an interception operation to make the length of the obtained target content meet the model input length limit, and then execute S316.
[0147] If the combined content length of the current function, similar functions, function declarations, and class variable definitions has exceeded the model input length limit, intercept the content before the cursor, the content after the cursor, similar functions, function declarations, and class variable definitions that meet the model input length limit to form the final target content.
[0148] S312. Add the class variable definitions to the target content.
[0149] S313. Determine whether the combined content length of the current function, similar functions, function declarations, class variable definitions, and reference package information exceeds the model input length limit. If so, execute S314; if not, execute S315.
[0150] S314. Perform an interception operation to make the length of the obtained target content meet the model input length limit, and then execute S316.
[0151] If the combined content length of the current function, similar functions, function declarations, class variable definitions, and reference package information has exceeded the model input length limit, intercept the content before the cursor, the content after the cursor, similar functions, function declarations, class variable definitions, and reference package information that meet the model input length limit to form the final target content.
[0152] S315. Add the reference package information to the target content, and then execute S316.
[0153] S316. Obtain the final target content.
[0154] It should be understood that Figure 2 and Figure 3 the schematic diagrams shown are merely exemplary rather than restrictive, and they are extensible. Those skilled in the art can make various obvious changes and / or substitutions based on the examples of Figure 2 and Figure 3 The obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0155] The embodiments of the present disclosure provide a code recommendation device, as shown in Figure 4As shown, the device may include: an instruction generation module 401, configured to generate a compression instruction according to a code file and a model input length limit; a content compression module 402, configured to execute the compression instruction to obtain a target content, where the length of the target content meets the model input length limit; a model recommendation module 403, configured to input the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model; wherein, the content compression module 402 includes: a function determination sub-module 4021, configured to determine the current function according to the code file and the cursor position; a content determination sub-module 4022, configured to determine the content before the cursor and the content after the cursor of the current function according to the current function and the cursor position; a similar function matching sub-module 4023, configured to determine a similar function of the current function according to the code file; a code structure parsing sub-module 4024, configured to determine code structure information according to the code file by the similar function matching sub-module; a content generation sub-module 4025, configured to compress the content before the cursor and the content after the cursor of the current function, as well as the similar function and the code structure information, to generate the target content.
[0156] In some embodiments, the instruction generation module 401 includes: an information acquisition sub-module, configured to acquire a code file and a model input length limit; a judgment and generation sub-module, configured to generate a compression instruction according to the number of lines of the code file and the model input length limit.
[0157] In some embodiments, the content determination sub-module 4022 is configured to: split the current function with the cursor position as the boundary; use the content before the cursor position as the content before the cursor; use the content after the cursor position as the content after the cursor.
[0158] In some embodiments, the similar function matching sub-module 4023 is configured to: calculate the similarity between each other function in the code file and the current function respectively; determine the similar function according to the similarity.
[0159] In some embodiments, the similar function matching sub-module 4023 is configured to: perform word segmentation processing on the current function according to the code file to generate a current function word list; perform word segmentation processing on other functions according to the code file to generate an auxiliary function word list; determine the similarity according to the current function word list and the auxiliary function word list.
[0160] In some embodiments, the similar function matching sub-module 4023 is configured to: sort each function in the code file according to the similarity to generate a similarity sorted list; extract at least one function with the largest similarity from the similarity sorted list according to a preset rule as the similar function.
[0161] In some embodiments, the code structure parsing sub-module 4024 is configured to: extract the function declaration information of each function in the code file; extract the class variable definition information of each class in the code file; extract the information of each referenced package in the code file; and use the function declaration information, the class variable definition information, and the referenced package information as the code structure information.
[0162] In some embodiments, the content generation sub-module 4025 is configured to: generate the subsequent content according to the content after the cursor; generate the target content according to the subsequent content, the content before the cursor, the similar functions, and the code structure information.
[0163] In some embodiments, the content generation sub-module 4025 is configured to: place the similar functions and the code structure information before the content before the cursor, and place the subsequent content after the content before the cursor to generate the target content.
[0164] In some embodiments, the content generation sub-module 4025 is configured to: generate the previous content according to the content before the cursor; generate the target content according to the previous content and the subsequent content; and in response to determining that the length of the target content exceeds the model input length limit, intercept the part that meets the model input length limit as the target content.
[0165] In some embodiments, the content generation sub-module 4025 is configured to: generate the previous content according to the content before the cursor and the similar functions; generate the target content according to the previous content and the subsequent content; and in response to determining that the length of the target content exceeds the model input length limit, intercept the part that meets the model input length limit as the target content.
[0166] In some embodiments, the content generation sub-module 4025 is configured to: generate the previous content according to the content before the cursor, the similar functions, and the function declaration information; generate the target content according to the previous content and the subsequent content; and in response to determining that the length of the target content exceeds the model input length limit, intercept the part that meets the model input length limit as the target content.
[0167] In some embodiments, the content generation sub-module 4025 is configured to: generate the previous content according to the content before the cursor, the similar functions, the function declaration information, and the class variable definition information; generate the target content according to the previous content and the subsequent content; and in response to determining that the length of the target content exceeds the model input length limit, intercept the part that meets the model input length limit as the target content.
[0168] In some embodiments, the content generation sub-module 4025 is configured to: generate the previous content according to the content before the cursor, the similar functions, the function declaration information, the class variable definition information, and the referenced package information; generate the target content according to the previous content and the subsequent content; and in response to determining that the length of the target content exceeds the model input length limit, intercept the part that meets the model input length limit as the target content.
[0169] For the specific functions and examples of each module and sub-module of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0170] In the code recommendation device according to the embodiments of the present disclosure, through the generation of compression instructions and target content, it can ensure that the input code content meets the model input length limit. Even if the code file is large, it can be reasonably trimmed to avoid problems such as the model being unable to process or performance degradation caused by too long input. By compressing the code content, the length of the input data is reduced, the computational amount during model processing is reduced, the processing efficiency is improved, and resource occupancy is reduced. During the compression process, the content before the cursor, the content after the cursor, similar functions, and code structure information are retained to ensure that the model input covers the key context and global background of the code. By extracting similar function and code structure information, the model can better understand the running logic and potential functional requirements of the current code, thereby generating more accurate code recommendation results that meet the actual requirements, and improving the accuracy of code recommendation.
[0171] The embodiments of the present disclosure provide a schematic diagram of the scenario of a code recommendation method, as Figure 5 shown.
[0172] As mentioned above, the code recommendation method provided by the embodiments of the present disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
[0173] Specifically, the electronic device can perform the following operations:
[0174] Generate compression instructions according to the code file and the model input length limit; execute the compression instructions to obtain the target content, and the length of the target content meets the model input length limit; input the target content into a pre-trained code recommendation model to obtain the code recommendation result output by the code recommendation model.
[0175] It should be understood that Figure 5 the shown scenario diagram is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 5 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.
[0176] In the technical solutions of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0177] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0178] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0179] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0180] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0181] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the code recommendation method. For example, in some embodiments, the code recommendation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the code recommendation method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the code recommendation method by any other suitable means (e.g., by means of firmware).
[0182] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0183] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0184] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0185] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0186] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: Local Area Network (LAN), Wide Area Network (WAN), and the Internet.
[0187] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.
[0188] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. No limitations are imposed herein.
[0189] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A code recommendation method, comprising: Generating a compression instruction according to a code file and a model input length limit; Executing the compression instruction to obtain target content, the length of the target content conforming to the model input length limit; Inputting the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model; Wherein, the executing the compression instruction to obtain target content includes: Determining a current function according to the code file and a cursor position; Determining content before the cursor and content after the cursor of the current function according to the current function and the cursor position; Determining a similar function of the current function according to the code file; Determining code structure information according to the code file; Compressing the content before the cursor, the content after the cursor, the similar function and the code structure information to generate the target content.
2. The method according to claim 1, wherein The generating a compression instruction according to a code file and a model input length limit includes: Obtaining the code file and the model input length limit; Generating the compression instruction according to the number of lines of the code file and the model input length limit.
3. The method according to claim 1, wherein The determining content before the cursor and content after the cursor of the current function according to the current function and the cursor position includes: Splitting the current function with the cursor position as a boundary; Taking the content before the cursor position as the content before the cursor; Taking the content after the cursor position as the content after the cursor.
4. The method according to claim 1, wherein, The determining a similar function of the current function according to the code file includes: Calculating the similarity between other functions in the code file and the current function respectively; Determining the similar function according to the similarity.
5. The method according to claim 4, wherein The calculating the similarity between other functions in the code file and the current function respectively includes: Performing word segmentation processing on the current function to generate a current function word list; Performing word segmentation processing on other functions in the code file to generate an auxiliary function word list; Determining the similarity according to the current function word list and the auxiliary function word list.
6. The method according to claim 4, wherein, The determining the similar function according to the similarity includes: Sorting other functions in the code file according to the similarity to generate a similarity sorted list; Extracting at least one function with the largest similarity from the similarity sorted list according to a preset rule as the similar function.
7. The method according to claim 1, wherein, The determining code structure information according to the code file includes: Extracting function declaration information of each function in the code file; Extracting class variable definition information of each class in the code file; Extracting information of each reference package in the code file; Taking the function declaration information, the class variable definition information and the reference package information as the code structure information.
8. The method according to claim 7, wherein The compressing the content before the cursor, the content after the cursor, the similar function and the code structure information to generate the target content includes: Generating subsequent content according to the content after the cursor; Generating the target content according to the subsequent content, the content before the cursor, the similar function and the code structure information.
9. The method according to claim 8, wherein Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Placing the similarity function and the code structure information before the content before the cursor, and placing the following content after the content before the cursor to generate the target content.
10. The method according to claim 8, wherein, Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Generating the previous content according to the content before the cursor; Generating the target content according to the previous content and the following content; In response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
11. The method according to claim 8, wherein Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Generating the previous content according to the content before the cursor and the similarity function; Generating the target content according to the previous content and the following content; In response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
12. The method according to claim 8, wherein, Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Generating the previous content according to the content before the cursor, the similarity function, and the function declaration information; Generating the target content according to the previous content and the following content; In response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
13. The method according to claim 8, wherein Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Generating the previous content according to the content before the cursor, the similarity function, the function declaration information, and the class variable definition information; Generating the target content according to the previous content and the following content; In response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
14. The method according to claim 8, wherein, Generating the target content according to the following content, the content before the cursor, the similarity function, and the code structure information includes: Generating the previous content according to the content before the cursor, the similarity function, the function declaration information, the class variable definition information, and the reference package information; Generating the target content according to the previous content and the following content; In response to determining that the length of the target content exceeds the model input length limit, intercepting the part that meets the model input length limit as the target content.
15. A code recommendation device, comprising: An instruction generation module, configured to generate a compression instruction according to a code file and a model input length limit; A content compression module, configured to execute the compression instruction to obtain a target content, where the length of the target content meets the model input length limit; A model recommendation module for inputting the target content into a pre-trained code recommendation model to obtain a code recommendation result output by the code recommendation model; Among them, the content compression module includes: A function determination sub-module for determining the current function according to the code file and the cursor position; A content determination sub-module for determining the content before the cursor and the content after the cursor of the current function according to the current function and the cursor position; A similar function matching sub-module for determining similar functions of the current function according to the code file; A code structure parsing sub-module for determining code structure information according to the code file; A content generation sub-module for compressing the content before the cursor, the content after the cursor, the similar functions and the code structure information to generate the target content.
16. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-14.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions for causing a computer to execute the method according to any one of claims 1-14.
18. A computer program product, comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1-14 when executed by a processor.
Citation Information
Cited By
Alarm information processing method, system and device and computer storage medium
CN121193586A