A code generation method, device and computing device cluster

By detecting and generating compliant code in real time during the code generation process, the plagiarism risk problem caused by the possibility of similarity to open source code in the big model generated code is solved, efficient and compliant code generation is achieved, and development efficiency is improved.

CN119396464BActive Publication Date: 2025-05-23SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510010713.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-23
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

When using large-model-based code generation tools, the generated code may be similar to open source code, resulting in possible plagiarism, bringing legal and commercial risks, requiring open source compliance detection of the generated code.

Method used

By outputting compliance code in real time during the code generation process, using the code generation model to output code line by line, and performing compliance detection after outputting one line of code at a time, ensuring that the output code complies with the compliance of the open source license.

Benefits of technology

It realizes real-time detection and generation of compliant code without interrupting user programming ideas, improves development efficiency and reduces legal and commercial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396464B_ABST
    Figure CN119396464B_ABST
Patent Text Reader

Abstract

A code generation method, the method comprising: obtaining codes that are serially output by a code generation model word by word; when the first line of code is obtained, the first line of code is matched with the code in the code library for similarity to obtain prompt information of the first line of code; when the prompt information of the first line of code does not have the same code identifier in the multiple prompt information corresponding to the second line of code, the first line of code and the second line of code are compliant codes. In the present application, after receiving a line of code each time, the method performs a compliance check on the received line of code while pausing the code generation model to generate the next line of code. If the received matching result indicates that the current line of code is non-compliant code, the code generation model can be instructed to regenerate compliant current line of code and other lines of code. During the whole process, the method continuously outputs compliant code without interrupting the user's programming ideas, thereby improving development efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of code generation, and in particular to a code generation method, apparatus and computing device cluster. Background Art

[0002] Code continuation and generation is a key application of intelligent R&D, which can significantly improve the work efficiency of developers. The code data for training these large models mostly comes from public channels, including a large amount of crawled open source code. Among these open source codes, some licenses are not friendly to commercial use. For example, the code requirements of using the general public license (GPL) are that any project that uses these codes must also open their source code and follow the GPL agreement.

[0003] Since the code generated by large models has a certain degree of randomness, developers may encounter situations similar to open source code when using code snippets generated by these models. If the proportion of similar code is too high, it may be regarded as plagiarism of open source code, which will bring legal and commercial risks. Therefore, when using code generation tools based on large models, it is necessary to perform open source compliance detection on the code generated by the large model. Summary of the invention

[0004] In order to solve the above problems, an embodiment of the present application provides a code generation method, which continuously outputs compliant code during the code generation process of a large model without interrupting the user's programming ideas, thereby improving development efficiency. In addition, the present application also provides a code generation device and a computing device cluster corresponding to the code generation method.

[0005] To this end, the following technical solutions are adopted in the embodiments of the present application:

[0006] In a first aspect, a code generation method is provided in an embodiment of the present application, the method is applied to a cloud management platform, the cloud management platform is used to manage infrastructure, the infrastructure includes at least one node, at least one node stores a code generation model, the code generation model is used to generate code, the method includes: obtaining the code that the code generation model outputs serially one by one; when the first line of code generated by the code generation model is obtained, the first line of code and the code in the code library are matched for similarity to obtain prompt information of the first line of code; the prompt information of the first line of code includes at least one code identifier, each code identifier is an identifier of a license to which the code that is the same or similar to the one line of code belongs; when the prompt information of the first line of code does not have the same code identifier in multiple prompt information corresponding to the prompt information of the second line of code, the first line of code and the second line of code are compliant codes.

[0007] In this implementation, each time a line of code is received, the method performs compliance detection on the received line of code and pauses the code generation model from generating the next line of code. If the received matching result indicates that the current line of code is non-compliant code, the code generation model can be instructed to regenerate the current line of code and other lines of code that are compliant. Throughout the entire process, the method continuously outputs compliant code without interrupting the user's programming ideas, thereby improving development efficiency.

[0008] In one embodiment, the method also includes: when the same code identifier exists in the prompt information of the first line number code and the multiple prompt information corresponding to the second line number code, based on the same code identifier, obtaining the code information corresponding to the same code identifier, the code information including one or more of the license of the open source software to which the code belongs, the address of the open source software to which the code belongs, and the file path of the open source software to which the code belongs; detecting whether the code information corresponding to the same code identifier is compliant code information; when the code information corresponding to the same code identifier is compliant code information, the first line number code and the second line number code are compliant codes.

[0009] In this embodiment, when determining that the prompt information of the first line number code and the multiple prompt information corresponding to the second line number code have the same code identifier, the method can detect whether the code information corresponding to the same code identifier is compliant code information to ensure the accuracy of the detection result.

[0010] In one embodiment, the method further includes: when the code information corresponding to the same code identifier is non-compliant code information, instructing the code generation model to regenerate the first line number code and the second line number code.

[0011] In this embodiment, when determining whether the code information corresponding to the same code identifier is non-compliant code information, the method may instruct the code generation model to regenerate compliant code to ensure that the output code is compliant code.

[0012] In one embodiment, when the first line of code generated by the code generation model is obtained, the first line of code is matched with the code in the code library for similarity to obtain prompt information of the first line of code, and the method also includes: creating an inverted index in the code library; the inverted index includes code features and code identifiers for each line of code in the code library; performing similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code, including: performing feature analysis on the first line of code to obtain the code features of the first line of code; based on the code features of the first line of code, obtaining at least one code identifier with the same code features as the code features of the first line of code from the inverted index to obtain prompt information of the first line of code.

[0013] In this embodiment, the method can use the inverted index to detect information about similar or identical codes. Since the inverted index records the association between code features and code identifiers, the method uses code features for matching to obtain code identifiers of similar or identical codes, without the need to compare one by one, thereby reducing the workload of detection and improving efficiency.

[0014] In one embodiment, the second number of lines of code is code generated by the code generation model before generating the first number of lines of code, and is continuous with the first number of lines of code, the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines, and the non-compliant code fragment refers to a code fragment with the same number of lines as the code in the code library being greater than or equal to the set number of lines.

[0015] In this embodiment, the method defines a rule for whether a code snippet is non-compliant, that is, if at least a set number of consecutive lines of code in the code snippet is identical or similar to the open source code, the code snippet is considered to be a non-compliant code snippet. When the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines and is identical or similar to the open source code, it can be determined that the first number of lines of code and the second number of lines of code are non-compliant code.

[0016] In one embodiment, the code generated by the code generation model is a code fragment missing in the code file, and the method also includes: obtaining the third line of code at a set position in the code file, the set position being a position after the position where the code generated by the code generation model is inserted into the code file; when the prompt information of the first line of code does not contain the same code information in multiple prompt information corresponding to the third line of code, the first line of code is a compliant code.

[0017] In this embodiment, when the method determines that the code generated by the code generation model is a missing code fragment in the code file, it is necessary to perform similarity matching on the prompt information of the first line of code with multiple prompt information corresponding to the third line of code in the original code, so as to improve the accuracy of compliance detection.

[0018] In one embodiment, when a first line of code generated by a code generation model is obtained, the first line of code is matched with the code in the code library for similarity to obtain prompt information of the first line of code. The method also includes: when the number of code lines stored in the code cache is equal to the second number of lines, obtaining multiple prompt information corresponding to the second number of lines of code from the code cache.

[0019] In the second aspect, a code generation device is provided in an embodiment of the present application, including: a first processing module, used to obtain the code that is serially output by the code generation model one by one; a second processing module, used to, when the first line of code generated by the code generation model is obtained, perform similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code; the prompt information of the first line of code includes at least one code identifier, each code identifier is an identifier of a license to which the code that is identical or similar to the one line of code belongs; a third processing module, used to, when the prompt information of the first line of code does not have the same code identifier in multiple prompt information corresponding to the second line of code, determine that the first line of code and the second line of code are compliant codes.

[0020] In one embodiment, the third processing module is also used to, when the same code identifier exists in the prompt information of the first line of code and the multiple prompt information corresponding to the second line of code, based on the same code identifier, obtain code information corresponding to the same code identifier, the code information including one or more of the license of the open source software to which the code belongs, the address of the open source software to which the code belongs, and the file path of the open source software to which the code belongs; detect whether the code information corresponding to the same code identifier is compliant code information; when the code information corresponding to the same code identifier is compliant code information, the first line of code and the second line of code are compliant codes.

[0021] In one implementation, the third processing module is further configured to instruct the code generation model to regenerate the first line of code and the second line of code when the code information corresponding to the same code identifier is non-compliant code information.

[0022] In one embodiment, the second processing module, when obtaining the first line of code generated by the code generation model, performs similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code, is also used to create an inverted index in the code library; the inverted index includes code features and code identifiers for each line of code in the code library; similarity matching is performed between the first line of code and the code in the code library to obtain prompt information of the first line of code, including: performing feature analysis on the first line of code to obtain the code features of the first line of code; based on the code features of the first line of code, obtaining at least one code identifier with the same code features as the code features of the first line of code from the inverted index to obtain prompt information of the first line of code.

[0023] In one embodiment, the second number of lines of code is code generated by the code generation model before generating the first number of lines of code, and is continuous with the first number of lines of code, the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines, and the non-compliant code fragment refers to a code fragment with the same number of lines as the code in the code library being greater than or equal to the set number of lines.

[0024] In one embodiment, the code generated by the code generation model is a missing code fragment in the code file, and the third processing module is also used to obtain a third line of code at a set position in the code file, where the set position is a position after the code generated by the code generation model is inserted into the code file; when the prompt information of the first line of code does not contain the same code information in multiple prompt information corresponding to the third line of code, the first line of code is a compliant code.

[0025] In one embodiment, the first processing module, when obtaining the first line of code generated by the code generation model, performs similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code, and is also used to obtain multiple prompt information corresponding to the second line of code from the code cache when the number of code lines stored in the code cache is equal to the second line.

[0026] In a third aspect, an embodiment of the present application provides a computing device, comprising: at least one memory; and at least one processor, the processor being used to execute instructions stored in the memory so that the computing device executes various possible implementations of the first aspect.

[0027] In a fourth aspect, a computer-readable storage medium is provided in an embodiment of the present application, comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes various possible implementations of the first aspect.

[0028] In a fifth aspect, a computer program product comprising instructions is provided in an embodiment of the present application, characterized in that the computer program product stores instructions, and when the instructions are executed by a computing device, the computing device implements each possible implementation embodiment of the first aspect.

[0029] In a sixth aspect, an embodiment of the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes various possible implementations of the first aspect.

[0030] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes various possible implementations of the first aspect.

[0031] In an eighth aspect, a computer program product comprising instructions is provided in an embodiment of the present application, characterized in that the computer program product stores instructions, and when the instructions are executed by a computing device cluster, the computing device cluster implements each possible implementation embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The following is a brief introduction to the drawings required for describing the embodiments or prior art.

[0033] Figure 1 A schematic diagram of the structure of a code generation system provided in an embodiment of the present application;

[0034] Figure 2 A schematic diagram of the process of code generation by the code matching module provided in an embodiment of the present application;

[0035] Figure 3 A schematic diagram of the time relationship for checking the compliance of N lines of code provided in an embodiment of the present application;

[0036] Figure 4 A schematic diagram of a process of generating code by a code generation system provided in an embodiment of the present application;

[0037] Figure 5 A schematic diagram of a scenario in which a user uses a code generation system provided in an embodiment of the present application;

[0038] Figure 6 A schematic diagram of the structure of a code generation device provided in an embodiment of the present application;

[0039] Figure 7 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0040] Figure 8 A schematic diagram of the architecture of a computing device cluster provided in an embodiment of the present application;

[0041] Fig. 9 A schematic diagram of the architecture of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0043] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.

[0044] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first response message and a second response message are used to distinguish different response messages rather than to describe a specific order of the response messages.

[0045] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0046] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0047] Before introducing the technical solution protected by this application, several professional terms involved in the technical solution protected by this application are explained in advance, namely:

[0048] A token is a basic processing unit in programming. When generating text, the model usually generates one token at a time. This process is serial, that is, the generation of each token depends on the previously generated token. The time delay for the model to generate each token is about 20-50 milliseconds (ms), which includes the time required for the model to calculate and generate the next token.

[0049] An inverted index is an indexing method that is widely used in information retrieval systems, especially in full-text search engines. The main idea of ​​an inverted index is to map words that appear in a document to a list of documents where the word appears. This index structure enables documents containing specific words to be quickly located during a query, thereby improving retrieval efficiency.

[0050] Next, the technical solution provided by this application is introduced.

[0051] In general, most of the current large-scale generative models are based on the Transformer architecture, which usually outputs serially token by token. On average, the output delay of each token is about 20-50 milliseconds, and a line of code contains about 10-20 tokens, so it takes about a few hundred milliseconds to one second to generate a line of code. Code generation based on such large models is either streamed output by token or output as a single line. Currently, overall output at the fragment level has not been implemented. The code output by the large model may be similar to existing open source code, so it is necessary to perform similarity detection on the output code.

[0052] Traditional open source similar code matching algorithms are based on the file level, and they find similarities by comparing fragments in the input file with fragments in the open source code base. These algorithms are not very real-time, and the scanning time for a single fragment takes at least several hundred milliseconds. If these algorithms are applied to compliance detection of large model-generated code, the scan can only be performed after the large model outputs the complete code fragment, and the user will be informed after several hundred milliseconds whether there is a license risk. For example, GitHub Copilot's Public code filter function is triggered after the user accepts the code completion suggestion, and the results are output in the form of asynchronous logs.

[0053] Methods like GitHub Copilot are relatively effective in identifying license issues in large model-generated code, but due to its asynchronous execution characteristics, it is triggered only after the user accepts the generated code. This means that when the scan results are returned, the user may have already written more new code based on them. If the user needs to review these scan results and make corresponding modifications, this may interrupt their programming train of thought, thus affecting development efficiency.

[0054] In view of this, an embodiment of the present application provides a code compliance detection method, which can utilize the serial output characteristics of the large model to generate code, and can perform compliance detection on the line of code after each line of code is output. While the method performs compliance detection on the received line of code, it pauses the code generation model to generate the next line of code. If the received test result indicates that the current line of code is non-compliant code, the method can instruct the code generation model to regenerate compliant current line of code and other lines of code. Throughout the process, the method continuously outputs compliant code without interrupting the user's programming ideas, thereby improving development efficiency.

[0055] Figure 1 Schematic diagram of the structure of a code generation system provided in an embodiment of the present application. Figure 1As shown, the code generation system 100 includes a code request module 110 , a code management module 120 , a code generation model 130 , a code matching module 140 , a code library 150 and a code cache 160 .

[0056] The code request module 110 can be deployed in the client to trigger code generation and code completion requests after receiving the operation instructions input by the user, so as to request the backend service (such as the code management module 120, etc.) to perform the code generation and code completion tasks. Among them, code generation refers to generating code to implement the corresponding function according to the text description. Code completion refers to supplementing the following code associated with the previous text according to the existing code context.

[0057] The code request module 110 may receive the generated code sent by the backend service and present the code to the user. The code request module 110 may render the non-compliant code, such as presenting the non-compliant code in red font, to remind the user that the current code is non-compliant.

[0058] The code management module 120 refers to a service deployed on the cloud management platform, which is used to generate corresponding code generation prompts based on the requests after receiving code generation and code completion requests, and input the prompts into the code generation model 130 to instruct the code generation model 130 to generate code that meets the prompts.

[0059] The code management module 120 can receive new codes generated by the code generation model 130. In the process of receiving codes, the code management module 120 can send the set number of lines of code to the code matching module 140 each time it receives the set number of lines of code, so that the code matching module 140 can detect the compliance of the set number of lines of code.

[0060] Normally, the time taken by the code generation model 130 to generate a line of code is longer than the time taken by the code matching module 140 to detect a line of code. Therefore, each time the code management module 120 receives a line of code, it sends the line of code to the code matching module 140, allowing the code matching module 140 to detect the compliance of the current line of code, thereby minimizing the delay in generating compliant code. Based on this, when introducing the protection scheme of this application below, the example of setting the number of lines as one line is used. Among them, a line of code refers to the code in the line where the enter key is located, and can also be a logical line of code after grammatical analysis.

[0061] The code management module 120 can receive the matching result sent by the code matching module 140. If the matching result shows that the current line of code is a compliant code, the code management module 120 can send the current line of code to the code request module 110 so that the user can get a new code. If the matching result shows that the current line of code is a non-compliant code, the code management module 120 can, based on the line number of the non-compliant code carried in the matching result or the user input to the instruction, allow the code generation model 130 to pause the generation of the next line of code, and modify the code generation prompt so that the code generation model 130 can regenerate compliant code. If the code management module 120 does not obtain a compliant line of code within a preset set time for generating a line of code, the last generated code can be sent to the code request module 110, and the code request module 110 can render the non-compliant code to let the user know that the code is non-compliant code.

[0062] The code generation model 130 refers to a large model that can automatically generate code, such as the Pangu large model, etc. After receiving the code generation prompt, the code generation model 130 can perform reasoning based on the code generation prompt to generate code that meets user needs. After each line of code generated by the code generation model 130, it can be sent to the code management module 120 in real time.

[0063] The code matching module 140 is a service deployed on the cloud management platform. The code matching module 140 is used to receive the code sent by the code management module 120, and after receiving the current line of code, the current line of code is matched with the open source code in the code library 150 for similarity, so as to obtain prompt information of the code that is the same or similar to the current line of code.

[0064] Exemplarily, the code matching module 140 may pre-build an inverted index and store it in the code library 150. The inverted index may mark different lines of open source code with different line numbers according to the order in which the code is generated, and may group open source code with different line numbers together according to the code features of each line of code. For example, the inverted index may be:

[0065] {Code feature 1: [(file 1, line 1), (file 2, line 2), ...], Code feature 2: [(file 3, line 3), (file 4, line 4), ...]}

[0066] The above inverted index shows that the inverted index includes code features and code identifiers, and the code identifier includes a file identifier and a code line number. The code feature is the information for feature extraction of the code, the file identifier is the file identifier where the relevant information of a line of code is located, and the code line number is the line number of a line of code. Among them, the code features of the code of line number 1 and the code of line number 2 are the same, and are code feature 1. The code features of the code of line number 3 and the code of line number 4 are the same, and are code feature 1. The code of line number 1 is associated with the file identifier of file 1, the code of line number 2 is associated with the file identifier of file 2, the code of line number 3 is associated with the file identifier of file 3, the code of line number 4 is associated with the file identifier of file 4, and so on.

[0067] The code base 150 may include an open source code base and a license base, wherein the open source code base is used to store a large amount of open source code, and the license base is used to store the license of the open source code. The code base 150 also stores a file index in advance. The file index may associate a file identifier with code information such as the license of the open source software to which a line of code belongs, the address of the open source software to which it belongs, and the file path of the open source software to which it belongs.

[0068] When receiving a line of code, the code matching module 140 may perform feature analysis on the received line of code to obtain the code feature of the current line of code. The code matching module 140 may search for at least one code identifier with the same code feature in the inverted index in the code library 150 based on the code feature of the current line of code.

[0069] If the code matching module 140 matches at least one code identifier with the same code feature as the current line code, prompt information of the current code may be generated based on the at least one code identifier. The prompt information includes at least one code identifier. Then, the code matching module 140 may cache the prompt information of the current line code in the code cache 160. If the code matching module 140 does not match a code identifier with the same code feature as the current line code, prompt information may not be generated, or prompt information indicating that the current line code does not have the same or similar code may be generated, and the current line code and the prompt information of the current line code are cached in the code cache 160.

[0070] The code matching module 140 can define rules for determining whether a code snippet is non-compliant. For example, if at least N consecutive lines of code in a code snippet are identical or similar to the open source code, the code snippet is considered to be a non-compliant code snippet. N is a positive integer greater than 1.

[0071] When the code cache 160 stores the prompt information of N-1 lines of code, the code matching module 140 can compare the prompt information of the N-1 lines of code to determine whether the prompt information of the N-1 lines of code all include the same code identifier. If not, the code matching module 140 can directly send the matching result to the code management module 120, and the matching result indicates that the N-1 lines of code are compliant codes. Subsequently, the code matching module 140 can clear the prompt information of the N-1 lines of code cached in the code cache 160.

[0072] If yes, the code matching module 140 can save the same code identifier. After the code matching module 140 receives the Nth line of code and generates the prompt information of the Nth line of code, it can compare the prompt information of the Nth line of code with the saved code identifier to determine again whether the same code identifier exists.

[0073] If not, the code matching module 140 can directly send a matching result to the code management module 120, and the matching result indicates that the N-line code is a compliant code. Subsequently, the code matching module 140 can clear the code identifier stored in the code cache 160 and the prompt information of the N-th line of code.

[0074] If yes, the code matching module 140 can search for the corresponding code information from the file index of the code library 150 based on the same code identifier. Optionally, the code matching module 140 can obtain the license from the license library based on the corresponding code information.

[0075] The code matching module 140 can detect whether the code information is compliant code information, such as whether the license is a compliant license, whether the address of the open source software is the address of the compliant open source software, etc. If the code information is compliant code information, the code matching module 140 can send a matching result to the code management module 120, and the matching result indicates that the N lines of code are compliant code. If the code information is non-compliant code information, the code matching module 140 sends a matching result to the code management module 120, and the matching result indicates that the N lines of code are non-compliant code. Optionally, when the code matching module 140 sends a matching result indicating that it is non-compliant, the line numbers of the current line of code and the N lines of code can be added to the matching result to instruct the code management module 120 to regenerate the corresponding lines of code.

[0076] The following is a flowchart to introduce the process of the code matching module 140 implementing the technical solution protected by the present application.

[0077] Figure 2 Schematic diagram of the process of code detection by the code matching module provided in the embodiment of the present application. Figure 2As shown, the method can be performed by the above-mentioned code matching module 140, and the specific implementation process is as follows:

[0078] Step S201, receiving the nth line of code.

[0079] Specifically, the nth line of code received by the code matching module 140 comes from the code management module 120. Each time the code management module 120 receives a line of code, it sends the current line of code to the code matching module 140, so that the code matching module 140 can detect the compliance of the current line of code. A line of code refers to the code in the line where a carriage return key is located, or it can be a logical line of code after syntax analysis.

[0080] Step S202 : performing similarity matching between the nth line of code and the open source code in the code library 150 , and obtaining prompt information of the nth line of code.

[0081] The code matching module 140 may perform similarity matching between the nth line of code and the open source code in the code library 150 to obtain prompt information of the code that is the same as or similar to the current line of code.

[0082] Exemplarily, the code matching module 140 can pre-build an inverted index and store it in the code base 150. The inverted index can mark different line numbers for different lines of open source code according to the order of code generation, and group open source codes with different line numbers together according to the code features of each line of code. The code base 150 can include an open source code base and a license base, the open source database is used to store a large amount of open source code, and the license base is used to store the licenses of the open source code. The code base 150 also pre-stores a file index. The file index can associate the file identifier with the code information such as the license of the open source software to which a line of code belongs, the address of the open source software to which it belongs, and the file path of the open source software to which it belongs.

[0083] When receiving the nth line of code, the code matching module 140 may perform feature analysis on the received nth line of code to obtain the code feature of the nth line of code. The code matching module 140 may search for at least one code identifier with the same code feature in the inverted index in the code library 150 based on the code feature of the nth line of code.

[0084] If the code matching module 140 matches at least one code identifier with the same code feature as the nth line of code, prompt information of the current code may be generated based on the at least one code identifier. The prompt information includes at least one code identifier. Then, the code matching module 140 may cache the prompt information of the nth line of code in the code cache 160. If the code matching module 140 does not match a code identifier with the same code feature as the current line of code, prompt information may not be generated, or prompt information indicating that the nth line of code does not have the same or similar code may be generated.

[0085] Step S203, checking whether the number of lines of code stored in the code cache 160 is less than N-1. If the number of lines of code is less than N-1, executing step S204. If the number of lines of code is equal to N-1, executing step S205.

[0086] Step S204 , cache the nth line of code and the prompt information of the nth line of code in the code cache 160 .

[0087] Step S205 , obtaining comparison information of N−1 lines from the code cache 160 .

[0088] Optionally, step S202 and step S203 may be performed in any order and may be performed simultaneously.

[0089] The code matching module 140 can detect the number of lines of code stored in the code cache 160 and determine whether the number of lines of code is less than N-1. In one case, the number of lines of code is less than N-1, and the code matching module 140 can cache the nth line of code and the prompt information of the nth line of code in the code cache 160.

[0090] In another case, the number of lines of code is equal to N-1, and the code matching module 140 may obtain the comparison information from the code cache 160. The comparison information is the same prompt information identified in the prompt information of the N-1 lines of code.

[0091] Optionally, when the code matching module 140 determines that the prompt information of N-1 lines of code is stored in the code cache 160, the prompt information of the N-1 lines of code can be compared to determine whether the prompt information of the N-1 lines of code all include the same code identifier. If not, the code matching module 140 can directly send the matching result to the code management module 120, and the matching result indicates that the N-1 lines of code are compliant codes. Subsequently, the code matching module 140 can clear the prompt information of the N-1 lines of code cached in the code cache 160, and cache the prompt information of the nth line of code and the nth line of code in the code cache 160. If yes, the code matching module 140 can store the same code identifier in the code cache 160.

[0092] Step S206, compare the prompt information of the nth line of code with the comparison information to determine whether there is an identical code identifier. If not, execute step S207. If yes, execute step S208.

[0093] Step S207: Send the first matching result to the code management module 120. The first matching result includes information indicating that the N lines of code are compliant codes.

[0094] After the code matching module 140 obtains the comparison information of the N-1th line, it can compare the prompt information of the nth line of code with the saved code identifier to determine whether there is an identical code identifier. If not, the code matching module 140 can directly send the matching result to the code management module 120, and the matching result indicates that the Nth line of code is a compliant code. Subsequently, the code matching module 140 can clear the code identifier saved in the code cache 160 and the prompt information of the Nth line of code.

[0095] Step S208: determining corresponding code information based on the same code identifiers of the N lines of code.

[0096] Step S209, determining whether the code information is non-compliant code information. If the code information is compliant code information, executing step S207. If the code information is non-compliant code information, executing step S210.

[0097] Step S210: Send the second matching result to the code management module 120. The second matching result includes information indicating that the N lines of code are non-compliant codes.

[0098] The code matching module 140 may search for corresponding code information from the file index of the code library 150 based on the same code identifier. Optionally, the code matching module 140 may obtain a license from a license library based on the corresponding code information.

[0099] The code matching module 140 can detect whether the code information is compliant code information, such as whether the license is a compliant license, whether the address of the open source software is the address of the compliant open source software, etc. If the code information is compliant code information, the code matching module 140 can send a matching result to the code management module 120, and the matching result indicates that the N lines of code are compliant code. If the code information is non-compliant code information, the code matching module 140 sends a matching result to the code management module 120, and the matching result indicates that the N lines of code are non-compliant code. Optionally, when the code matching module 140 sends a matching result indicating that it is non-compliant, the line numbers of the current line of code and the N lines of code can be added to the matching result to instruct the code management module 120 to regenerate the corresponding lines of code.

[0100] Optionally, when the code matching module 140 determines that the code request module 110 triggers a request for code completion, the N-1 lines of original code after the position to be completed can be obtained from the code file to be completed. The code matching module 140 performs similarity matching on the N-1 lines of original code with the code in the code library 150 to obtain prompt information of the N-1 lines of original code. The code matching module 140 compares the prompt information of the N-1 lines of original code to obtain comparison information of the N-1 lines of original code.

[0101] After determining that the nth line of code is a compliant code, the code matching module 140 can compare the prompt information of the current line of code with the comparison information of the N-1 line of original code to determine whether there is an identical code identifier in the prompt information of the N lines of code. If not, the code matching module 140 can send a matching result to the code management module 120, and the matching result indicates that the nth line of code is a compliant code. If so, the code matching module 140 sends a matching result to the code management module 120, and the matching result indicates that the nth line of code is an uncompliant code.

[0102] Assume that the time it takes for the code matching module 140 to search for a line of code in the inverted index is t1. The time it takes for the code matching module 140 to find a code information (taking a license as an example) from the code library 150 is t2. The time it takes for the code matching module 140 to compare the prompt information of a line of code with the prompt information of another line of code is t3. The number of code information (taking a license as an example) that the code matching module 140 searches for in the code library 150 is K. The time it takes for the code generation model 130 to generate a line of code is t4. The total number of lines of code that the code generation model 130 needs to generate is M, where M is greater than N and is a positive integer greater than 1.

[0103] like Figure 3 As shown in the figure, the time relationship for checking the compliance of N lines of code is:

[0104] The time T1=t4 for the code generation model 130 to generate a line of code.

[0105] In the related art, after the code generation model generates M lines of code, a sliding window of length N can be used to slide on the M lines of code, and each slide will move one line of code. The time for the code matching module to detect the compliance of N lines of code each time is T2. That is, the time for simultaneously detecting the compliance of N lines of code in the related art is T2=N*t1+K*t2+(N-1)*t3.

[0106] Since M lines of code have been generated, the code matching module detects that one or several N lines of code among the M lines of code are non-compliant codes. The non-compliant codes need to be modified again, and the subsequent codes of the non-compliant codes need to be regenerated, which results in a longer time and interrupts the user's programming ideas.

[0107] In the technical solution of the present application, in the process of the code generation model 130 generating N lines of code, when the code generation model 130 generates the first line of code, the code matching module 140 will not detect. When the code generation model 130 generates the second line of code, the code matching module 140 searches the first line of code in the inverted index at time t1. When the code generation model 130 generates the third line of code, the code matching module 140 searches the second line of code in the inverted index at time t1. Similarly, when the code generation model 130 generates the Nth line of code, the code matching module 140 searches the N-1th line of code in the inverted index at time t1. At this time, since the code cache 160 stores N-1 lines of code, the code matching module 140 can compare the prompt information of the N-1 line of code, and the time is (N-2)*t3. During the whole process, the code generation model 130 executes time T4. That is, the time T4 = (N-1)*t1+ (N-2)*t3 for the code matching module 140 to detect the compliance of the N-1 codes before the Nth row.

[0108] After the code generation model 130 generates the Nth line of code, the code generation can be paused. At this time, the code matching module 140 can search the Nth line of code in the inverted index, and the time is t1. The code matching module 140 can compare the prompt information of the Nth line of code with the comparison information of the N-1th line, and the time is t3. The code matching module 140 searches for K licenses from the code library 150, and the time is K*t2. During the whole process, the code generation model 130 executes for T3. That is, the time when the code matching module 140 detects the compliance of the Nth line of code is T3=t1+K*t2+t3.

[0109] If the matching result received by the code management module 120 indicates that the current line of code is non-compliant code, the code generation model 130 can be instructed to pause generating the next line of code, and to regenerate the compliant current line of code and other lines of code. Throughout the process, the code generation system 100 continuously outputs compliant code without interrupting the user's programming ideas, thereby improving development efficiency.

[0110] The following is a flowchart to introduce the process of the code matching module 140 implementing the technical solution protected by the present application.

[0111] Figure 4Schematic diagram of the process of generating code by the code generation system provided in the embodiment of the present application. Figure 4 As shown, the method can be jointly executed by the code request module 110, the code management module 120, the code generation model 130 and the code matching module 140 in the above-mentioned code generation system 100, and the specific implementation process is as follows:

[0112] Step S401 : the code request module 110 sends a code generation request to the code management module 120 .

[0113] Specifically, after receiving the operation instruction input by the user, the code request module 110 triggers the request for code generation and code completion, and sends the code generation request to the code management module 120. The code generation request is used to request the backend service (such as the code management module 120, etc.) to perform the code generation and code completion tasks. Among them, code generation refers to generating code to implement the corresponding function according to the text description. Code completion refers to supplementing the following code associated with the previous text according to the existing code context.

[0114] Step S402: the code management module 120 generates a code generation prompt based on the code generation request.

[0115] Step S403 : the code management module 120 sends a code generation prompt to the code generation model 130 .

[0116] Specifically, after receiving the code generation request, the code management module 120 may generate corresponding code generation hints based on the code generation request, and input the hints into the code generation model 130 to instruct the code generation model 130 to generate code that satisfies the hints.

[0117] Step S404 : after generating a line of code each time, the code generation model 130 sends the generated line of code to the code management module 120 .

[0118] Step S405 , after receiving a line of code, the code management module 120 sends the generated line of code to the code matching module 140 .

[0119] Step S406 : the code matching module 140 sends the matching result to the code management module 120 .

[0120] Specifically, the code management module 120 can receive the code generated by the code generation model 130. In the process of receiving the code, the code management module 120 can send a line of code to the code matching module 140 each time it receives a line of code, so that the code matching module 140 can detect the compliance of the line of code. The process of the code matching module 140 detecting the compliance of a line of code can be referred to Figure 2 and Figure 2Related content described.

[0121] Step S407: The code management module 120 determines whether a generated line of code is compliant code based on the matching result. If the generated line of code is compliant code, step S408 is executed. If the generated line of code is non-compliant code, step S409 is executed.

[0122] Step S408: The code management module 120 sends the generated line of code to the code request module 110.

[0123] Specifically, the code management module 120 can receive the matching result sent by the code matching module 140 and determine whether a generated line of code is compliant code. If the matching result indicates that the current line of code is compliant, the code management module 120 can send the current line of code to the code request module 110 so that the user can obtain a line of code.

[0124] At this time, the code management module 120 can continue to wait. After the time reaches the set time for generating a line of code, it will receive the next line of code sent by the code generation model 130, and then execute the process of steps S405 - S408, or steps S405 - S410, or steps S405 - S413 again.

[0125] Step S409: The code management module 120 determines whether the generation of a line of code exceeds the set time. If the generation of a line of code exceeds the set time, step S410 is executed. If the generation of a line of code does not exceed the set time, step S411 is executed.

[0126] Step S410: The code management module 120 sends the generated line of code to the code request module 110.

[0127] Specifically, if the matching result indicates that the current line of code is non-compliant code, the code management module 120 determines whether the time for generating the current line of code exceeds the set time. If it exceeds the set time, the code management module 120 can send the current line of code to the code request module 110 and let the code request module 110 render the non-compliant code, so that the user knows that the current line of code is non-compliant code.

[0128] At this time, the code management module 120 will receive the next line of code sent by the code generation model 130, and then execute the process of steps S405 - S408, or steps S405 - S410, or steps S405 - S413 again.

[0129] Step S411: The code management module 120 sends a pause instruction to the code generation model 130.

[0130] Step S412: the code management module 120 modifies the code generation prompt to obtain a modified code generation prompt.

[0131] Step S413 : the code management module 120 sends the modified code generation hint to the code generation model 130 .

[0132] Specifically, if the set time has not been exceeded, the code management module 120 first sends a pause instruction to the code generation model 130, instructing the code generation model 130 to pause generating the next line of code. Then, the code management module 120 can modify the code generation prompt based on the line number of the non-compliant code carried in the matching result or the user input to the instruction, and send the modified code generation prompt to the code generation model 130, so that the code generation model 130 regenerates a line of code, and then repeats the process of steps S404-S413.

[0133] In the embodiment of the present application, each time the code management module 120 receives a line of code, it sends the received line of code to the code matching module 140. While the code matching module 140 performs compliance detection on the received line of code, it suspends the code generation model 130 from generating the next line of code. If the matching result received by the code management module 120 indicates that the current line of code is non-compliant code, the code generation model 130 can be instructed to regenerate the compliant current line of code and other lines of code. During the whole process, the code generation system 100 continuously outputs compliant code without interrupting the user's programming ideas, thereby improving development efficiency.

[0134] It should be understood that the functional modules, functional devices, etc. involved in the above-mentioned code generation system 100 can also be implemented by software or hardware, which can be determined according to actual conditions and is not limited here. In addition, the functional modules, functional devices, etc. involved in the above-mentioned code generation system 100 can be arranged separately or integrated, which is not limited here.

[0135] The above is an introduction to the code generation system 100 provided in the embodiment of the present application. It is understandable that the above-mentioned code generation system 100 can be configured on a cloud management platform, for example, deployed on at least one virtual machine or container instance, so that the cloud management platform can provide code generation services. Of course, the code generation system 100 can also be configured on a node other than the cloud management platform. For example, it can be deployed in at least one data center, or deployed on at least one server. The specific situation can be determined according to the actual situation and is not limited here. Among them, the cloud management platform can provide pages related to public cloud services for users to remotely access public cloud services. In this embodiment, users can purchase the code generation service that the code generation system 100 can provide on the cloud management platform in advance. For ease of understanding, the interaction between the user and the cloud management platform is described below.

[0136] like Figure 5 As shown, the interaction between the user and the cloud management platform mainly includes: the user logs in to the cloud management platform 500 through the web page of the client, selects and purchases the cloud service (i.e., the code generation service) related to the code generation system 100 in the cloud management platform 500, and after the purchase, the user can generate the code generation system 100 on the cloud management platform 500 based on the functions provided by the code generation service. Among them, the cloud management platform 500 is mainly used to manage the infrastructure for running the code generation service. Exemplarily, the infrastructure of the code generation service may include multiple data centers set up in different regions, each data center including multiple servers. The data center can provide basic resources for the code generation service, such as computing resources, storage resources, etc. Therefore, when the user purchases and uses the code generation service, the user mainly pays for the resources used. When using the code generation service, the user can input its demand for the code generation service through the configuration interface, application program interface (API) or interface for interacting with the user provided by the cloud management platform 500, and the cloud management platform 500 can generate a code generation service that matches the user's demand according to the demand input by the user (or other software / hardware, etc.).

[0137] In addition, a part of the modules in the code generation system 100 can also be configured on the cloud side, and another part can be configured on the terminal side, thereby realizing the code generation service through the terminal-cloud collaboration. In addition, the code generation system 100 can also be configured entirely on the terminal side, which can be determined according to the actual situation and is not limited here.

[0138] Based on the above description, the embodiment of the present application provides a code generation device 600. The device is connected to a storage end and a large model. The storage end is used to store a large amount of open source code, and the large model is used to generate code. Figure 6 As shown, the device 600 includes:

[0139] The first processing module 610 is used to obtain the code that the code generation model outputs token by token in series; the second processing module 620 is used to, when obtaining the first line of code generated by the code generation model, perform similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code; the prompt information of the first line of code includes at least one code identifier, each code identifier is an identifier of a license to which the code that is identical or similar to the one line of code belongs; the third processing module 630 is used to determine that the first line of code and the second line of code are compliant codes when there is no identical code identifier in the prompt information of the first line of code and the multiple prompt information corresponding to the second line of code.

[0140] In one embodiment, the third processing module 630 is also used to, when the same code identifier exists in the prompt information of the first line of code and the multiple prompt information corresponding to the second line of code, based on the same code identifier, obtain code information corresponding to the same code identifier, the code information including one or more of the license of the open source software to which the code belongs, the address of the open source software to which the code belongs, and the file path of the open source software to which the code belongs; detect whether the code information corresponding to the same code identifier is compliant code information; when the code information corresponding to the same code identifier is compliant code information, the first line of code and the second line of code are compliant codes.

[0141] In one implementation, the third processing module 630 is further configured to instruct the code generation model to regenerate the first line of code and the second line of code when the code information corresponding to the same code identifier is non-compliant code information.

[0142] In one embodiment, when the first line of code generated by the code generation model is obtained, the second processing module 620 performs similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code, and is also used to create an inverted index in the code library; the inverted index includes code features and code identifiers for each line of code in the code library; similarity matching is performed between the first line of code and the code in the code library to obtain prompt information of the first line of code, including: performing feature analysis on the first line of code to obtain code features of the first line of code; based on the code features of the first line of code, obtaining at least one code identifier with the same code features as the code features of the first line of code from the inverted index to obtain prompt information of the first line of code.

[0143] In one embodiment, the second number of lines of code is code generated by the code generation model before generating the first number of lines of code, and is continuous with the first number of lines of code, the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines, and the non-compliant code fragment refers to a code fragment with the same number of lines as the code in the code library being greater than or equal to the set number of lines.

[0144] In one embodiment, the code generated by the code generation model is a missing code fragment in the code file, and the third processing module 630 is also used to obtain a third line of code at a set position in the code file, where the set position is a position after the code generated by the code generation model is inserted into the code file; when the prompt information of the first line of code does not contain the same code information in multiple prompt information corresponding to the third line of code, the first line of code is a compliant code.

[0145] In one embodiment, when the first line of code generated by the code generation model is obtained, the first processing module 610 performs similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code. When the number of code lines stored in the code cache is equal to the second number of lines, multiple prompt information corresponding to the second number of lines of code are obtained from the code cache.

[0146] Among them, the first processing module 610, the second processing module 620 and the third processing module 630 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the first processing module 610 is introduced below by taking the first processing module 610 as an example. Similarly, the implementation of the second processing module 620 and the third processing module 630 can refer to the implementation of the first processing module 610.

[0147] As an example of a software functional unit, the first processing module 610 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the first processing module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Generally, a region may include multiple AZs.

[0148] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0149] As an example of a hardware functional unit, the first processing module 610 may include at least one computing device, such as a server, etc. Alternatively, the first processing module 610 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0150] The multiple computing devices included in the first processing module 610 can be distributed in the same region or in different regions. The multiple computing devices included in the first processing module 610 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the first processing module 610 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0151] It should be noted that, in other embodiments, the first processing module 610 can be used to perform the following steps: Figure 2 In any step of the method shown in FIG. 1 , the second processing module 620 can be used to perform the following steps: Figure 2 In any step of the method shown in FIG. 1 , the third processing module 630 can be used to perform the following steps: Figure 2 In any step of the method shown in FIG. 1 , the first processing module 610, the second processing module 620, and the third processing module 630 are responsible for implementing the steps that can be specified as needed, and the first processing module 610, the second processing module 620, and the third processing module 630 are respectively implemented as follows: Figure 2The different steps in the method shown are used to implement the overall functions of the apparatus 600 .

[0152] Figure 7 Schematic diagram of a computing device provided in an embodiment of the present application. Figure 7 As shown, the computing device 700 includes a bus 710, a processor 720, a memory 730, and a communication interface 740. The processor 720, the memory 730, and the communication interface 740 communicate with each other through the bus 710. The computing device 700 can be a server, a computer, a portable notebook, a cabinet, etc. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.

[0153] The bus 710 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 710 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 710 may include a path for transmitting information between various components of the computing device 700 (eg, the processor 720, the memory 730, and the communication interface 740).

[0154] The processor 720 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0155] The memory 730 may include a volatile memory, such as a random access memory (RAM). The memory 730 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0156] The memory 730 stores executable program codes, and the processor 720 executes the executable program codes to respectively implement the functions of the aforementioned multiple modules, such as the first processing module 610, the second processing module 620, and the third processing module 630, so as to implement the following. Figure 2 That is, the memory 730 stores the method for executing Figure 2 Instructions for the method shown.

[0157] Alternatively, the memory 730 stores executable codes, and the processor 720 executes the executable codes to respectively implement the functions of the aforementioned modules, thereby implementing the following: Figure 2 That is, the memory 730 stores the method for executing Figure 2 Instructions for the method shown.

[0158] The communication interface 740 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.

[0159] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0160] like Figure 8 As shown, the computing device cluster includes at least one computing device 700. The memory 730 in one or more computing devices 700 in the computing device cluster may store the same Figure 2 Instructions for the method shown.

[0161] In some possible implementations, the memory 730 of one or more computing devices 700 in the computing device cluster may also store a program for executing the following steps: Figure 2 In other words, the combination of one or more computing devices 700 can jointly execute instructions for performing the method as shown. Figure 2 Instructions for the method shown.

[0162] It should be noted that the memory 730 in different computing devices 700 in the computing device cluster may store different instructions, which are respectively used to execute part of the functions of the first processing module 610, the second processing module 620, and the third processing module 630. That is, the instructions stored in the memory 730 in different computing devices 700 may implement the functions of one or more of the first processing module 610, the second processing module 620, and the third processing module 630.

[0163] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig. 9 A possible implementation is shown. Fig. 9 As shown, two computing devices are connected via a network, namely computing device 700A and computing device 700B. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 730 in the computing device 700A stores instructions for executing the functions of some modules in the first processing module 610 and the second processing module 620. At the same time, the memory 730 in the computing device 700B stores instructions for executing the functions of another part of the modules in the third processing module 630.

[0164] Fig. 9 The connection mode between the computing device clusters shown can be considered as provided in the present application. Figure 2 The method shown requires a large amount of data storage, so it is considered that the functions implemented by another part of the first processing module 610, the second processing module 620 and the third processing module 630 are handed over to the computing device 700B for execution.

[0165] It should be understood that Fig. 9 The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.

[0166] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 7 and Figure 8 The difference is that the memory 730 in one or more computing devices 700 in the computing device cluster may store the same memory 730 for executing the following operations: Figure 2 Instructions for the method shown.

[0167] In some possible implementations, the memory 730 of one or more computing devices 700 in the computing device cluster may also store a program for executing the following steps: Figure 2 In other words, the combination of one or more computing devices 700 can jointly execute instructions for performing the method as shown. Figure 2 Instructions for the method shown.

[0168] It should be noted that the memory 730 in different computing devices 700 in the computing device cluster may store different instructions for executing partial functions of the computing device 700. That is, the instructions stored in the memory 730 in different computing devices 700 may implement the functions of one or more of the first processing module 610, the second processing module 620, and the third processing module 630 described above.

[0169] The present application also provides a computer program product including instructions. The computer program product may be a software or program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the following steps: Figure 2 The method shown.

[0170] The present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the following: Figure 2 The method shown.

[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A code generation method, characterized in that: The method is applied to a cloud management platform, the cloud management platform is used to manage infrastructure, the infrastructure includes at least one node, the at least one node stores a code generation model, the code generation model is used to generate code, the method includes: Obtain the code that the code generation model outputs word by word in series; When the first line of code generated by the code generation model is obtained, the first line of code is matched with the code in the code library for similarity to obtain prompt information of the first line of code; the prompt information of the first line of code includes at least one code identifier, each code identifier is an identifier of a license to which a code identical or similar to the first line of code belongs; When the prompt information of the first number of lines of code and the multiple prompt information corresponding to the second number of lines of code do not have the same code identifier, the first number of lines of code and the second number of lines of code are compliant codes, the second number of lines of code is the code generated by the code generation model before generating the first number of lines of code, and is continuous with the first number of lines of code, the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines, and an non-compliant code fragment refers to a code fragment in which the number of lines of code that is the same as the code in the code library is greater than or equal to the set number of lines.

2. The method according to claim 1, characterized in that The method further comprises: When the prompt information of the first number of lines of code and the plurality of prompt information corresponding to the second number of lines of code have the same code identifier, based on the same code identifier, obtaining code information corresponding to the same code identifier, the code information including one or more of a license of the open source software to which the code belongs, an address of the open source software to which the code belongs, and a file path of the open source software to which the code belongs; Detecting whether the code information corresponding to the same code identifier is compliant code information; When the code information corresponding to the same code identifier is the compliant code information, the first line number code and the second line number code are compliant codes.

3. The method according to claim 2, characterized in that The method further comprises: When the code information corresponding to the same code identifier is non-compliant code information, the code generation model is instructed to regenerate the first line number of code and the second line number of code.

4. The method according to any one of claims 1 to 3, characterized in that: When the first line of code generated by the code generation model is obtained, the first line of code is matched with the code in the code library for similarity, and before obtaining prompt information of the first line of code, the method further includes: Creating an inverted index in the code base; the inverted index includes code features and code identifiers of each line of code in the code base; The performing similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code includes: Performing feature analysis on the first line code to obtain a code feature of the first line code; Based on the code feature of the first line of code, at least one code identifier of a code feature that is the same as the code feature of the first line of code is obtained from the inverted index to obtain prompt information of the first line of code.

5. The method according to any one of claims 1 to 3, characterized in that: The code generated by the code generation model is a code fragment missing from the code file, and the method further includes: Obtaining a third line of code at a set position in the code file, wherein the set position is a position after the code generated by the code generation model is inserted into the code file; When the prompt information of the first line number code and the plurality of prompt information corresponding to the third line number code do not contain the same code information, the first line number code is a compliant code.

6. The method according to any one of claims 1 to 3, characterized in that: When the first line of code generated by the code generation model is obtained, the first line of code is matched with the code in the code library for similarity, and before obtaining prompt information of the first line of code, the method further includes: When the number of code lines stored in the code cache is equal to the second number of lines, multiple prompt information corresponding to the second number of code lines is obtained from the code cache.

7. A code generating device, characterized in that: include: A first processing module is used to obtain the code that is serially output by the code generation model word by word; The second processing module is used for, when obtaining the first line of code generated by the code generation model, combining the first line of code with the code generation model The codes in the code library are matched for similarity to obtain prompt information of the first line of code; The prompt information of the first line of code includes at least one code identifier, each code identifier being an identifier of a license to which a code identical or similar to the first line of code belongs; The third processing module is used for, when there is no identical code identifier in the prompt information of the first number of lines of code and the multiple prompt information corresponding to the second number of lines of code, the first number of lines of code and the second number of lines of code are compliant codes, the second number of lines of code is the code generated by the code generation model before generating the first number of lines of code, and is continuous with the first number of lines of code, the sum of the first number of lines and the second number of lines is greater than or equal to the set number of lines, and the non-compliant code fragment refers to a code fragment with the same number of lines as the code in the code library being greater than or equal to the set number of lines.

8. The device according to claim 7, characterized in that The third processing module is also used When the prompt information of the first number of lines of code and the plurality of prompt information corresponding to the second number of lines of code have the same code identifier, based on the same code identifier, obtaining code information corresponding to the same code identifier, the code information including one or more of a license of the open source software to which the code belongs, an address of the open source software to which the code belongs, and a file path of the open source software to which the code belongs; Detecting whether the code information corresponding to the same code identifier is compliant code information; When the code information corresponding to the same code identifier is the compliant code information, the first line number code and the second line number code are compliant codes.

9. The device according to claim 8, characterized in that The third processing module is also used When the code information corresponding to the same code identifier is non-compliant code information, the code generation model is instructed to regenerate the first line number of code and the second line number of code.

10. The device according to any one of claims 7 to 9, characterized in that: The second processing module, when acquiring the first line of code generated by the code generation model, performs similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code, is further used to Creating an inverted index in the code base; the inverted index includes code features and code identifiers of each line of code in the code base; The performing similarity matching between the first line of code and the code in the code library to obtain prompt information of the first line of code includes: Performing feature analysis on the first line code to obtain a code feature of the first line code; Based on the code feature of the first line of code, at least one code identifier of a code feature that is the same as the code feature of the first line of code is obtained from the inverted index to obtain prompt information of the first line of code.

11. A computing device cluster, characterized in that: include: at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 6.

12. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device performs the method according to any one of claims 1 to 6.

13. A computer program product comprising instructions, characterized in that The computer program product stores instructions, which, when executed by a computing device cluster, enable the computing device cluster to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for converting source codes into numeric identifiers and comparison against data sets

    CN110941726A

  • Software license-based code suggestions

    US20240111843A1