Code generation method and device based on large model, medium and equipment

By syntax analysis and editing of the business logic process, and then generating the secondary expression text, using the large language model to generate code, the accuracy and reliability problems of the large language model when generating code in specific fields are solved, and the standardization and accuracy of code generation are achieved.

CN120540644AActive Publication Date: 2025-08-26ZHEJIANG ANT MISUAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511039367.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-26
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

When generating code in specific domains, large language models face problems such as lack of knowledge, errors in knowledge or expired knowledge, resulting in reduced accuracy and reliability of performing tasks.

Method used

By performing grammatical analysis of the business logic process described by natural language, atomic task units and their logical relationships are generated, re-edited into quadratic expression text, and code is generated using a pre-trained large language model to avoid logical confusion and ambiguity.

Benefits of technology

Improves the accuracy and reliability of generated code, ensuring that the standardization of input text is guaranteed regardless of the quality of business logic process writing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540644A_ABST
    Figure CN120540644A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a code generation method based on a large model, and the method comprises the steps: carrying out the grammatical analysis of a business logic process described by a natural language, and obtaining atomic task units and a logic relation between the atomic task units; and then re-editing according to the logical relationship to obtain a secondary expression text, and generating a code through LLM. And the problem of subsequent code generation through LLM due to logic relation chaos caused by natural language description for a business logic process is avoided. In the syntactic analysis process, the simplest expression corresponding to the natural language can be determined, ambiguity of description of steps to be executed in the business logic process is reduced, secondary expression texts are obtained through re-editing, the influence of logic chaos possibly occurring in the business logic process can be avoided, and the business logic processing efficiency is improved. Therefore, the normalization of the input LLM text is not affected regardless of the writing degree of the business logic process, and the accuracy and reliability of code generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a large model reasoning process detection method, device, storage medium and equipment. Background Art

[0002] With the development of artificial intelligence (AI) technology, large language models (LLM) have been widely used in various fields.

[0003] However, when LLMs are applied to generate code for specific domains, they often face numerous challenges. This is because specialized domains contain a vast, highly complex, and constantly evolving body of knowledge. LLMs, due to limited training data, can suffer from knowledge deficiencies, errors, or outdated knowledge. These issues can cause LLMs to experience "hallucinations" when performing tasks, reducing their accuracy and reliability.

[0004] Based on this, how to improve the accuracy and reliability of LLM execution tasks has become an urgent problem to be solved. Therefore, this specification provides a code generation method based on a large model. Summary of the Invention

[0005] The embodiments of this specification provide a code generation method, device, storage medium and electronic device based on a large model to partially solve the problems existing in the above-mentioned prior art.

[0006] The embodiments of this specification adopt the following technical solutions: This specification provides a code generation method based on a large model, the method comprising: Obtain business logic processes described in natural language; Performing grammatical analysis on the business logic process to generate atomic task units corresponding to the business logic process and logical relationships between the atomic task units, wherein the atomic task units are described in natural language; Re-editing the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process; Based on the secondary expression text, a knowledge high-level program for executing the business logic process is generated through a pre-trained large language model LLM.

[0007] This specification provides a code generation device based on a large model, the device comprising: An acquisition module is used to acquire the business logic process described in natural language; Parsing module, used to obtain business logic processes described in natural language; An editing module, configured to re-edit the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process; A generation module is used to generate a knowledge high-level program for executing the business logic process based on the secondary expression text through a pre-trained large language model LLM.

[0008] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned large model-based code generation method.

[0009] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned large model-based code generation method is implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: The embodiment of this specification discloses a code generation method based on a large model, which performs grammatical analysis on the business logic process described in natural language to obtain atomic task units and the logical relationship between atomic task units. It is then re-edited according to the logical relationship to obtain a secondary expression text, and the code is generated through LLM. The logical relationship confusion caused by the natural language description of the business logic process is avoided, which leads to problems when the code is subsequently generated through LLM. Among them, the simplest expression corresponding to the natural language can be determined in the grammatical analysis process, reducing the ambiguity in the description of the steps to be executed in the business logic process, and re-editing to obtain a secondary expression text can avoid the influence of logical confusion that may occur in the business logic process, so that no matter how high or low the business logic process is written, it will not affect the standardization of the input LLM text, thereby improving the accuracy and reliability of the generated code. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 A code generation flow chart based on a large model provided in the embodiments of this specification; Figure 2 A flow chart of a method for generating a HOP provided in an embodiment of this specification; Figure 3 A schematic diagram of a code generation device based on a large model provided in an embodiment of this specification; Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0012] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0013] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0014] Figure 1 The code generation flow chart based on the large model provided in the embodiment of this specification specifically includes the following steps: S100: Acquire a business logic process described in natural language.

[0015] In the embodiments of this specification, Figure 1 The device for generating codes using the method shown can be any electronic device, such as a computer, a server, or a server cluster consisting of multiple servers, etc. For ease of description, the following description will only take a server as an example.

[0016] In order to generate code for executing a certain business, the server can first obtain the business logic process of the business. Specifically, the business logic process described in this specification includes a standard operating procedure (SOP). The SOP is a standard operating procedure for the business described in natural language. It is a document that details the work, workflow, and steps in the business. It provides standard process specifications within the organization of the business for reference and operation by internal personnel of the organization. In addition to containing the standardized process of the business, the SOP also includes clear responsibilities and authorities of each role in the business, emergency measures, quality standards and assessments, document management, and other content.

[0017] Since the business logic process described in natural language in this specification includes SOP, the business logic process includes business operation steps and can reflect the logical relationship between the steps. Therefore, by processing the subsequent steps of the business logic process, code for implementing the business logic process can be generated. Whether it is a business logic process or an SOP, it is generally based on expert experience and is a manually written text. Therefore, it is difficult to avoid problems such as logical confusion or inconsistent descriptions. Although the accuracy of the logical description of the final generated code can be guaranteed by relying on traditional programming languages ​​(such as Python, Java, etc.), the conversion between the text described in natural language and the accurate code, as well as the possible grammatical and semantic problems of natural language, cannot be solved by programming language. For this reason, the embodiment of this specification processes the business logic process through the following steps and pre-processes the business logic process to avoid the above-mentioned situation causing errors in the final generated code.

[0018] S102: Performing grammatical analysis on the business logic process to generate atomic task units corresponding to the business logic process and logical relationships between the atomic task units, wherein the atomic task units are described in natural language.

[0019] After the server obtains the above-mentioned business logic process, it can perform a grammatical analysis on the business logic process described in natural language to determine how many steps are included in the business logic process, where each step generally corresponds to a task that needs to be performed. Therefore, a grammatical analysis is performed on each step to determine the sentence structure of the step, construct the atomic task unit corresponding to each step, and determine the logical relationship between each atomic task unit. The atomic task unit refers to a text that uses the minimum sentence elements to express the step content in natural language, so the atomic task unit is also described in natural language. The main purpose of performing grammatical analysis here is, on the one hand, to determine the logical relationship between the steps included in the business logic process, and on the other hand, to eliminate the ambiguous content in the step and retain only the necessary sentence elements to form the atomic task unit.

[0020] Specifically, the server may first determine the various steps included in the business logic process, where the steps are texts described in natural language. The server may utilize a trained semantic analysis model to input the business logic process into the semantic analysis model, and obtain classification results belonging to different statements given by the semantic analysis model. In this way, the business logic process is first split into several different steps from the perspective of statements and semantics. Therefore, the process can only be split, so the content of the steps is still the content of the original business logic process, that is, the steps described in natural language. Alternatively, the server directly determines the different statements in the business logic process, as well as the sentence structure and elements that make up the sentence structure in the statements through grammatical analysis. Compared with semantic analysis performed by a semantic analysis model, grammatical analysis is only based on grammar analysis, so there is a difference in the statement splitting capability. Of course, this manual does not limit the method to be used, and the following explanation will be given by first splitting the statements through a semantic analysis model and then determining the sentence structure through grammatical analysis.

[0021] The main purpose of this decomposition is to subsequently analyze each step individually and filter out ambiguous content. Furthermore, because the business logic process is manually written, the accuracy of the text description varies, and the description may not even follow the order of the steps. Therefore, to simplify analysis, the business logic process is first decomposed to determine the corresponding tasks for each step, that is, to identify atomic task units.

[0022] Afterwards, for each step, the server can perform a grammatical analysis on the step to determine the sentence structure of the step and the elements that make up the sentence structure. That is, a step is treated as a statement, and a grammatical analysis is performed to determine the sentence structure of the step and the elements in the sentence under the sentence structure. For example, a grammatical analysis is performed on step A to determine that step A is a sentence format with the object placed in front. Then, according to the sentence format, the text of the step is determined as follows: what text is the subject, what text is the predicate, and what text is the object. Moreover, the text used to describe the step is not necessary content for the atomic task unit and does not need to be determined. For example, the step description is "the content of the cmd field needs to be analyzed urgently". After splitting, the sentence elements predicate "analysis" and object "cmd field" are determined, and other content is interference content.

[0023] Finally, the elements that make up the sentence structure of this step are reconstructed according to the preset sentence format as atomic task units. Generally, to facilitate code generation in subsequent steps, the preset sentence format here uses a subject-verb-object format. Choosing other sentence formats, such as object placement, subject-verb inversion, or adverbial placement, may increase the complexity of subsequent steps. Of course, this manual does not limit the specific method used here, and it can be set as needed.

[0024] As can be seen from the last step, the obtained atomic task unit expresses the step content in the simplest form of natural language, which can reduce ambiguity caused by personal writing habits or non-standard writing habits in manual writing.

[0025] In addition, in the embodiments of this specification, the business logic process should be a text that is consistent with the context. However, manual writing may still contain errors and omissions. Therefore, after determining each atomic task unit, the server can also perform a similarity judgment on the elements that make up the sentence structure contained therein to determine whether the same concept is expressed in different texts. If synonyms are determined to exist and, based on the semantic analysis of the steps, it is determined that the synonyms do express the same concept, these synonyms are unified with a single text expression.

[0026] In the embodiments of this specification, when determining the steps included in the business logic process, the server can obtain the logical relationship between the steps through the aforementioned semantic analysis model. However, this is only the case when there is a description of the relationship between the steps in the business logic process, such as a description of conjunctions, step numbers, etc. However, the level of textual description of the business logic process is difficult to guarantee, and there may be situations where the logical relationship between the steps cannot be determined through semantic analysis. In this case, the logical relationship between the steps (which is also an atomic task unit) can be determined through a pre-configured composite task template. The specific method can be referred to in the subsequent content.

[0027] It should be noted that while the embodiments of this specification distinguish between grammatical analysis and semantic analysis, in mature natural language solutions, these two are often confused, that is, grammatical analysis and semantic analysis are performed together without distinction. Therefore, the semantic analysis model or grammatical analysis in the embodiments of this specification can be implemented using the same technical means to achieve the desired effects at different stages, and therefore, the subsequent description will not make a specific distinction between them.

[0028] S104: Re-edit the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process.

[0029] After step S102, the server determines each atomic task unit and the logical relationship between the atomic task units. In order to avoid the influence of possible description logic confusion or description ambiguity in the original description of the business logic process, the server can re-edit the business logic process and generate a secondary expression text of the business logic process based on the determined logical relationship and the atomic task unit as the re-editing object.

[0030] Because different business logic processes may use different methods to describe the same logical relationship, even if the business logic process describes the logical relationship correctly, subsequent steps still require a certain amount of effort to parse and determine the logical relationship. To further reduce the difficulty of logical relationship parsing, the representation of logical relationships is unified when generating secondary expression text. This unified expression of logical relationships reduces the difficulty of LLM processing in subsequent steps.

[0031] Specifically, the server first obtains preset logical symbols for expressing logical relationships, for example, using propositional logical symbols to express logical relationships of implication, equivalence, and negation.

[0032] Afterwards, the logical symbols corresponding to the logical relationships between the atomic task units are determined based on the preset logical symbols expressing the logical relationships. Each logical symbol corresponds to a description of the logical relationship, and the logical symbols corresponding to the logical relationships between the atomic task units are determined based on the description.

[0033] Finally, using the identified logical symbols as connectors between atomic task units, the business logic process is re-edited to create a secondary representation of the business logic process. For example, consider three atomic task units that originally had a judgment relationship. Step A (corresponding to atomic task unit A) is executed for judgment. If the judgment is true, step B (corresponding to atomic task unit B) is executed. If not, step C (corresponding to atomic task unit C) is executed. The resulting secondary representation is "Atomic task unit A → Atomic task unit B; ¬Atomic task unit A → Atomic task unit C." Here, "→" indicates the sequential execution of the atomic task units, and "¬" indicates a negative judgment result.

[0034] By obtaining the secondary expression text, no matter what kind of business logic process is obtained in step S100, whether it is a high-quality text or a low-quality text, whether the logic is clear, or whether there is useless description content in the middle, it can all be unified into a formatted and standardized expression text through the operations of the above steps.

[0035] S106: Generate a knowledge high-level program for executing the business logic process based on the secondary expression text through a large language model (LLM).

[0036] In the embodiment of the present specification, after obtaining the secondary expression text, the secondary expression text can be input into the trained LLM, so that the LLM generates a knowledge high-order program (HOP) for executing the business logic process.

[0037] Specifically, the server first inputs the secondary expression text into the pre-trained large language model LLM; then, based on the logical relationship contained in the secondary expression text, it generates a logic code for expressing the logical relationship between each step, and for each atomic task unit, generates a functional code for executing the atomic task unit; finally, based on the logical code and the functional code, it generates a knowledge high-level program for executing the business logic process.

[0038] based on Figure 1 The code generation method based on the large model shown in the figure performs grammatical analysis on the business logic process described in natural language to obtain the atomic task units and the logical relationship between the atomic task units. It is then re-edited according to the logical relationship to obtain a secondary expression text, and the code is generated through LLM. Avoid the confusion of logical relationships caused by natural language descriptions for business logic processes, which leads to problems when generating code through LLM later. Among them, the simplest expression corresponding to the natural language can be determined in the grammatical analysis process, reducing the ambiguity in the description of the steps to be executed in the business logic process, and re-editing to obtain a secondary expression text can avoid the influence of logical confusion that may occur in the business logic process, so that no matter how high or low the business logic process is written, it will not affect the standardization of the input LLM text, thereby improving the accuracy and reliability of the generated code.

[0039] In addition, in the embodiment of the present specification, when executing step S102, the server may also split the business logic process based on a preset composite task template to determine the atomic task units and the logical relationships between the atomic task units.

[0040] Specifically, the server can match the business logic process with each preset compound task template respectively, and determine the compound task template that matches the business logic process. Among them, the compound task template is a natural language expression template, and the compound task template that matches the business logic process can be determined by the simplest similarity calculation and other methods. In the embodiment of this specification, the so-called compounding of the compound task template refers to a template that combines different sub-business logic processes and executes them with a certain logic. Therefore, the compound task template naturally contains the logical relationship between the execution of the sub-business logic processes.

[0041] Therefore, the business logic process can be split according to the matching composite task template to determine several sub-business logic processes.

[0042] A sub-business logic process can be a single atomic task unit or composed of multiple atomic task units, forming a nested relationship. For example, a business logic process may include three sub-business logic processes, one of which contains multiple atomic task units. Sub-business logic processes containing multiple atomic task units can be further split using composite task templates until all the sub-processes are atomic task units.

[0043] Finally, the resulting sub-business logic process is parsed to obtain the atomic task units corresponding to the business logic process and the logical relationships between them. Based on the logical relationships between the sub-business logic processes during the aforementioned split and the logical relationships between the atomic task units in each sub-business logic process, all atomic task units corresponding to the business logic process and the logical relationships between them are determined.

[0044] In addition, when matching a composite task template, the server may also first use the splitting method described in step S102 to split out each atomic task unit. Then, similarity calculation is performed between the task descriptions contained in the composite task template and the split atomic task units. The composite task template with the highest similarity is determined and matched with the business logic process. Of course, the similarity calculation requires taking the task descriptions contained in the composite task template as a whole and the split atomic task units as a whole, and performing similarity calculation between the two wholes.

[0045] If the step description is relatively brief, the atomic task units identified through analysis still struggle to generate accurate code directly using LLM. For example, the steps include "take out a pen from a drawer," but implementing these steps also requires "walking to the drawer," "opening the drawer," "searching for a pen in the drawer," and "taking out a pen." This process can be further broken down using composite task templates until the smallest divisible atomic task unit is reached.

[0046] Furthermore, as mentioned in step S102, the level of textual description of the business logic process is difficult to guarantee, and semantic analysis may not be able to determine the logical relationship between steps. In this case, the server can use a composite task template to determine the logical relationship between each step (also an atomic task unit). For specific methods, please refer to the aforementioned process.

[0047] In the embodiment of this specification, the process of generating HOP in step S106 can be as follows: Figure 2 shown.

[0048] Figure 2 The flow chart of the HOP generation method provided in the embodiment of this specification specifically includes the following steps: S200: Obtaining a secondary expression text of a business logic process described in natural language.

[0049] In the embodiment of this specification, the server uses the regenerated secondary expression text of the business logic process as the input of the LLM. The embodiment of this specification below only takes the above-mentioned business as the risk control business as an example to illustrate the code generation method provided in this specification.

[0050] When the above-mentioned business is a risk control business based on the recorded risk log, the business logic process of the risk control business (that is, the risk control logic process) can be described in the following natural language: The first step is to analyze the cmd field to determine whether the file is granted executable permissions. If the cmd command line does not grant the file executable permissions, the conclusion is that no action is required. The second step is to determine whether the installation package is a commonly used software installation package. If it is a commonly used installation software, the conclusion is that no action is required; The third step is to determine whether the cmd field has been manually analyzed by security operations personnel within N days and determined to require no action. For example, if the chmod command line is determined to require no action, then no action is required for this risk log. The fourth step is to determine whether the risk log does not need to be handled or generates an attack alarm based on the results of the above three steps.

[0051] Converted to quadratic expression: Analyze the cmd field → The cmd field does not grant the file executable permission → No action is required; The cmd field does not grant executable permissions to the file. → The installation package is for commonly used software. → No action is required. The installation package is not a commonly used software installation package. The cmd field has been manually analyzed by security operations personnel within N days and determined to require no action. This risk log does not require any action. The ¬cmd field has been manually analyzed by security operations personnel within N days and determined to require no action. → An attack alarm is generated.

[0052] The risk log is a risk log of other businesses recorded by the user when performing other businesses, and the risk control business is a business that requires risk control of the other businesses based on the risk log.

[0053] S202: Input the secondary expression text into the pre-trained large language model LLM.

[0054] After the server obtains the secondary representation text, it can input the natural language secondary representation text into a pre-trained LLM. This LLM can be pre-deployed on the server, or it can be deployed on other devices. If deployed on other devices, the server can send the business logic process to the other device, allowing the other device to input the received business logic process into the LLM. The following description uses the LLM deployed on the server as an example.

[0055] In order to avoid the LLM from generating "hallucinations" and code that is inconsistent with the actual business logic process, and to improve the accuracy and reliability of the LLM generated code, in the embodiment of this specification, the server directly outputs the secondary expression text obtained in step S200 to the LLM, so that the LLM generates code for executing the business based on the secondary expression text.

[0056] S204: Identify the atomic task units contained in the secondary expression text and the logical relationships between the atomic task units through the LLM.

[0057] Generally, a business logic process consists of several steps arranged according to certain logical relationships. For example, in the example above, the first step is executed first. If the first step evaluates to negative, the output is output directly without any action required. If the evaluation result is positive, the second step is executed. If the second step evaluates to positive, the output is output directly without any action required. Otherwise, the third step is executed. If the third step evaluates to positive, the output is output directly without any action required. Otherwise, an attack alarm is generated. Therefore, the logical relationships between the different steps in a business logic process constitute the basic framework of the process. By constructing secondary expression text, atomic task units are used to replace the step content, reducing the ambiguity of the step content. By using logical symbols to unify the logical relationships between atomic task units, LLM can directly identify logical symbols and determine the logical relationships between steps.

[0058] S206: Based on the logical relationship, generate a computer program code of a preset type for representing the logical relationship between the atomic task units as a logical code.

[0059] The server may generate a preset type of computer program code for representing the logical relationship between the atomic task units as the logic code based on the logic symbols between the atomic task units identified in step S204 through the LLM.

[0060] The preset type of computer program code may be any type of computer program code, including but not limited to Python code. The following description will only take Python code as an example.

[0061] S208: For each atomic task unit, generate at least one operator for executing the step as a function code based on the natural language used to describe the atomic task unit in the secondary expression text and the knowledge base corresponding to the preset business logic process.

[0062] In the embodiment of this specification, the execution order of step S206 and step S208 may not be particular. Specifically, the server may execute step S206 through the LLM while also executing step S208 through the LLM.

[0063] As for the knowledge required by each atomic task unit in the business logic process, natural language is more natural and more understandable than computer programming language. Therefore, in the embodiments of this specification, an atomic task unit is represented by an operator defined in natural language.

[0064] The computational logic of the operators described in the embodiments of this specification is defined in natural language, but the operator names of these operators still conform to the function name format requirements of the aforementioned predefined type of computer program code. Therefore, the operators defined in natural language in this specification are essentially custom functions within the predefined type of computer program code, which are high-level abstract expressions of the logic that implements the atomic task unit.

[0065] Specifically, LLM can determine the knowledge required to execute each atomic task unit identified from the business logic process, such as one or several entities in the knowledge graph, one or several attributes of the entity, and the relationship between the entity and other entities, and / or the calling interface of the tool required by the atomic task unit, based on the semantics of the natural language used to describe the atomic task unit in the identified business logic process and the knowledge base corresponding to the above-mentioned business logic process. After determining the knowledge required to execute the atomic task unit, at least one operator for executing the atomic task unit can be generated based on it. The operator is defined in natural language, and the operator name meets the function name format requirements of the above-mentioned preset type of computer program code.

[0066] Continuing with the previous example, the first step is to determine whether the cmd field grants the file executable permission. Therefore, LLM can generate an operator name hop_judge that meets the Python code function name requirements and define the operator hop_judge as follows: def hop_judge Determine whether {cmd} grants file executable permissions As can be seen, the definition of the operator hop_judge is completely defined using natural language, "Judge whether {cmd} grants file executable permissions," making the operator more understandable. The operators whose computational logic is defined in natural language are special operators in HOP, serving as the functional code for this atomic task unit.

[0067] The above is just an example of expressing an atomic task unit by one operator. Of course, after the splitting in step S104, the atomic task unit is generally expressed by only one operator. However, those skilled in the art should understand that in actual applications, it is not ruled out that for more complex atomic task units, more operators (such as more than two operators) can be used to express the atomic task unit.

[0068] S210: Generate a knowledge high-level program for executing the business logic process based on the logic code and the function code generated for each atomic task unit.

[0069] As explained above, the secondary expression text can be regarded as consisting of the logical relationship between multiple atomic task units and the knowledge required to execute each atomic task unit. The logical relationship has been expressed by the generated logic code in step S206, and the knowledge required to execute each step has also been expressed by the operator defined by natural language in step S208, and the operator name still conforms to the preset type of computer program code function name format requirements. Therefore, LLM can combine the logic code and the functional code (that is, the operator) corresponding to each step to generate a knowledge high-level program (HOP) for executing the business logic process.

[0070] Continuing with the above example, LLM combines the operators corresponding to each step into the framework code to generate the following HOP for executing the above risk control logic process: def evaluate(cmd): if hop_judge_a("Judge (cmd) whether to grant file executable permissions") #High-level judgment pkg = hop_get("Get the installation package used by (cmd)") if not pkg in hop_knowledge_retrieve("Common Installation Package") #High-level knowledge concept matching if not hop_judge_b("Check the history to see if (cmd) has been manually judged as not requiring action") return "Attack Alarm" return "No need to dispose" In the HOP above, the content enclosed by ("") is the natural language definition of the corresponding operator. These operators include hop_judge_a, hop_get, hop_knowledge_retrieve, and hop_judge_b. The framework code expressed in Python above shows the entire risk control logic process as follows: first, execute the hop_judge_a operator to determine whether (cmd) has file executable permissions. If so, execute the hop_get operator, obtain the installation package used by (cmd), and assign it to pkg. If not, directly output "No action required." After obtaining the installation package used by (cmd) and assigning it to pkg, execute the hop_knowledge_retrieve operator to obtain the commonly used installation package and determine whether pkg is not among the commonly used installation packages. If so, execute the hop_judge_b operator. If not, directly output "No action required." Execute the hop_judge_b operator to determine whether (cmd) has not been manually determined to require action in the past. If so, output "Attack Alarm"; otherwise, output "No action required."

[0071] It can be seen that the HOP expressed in the Python code is completely consistent with the original intention of the risk control logic process described in natural language. Therefore, through the HOP, even if there is a change in the personnel in the department performing the risk control business, when the HOP is read by other personnel (other than the personnel who wrote the original risk control logic process), the logical confusion of each step caused by the ambiguity and ambiguity of natural language itself can be avoided.

[0072] In addition, in the embodiment of this specification, after obtaining the HOP, the server can also verify the HOP and execute the HOP after passing the verification. Specifically, the correctness of the HOP code level of the atomic task unit can be verified. Or the correctness of the HOP logic between atomic tasks can be verified. Or other LLMs can be used. Figure 2 The HOP is regenerated through the process and the two HOPs are cross-checked. Alternatively, the HOP is verified in principle, specifically identifying code that does not conform to the principle based on the knowledge base. For example, code that is clearly inconsistent with the actual situation, such as "obtaining user accounts over 200 years old," is verified.

[0073] The above is a code generation method based on a large model provided in an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media and electronic devices.

[0074] Figure 3 A schematic diagram of a code generation device based on a large model provided in an embodiment of this specification, the device comprising: An acquisition module 301 is used to acquire a business logic process described in a natural language; A parsing module 302 is configured to perform syntax analysis on the business logic process to generate atomic task units corresponding to the business logic process and logical relationships between the atomic task units, wherein the atomic task units are described in natural language; An editing module 303 is configured to re-edit the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process; The generation module 304 is used to generate a knowledge high-level program for executing the business logic process based on the secondary expression text through a pre-trained large language model LLM.

[0075] Optionally, the parsing module 302 is used to match the business logic process with each preset composite task template respectively to determine the composite task template that matches the business logic process; split the business logic process according to the matched composite task template to determine several sub-business logic processes; perform grammatical analysis on the obtained sub-business logic processes to obtain the atomic task units corresponding to the business logic process, as well as the logical relationship between the atomic task units.

[0076] Optionally, the parsing module 302 is used to perform grammatical analysis on each sub-business logic process to determine the atomic task units corresponding to the sub-business logic process and the logical relationship between the atomic task units; determine the logical relationship between the sub-business logic processes based on the composite task template; determine the logical relationship between the atomic task units obtained by processing the business logic process based on the logical relationship between the atomic task units contained in the sub-business logic process and the logical relationship between the sub-business logic processes.

[0077] Optionally, the parsing module 302 is used to determine the steps included in the business logic process, where the steps are texts described in natural language; for each step, the step is subjected to grammatical analysis to determine the sentence structure of the step and the elements that constitute the sentence structure; and the elements that constitute the sentence structure of the step are reconstructed according to a preset sentence format as atomic task units.

[0078] Optionally, the editing module 303 is used to determine the logical symbols corresponding to the logical relationships between the atomic task units based on preset logical symbols for expressing logical relationships; and using the determined logical symbols as connectors between the atomic task units, re-edit the atomic task units included in the business logic process to obtain a secondary expression text of the business logic process.

[0079] Optionally, the generation module 304 is used to input the secondary expression text into a pre-trained large language model LLM; generate a logic code for representing the logical relationship between the atomic task units based on the logical relationship contained in the secondary expression text, and generate a function code for executing the atomic task unit for each atomic task unit; and generate a knowledge high-level program for executing the business logic process based on the logic code and the function code.

[0080] Optionally, the business logic process includes a risk control logic process; the high-level knowledge program is used to execute risk control business.

[0081] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the large model-based code generation method provided above.

[0082] based on Figure 1 The code generation method based on the large model shown in the embodiment of this specification also provides Figure 4 The structural diagram of the electronic device shown in FIG. Figure 4 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it, implementing the aforementioned large model-based code generation method.

[0083] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A code generation method based on a large model, the method comprising: Obtain business logic processes described in natural language; Performing grammatical analysis on the business logic process to generate atomic task units corresponding to the business logic process and logical relationships between the atomic task units, wherein the atomic task units are described in natural language; Re-editing the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process; Based on the secondary expression text, a knowledge high-level program for executing the business logic process is generated through a pre-trained large language model LLM.

2. The method according to claim 1, further comprising: performing grammatical analysis on the business logic process to generate atomic task units corresponding to the business logic process and logical relationships between the atomic task units, specifically comprising: Matching the business logic process with each preset composite task template respectively, and determining the composite task template that matches the business logic process; According to the matched composite task template, the business logic process is split to determine a number of sub-business logic processes; The obtained sub-business logic process is subjected to syntax analysis to obtain the atomic task units corresponding to the business logic process and the logical relationships between the atomic task units.

3. The method according to claim 2, further comprising: performing grammatical analysis on the obtained sub-business logic process to obtain the atomic task units corresponding to the business logic process and the logical relationships between the atomic task units, specifically comprising: For each sub-business logic process, perform syntax analysis on the sub-business logic process to determine the atomic task units corresponding to the sub-business logic process and the logical relationship between the atomic task units; Determining the logical relationship between the sub-business logic processes according to the composite task template; According to the logical relationship between the atomic task units included in the sub-business logic process and the logical relationship between the sub-business logic processes, the logical relationship between the atomic task units obtained by processing the business logic process is determined.

4. The method according to claim 1, wherein the business logic process is parsed to generate atomic task units corresponding to the business logic process, specifically comprising: Determine each step included in the business logic process, wherein the steps are texts described in natural language; For each step, perform grammatical analysis on the step to determine the sentence structure of the step and the elements that constitute the sentence structure; According to the preset sentence format, the elements that make up the sentence structure of this step are reconstructed as atomic task units.

5. The method according to claim 1, wherein the generated atomic task units are re-edited according to the logical relationship to obtain a secondary expression text of the business logic process, specifically comprising: Determine the logical symbols corresponding to the logical relationships between the atomic task units based on the preset logical symbols expressing the logical relationships; The determined logical symbols are used as connectors between the atomic task units, and the atomic task units included in the business logic process are re-edited to obtain a secondary expression text of the business logic process.

6. The method according to claim 1, wherein generating a knowledge high-level program for executing the business logic process based on the secondary expression text through a large language model (LLM) specifically comprises: Inputting the secondary expression text into a pre-trained large language model LLM; Generate logic codes for representing the logic relationships between the atomic task units according to the logic relationships contained in the secondary expression text, and generate function codes for executing the atomic task units for each atomic task unit; A knowledge high-level program for executing the business logic process is generated according to the logic code and the function code.

7. The method according to claim 1, wherein the business logic process includes a risk control logic process; and the high-level knowledge program is used to execute risk control business.

8. A code generation device based on a large model, the device comprising: An acquisition module is used to acquire the business logic process described in natural language; A parsing module, configured to perform syntax analysis on the business logic process, generate atomic task units corresponding to the business logic process, and logical relationships between the atomic task units, wherein the atomic task units are described in natural language; An editing module, configured to re-edit the generated atomic task units according to the logical relationship to obtain a secondary expression text of the business logic process; A generation module is used to generate a knowledge high-level program for executing the business logic process based on the secondary expression text through a pre-trained large language model LLM.

9. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

Citation Information

Patent Citations

  • Method and device for establishing logic expression by using large language model and query system

    CN117076607A

  • Combined event logic extraction method and related device

    CN117648395A

  • Model-based code generation method and device, storage medium and electronic equipment

    CN117873453A

  • Structured common sense reasoning method based on code language

    CN118396109A

  • Visual intelligent programming method based on large language model

    CN119045806A

Cited By

  • Code generation method and system based on large model

    CN120763072A

  • Evaluation model training method, data processing method and device

    CN121094052A

  • Business program code generation method and device, medium and electronic equipment

    CN121614122A