Method, device and medium for automatically generating program code driven by Chinese language

By designing abstract grammar rulesets and small-scale neural network models, sequence generation and abstract grammar tree analysis of Chinese pseudo-code, combined with human-computer dialogue technology, the problems of high code generation cost and difficult to guarantee grammar accuracy in the existing technology are solved, and efficient and accurate automatic code generation is achieved.

CN117971178BActive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410011848.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-05-06
Estimated Expiration
2044-01-02

AI Technical Summary

Technical Problem

In the prior art, Chinese language-driven program code automatic generation tasks rely on large-scale training data and model parameters, the training cost is high, and the syntax accuracy of the generated code cannot be guaranteed, it requires manual evaluation, which is time-consuming and labor-consuming.

Method used

Design a programming language-independent abstract grammar rule set, sequence generation of Chinese pseudocode through a small-scale programming-decoder neural network model, generate abstract grammar rule sequences, and analyze word element legality, subtree structure legality and overall legality through an abstract grammar tree, and complete information completion with human-computer dialogue technology to generate high-quality code.

Benefits of technology

Implementing program-level code generation on small-scale models and data sets reduces training and usage costs, ensures the lexical, syntactic and semantic correctness of generated code, simplifies code quality evaluation, and improves development efficiency and ease of use of code maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117971178B_ABST
    Figure CN117971178B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and medium for automatic generation of program code driven by Chinese language, and relates to the code generation technology of computer software engineering. The method includes: designing an abstract grammar rule set that is independent of programming language; performing sequence generation based on the abstract grammar rule set on the input Chinese pseudocode to obtain an abstract grammar rule sequence; for the abstract grammar rule sequence, expanding the tree structure according to the node category in the sequence, defining the subtree range, and abstracting the rule sequence into an abstract syntax tree; for the abstract syntax tree, performing word unit legitimacy analysis, subtree structure legitimacy analysis and abstract syntax tree overall legitimacy analysis, and generating a quality assessment result based on the abstract grammar rule set; according to the quality assessment result, completing information completion in combination with human-computer dialogue technology; when it is detected that the quality assessment result is qualified, generating the final high-quality code. The present invention can realize the automatic generation of high-quality Chinese language-driven program code with small-scale resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to code generation technology for computer software engineering, and in particular to a method, device and medium for automatically generating program codes driven by a Chinese language. Background Art

[0002] With the rapid development of computer technology and the iterative update of software and hardware technology, the way and language of writing programs that can be understood and executed by computers are constantly changing and simplifying, but the code writing task is still one of the necessary tasks for developers. The pseudocode generation code task usually appears as the last step of program development. Developers manually write corresponding codes based on pseudocodes. There are a lot of repeated writing of simple codes and codes with the same functions, as well as the repeated debugging and solving the defects of the written codes. These tasks greatly reduce the development efficiency and increase the maintenance cost of the code in the later stage. The pseudocode automatic code generation method effectively solves the problem of repeated code writing and maintenance, and effectively reduces the learning threshold from mastering computational thinking to realizing specific program code writing. With the large-scale application of deep learning models, especially large-scale pre-trained models, in the pseudocode automatic code generation task, the pseudocode automatic code generation task has developed rapidly.

[0003] Currently, for the task of automatic program code generation driven by the Chinese language, previous work has designed many large-scale pre-training models to learn the intrinsic grammatical structure information and logical structure information of the code language. These large-scale pre-training models are extremely dependent on large amounts of training data and model parameters, and the training cost is extremely high. At the same time, the existing code generation methods cannot guarantee the grammatical correctness of the generated code, and the actual value of the generated code needs to be manually evaluated, which is time-consuming and labor-intensive. Summary of the invention

[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the object of the present invention is to provide a method, device and medium for automatically generating program code driven by Chinese language.

[0005] The technical solution adopted by the present invention is:

[0006] A method for automatically generating program code driven by a Chinese language comprises the following steps:

[0007] S1. Design an abstract grammar rule set that is independent of programming languages ​​to capture the minimum data information required for code construction; generate a sequence based on the abstract grammar rule set for the input Chinese pseudocode to obtain an abstract grammar rule sequence corresponding to the Chinese pseudocode sample;

[0008] S2, for the abstract syntax rule sequence obtained in step S1, expand the tree structure according to the node categories in the sequence, define the subtree range, and abstract the rule sequence into an abstract syntax tree;

[0009] S3, for the abstract syntax tree obtained in step S2, perform word unit legitimacy analysis, subtree structure legitimacy analysis and abstract syntax tree overall legitimacy analysis to generate a quality assessment result based on the abstract syntax rule set;

[0010] S4. Complete the information completion based on the quality assessment results based on the abstract grammar rule set and combined with human-computer dialogue technology;

[0011] S5. When it is detected that the quality assessment result is qualified, the final high-quality code is generated.

[0012] Furthermore, in step S1, the designing of a programming language-independent abstract grammar rule set includes:

[0013] By designing the rule set to eliminate redundant information recognition and capture only the minimum data information required for code construction, the corresponding generated abstract syntax tree can generate code in multiple programming languages. At the same time, specific rules are designed for different types of basic code statements to meet the corresponding constraints, reducing the number of categories when the model makes multiple classifications of a single node, and providing a basis for structural disassembly and error detection of the quality assessment model.

[0014] Furthermore, in step S1, the input Chinese pseudocode is sequenced based on an abstract grammar rule set to obtain an abstract grammar rule sequence corresponding to the Chinese pseudocode sample, including:

[0015] A small-scale encoder-decoder neural network model is used to generate a sequence of the input Chinese pseudocode based on the abstract grammar rule set, and an abstract grammar rule sequence corresponding to the Chinese pseudocode sample is obtained; the abstract grammar rule sequence includes rule nodes, terminal nodes and word nodes.

[0016] Furthermore, the step S2 comprises:

[0017] By making category judgment and position recognition on the rule nodes, terminal nodes and word-unit nodes of the abstract grammar rule sequence generated in step S1, the sequence is expanded and abstracted into an abstract syntax tree, and the subtree range and margin are divided to prepare for the subsequent abstract syntax tree structure analysis.

[0018] Furthermore, the step S3 comprises:

[0019] The abstract syntax tree is subjected to word-unit validity analysis, subtree structure validity analysis and abstract syntax tree overall validity analysis. Among them, word-unit validity analysis is used to determine whether the word at a specific position satisfies the constraints or context conditions assigned by its subtree or parent node. Subtree structure validity analysis is used to determine whether the node type of each node of the subtree and the number of node branches meet the constraints. The abstract syntax tree overall validity analysis is used to determine whether the abstract syntax tree is legal based on the overall logical structure and context information of the basic code statements.

[0020] Furthermore, the step S4 comprises:

[0021] The abstract grammar rule set, quality assessment model and human-computer dialogue are combined to complete the missing or ambiguous information in the Chinese pseudocode description through human-computer dialogue. Depending on the type of error, quality re-evaluation or sequence re-prediction methods are used to form an information completion and interaction method for high-quality code generation.

[0022] Furthermore, according to different error types, the quality re-evaluation method or sequence re-prediction method is respectively adopted to form an information completion and interaction method for high-quality code generation, including:

[0023] If a missing word or illegal error occurs, the human-computer dialogue interacts with the user and supplements the word information to the abstract syntax tree, and then returns to step S3; if a structural error or illegal rule node error occurs, the human-computer dialogue interacts with the user and supplements the structural information, and then returns to steps S1-S3.

[0024] Furthermore, in step S2, the tree structure is expanded according to the node categories in the sequence, the subtree range is defined, and the rule sequence is abstracted into an abstract syntax tree, including:

[0025] Use a quality assessment model based on an abstract syntax rule set to expand the tree structure and define the subtree range according to the node categories in the sequence, and abstract the rule sequence into an abstract syntax tree;

[0026] In step S3, the legitimacy analysis of the word unit, the legitimacy analysis of the subtree structure and the overall legitimacy analysis of the abstract syntax tree include:

[0027] A quality assessment model based on an abstract grammar rule set is used to perform word legitimacy analysis, subtree structure legitimacy analysis, and abstract syntax tree overall legitimacy analysis.

[0028] Another technical solution adopted by the present invention is:

[0029] A Chinese language driven program code automatic generation device, comprising:

[0030] at least one processor;

[0031] at least one memory for storing at least one program;

[0032] When the at least one program is executed by the at least one processor, the at least one processor implements the method described above.

[0033] Another technical solution adopted by the present invention is:

[0034] A computer-readable storage medium stores a program executable by a processor, wherein the program executable by the processor is used to execute the method described above when executed by the processor.

[0035] Compared with the prior art, the present invention has the following advantages and technical effects:

[0036] (1) The method of the present invention realizes program-level code generation on a small-scale model and data set, greatly reducing the training cost and usage cost; at the same time, it ensures the lexical, grammatical and semantic correctness of the generated code, without spending a lot of manpower and material resources on code quality evaluation.

[0037] (2) The method of the present invention has simple steps and is easy to implement. The user only needs to input pseudocode according to the needs and obtain the required code according to the human-computer dialogue instructions. The user can operate and use it without understanding the internal principles. It has extremely strong ease of use and has a broad application space. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0039] Figure 1 It is a flowchart of the steps of a method for automatically generating program code driven by a Chinese language according to an embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of some abstract grammar rules of an embodiment of the present invention;

[0041] Figure 3 It is a schematic diagram of an abstract grammar rule sequence and a grammar tree corresponding to the pseudo code “create an integer flag with a value of 0” in an embodiment of the present invention;

[0042] Figure 4 It is a schematic diagram of a Chinese pseudo code-code example of an embodiment of the present invention;

[0043] Figure 5 The present invention is a flowchart of a method for automatically generating program codes driven by the Chinese language. DETAILED DESCRIPTION

[0044] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0045] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0046] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0047] In addition, in the description of the present invention, unless otherwise specified, "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0048] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0049] In response to the existing technical problems, the present invention focuses on structured modeling of the abstract grammar rules required for code generation, establishes an abstract grammar rule set as the prediction basis of the neural network model, generates an abstract grammar rule sequence corresponding to the code, and combines the quality assessment model to perform lexical and grammatical correctness detection. In this way, the program-level target code can be generated and evaluated on a small-scale model, thereby improving the model prediction performance and ensuring the lexical and grammatical correctness of the generated target code, thereby improving the accuracy of automatic generation of Chinese language-driven program code.

[0050] like Figure 1 As shown, this embodiment provides a method for automatically generating program code driven by a Chinese language, comprising the following steps:

[0051] S1. Design an abstract grammar rule set that is independent of programming languages; generate a sequence based on the abstract grammar rule set for the input Chinese pseudocode to obtain an abstract grammar rule sequence corresponding to the Chinese pseudocode sample.

[0052] In this embodiment, the abstract grammar rule set corresponding to the design code is first converted into an abstract grammar rule sequence based on the image grammar rule set through a neural network model in the subsequent processing of the pseudo code, so that the quality assessment model can perform structural decomposition and error detection, and then converted into the final target code after passing the quality assessment. The rule design of the abstract grammar rule set and the inherent constraint relationship between the rules serve as the basis for the label predicted by the neural network model and the analysis of the quality assessment model.

[0053] S2. For the obtained abstract syntax rule sequence, expand the tree structure according to the node categories in the sequence, define the subtree range, and abstract the rule sequence into an abstract syntax tree.

[0054] As an optional implementation, a small-scale encoder-decoder neural network model is used for sequence generation to obtain an abstract grammar rule sequence, which includes rule nodes, termination nodes and word-unit nodes. The rule nodes are non-leaf nodes of the abstract syntax tree, representing the structural information in the code; the word-unit nodes are leaf nodes of the abstract syntax tree, representing specific code snippets at specific locations; the termination nodes appear in the leaf nodes of the abstract syntax tree, and as the terminator of the abstract syntax tree, they represent the termination of a subtree structure or a nested structure.

[0055] S3. Perform word unit validity analysis, subtree structure validity analysis and abstract syntax tree overall validity analysis on the obtained abstract syntax tree to generate a quality assessment result based on the abstract syntax rule set.

[0056] Specifically, the quality assessment model based on the image grammar rule set identifies and extracts the tree structure information in the sequence according to the node features in the abstract grammar rule sequence, and performs analysis and inspection step by step to generate the final assessment result.

[0057] S4. Complete the information completion based on the quality assessment results based on the abstract grammar rule set and combined with human-computer dialogue technology.

[0058] Specifically, when there is an error in the quality assessment result generated in step S3, the corresponding operation is performed according to the error type: if there is a missing or illegal word, the human-computer dialogue communicates with the user, supplements the word information, and then re-evaluates the quality; if a structural error or an illegal rule node error occurs, such as when the user does not provide the sentence structure information of the current pseudocode, the human-computer dialogue communicates with the user, clearly supplements the current sentence structure information, and then re-predicts the abstract grammar rule sequence.

[0059] S5. When it is detected that the quality assessment result is qualified, the final high-quality code is generated.

[0060] The above method is explained in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 5 As shown, this embodiment provides a method for automatically generating program code driven by a Chinese language, comprising the following steps:

[0062] The first step is to design the abstract grammar rule set by yourself.

[0063] like Figure 2 As shown, each rule is expressed in the form of an stmt expression. The left side of the "=" symbol represents the starting node or the parent node, and the right side of the "=" symbol represents the rule form of the currently selected specific node and the corresponding subtree members. The generation is completed by continuously selecting and expanding the subtree structure until all subtrees reach the word node or the terminal node.

[0064] In the second step, based on the abstract grammar rule set, a small-scale encoder-decoder neural network model is used to generate sequences for the input Chinese pseudocode to obtain the abstract grammar rule sequence corresponding to the Chinese pseudocode sample. The pseudocode "create an integer flag with a value of 0" is generated by the neural network model corresponding to the abstract grammar rule sequence and abstract syntax tree structure, as shown in Figure 3 As shown, in the prediction process, the expandable nodes are processed in a dynamic stack. When the node is selected, the node is pushed into the stack. When all subtrees under the node reach the word node or the terminal node, the node is popped out of the stack. When the dynamic stack is empty, it means that the generation is completed.

[0065] In the third step, the quality assessment model based on the abstract syntax rule set abstracts the rule sequence into an abstract syntax tree, and performs word unit validity analysis, subtree structure validity analysis, and abstract syntax tree overall validity analysis in sequence. Word unit validity analysis focuses on whether the word unit or keyword information is legal; subtree structure validity analysis and abstract syntax tree overall validity analysis focus on whether it contains sufficient grammatical structure information. Figure 3 As shown, the word legitimacy analysis determines whether the type defined by the variable in the definition statement is a legal type in the type library, and whether the types of the child nodes "flag", "int", and "0" of each subtree tree are legally matched; the subtree structure legitimacy analysis and the overall legitimacy analysis of the abstract syntax tree determine whether the node selection in the rule node and the overall tree structure are legal.

[0066] The fourth step is to evaluate the quality of the model based on the abstract grammar rule set. If a word is missing or illegal, the human-computer dialogue will supplement the word information and re-evaluate the quality. If a structural error or illegal rule node error occurs, the human-computer dialogue will supplement the structural information and re-generate the sequence.

[0067] Complete Chinese pseudocode - code example Figure 4 As shown in the figure, the quality assessment result of the statement "define variable fast with value 2" is "variable type undefined", which is generated after the type description is supplemented through human-computer dialogue; the quality assessment result of the statement "nums is assigned to nums[fast]" is type mismatch, and the statement needs to be changed through human-computer dialogue to make the left value and the right value match the type.

[0068] From the example results, the target code corresponding to the Chinese pseudocode can be successfully generated by the method of this application. For other Chinese pseudocodes, it is only necessary to input the Chinese pseudocode description into the system and finally obtain the target code according to the human-computer dialogue instructions.

[0069] This embodiment also provides a Chinese language driven program code automatic generation device, comprising:

[0070] at least one processor;

[0071] at least one memory for storing at least one program;

[0072] When the at least one program is executed by the at least one processor, the at least one processor implements the following Figure 1 or Figure 5 The method shown.

[0073] A Chinese language driven program code automatic generation device of this embodiment can execute a Chinese language driven program code automatic generation method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0074] The present application also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes Figure 1 or Figure 5 The method shown.

[0075] This embodiment also provides a storage medium, which stores instructions or programs that can execute a Chinese language-driven program code automatic generation method provided by the method embodiment of the present invention. When the instructions or program are run, any combination of implementation steps of the method embodiment can be executed, and the corresponding functions and beneficial effects of the method can be obtained.

[0076] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0077] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0078] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0079] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0080] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0081] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0082] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0083] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0084] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for automatically generating program code driven by Chinese language, characterized in that: The following steps are involved: S1. Design an abstract grammar rule set that is independent of programming languages; generate a sequence based on the abstract grammar rule set for the input Chinese pseudocode to obtain an abstract grammar rule sequence corresponding to the Chinese pseudocode sample; S2. For the obtained abstract syntax rule sequence, expand the tree structure according to the node categories in the sequence, define the subtree range, and abstract the rule sequence into an abstract syntax tree; S3. Perform word-unit legitimacy analysis, subtree structure legitimacy analysis, and overall legitimacy analysis on the obtained abstract syntax tree to generate a quality assessment result based on the abstract syntax rule set; S4. Complete the information completion based on the quality assessment results based on the abstract grammar rule set and combined with human-computer dialogue technology; S5. When the quality assessment result is detected to be qualified, the final high-quality code is generated; In step S1, the design of a programming language-independent abstract grammar rule set includes: By eliminating redundant information recognition through the design of the rule set, only the minimum data information required for code construction is captured, and the corresponding generated abstract syntax tree can generate code in multiple programming languages; at the same time, specific rules are designed for different types of basic code statements to meet the corresponding constraints, reduce the number of categories when the model performs multiple classifications on a single node, and provide a basis for the structural disassembly and error detection of the quality assessment model; In step S1, the input Chinese pseudocode is sequenced based on an abstract grammar rule set to obtain an abstract grammar rule sequence corresponding to the Chinese pseudocode sample, including: A small-scale encoder-decoder neural network model is used to generate a sequence based on an abstract grammar rule set for the input Chinese pseudocode, and an abstract grammar rule sequence corresponding to the Chinese pseudocode sample is obtained; wherein the abstract grammar rule sequence includes rule nodes, termination nodes, and word-unit nodes; The step S2 comprises: By making category judgment and position recognition on the rule nodes, terminal nodes and word-element nodes of the abstract grammar rule sequence generated in step S1, the sequence is expanded and abstracted into an abstract syntax tree, and the subtree range and margin are divided to prepare for the subsequent abstract syntax tree structure analysis; The step S3 comprises: Perform word-unit validity analysis, subtree structure validity analysis and abstract syntax tree overall validity analysis on the abstract syntax tree; among them, word-unit validity analysis is used to determine whether the word at a specific position satisfies the constraints or context conditions assigned by its subtree or parent node; subtree structure validity analysis is used to determine whether the node type of each node in the subtree and the number of node branches meet the constraints; and the abstract syntax tree overall validity analysis is used to determine whether the abstract syntax tree is legal in terms of the overall logical structure and context information based on the basic code statements; The step S4 comprises: The abstract grammar rule set, quality assessment model and human-computer dialogue are combined to complete the missing or ambiguous information in the Chinese pseudocode description through human-computer dialogue. Depending on the type of error, quality re-evaluation or sequence re-prediction methods are used to form an information completion and interaction method for high-quality code generation.

2. The method for automatically generating program code driven by Chinese language according to claim 1, characterized in that: According to different error types, the quality re-evaluation method or sequence re-prediction method is adopted to form an information completion and interaction method for high-quality code generation, including: If a missing word or illegal error occurs, the human-computer dialogue interacts with the user and supplements the word information to the abstract syntax tree, and then returns to step S3; if a structural error or illegal rule node error occurs, the human-computer dialogue interacts with the user and supplements the structural information, and then returns to steps S1-S3.

3. The method for automatically generating program code driven by Chinese language according to claim 1, characterized in that: In step S2, the tree structure is expanded according to the node categories in the sequence, the subtree range is defined, and the rule sequence is abstracted into an abstract syntax tree, including: Use a quality assessment model based on an abstract syntax rule set to expand the tree structure and define the subtree range according to the node categories in the sequence, and abstract the rule sequence into an abstract syntax tree; In step S3, the legitimacy analysis of the word unit, the legitimacy analysis of the subtree structure and the overall legitimacy analysis of the abstract syntax tree include: A quality assessment model based on an abstract grammar rule set is used to perform word legitimacy analysis, subtree structure legitimacy analysis, and abstract syntax tree overall legitimacy analysis.

4. A Chinese language driven program code automatic generation device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 3.

5. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 3 when executed by the processor.

Citation Information

Patent Citations

  • Automatic programming method based on human-computer interaction

    CN115185497A

  • Code generation method and device and storage medium

    CN115878120A