Component code generation method and device, storage medium and processor

By converting natural language information into a programming language, building a component syntax tree and extracting key information, inputting it into a deep generation model to generate component code, solving the problem of high cost of generating component codes in a large language model, and achieving efficient and low-cost component code generation.

CN120010822APending Publication Date: 2025-05-16中国农业银行股份有限公司宁夏回族自治区分行
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510093244.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

Smart Images

  • Figure CN120010822A_ABST
    Figure CN120010822A_ABST
Patent Text Reader

Abstract

The invention discloses a component code generation method and device, a storage medium and a processor. According to the scheme, natural language information describing a to-be-generated component is obtained; converting the natural language information into a programming language of the to-be-generated component; constructing a component syntax tree based on the programming language of the to-be-generated component, and extracting key information in the component syntax tree; and inputting the key information into a depth generation model, performing prediction by the depth generation model according to the key information, and outputting a code of the to-be-generated component. Different from the mode of generating component codes based on a large language model in the prior art, the method has the advantages that the subsequent work of the depth generation model is accurately guided in the mode of constructing the component syntax tree and extracting key information of the component syntax tree, unnecessary exploration is avoided, the used depth generation model aims at a specific task of code generation, and the efficiency of code generation is improved. Related features are learned more efficiently, so that training time and resource consumption are reduced. Compared with the prior art that the cost is high, the method has obvious advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, storage medium and processor for generating component code. Background Art

[0002] In software development, a component is a modular unit that can be designed, implemented, and used independently. It is usually used to build user interfaces or perform specific functions, such as buttons, input boxes, and dialog boxes. Components break down complex applications into smaller and easier-to-understand fragments. This modular architecture improves the readability and maintainability of the code.

[0003] At present, the technical solutions for component code generation tasks are mainly based on large language models (such as the GPT series, BERT, etc.). The large language model can learn the grammatical structure, semantic information, and programming logic of the code through training on massive code data sets. In the code generation task, the large language model will automatically generate logically and semantically correct code based on the input natural language description or specific requirements. In order to improve the accuracy and efficiency of generation, some technical means will also be used, such as the introduction of contextual information, gradual generation and correction, and combination with external retrieval. Through these technical solutions, the large language model has achieved remarkable results in component code generation tasks, bringing more convenience and possibilities to software development.

[0004] However, there are drawbacks in generating component code based on large language models: large language models need to process a large amount of code data, which requires high-performance computing devices and large-scale storage space. At the same time, in order to obtain high-quality generation results, large language models need to be trained and adjusted for a long time, which also increases manpower and time costs. These high costs limit the promotion and application of code generation technology in some application scenarios with limited resources or tight budgets.

[0005] Faced with the high cost of generating component code through large language models, how to reduce the cost while ensuring the accuracy of the generated component code is a technical problem that needs to be solved urgently. Summary of the invention

[0006] Based on the above problems, the present application provides a method, device, storage medium and processor for generating component codes, with the aim of reducing costs while ensuring the accuracy of the generated component codes.

[0007] The embodiments of the present application disclose the following technical solutions:

[0008] The first aspect of the present application provides a method for generating component code, the method comprising:

[0009] Obtain natural language information describing the component to be generated;

[0010] Converting the natural language information into the programming language of the component to be generated;

[0011] Constructing a component syntax tree based on the programming language of the component to be generated, and extracting key information in the component syntax tree; the key information includes key nodes or key subtrees;

[0012] The key information is input into a deep generative model, and the deep generative model makes predictions based on the key information and outputs the code of the component to be generated.

[0013] Optionally, before inputting the key information into a deep generative model, and having the deep generative model predict according to the key information and outputting the code of the component to be generated, the method further includes:

[0014] Analyzing the key information using a saliency map algorithm;

[0015] The component syntax tree is adjusted based on the analysis result, and the key information is updated.

[0016] Optionally, after inputting the key information into a deep generative model, and the deep generative model predicts based on the key information and outputs the code of the component to be generated, the method further includes:

[0017] The self-attention mechanism is integrated with the component syntax tree to calculate the weight of each node or subtree in the component syntax tree;

[0018] Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

[0019] Optionally, constructing a component syntax tree based on the programming language of the component to be generated, and extracting key information in the component syntax tree includes:

[0020] Static code analysis technology is used to extract key syntax information contained in the component syntax tree; the key syntax information includes component interface information, dependency relationships between components, or internal logic information of components.

[0021] Optionally, before inputting the key information into a deep generative model, and having the deep generative model make predictions based on the key information and outputting the code of the component to be generated, the method further includes:

[0022] The key information is converted into a form that matches the processing form of the deep generative model; the form includes an embedded vector form or a path expression form.

[0023] A second aspect of the present application provides a device for generating a component code, the device comprising:

[0024] A natural language information acquisition module, used to acquire natural language information describing the component to be generated;

[0025] A natural language processing module, used for converting the natural language information into the programming language of the component to be generated;

[0026] A component syntax tree key information extraction module, used to construct a component syntax tree based on the programming language of the component to be generated, and extract key information in the component syntax tree; the key information includes key nodes or key subtrees;

[0027] The component code generation module is used to input the key information into the deep generation model, and the deep generation model predicts according to the key information and outputs the code of the component to be generated.

[0028] Optionally, the device further comprises: a key information verification module;

[0029] The key information verification module is used to analyze the key information using a saliency map algorithm;

[0030] The component syntax tree is adjusted based on the analysis result, and the key information is updated.

[0031] Optionally, the device further comprises: a component code checking module;

[0032] The component code checking module is used to integrate the self-attention mechanism with the component syntax tree and calculate the weight of each node or subtree in the component syntax tree;

[0033] Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

[0034] A third aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the program is executed by a processor, a method for generating component code as provided in any implementation of the first aspect is implemented.

[0035] A fourth aspect of the present application provides a processor for running a computer program, wherein when the program is run, the method for generating component code provided in any implementation of the first aspect is executed.

[0036] Compared with the prior art, this application has the following beneficial effects:

[0037] The component code generation method provided in the present application converts the acquired natural language information of the component to be generated into the programming language of the component to be generated, constructs a component syntax tree based on the programming language of the component to be generated, extracts key information in the component syntax tree, and more accurately guides the work of the subsequent deep generation model by constructing a component syntax tree and extracting its key information, thereby avoiding unnecessary exploration and improving the efficiency of code generation. Compared with using a large language model to generate component code, although a certain training cost is also required, since the deep generation model used in this application is for the specific task of code generation, it can learn relevant features more efficiently, thereby reducing training time and resource consumption, and achieving the purpose of reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0039] Figure 1 A flowchart of a method for generating component code provided in an embodiment of the present application;

[0040] Figure 2 A flowchart of another method for generating component code provided in an embodiment of the present application;

[0041] Figure 3 A schematic diagram of the structure of a device for generating component code provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] As described above, the current technical solutions for component code generation tasks are mainly based on large language models (such as the GPT series, BERT, etc.). However, there are defects in generating component code based on large language models: large language models need to process a large amount of code data for component code generation, which requires high-performance computing devices and large-scale storage space. At the same time, in order to obtain high-quality generation effects, large language models require long-term training and parameter adjustment, which also increases manpower and time costs. These high costs limit the promotion and application of code generation technology in some application scenarios with limited resources or tight budgets.

[0043] In view of the above problems, the inventors have proposed a method, device, storage medium and processor for generating component code after research, which obtain natural language information describing the component to be generated; convert the natural language information into the programming language of the component to be generated; construct a component syntax tree based on the programming language of the component to be generated, and extract key information in the component syntax tree; the key information includes key nodes or key subtrees; the key information is input into a deep generation model, and the deep generation model makes predictions based on the key information and outputs the code of the component to be generated.

[0044] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0045] See also Figure 1 , which is a flow chart of a method for generating component code provided by an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:

[0046] S101. Obtain natural language information describing the component to be generated.

[0047] This step involves obtaining a natural language description provided by the user about the component they want to create, which can be a functional requirement, a behavioral specification, or any other form of description. This can be obtained by receiving text input through a user interface or transcribing a voice command.

[0048] Obtaining natural language information from users describing the components to be generated is the first step in the component code generation process, which determines the quality and accuracy of subsequent steps.

[0049] S102: Convert the natural language information into the programming language of the component to be generated.

[0050] Natural language processing (NLP) techniques, such as named entity recognition (NER) and dependency parsing, are used to convert unstructured natural language into structured programming language. This conversion process usually involves NLP tasks such as semantic parsing and intent recognition to understand user requests and map them to corresponding programming concepts.

[0051] S103: construct a component syntax tree based on the programming language of the component to be generated, and extract key information in the component syntax tree.

[0052] The key information includes key nodes or key subtrees.

[0053] Using an Abstract Syntax Tree (AST) parser to convert the programming language of the component to be generated into a structured AST representation is a core method in the field of software engineering and programming language processing. AST clearly displays the grammatical elements and their interrelationships in the programming language in a tree structure. During the construction of AST, the parser will parse the character sequence in the programming language into a series of nodes with clear semantics based on the grammatical rules of the programming language. These nodes are connected together through parent-child relationships to form a hierarchical tree structure. AST not only retains key grammatical information in the source code, such as function definitions, loop structures, conditional branches, etc., but also shields grammatical details such as brackets and semicolons, making subsequent analysis and processing more efficient and intuitive.

[0054] S104: Input the key information into a deep generative model, and the deep generative model makes predictions based on the key information and outputs the code of the component to be generated.

[0055] Use a trained deep generative model, such as a model based on the Transformer architecture, to analyze the extracted key information and generate the final code based on it. Deep generative models are usually trained on a large number of labeled data sets and can learn how to infer the correct code snippets from a given context. The generated code should meet the expected functional requirements and be syntactically correct.

[0056] The component code generation method provided in this embodiment converts the acquired natural language information of the component to be generated into the programming language of the component to be generated, constructs a component syntax tree based on the programming language of the component to be generated, extracts key information in the component syntax tree, and guides the subsequent deep generation model by constructing the component syntax tree and extracting its key information, thereby avoiding unnecessary exploration and improving the code generation efficiency. Compared with using a large language model to generate component code, although a certain training cost is also required, since the deep generation model used in this application is for the specific task of code generation, it can learn relevant features more efficiently, thereby reducing training time and resource consumption, and achieving the purpose of reducing costs.

[0057] In order to further improve the accuracy of component code generation, a saliency map algorithm is introduced on the basis of the previous embodiment to provide more precise guidance for the deep generation model.

[0058] See also Figure 2 , which is a flowchart of another method for generating component code provided in an embodiment of the present application. Figure 2As shown, the method comprises the following steps:

[0059] S201. Obtain natural language information describing the component to be generated.

[0060] For example, assuming that a user is developing a tool for generating a Web API, the user may need to provide the following natural language description: "I need an API endpoint that accepts two integer parameters and returns the sum of the two numbers."

[0061] At this step, it is important for the user to provide as many details as possible, such as data types, boundary conditions, error handling methods, etc., because these will affect the quality of the final generated code.

[0062] S202: Convert the natural language information into the programming language of the component to be generated.

[0063] Exemplarily, semantic understanding is performed through the TinyBERT model to further refine and abstract the natural language information to improve the accuracy of subsequent steps. TinyBERT is a knowledge distillation method for transformer-based models jointly proposed by Huazhong University of Science and Technology and Huawei Noah's Ark Laboratory. It compresses and optimizes large pre-trained models such as BERT to create a smaller, faster model with similar performance to solve the problems of too many parameters, large size, and long inference time of models such as BERT. Training TinyBERT usually involves two main stages: knowledge distillation and data fine-tuning. In the first stage, the knowledge distillation process, TinyBERT uses a large pre-trained model (such as BERT) as a teacher model and learns by imitating its output behavior. In the second stage, the data fine-tuning stage, TinyBERT is trained on labeled datasets for specific tasks (such as text classification, named entity recognition, etc.) to further adjust its parameters to better adapt to downstream tasks. Through these two stages of training, TinyBERT is able to achieve performance similar to or even surpass that of large models in some tasks while keeping the model small.

[0064] S203: construct a component syntax tree based on the programming language of the component to be generated, and extract key information in the component syntax tree.

[0065] The key information includes key nodes or key subtrees.

[0066] Static code analysis technology is used to extract key syntax information contained in the component syntax tree; the key syntax information includes component interface information, dependency relationships between components, or internal logic information of components.

[0067] Static code analysis is a technology that analyzes and processes program code without executing it. By traversing and parsing the AST, you can accurately identify components such as classes, methods, variables, and their calling relationships, dependencies, etc. When extracting component interfaces, you can focus on the nodes in the AST that represent method declarations, as well as the parameter lists and return types of these nodes. Similarly, by analyzing the call expressions and reference relationships in the AST, you can accurately identify the dependencies between components. In addition, you can also extract the internal logic information of a component by analyzing the control flow nodes in the AST (such as conditional statements, loop statements, etc.).

[0068] The method of combining AST parser and static code analysis technology can efficiently extract key information from source code and provide strong support for subsequent software development, testing, maintenance, etc. This method has broad application prospects in software quality assurance, code refactoring, security vulnerability detection and other fields.

[0069] S204: Analyze the key information using a saliency graph algorithm, adjust the component syntax tree based on the analysis result, and update the key information.

[0070] The saliency map algorithm is an algorithm for identifying key areas in images or data. In the code generation task, "saliency" can refer to the importance of natural language descriptions or abstract syntax tree nodes to the final code generation. Exemplarily, the key information is analyzed using the saliency map algorithm, and the component syntax tree is adjusted based on the analysis results to update the key information, including:

[0071] Based on the back-propagation algorithm, for each key information item, calculate its gradient relative to the model output (i.e., the generated code);

[0072] Generate a saliency map based on the gradient values ​​to show which parts of the information are most important for code generation;

[0073] For AST nodes or subtrees with higher significance, make necessary adjustments to the AST, such as optimizing control flow, streamlining expressions, etc. For parts with lower significance, consider whether they can be simplified or automated.

[0074] These key nodes or subtrees are usually closely related to the core logic, input and output interfaces, etc. of the component, and are the areas that need to be focused on when generating new component instances. Based on the guidance provided by the saliency map, the key information is continuously adjusted iteratively until the requirements for high-quality code generation are met, so that the deep generative model can focus more on these key areas during the playback process, thereby generating more accurate and useful new component instances.

[0075] The updated key information is converted into a form that matches the processing form of the deep generative model; the form includes an embedded vector form or a path expression form.

[0076] S205: Input the updated key information into the deep generation model, and the deep generation model performs prediction based on the key information and outputs the code of the component to be generated.

[0077] For example, if the deep generative model is a Generative Adversarial Network (GANs), generating new component instances from historical data is an innovative and effective method to expand the training data set. GANs consists of two neural networks, a generator and a discriminator. Through the process of mutual competition and optimization, the generator can learn the potential distribution of the data set and generate realistic new data accordingly. In the code generation scenario, component instances in the historical code base can be used as training data to generate new component instances with similar grammatical and semantic structures through GANs.

[0078] After inputting the key information into the deep generative model, and the deep generative model predicts according to the key information and outputs the code of the component to be generated, the method further includes:

[0079] The self-attention mechanism is integrated with the component syntax tree to calculate the weight of each node or subtree in the component syntax tree;

[0080] Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

[0081] The self-attention mechanism is a neural network architecture that captures the interdependencies between different positions in the input sequence and highlights important information by calculating attention weights. In the context of AST, each node or subtree can be regarded as an information unit, and the self-attention mechanism can calculate the correlation between these units and assign importance weights to them accordingly.

[0082] To further improve the accuracy of code inspection and analysis, the attention weights are fused with the representation of component language models such as CodeBERT. CodeBERT is based on the Transformer architecture and can capture grammatical and semantic information in the code. An enhanced inspection model is formed by combining the weights calculated by the self-attention mechanism with the representation of CodeBERT. This inspection model can not only capture the key grammatical and semantic features in the code, but also weight them according to the importance of nodes or subtrees, so as to more accurately identify potential problems or optimization points in the code. By utilizing this enhanced inspection model, we can perform more in-depth analysis and inspection of the code to improve code quality, security, and maintainability.

[0083] Another component code generation method provided in this embodiment combines the powerful generation capability of deep learning and the precise positioning characteristics of saliency maps, and can continuously learn and update itself from new data. This method not only reduces the cost of traditional generation methods and improves the generation efficiency of component codes, but also has the ability of continuous learning and can continuously optimize the deep generation model to adapt to ever-changing and complex application scenarios.

[0084] Based on the component code generation method introduced in the above embodiment, the present application also provides a component code generation device. Figure 3 Figure 1 is a schematic diagram of the structure of the device. Figure 3 As shown, the component code generation device includes:

[0085] A natural language information acquisition module 301 is used to acquire natural language information describing the component to be generated;

[0086] A natural language processing module 302, used to convert the natural language information into the programming language of the component to be generated;

[0087] A component syntax tree key information extraction module 303 is used to construct a component syntax tree based on the programming language of the component to be generated, and extract key information in the component syntax tree; the key information includes key nodes or key subtrees;

[0088] The component code generation module 304 is used to input the key information into the deep generation model, and the deep generation model makes predictions based on the key information and outputs the code of the component to be generated.

[0089] In an optional implementation, constructing a component syntax tree based on the programming language of the component to be generated, and extracting key information in the component syntax tree includes:

[0090] Static code analysis technology is used to extract key syntax information contained in the component syntax tree; the key syntax information includes component interface information, dependency relationships between components, or internal logic information of components.

[0091] In an optional implementation, before the key information is input into a deep generative model, and the deep generative model predicts based on the key information and outputs the code of the component to be generated, the method further includes:

[0092] The key information is converted into a form that matches the processing form of the deep generative model; the form includes an embedded vector form or a path expression form.

[0093] In an optional embodiment, the device further comprises: a key information verification module;

[0094] The key information verification module is used to analyze the key information using a saliency map algorithm;

[0095] The component syntax tree is adjusted based on the analysis result, and the key information is updated.

[0096] In an optional embodiment, the device further comprises: a component code checking module;

[0097] The component code checking module is used to integrate the self-attention mechanism with the component syntax tree and calculate the weight of each node or subtree in the component syntax tree;

[0098] Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

[0099] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the program is executed by a processor, the component code generation method described in any of the method embodiments is implemented.

[0100] In addition, an embodiment of the present application further provides a processor, which is used to run a computer program, and when the program is run, the method for generating component code as described in any implementation manner of the aforementioned method embodiment is executed.

[0101] It should be noted that each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely schematic, in which the unit described as a separate component may or may not be physically separated, and the component prompted as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.

[0102] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for generating component code, characterized in that: include: Obtain natural language information describing the component to be generated; Converting the natural language information into the programming language of the component to be generated; Building a component syntax tree based on the programming language of the component to be generated, and extracting key information from the component syntax tree; The key information includes key nodes or key subtrees; The key information is input into a deep generative model, and the deep generative model makes predictions based on the key information and outputs the code of the component to be generated.

2. The method according to claim 1, characterized in that The key information is input into a deep generative model, and the deep generative model predicts based on the key information, and before outputting the code of the component to be generated, the method further includes: Analyzing the key information using a saliency map algorithm; The component syntax tree is adjusted based on the analysis result, and the key information is updated.

3. The method according to claim 1, characterized in that After inputting the key information into the deep generative model, and the deep generative model predicts according to the key information and outputs the code of the component to be generated, the method further includes: The self-attention mechanism is integrated with the component syntax tree to calculate the weight of each node or subtree in the component syntax tree; Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

4. The method according to claim 1, characterized in that: Constructing a component syntax tree based on the programming language of the component to be generated, and extracting key information in the component syntax tree, including: Static code analysis technology is used to extract key syntax information contained in the component syntax tree; the key syntax information includes component interface information, dependency relationships between components, or internal logic information of components.

5. The method according to claim 1, characterized in that The key information is input into a deep generative model, and the deep generative model predicts based on the key information, and before outputting the code of the component to be generated, the method further includes: The key information is converted into a form that matches the processing form of the deep generative model; the form includes an embedded vector form or a path expression form.

6. A device for generating a component code, characterized in that: include: A natural language information acquisition module, used to acquire natural language information describing the component to be generated; A natural language processing module, used for converting the natural language information into the programming language of the component to be generated; A component syntax tree key information extraction module, used to construct a component syntax tree based on the programming language of the component to be generated, and extract key information from the component syntax tree; The key information includes key nodes or key subtrees; The component code generation module is used to input the key information into the deep generation model, and the deep generation model predicts according to the key information and outputs the code of the component to be generated.

7. The device according to claim 6, characterized in that Also includes: Key information verification module; The key information verification module is used to analyze the key information using a saliency map algorithm; The component syntax tree is adjusted based on the analysis result, and the key information is updated.

8. The device according to claim 6, characterized in that Also includes: Component code inspection module; The component code checking module is used to integrate the self-attention mechanism with the component syntax tree and calculate the weight of each node or subtree in the component syntax tree; Based on the calculated weight of each node or subtree in the component syntax tree, a component language model is used to perform a check to identify problems existing in the output code of the component to be generated.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the program is executed by a processor, the method for generating component code according to any one of claims 1 to 5 is implemented.

10. A processor, characterized in that: Used to run a computer program, which executes the method for generating component code according to any one of claims 1 to 5 when the program is run.

Citation Information

Cited By

  • Code generation method based on grammar discriminator and related equipment

    CN120371274A