Code generation method and device based on AI agent, equipment and medium
Through the AI agent-based code generation method, large language models are used to conduct interactive query and generate target code, which solves the problems of difficult and inefficient code writing in the existing technology, and achieves more efficient code generation.
Patent Information
- Application Number
- CN202510224044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The existing code generation process relies on programmers' writing techniques and domain knowledge, making code writing difficult and inefficient.
Using the code generation method based on AI agents, the initial description information of the object code generation task is obtained, the initial task element components are generated, and the related problem list is generated through a large language model for interactive inquiry, and supplementary description information is obtained, and the object code is finally generated based on the component combination logic.
It reduces the difficulty of code writing, improves code generation efficiency, and improves user experience.
Smart Images

Figure CN120066477A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology and dialogue / question-answering systems, and is applicable to the medical and financial fields. In particular, it relates to a method, device, computer device, and computer-readable storage medium for code generation based on an AI agent. Background Art
[0002] Currently, in the code generation process, programmers need to disassemble a given code generation task, and then write the disassembled tasks through computer language prompts. After continuously iterating and optimizing the prompts, the final code corresponding to the code generation task is obtained. However, the current code generation process mainly depends on programmers' code writing skills and their understanding of domain knowledge. Not only is the code writing difficult, but the code generation efficiency is also low.
[0003] Therefore, how to solve the problem of reducing the difficulty of code writing and thus improving the code generation efficiency has become an urgent technical problem to be solved at present. Summary of the Invention
[0004] The main objective of the present invention is to provide a method, device, equipment, and computer-readable storage medium for code generation based on an AI agent, aiming to reduce the difficulty of code writing and thus improve the code generation efficiency.
[0005] To achieve the above objective, the present invention provides a method for code generation based on an AI agent, and the code generation method includes the following steps:
[0006] Obtain the initial description information of the target code generation task, and generate the initial task element components corresponding to the initial description information based on the target AI agent;
[0007] Based on the target large language model, generate a list of relevant questions corresponding to the initial description information and the initial task element components, and execute an interactive code generation task question inquiry to the target user based on the list of relevant questions to obtain the supplementary description information corresponding to the target user;
[0008] Generate supplementary task element components corresponding to the supplementary description information based on the target AI agent, and determine the component combination logic corresponding to the initial task element components and the supplementary task element components based on the target AI agent;
[0009] Based on the initial task element components, supplementary task element components, and the component combination logic, combine and generate the target code corresponding to the target code generation task.
[0010] In addition, to achieve the above objective, the present invention also provides a device for code generation based on an AI agent, and the device for code generation based on an AI agent includes:
[0011] An initial component generation module, configured to obtain initial description information of a target code generation task, and generate an initial task element component corresponding to the initial description information based on a target AI agent;
[0012] A supplementary information acquisition module, configured to generate a list of related questions corresponding to the initial description information and the initial task element component based on a target large language model, and perform interactive code generation task question inquiries to a target user based on the list of related questions to obtain supplementary description information corresponding to the target user;
[0013] A component information supplement module, configured to generate a supplementary task element component corresponding to the supplementary description information based on the target AI agent, and determine a component combination logic corresponding to the initial task element component and the supplementary task element component based on the target AI agent;
[0014] A target code generation module, configured to generate a target code corresponding to the target code generation task by combining the initial task element component, the supplementary task element component, and the component combination logic.
[0015] In addition, to achieve the above object, the present invention further provides a computer device, the computer device includes a processor, a memory, and an AI agent-based code generation program stored on the memory and executable by the processor, wherein when the AI agent-based code generation program is executed by the processor, the steps of the above-mentioned AI agent-based code generation method are implemented.
[0016] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which an AI agent-based code generation program is stored, wherein when the AI agent-based code generation program is executed by a processor, the steps of the above-mentioned AI agent-based code generation method are implemented.
[0017] The present application provides a code generation method based on an AI agent. The method obtains initial description information of a target code generation task, and generates initial task element components corresponding to the initial description information based on the target AI agent; based on a target large language model, generates a list of related questions corresponding to the initial description information and the initial task element components, and executes an interactive code generation task question inquiry to a target user based on the list of related questions to obtain supplementary description information corresponding to the target user; generates supplementary task element components corresponding to the supplementary description information based on the target AI agent, and determines a component combination logic corresponding to the initial task element components and the supplementary task element components based on the target AI agent; combines and generates target code corresponding to the target code generation task based on the initial task element components, the supplementary task element components, and the component combination logic. Through the above method, the present application first analyzes the initial description information through an AI agent to determine the initial task element components corresponding to the target code, and further generates a list of related questions for the target code based on the initial description information and the initial task element components through a target large language model, and executes an interactive inquiry to the target user based on the list of related questions to obtain supplementary description information of the target code. Then, the supplementary task element components corresponding to the target code are determined by analyzing the supplementary description information through the AI agent. Finally, the target code is combined and generated based on the initial task element components, the supplementary task element components, and the component combination logic. Thus, through the AI agent and the target large language model, the automatic generation of code is realized, the code generation efficiency is improved, and the user experience is enhanced. Description of the Drawings
[0018] Figure 1 It is a schematic flowchart of the first embodiment of the code generation method based on an AI agent of the present invention;
[0019] Figure 2 It is a schematic diagram of task parsing based on an AI agent provided by an embodiment of the present invention;
[0020] Figure 3 It is a schematic diagram of the code generation device based on an AI agent provided by an embodiment of the present invention;
[0021] Figure 4 It is a schematic diagram of the hardware structure of a computer device involved in the solution of an embodiment of the present invention.
[0022] The realization, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0025] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0026] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] The code generation method based on an AI agent involved in the embodiments of the present invention is mainly applied to a computer device, which can be a device with display and processing functions such as a PC, a portable computer, a mobile terminal, etc.
[0028] Next, some embodiments of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0029] Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the code generation method based on an AI agent of the present invention.
[0030] As Figure 1 shown, the embodiments of the present invention provide a code generation method based on an AI agent, and the code generation method based on an AI agent includes steps S10 to S40.
[0031] Step S10: Obtain the initial description information of the target code generation task, and generate an initial task element component corresponding to the initial description information based on the target AI agent;
[0032] The code generation method in this embodiment can be applied to medical scenarios, such as generating codes corresponding to physical examination item recommendation tasks, medical insurance payment tasks, etc. Specifically, the physical examination item recommendation task can be split into the following task element components: a basic information acquisition module for acquiring the basic health information of the target user; a dynamic information acquisition module for generating a list of relevant questions corresponding to the basic health information based on the target large language model, and performing interactive physical examination question inquiries on the target user based on the list of relevant questions to obtain the dynamic health information corresponding to the target user; a physical examination item recommendation module for generating a code corresponding to the physical examination item recommendation task based on the physical examination item recommendation logic, the basic information acquisition module, and the dynamic information acquisition module.
[0033] The code generation method in this embodiment can also be used in the financial field, such as generating codes corresponding to product detail page generation tasks, risk control data analysis tasks, etc. in the e-commerce field. Specifically, the product detail page generation task in the e-commerce field can be split into the following task element components: a product basic information module, a payment module, and a comment module.
[0034] Receive the code task requirement input by the target user (such as "I want to create a product detail page"), and obtain the initial description information from the code task requirement. The initial description information includes the product basic information, payment information, comment information, etc. corresponding to the product page. That is, as Figure 2 shown, through the chain of thought and planning ability of the large language model (LLM) of the AI agent (i.e., the AI programming agent shown in Figure 2 ), complex tasks are reasoned. For example, in the e-commerce field, for the several major element components that must be included in the product detail page component, namely the product basic information module, the payment module, and the comment module. Based on the analysis of the initial description information by the target AI agent, each disassembled information for realizing each target function in the initial description information is disassembled, so as to determine each initial task element component based on each disassembled information for each target function. Specifically, the target AI agent can determine the element component that can achieve this effect based on the role of each disassembled information in the initial description information as the initial task element component.
[0035] In a specific embodiment, based on the corresponding relationship diseasefactors between the target function of each disassembled information and each element component, calculate the association degree score DR(i) between a certain disassembled information and the element component i. The calculation formula is as follows:
[0036]
[0037] Among them, diseasefactors(i) is used to represent all functions corresponding to the element component i, and the denominator (|diseasefactors(i)|) is used to represent the total number of functions corresponding to the element component i. user_risk represents the overlapping effect of the element component i and a certain disassembly information, and the numerator is used to represent the number of overlapping effects of the element component i and a certain disassembly information. DR(i) is used to represent the correlation score between the element component i and a certain disassembly information, and the correlation score ranges from 0 to 1. The closer it is to 1, the greater the correlation.
[0038] According to the calculated correlation scores between each disassembly information and the element components, the higher the correlation score, the greater the correlation score between the element component i and a certain disassembly information. The element component with the highest correlation score is determined as the task element component corresponding to the disassembly information, and thus the initial task element component set is obtained.
[0039] Step S20: Based on the target large language model, generate the initial description information and the relevant question list corresponding to the initial task element components, and perform an interactive code generation task question inquiry to the target user based on the relevant question list to obtain the supplementary description information corresponding to the target user;
[0040] In this embodiment, the initial description information and the initial task element components are input into the target large language model, and the large language model generates questions related to the keywords in the initial description information or questions related to the functions of each component in the initial task element components.
[0041] Specifically, task element components with corresponding functions can be pre-generated according to each code task encapsulation, and each task element component and its corresponding function are stored as keys.
[0042] Furthermore, count the total number of occurrences of each element component in each code task, and determine the correlation between each element component according to the number of occurrences of the element component.
[0043] Input the initial description information into the target large language model, analyze the initial description information through the target large language model, and determine the element components corresponding to each information. Input the initial task element components into the target large language model, and determine the associated components of each element component in the initial task element components through the target large language model. And generate a relevant question list based on the element components corresponding to each information and / or the associated components of each element component in the initial task element components.
[0044] Exemplarily, the generating the relevant question list corresponding to the initial description information and the initial task element components based on the target large language model specifically includes:
[0045] Based on the target knowledge graph, obtain the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element components;
[0046] Input the initial description information and its corresponding domain context information, the initial task element components and their corresponding domain context information into the target large language model to obtain the list of relevant questions.
[0047] In this embodiment, based on the list of relevant questions, questions are asked to the target user to discover supplementary description information for the target code generation task. Specifically, the DR of each element component score can be obtained, and the associated component i is sorted in descending order of the element component score, and the element components with higher score rankings are selected for questioning.
[0048] Specifically, based on the function description of the element component and the component knowledge graph, obtain relevant information about the element component and its corresponding function; for the element component i, execute the GraphRAG retrieval algorithm to obtain relevant domain knowledge such as community reports and component functions. The target large language model combines the initial description information, the initial task element components, and the extracted domain context knowledge, and analyzes other element components that may be required for the target code generation task with the help of a predefined Prompt template, and generates a list of relevant questions based on the above other element components. Specifically include:
[0049] 1. Information integration and filling: Fill the initial description information, the initial task element components, and the domain context knowledge extracted by the GraphRAG method into the Prompt template together to provide comprehensive data support for the analysis. That is, based on the large model, use the Chain-of-Thought (COT) reasoning and In-Context Learning (ICL) methods to guide the model to gradually reason and analyze the initial description information, the initial task element components, and the domain context, and generate other element components related to the initial description information and the initial task element components. And comprehensively consider the interaction of domain context knowledge and various factors during the analysis process.
[0050] 2. Candidate list generation: Output a structured list of other element components through the large model.
[0051] 3. Confidence score: The target large language model gives a relevance score for each element component. The relevance score is based on the analysis result of the total number of occurrences of each element component in the historical samples.
[0052] For example, the initial description information corresponding to the target code generation task is "I want to create a product details page". The initial task element components include a product information display component, a product payment component, a purchase link jump component, etc. Other element components include a comment component, a payment status component, an associated product recommendation component, etc., and specific relevant questions are generated based on the other element components. For example:
[0053] If the other element component is a comment component, the relevant question generated is: "Is it necessary to provide a product evaluation function through the comment component?"
[0054] If the other element component is a payment status display component, the relevant question generated is: "Is it necessary to provide a product payment status display function through the payment status component?"
[0055] Further, after performing an interactive code generation task question inquiry to the target user based on the relevant question list to obtain supplementary description information corresponding to the target user, it further includes:
[0056] Count the number of dialogue turns corresponding to the target user, and calculate the perfection degree of the element components of the target code generation task based on the initial task element components and the supplementary task element components of the target code generation task;
[0057] When the number of dialogue turns corresponding to the target user exceeds a preset number threshold or when the perfection degree of the element components exceeds a preset percentage threshold, generate an inquiry stop instruction to terminate the code generation task question inquiry corresponding to the target user.
[0058] In this embodiment, after asking questions to the target user, the target user supplements the description information based on the relevant questions.
[0059] Furthermore, the AI agent is responsible for managing the entire interactive Q&A process and generating an inquiry stop instruction to end the code generation task interactive inquiry process.
[0060] Specifically, based on the initial task element components and the target user's answers currently collected, the AI agent judges whether the current task element components are used to form a runnable program and the degree of program perfection based on predefined Prompts and domain-related knowledge extracted by GraphRAG. If it is not runnable or the program function is not perfect enough, and the current dialogue turn is less than or equal to 10, the AI agent can continue to ask questions to the target user for this code generation task; if the task element components can be used to form a runnable program and the program perfection degree is relatively high, or the dialogue turn exceeds 10, an inquiry stop instruction is generated to terminate the question inquiry corresponding to the target user.
[0061] Step S30: Generate a supplementary task element component corresponding to the supplementary description information based on the target AI agent, and determine a component combination logic corresponding to the initial task element component and the supplementary task element component based on the target AI agent;
[0062] In this embodiment, after determining the supplementary task element component corresponding to the supplementary function based on the supplementary description information, based on the pre-stored component combination logic (i.e., component operation logic) between each component. The component combination logic includes the call order, operation order of each component, and parameters that need to be adjusted correspondingly in the component, etc.
[0063] Step S40: Combine and generate a target code corresponding to the target code generation task based on the initial task element component, the supplementary task element component, and the component combination logic.
[0064] Further, after the step S40, it further includes:
[0065] Obtain the script language type and the front-end framework type corresponding to the target code generation task, and generate and display the corresponding front-end interface code based on the target code, the script language type, and the front-end framework type.
[0066] In this embodiment, based on the initial task element component, the supplementary task element component, and the component combination logic, perform a front-end architecture design on the target code generation task to generate a task skeleton corresponding to the target code generation task. Among them, the task skeleton corresponds to several simple code generation subtasks. For example, each element component corresponds to a task code generation subtask. Further, through front-end technology selection, perform corresponding code generation and redundant code merging respectively to obtain the target code.
[0067] This embodiment provides a code generation method based on an AI agent. The method obtains the initial description information of the target code generation task, and generates the initial task element components corresponding to the initial description information based on the target AI agent; based on the target large language model, generates a list of related questions corresponding to the initial description information and the initial task element components, and executes an interactive code generation task question inquiry to the target user based on the list of related questions to obtain the supplementary description information corresponding to the target user; generates the supplementary task element components corresponding to the supplementary description information based on the target AI agent, and determines the component combination logic corresponding to the initial task element components and the supplementary task element components based on the target AI agent; based on the initial task element components, the supplementary task element components, and the component combination logic, combines and generates the target code corresponding to the target code generation task. In the above manner, this application first analyzes the initial description information through the AI agent to determine the initial task element components corresponding to the target code, and further generates a list of related questions for the target code based on the initial description information and the initial task element components through the target large language model, and executes an interactive inquiry to the target user based on the list of related questions to obtain the supplementary description information of the target code, and then analyzes the supplementary description information through the AI agent to determine the supplementary task element components corresponding to the target code. Finally, the target code is combined and generated based on the initial task element components, the supplementary task element components, and the component combination logic. Thus, through the AI agent and the target large language model, the automatic generation of code is realized, the code generation efficiency is improved, and the user experience is enhanced.
[0068] Further, before obtaining the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element components based on the target knowledge graph, it further includes:
[0069] Collect the relevant information of the historical code generation task, and split the relevant information into text units of each preset unit;
[0070] Based on the target large language model, respectively extract each element component and the combined logical relationship corresponding to each element component in each text unit;
[0071] Based on each element component and each combined logical relationship, construct a homogeneous undirected weighted graph, where the nodes and edges in the homogeneous undirected weighted graph respectively represent each element component and the combined logical relationship corresponding to each element component, and the weight of each edge represents the standardized count of the detected combined logical relationship instances;
[0072] Merge the subgraphs corresponding to different text units with the same element components in the homogeneous undirected weighted graph, and perform hierarchical clustering on the merged homogeneous undirected weighted graph to obtain each community;
[0073] Extract the report-style summaries corresponding to each community, and generate the target knowledge graph based on each community and its corresponding report-style summary.
[0074] In this embodiment, based on the historical code sample data, the GraphRAG indexing technology is used to construct a relevant knowledge base.
[0075] It can be understood that the collected data is preprocessed in advance, including text cleaning, removing irrelevant information, and term standardization.
[0076] Split the relevant information into a series of text units (TextUnits), and the above units are the basic units for subsequent analysis. For example, the unit size is 300 tokens.
[0077] Furthermore, in order to avoid sentences and paragraphs from being truncated, the sequential chunking method can be adopted, introducing a sliding window to provide context for each chunk.
[0078] Use a large language model (LLM) to extract all entities, relationships, and key statements from different text units respectively. The key statement is a statement that describes the interaction between the subject entity and the object entity (that is, the correspondence between each code generation task and each element component, such as optional components or mandatory components). Combine the predefined prompt and examples, clarify the extraction objectives and relationships, and then use the large language model to output the extraction results in a formalized structure.
[0079] Construct a homogeneous undirected weighted graph using the extracted entities and relationships. The nodes in the homogeneous undirected weighted graph are used to represent each entity, the edges in the homogeneous undirected weighted graph are used to represent the corresponding relationships of each entity, and the weight of the edge represents the normalized count of the detected relationship instances. Merge the subgraphs corresponding to different text units according to the same entity (that is, the name and type are the same).
[0080] To ensure the connectivity within the community and further improve the code generation efficiency, after the merging is completed, perform hierarchical clustering on the merged homogeneous undirected weighted graph through the Leiden algorithm to form different communities, so that the community division is more reasonable.
[0081] It can be understood that the traditional Louvain algorithm can also be used to perform hierarchical clustering on the merged homogeneous undirected weighted graph.
[0082] For each community, generate summaries of the community and its members from bottom to top. Create a report-style summary for each community in the Leiden hierarchy. In this way, it is convenient to understand the global structure and semantics of the dataset through the above summaries.
[0083] Generate the target knowledge graph based on each community and its corresponding report-style summary. And store the constructed knowledge graph in the graph database Neo4j for subsequent query and analysis.
[0084] Further, before the step S20, it further includes:
[0085] Initialize the weight parameters and bias parameters of each network layer in the initial large language model, and generate the survival probability parameters of each network layer based on a preset rule;
[0086] Based on the survival probability parameters of each network layer, determine the target network layer and the optimized network layer in each network layer. The target network layer is the network layer to be trained, and the optimized network layer is the network layer to be skipped during training;
[0087] Based on a preset training sample set, initial weight parameters, and initial bias parameters, skip the optimized network layer and train the target network layer until a preset convergence condition is reached to obtain a trained model as the target large language model.
[0088] In this embodiment, in order to prevent the target large language model from overfitting or underfitting during training, resulting in a long model training time, a stochastic depth mechanism is provided based on the survival probability corresponding to each network layer in the model. Based on this stochastic depth mechanism, some network layers can be randomly skipped during the training process, reducing the computational amount of the forward and backward propagation during model training, thereby reducing the computational complexity during model training; moreover, only some network layers are used in each training, so that while significantly shortening the training time, the convergence of the model is ensured, overfitting of the model is prevented, and the robustness of the model to abnormal data and noise data is enhanced. In addition, the target large language model trained based on the stochastic depth mechanism can further improve the model training efficiency and label generation accuracy on the premise of ensuring the model performance.
[0089] Specifically, first, initialize the network parameters and the survival probability of each layer. Then, in each forward propagation process, randomly determine whether to skip the layer according to the survival probability of each layer. For the network layers participating in training, perform forward and backward propagation to update the parameters normally. For the network layers not participating in training, do not update the parameters, thereby realizing multiple iterative trainings of the model to ensure the convergence of the model and achieve the optimal performance.
[0090] Exemplarily, based on the He initialization method or the Xavier initialization method, initialize the weight parameters and bias parameters of each network layer to obtain the initial weight parameters and the initial bias parameters;
[0091] Based on a linear decay strategy, survival probability parameters for each network layer are generated, where the level of the network layer is inversely proportional to the value of its corresponding survival probability parameter.
[0092] In this embodiment, by introducing a stochastic depth mechanism during training, efficient training of a deep neural network is achieved, specifically including:
[0093] First, initialize network parameters, that is, initialize the weight parameters and bias parameters for each layer of the stochastic depth neural network. Specifically, methods such as He initialization or Xavier initialization can be used to initialize the model parameters of the stochastic depth neural network to ensure the stability and effectiveness of training.
[0094] Second, initialize the survival probability, that is, assign an initial survival probability p to each layer of the network l 。
[0095] Specifically, a linear decay strategy can be adopted, that is, the level of the network layer is inversely proportional to the value of its corresponding survival probability parameter. By adopting a linear decay survival probability design, the effective depth of the network during training is ensured, while overly complex calculations are avoided.
[0096] That is to say, the survival probability near the input layer is higher, while the survival probability near the output layer is lower. For example, the survival probability can be set as:
[0097]
[0098] where l is the current network layer, L is the total number of network layers, and p L is the survival probability of the output layer in the previous round. The initial value can be taken in the first round, and this value is generally relatively low, such as less than 50%.
[0099] Specifically, a preset probability threshold can be set. When the survival probability parameter is not less than this preset probability threshold, this network layer can be marked as the target network layer for training. When the survival probability parameter is less than this preset probability threshold, this network layer can be marked as the optimized network layer to skip this network layer.
[0100] The initial model is a random depth network model with the initial weight parameters and initial bias parameters as the current model parameters. In each forward propagation process, the description information and element components in the preset historical task code sample set are input into each target network layer in the initial random depth network model, that is, the output results of each target network layer are directly passed to the next layer of the target network layer, skipping the next layer of the optimized network layer. After completing the training of all target network layers in the random depth network model, it is judged whether the current random depth network model reaches the convergence condition. If not, the target network layers in the random depth network model are iteratively trained again, still skipping the optimized network layer, until the random depth network model reaches the preset convergence condition. Stop the model training to obtain the trained target large language model.
[0101] Further, determining the target network layer and the optimized network layer based on the survival probability parameters of each network layer specifically includes:
[0102] Obtain any value within the preset numerical range as the random parameter of a network layer;
[0103] Compare the survival probability parameter of each network layer with its corresponding random parameter;
[0104] Based on the comparison results of each network layer, determine the network layer with the survival probability parameter not less than the random parameter as the target network layer, and determine the network layer with the survival probability parameter less than the random parameter as the optimized network layer.
[0105] First, input the description information and element component samples in the historical task code sample set into the first-level network layer.
[0106] Then, process each network layer outside the first-level network layer in the random depth network model layer by layer, and perform forward propagation on each network layer of the initialized random depth network model:
[0107] For each network layer l, generate a random number r between [0, 1] l , as the random parameter of each network layer;
[0108] The random parameter r of each network layer l is compared with its corresponding survival probability parameter p l as follows:
[0109] 1. If r l ≤p l , then this layer participates in the calculation and performs normal forward propagation:
[0110] x l+1 =f l (x l ,Wl , b l )
[0111] Among them, x l is the input of the l-th layer, W l and b l are the weight and bias of the l-th layer respectively, and f l is the activation function of this layer.
[0112] 2. If r l > p l , then skip this layer and directly pass the current input to the next layer:
[0113] x l+1 = x l
[0114] In the above way, this embodiment generates corresponding random parameters for each network layer to ensure the randomness of the skipped layers. Then, the random parameters of this layer are compared with their corresponding survival probability parameters, and it is determined whether to skip this layer during training according to the comparison result. Thus, based on the random parameters, the randomness of the optimized network layer is ensured, avoiding a fixed optimized network layer caused by a fixed probability parameter threshold, and further improving the prediction result accuracy of the random depth network model.
[0115] After completing one round of the forward propagation of the model, backpropagation is performed to update the parameters of the network layers participating in the forward propagation. Specifically:
[0116] 1. Calculate the loss function of the target-related problem sample (i.e., the target value y true ) corresponding to the description information and the element component sample, and the current related problem (network output y):
[0117] Calculate the loss function L(y, y true ) according to the network output y and the target value y true ), using cross-entropy loss or mean squared error loss.
[0118] 2. Backpropagate layer by layer:
[0119] Starting from the output layer, calculate the gradient layer by layer forward.
[0120] For each layer l, if this layer participates in the forward propagation (i.e., r l ≤ p l ), then calculate the gradient normally and update the parameters:
[0121]
[0122] Among them, η is the preset learning rate.
[0123] If this layer is skipped (i.e., rl >p l ) If not, no parameter update is performed, and the gradient is directly passed to the previous layer:
[0124]
[0125] Through the above method, in this embodiment, based on the above forward propagation and backward propagation steps, it will iterate multiple times on the entire training set until the loss function converges or reaches the preset number of training rounds. During each iteration, the survival probability p of each network layer in this round l remains unchanged to ensure the stability and effect of model training.
[0126] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of a code generation device based on an AI agent provided by an embodiment of the present application.
[0127] As Figure 3 shown, the code generation device 300 based on an AI agent includes:
[0128] An initial component generation module 310, configured to obtain initial description information of a target code generation task, and generate initial task element components corresponding to the initial description information based on a target AI agent;
[0129] A supplementary information acquisition module 320, configured to generate a list of related questions corresponding to the initial description information and the initial task element components based on a target large language model, and perform interactive code generation task question inquiries to a target user based on the list of related questions to obtain supplementary description information corresponding to the target user;
[0130] A component information supplement module 330, configured to generate supplementary task element components corresponding to the supplementary description information based on the target AI agent, and determine component combination logics corresponding to the initial task element components and the supplementary task element components based on the target AI agent;
[0131] A target code generation module 340, configured to combine and generate a target code corresponding to the target code generation task based on the initial task element components, supplementary task element components, and the component combination logic.
[0132] Furthermore, the code generation device 300 based on an AI agent further includes a model training module, configured to:
[0133] Initialize the weight parameters and bias parameters of each network layer in an initial large language model, and generate survival probability parameters of each network layer based on a preset rule;
[0134] Based on the survival probability parameters of each network layer, determine a target network layer and an optimized network layer in each network layer, where the target network layer is the network layer to be trained, and the optimized network layer is the network layer to be skipped from training;
[0135] Based on a preset training sample set, initial weight parameters, and initial bias parameters, skip the optimized network layer and train the target network layer until a preset convergence condition is reached, and obtain the trained model as the target large language model.
[0136] Further, the model training module is further configured to:
[0137] Obtain any value within a preset numerical range as a random parameter of a network layer;
[0138] Compare the survival probability parameter of each network layer with its corresponding random parameter;
[0139] Based on the comparison results of each network layer, determine that the network layer with a survival probability parameter not less than the random parameter is the target network layer, and determine that the network layer with a survival probability parameter less than the random parameter is the optimized network layer.
[0140] Further, the supplementary information acquisition module is further configured to:
[0141] Based on the target knowledge graph, obtain the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element components;
[0142] Input the initial description information and its corresponding domain context information, the initial task element components and their corresponding domain context information into the target large language model to obtain the relevant question list.
[0143] Further, the supplementary information acquisition module is further configured to:
[0144] Collect relevant information on historical code generation tasks and split the relevant information into text units of each preset unit;
[0145] Based on the target large language model, respectively extract each element component and the combined logical relationship corresponding to each element component in each text unit;
[0146] Based on each element component and each combined logical relationship, construct a homogeneous undirected weighted graph, where the nodes and edges in the homogeneous undirected weighted graph respectively represent each element component and the combined logical relationship corresponding to each element component, and the weights of each edge represent the standardized count of the detected combined logical relationship instances;
[0147] Merge the subgraphs corresponding to different text units with the same element components in the homogeneous undirected weighted graph, and perform hierarchical clustering on the merged homogeneous undirected weighted graph to obtain each community;
[0148] Extract the report-style summaries corresponding to each community, and generate the target knowledge graph based on each community and its corresponding report-style summary.
[0149] Further, the code generation device 300 based on the AI agent further includes an inquiry termination module for:
[0150] Count the number of dialogue turns corresponding to the target user, and calculate the perfection degree of the element components of the target code generation task based on the initial task element components and the supplementary task element components of the target code generation task;
[0151] When the number of dialogue turns corresponding to the target user exceeds the preset number threshold or when the perfection degree of the element components exceeds the preset ratio threshold, generate an inquiry stop instruction to terminate the code generation task question inquiry corresponding to the target user.
[0152] Further, the code generation device 300 based on the AI agent further includes a front-end code display module for:
[0153] Obtain the script language type and front-end framework type corresponding to the target code generation task, and generate and display the corresponding front-end interface code based on the target code, the script language type, and the front-end framework type.
[0154] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0155] The above device can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 4 shown.
[0156] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. This computer device can be a server.
[0157] Refer to Figure 4 , this computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0158] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, enable the processor to execute any one of the code generation methods based on an AI agent.
[0159] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0160] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it enables the processor to execute any one of the code generation methods based on an AI agent.
[0161] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0162] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0163] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:
[0164] Obtain the initial description information of the target code generation task, and generate the initial task element components corresponding to the initial description information based on the target AI agent;
[0165] Based on the target large language model, generate a list of related questions corresponding to the initial description information and the initial task element components, and execute an interactive code generation task question inquiry to the target user based on the list of related questions to obtain supplementary description information corresponding to the target user;
[0166] Generate a supplementary task element component corresponding to the supplementary description information based on the target AI agent, and determine a component combination logic corresponding to the initial task element component and the supplementary task element component based on the target AI agent;
[0167] Based on the initial task element component, the supplementary task element component, and the component combination logic, generate a target code corresponding to the target code generation task.
[0168] In one embodiment, before the processor is used to process generating a relevant question list corresponding to the initial description information and the initial task element component based on the target large language model, and performing an interactive code generation task question inquiry on the target user based on the relevant question list to obtain supplementary description information corresponding to the target user, the processor is further used to implement:
[0169] Initialize the weight parameters and bias parameters of each network layer in the initial large language model, and generate survival probability parameters for each network layer based on a preset rule;
[0170] Based on the survival probability parameters of each network layer, determine a target network layer and an optimized network layer in each network layer, where the target network layer is the network layer to be trained, and the optimized network layer is the network layer to be skipped from training;
[0171] Based on a preset training sample set, initial weight parameters, and initial bias parameters, skip the optimized network layer and train the target network layer until a preset convergence condition is reached, and obtain a trained model as the target large language model.
[0172] In one embodiment, when the processor is used to process determining a target network layer and an optimized network layer in each network layer based on the survival probability parameters of each network layer, the processor is further used to implement:
[0173] Obtain any value within a preset numerical range as a random parameter of a network layer;
[0174] Compare the survival probability parameter of each network layer with its corresponding random parameter;
[0175] Based on the comparison results of each network layer, determine that the network layer with a survival probability parameter not less than the random parameter is the target network layer, and determine that the network layer with a survival probability parameter less than the random parameter is the optimized network layer.
[0176] In one embodiment, when the processor is used to process generating a relevant question list corresponding to the initial description information and the initial task element component based on the target large language model, the processor is further used to implement:
[0177] Based on the target knowledge graph, obtain the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element components;
[0178] Input the initial description information and its corresponding domain context information, and the initial task element components and their corresponding domain context information into the target large language model to obtain the list of relevant questions.
[0179] In one embodiment, before the processor is used to process obtaining the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element components based on the target knowledge graph, it is further used to implement:
[0180] Collect information related to historical code generation tasks, and segment the related information into text units of each preset unit;
[0181] Based on the target large language model, respectively extract each element component and the combined logical relationship corresponding to each element component in each text unit;
[0182] Based on each element component and each combined logical relationship, construct a homogeneous undirected weighted graph, where the nodes and edges in the homogeneous undirected weighted graph respectively represent each element component and the combined logical relationship corresponding to each element component, and the weights of each edge represent the standardized count of the detected combined logical relationship instances;
[0183] Merge the subgraphs corresponding to different text units with the same element component in the homogeneous undirected weighted graph, and perform hierarchical clustering on the merged homogeneous undirected weighted graph to obtain each community;
[0184] Extract the report-style summary corresponding to each community, and generate the target knowledge graph based on each community and its corresponding report-style summary.
[0185] In one embodiment, after the processor is used to process asking questions about the interactive code generation task to the target user based on the list of relevant questions to obtain the supplementary description information corresponding to the target user, it is further used to implement:
[0186] Count the number of dialogue turns corresponding to the target user, and calculate the perfection degree of the element components of the target code generation task based on the initial task element components and the supplementary task element components of the target code generation task;
[0187] When the number of dialogue turns corresponding to the target user exceeds the preset number threshold or when the perfection degree of the element components exceeds the preset proportion threshold, generate an inquiry stop instruction to terminate the question inquiry of the code generation task corresponding to the target user.
[0188] In one embodiment, after the processor is used to process the combination of the initial task element components, supplementary task element components, and the component combination logic to generate the target code corresponding to the target code generation task, it is further used to implement:
[0189] Obtain the script language type and front-end framework type corresponding to the target code generation task, and generate and display the corresponding front-end interface code based on the target code, the script language type, and the front-end framework type.
[0190] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the code generation methods based on the AI agent provided in the embodiments of the present application.
[0191] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0192] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A code generation method based on AI agent, characterized in that: The code generation method comprises the following steps: Obtaining initial description information of the target code generation task, and generating initial task element components corresponding to the initial description information based on the target AI agent; Based on the target large language model, generate the initial description information and a list of related questions corresponding to the initial task element components, and perform an interactive code generation task question query to the target user based on the list of related questions to obtain supplementary description information corresponding to the target user; Generate a supplementary task element component corresponding to the supplementary description information based on the target AI agent, and determine the component combination logic corresponding to the initial task element component and the supplementary task element component based on the target AI agent; Based on the initial task element components, the supplementary task element components and the component combination logic, the target code corresponding to the target code generation task is generated in combination.
2. The code generation method based on AI agent according to claim 1, characterized in that: Before generating the initial description information and a list of related questions corresponding to the initial task element components based on the target large language model, and performing an interactive code generation task question inquiry to the target user based on the list of related questions to obtain the supplementary description information corresponding to the target user, the method further includes: Initialize the weight parameters and bias parameters of each network layer in the initial large language model, and generate the survival probability parameters of each network layer based on preset rules; Based on the survival probability parameters of each network layer, a target network layer and an optimized network layer are determined in each network layer, wherein the target network layer is a network layer to be trained, and the optimized network layer is a network layer for which training needs to be skipped; Based on the preset training sample set, initial weight parameters and initial bias parameters, the optimization network layer is skipped, and the target network layer is trained until a preset convergence condition is reached, and a trained model is obtained as the target large language model.
3. The code generation method based on AI agent according to claim 2, characterized in that: The determining of the target network layer and optimizing the network layer at each network layer based on the survival probability parameters of each network layer specifically includes: Obtain any value within a preset value range as a random parameter of a network layer; Compare the survival probability parameter of each network layer with its corresponding random parameter; Based on the comparison results of each network layer, the network layer whose survival probability parameter is not less than the random parameter is determined as the target network layer, and the network layer whose survival probability parameter is less than the random parameter is determined as the optimized network layer.
4. The code generation method based on AI agent according to claim 1, characterized in that: The generating of the initial description information and the list of related questions corresponding to the initial task element components based on the target large language model specifically includes: Based on the target knowledge graph, obtain the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element component; The initial description information and its corresponding domain context information, the initial task element components and their corresponding domain context information are input into the target large language model to obtain the related question list.
5. The code generation method based on AI agent according to claim 4, characterized in that: Before acquiring the domain context information corresponding to the initial description information and the domain context information corresponding to the initial task element component based on the target knowledge graph, the method further includes: Collecting relevant information of historical code generation tasks, and dividing the relevant information into text units of various preset units; Based on the target large language model, extracting each element component and the combination logic relationship corresponding to each element component in each text unit; Based on each element component and each combinatorial logical relationship, a homogeneous undirected weighted graph is constructed, wherein the nodes and edges in the homogeneous undirected weighted graph represent each element component and the combinatorial logical relationship corresponding to each element component, respectively, and the weight of each edge represents the standardized count of the detected combinatorial logical relationship instances; Merging subgraphs corresponding to different text units with the same element components in the homogeneous undirected weighted graph, and performing hierarchical clustering on the merged homogeneous undirected weighted graph to obtain various communities; The report-style summaries corresponding to each community are extracted, and the target knowledge graph is generated based on each community and its corresponding report-style summaries.
6. The code generation method based on AI agent according to claim 1, characterized in that: After performing interactive code generation task question query to the target user based on the relevant question list to obtain the supplementary description information corresponding to the target user, the method further includes: Counting the conversation turns corresponding to the target user, and calculating the element component perfection of the target code generation task based on the initial task element component and the supplementary task element component of the target code generation task; When the dialogue round corresponding to the target user exceeds a preset number threshold or the perfection of the element component exceeds a preset ratio threshold, an inquiry stop instruction is generated to terminate the code generation task question inquiry corresponding to the target user.
7. The AI agent-based code generation method according to any one of claims 1 to 6, characterized in that: After the target code corresponding to the target code generation task is generated based on the initial task element component, the supplementary task element component and the component combination logic, the method further includes: The scripting language type and the front-end framework type corresponding to the target code generation task are obtained, and based on the target code, the scripting language type and the front-end framework type, corresponding front-end interface code is generated and displayed.
8. A code generation device based on AI agent, characterized in that: The code generation device based on AI agent includes: An initial component generation module, used to obtain initial description information of a target code generation task, and generate initial task element components corresponding to the initial description information based on a target AI agent; A supplementary information acquisition module, used to generate the initial description information and a list of related questions corresponding to the initial task element components based on the target large language model, and perform an interactive code generation task question inquiry to the target user based on the list of related questions to obtain supplementary description information corresponding to the target user; A component information supplementation module, used to generate a supplementary task element component corresponding to the supplementary description information based on the target AI agent, and determine the component combination logic corresponding to the initial task element component and the supplementary task element component based on the target AI agent; The target code generation module is used to generate the target code corresponding to the target code generation task based on the initial task element components, the supplementary task element components and the component combination logic.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and an AI agent-based code generation program stored in the memory and executable by the processor, wherein when the AI agent-based code generation program is executed by the processor, the steps of the AI agent-based code generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an AI agent-based code generation program, wherein when the AI agent-based code generation program is executed by a processor, the steps of the AI agent-based code generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Code generation method and device, equipment and storage medium
CN120743244A
Task description optimization method and device, equipment, storage medium and product
CN120764579A
Code generation method and computing device
CN121166525A
Method and apparatus for sharing file in association with large language model
KR102956211B1