One-key code generation method and system for an arrangement model
By decoupling neural network models and orchestrating components, building data sets for large-scale training, the problems of computing resource consumption and deployment difficulty in the existing technology are solved, and high-precision and high-efficiency code generation are achieved.
Patent Information
- Application Number
- CN202510413196.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing code generation methods based on large language models have challenges in computing resource consumption and deployment difficulty, and the accuracy and reliability of code generation are insufficient.
By decoupling multiple neural network models, encapsulating them into different components, and simulating the component arrangement method, building a data set for large-model training, and obtaining the trained large-model for code generation in one-click.
This reduces the demand for computing resources, improves the accuracy and efficiency of code generation, and makes the generated code easier to apply and deploy on different edges and ends.
Smart Images

Figure CN119938032B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of code generation, and particularly to a method and system for one-key code generation of an orchestration model. Background Art
[0002] CN118409741A discloses a code generation method based on a large language model, including the following steps: First, receive a requirement document input by a user and a data model layer generated from database fields, and use the received data as source text; form prompt words according to the source text, and the ChatGLM2 model obtains function signatures and outputs text content; form prompt word groups according to the source text, and the WizardCoder model obtains file paths and code content and outputs text content; construct folders and files according to the code paths, write the code content, and obtain the conversion of the target programming language. This invention uses two large models, ChatGLM2 and WizardCoder. ChatGLM2 is used to understand user requirements and generate text, and WizardCoder generates code according to the text. Although the accuracy is improved, it relies on two large models, requires too much computing resources, is difficult to apply on the edge side, and is difficult to deploy.
[0003] CN118210489A discloses a code generation method and system based on a large language model, wherein the code generation method includes: generating a workflow from user requirements described in natural language through the large language model; converting the workflow into a flowchart and performing verification based on the flowchart; generating executable code from the verified workflow through the large language model. This invention depends relatively much on the accuracy of workflow generation. If there are deviations during the workflow generation process, the results of the entire code generation are unreliable. Moreover, this invention divides the process of using the large model into two parts. One part generates a workflow according to user requirements, and the other part generates code according to the workflow, increasing the instability of the large model generation results.
[0004] Therefore, there is an urgent need for a code generation method that requires low computing resources and has high code generation accuracy. Summary of the Invention
[0005] The present invention provides a method and system for one-key code generation of an orchestration model to solve the technical problems mentioned in the background art.
[0006] To achieve the above object, the technical solution of the present invention is realized as follows:
[0007] The present invention provides a method for one-key code generation of an orchestration model, including the following steps:
[0008] S1. Decouple multiple neural network models to obtain multiple single data processing logics respectively, and encapsulate the multiple single data processing logics into different components respectively;
[0009] S2. Simulate different orchestration connection methods for all components, collect orchestration cases in multiple different scenarios, generate text descriptions for each orchestration case and each component, and use the multiple text descriptions to construct a dataset;
[0010] S3. Use the dataset to train a large model, obtain the trained large model, and deploy the trained large model to the device side;
[0011] S4. Orchestrate the required model using all components, and then use the trained large model to generate the code of the required model obtained by orchestration with one key.
[0012] Further, the components at least include data input / output, preprocessing, model selection, hyperparameter setting, deployment framework, deployment platform, and postprocessing.
[0013] Further, among the multiple components, there are components with the same component structure but inconsistent weight parameters.
[0014] Further, the S2 specifically includes the following steps:
[0015] S21. Simulate different orchestration connection methods for all components multiple times, and collect orchestration cases in multiple different scenarios;
[0016] S22. Establish mapping rules from function to text description for each component and each orchestration case;
[0017] S23. Describe each component and each orchestration case in natural language, and unify the sentence structures of all components and all orchestration cases into templates;
[0018] S24. According to the mapping rules and sentence structures of each component, use a program to automatically generate the text description of each component, and then add the dynamic parameters in each component to the text description of each component respectively;
[0019] S25. According to the mapping rules and sentence structures of each orchestration, use a program to automatically generate the text description of each orchestration, and then add the dynamic parameters in all components within each orchestration to the text description of each orchestration respectively;
[0020] S26. Use the text descriptions generated by each component and each orchestration to construct a dataset.
[0021] Further, the S2 also includes the following steps:
[0022] S27. Add label information to each text description in the data set;
[0023] S28. Review the label information of each text description in the data set to improve the accuracy of the samples in the data set.
[0024] Further, the label information in S27 includes at least one or more of task type, programming language used, and framework version.
[0025] Further, the large model is the Jiuge 2B large model.
[0026] Further, S3 specifically includes the following steps:
[0027] S31. Divide the data set into a training set and a validation set according to a preset ratio;
[0028] S32. Set a loss function, use the training set to perform LoRA fine-tuning training on the large model, and during each round of training, calculate the loss value according to the label information and the set loss function. Train for multiple rounds until the loss function converges or the number of training times reaches the set number. Then use the validation set for verification, and select the set of weights with the highest accuracy as the weights of the large model to obtain the trained large model;
[0029] S33. Deploy the trained large model to the device side to achieve one-key code generation.
[0030] Further, S4 specifically includes the following steps:
[0031] S41. Orchestrate the required model using all components, then establish the relationship between the front and rear components using connection lines, and form a complete thought chain based on the relationship between the front and rear components;
[0032] S42. Generate the text description of the complete thought chain for each component in the complete thought chain through their respective mapping rules and sentence structures;
[0033] S43. Input the text description of the complete thought chain obtained in S42 into the trained large model for code generation to obtain the code DM1 for the entire complete process;
[0034] S44. Isolate and split the complete thought chain obtained in S41 according to components into multiple single thought chains; then generate the text description of the single thought chain for each component in the single thought chain through their respective mapping rules and sentence structures;
[0035] S45. Input the text description of the single thought chain into the trained large model one by one in the front-to-back order of the single thought chain to obtain the corresponding code;
[0036] S46. Perform data preprocessing or data space alignment processing on the output result of the previous single thought chain, convert it into a text description, and add it to the input text of the next single thought chain to guide the large model to generate input code that matches the data type and scale of the output result of the previous single thought chain, ensuring that the interfaces between the front and back single thought chains are completely consistent;
[0037] S47. Loop S46 multiple times until the codes output by each single thought chain are obtained, and merge the codes output by each single thought chain to form the code DM2 for completing the processing flow;
[0038] S48. Compare code DM1 with code DM2, identify the parts where the code differences reach the set differences, analyze and run the verification on the parts that reach the set differences, and make a choice on the parts that reach the set differences in code DM1 according to the verification results. After the choice is completed, obtain code DM3;
[0039] S49. Perform a demonstration run on code DM3, display the input and output, and make a final confirmation.
[0040] On the other hand, the present invention also provides a one - key code generation system for an orchestration model, including a computer device, which is programmed or configured to execute the above - mentioned one - key code generation method.
[0041] Advantages of the present invention:
[0042] 1. The present invention discloses a one - key code generation method for an orchestration model. Compared with the prior art, the prior art mainly refers to CN118409741A, a code generation method based on a large - language model. The present invention has low requirements for computing resources. At the same time, the present invention and the codes of the required models orchestrated by the present invention are both easy to be applied and deployed on different edge - side devices, with a wide range of applications. In addition, the present invention is mainly deployed on a model orchestration software platform for one - key code generation of orchestrated models.
[0043] 2. The one - key code generation method disclosed by the present invention not only improves the accuracy of code generation but also significantly reduces the time consumption of model calculation, with higher code generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the one - key code generation method in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many other different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present invention more thorough and comprehensive.
[0046] Referring to Figure 1 , an embodiment of the present application provides a method for one-key generating the code of an orchestration model, including the following steps:
[0047] S1. Decouple multiple neural network models to obtain multiple single data processing logics respectively, and encapsulate the multiple single data processing logics into different components respectively;
[0048] S2. Simulate different orchestration connection methods of all components, collect orchestration cases in multiple different scenarios, generate text descriptions for each orchestration case and each component, and construct a data set using the multiple text descriptions;
[0049] S3. Train a large model using the data set to obtain the trained large model, and deploy the trained large model to the device side;
[0050] S4. Orchestrate the required model using all components, and then use the trained large model to generate the code of the required model obtained by orchestration with one key.
[0051] The present invention has low requirements for computing resources. At the same time, the present invention and the code of the required model orchestrated by the present invention are both easy to be applied and deployed on different edge sides, with a wide range of applications. In addition, the present invention is mainly deployed on the model orchestration software platform for one-key generating the code of the orchestrated model.
[0052] In some embodiments, the components at least include data input / output, preprocessing, model selection, hyperparameter setting, deployment framework, deployment platform, and postprocessing.
[0053] In some embodiments, among the multiple components, there are components with the same component structure but inconsistent weight parameters.
[0054] In some embodiments, S2 specifically includes the following steps:
[0055] S21. Simulate different orchestration connection methods of all components multiple times to collect orchestration cases in multiple different scenarios;
[0056] S22. Establish a mapping rule from function to text description for each component and each orchestration case;
[0057] S23. Describe each component and each orchestration case in natural language, and unify and template the sentence structures of all components and all orchestration cases;
[0058] S24. According to the mapping rules and sentence structures of each component, and use a programming language to automatically generate the text descriptions of each component, and then add the dynamic parameters in each component to the text descriptions of each component respectively;
[0059] S25. According to the mapping rules and sentence structures of each orchestration, and use a programming language to automatically generate the text descriptions of each orchestration, and then add the dynamic parameters in all components within each orchestration to the text descriptions of each orchestration respectively;
[0060] S26. Use the text descriptions generated for each component and each orchestration to construct a dataset.
[0061] In some embodiments, step S2 further includes the following steps:
[0062] S27. Add label information to each text description in the dataset;
[0063] S28. Review the label information of each text description in the dataset to improve the accuracy of the samples in the dataset.
[0064] In some embodiments, the label information in S27 includes at least one or more of task type, programming language used, and framework version.
[0065] In some embodiments, the large model is the Nine Grid 2B large model. The present invention uses the Nine Grid 2B large model for LoRA fine-tuning training. LoRA fine-tuning is an efficient fine-tuning method for large pre-trained models, which allows learning incremental knowledge for specific tasks by introducing low-rank matrices without changing the weights of the original model, helps reduce the number of parameter updates, lower the computational cost and memory occupancy, and at the same time can maintain good performance.
[0066] Before LoRA fine-tuning training, first place the dataset file (.json) in the data folder under the project, and write the dataset file information into data / dataset_info.json. dataset_info.json is mainly used to record the file names of each dataset under data. file_name is the dataset file name, and file_sha1 can be left blank (the content is an empty string, i.e., "file_sha1": "").
[0067] LoRA fine-tuning parameter configuration. Taking single-card training as an example, the following are the descriptions of each parameter:
[0068] # Model Configuration
[0069] ## Model
[0070] - `model_name_or_path`: The path points to the location where the pre-trained model or dialogue model is stored.
[0071] ## Training Method
[0072] - `stage`: The training stage or mode, set to `sft` (Supervised Fine-Tuning) here, which means fine-tuning.
[0073] - `do_train`: Whether to execute the training process, `true` means yes.
[0074] - `finetuning_type`: The fine-tuning type, using `lora` (Low-Rank Adaptation) here, which is a lightweight fine-tuning method that only adjusts the weights of some layers.
[0075] - `lora_target`: Specify the specific model layers to which LoRA adaptation is applied. `q_proj`, `v_proj` usually correspond to the query (Query) and value (Value) projection layers in the attention mechanism of the Transformer model, or it can also be all layers for lora adaptation, `all`.
[0076] ## Dataset Configuration
[0077] - `dataset`: A list of datasets, separated by commas, such as `alpaca_en`, `alpaca_zh`, using the key names of the corresponding dataset dictionaries in data / dataset_info.json.
[0078] - `template`: The template type, using `9g` here, which is a data processing template for the structure of the model to be trained.
[0079] - `cutoff_len`: The maximum length limit of the input sequence. Inputs longer than this length will be truncated.
[0080] - `max_samples`: The maximum number of samples, which limits the total number of samples loaded from the dataset.
[0081] - `val_size`: The proportion of the validation set in the total dataset, which is 10% here.
[0082] - `overwrite_cache`: If `true`, the cache for loaded data will be overwritten before each run.
[0083] - `preprocessing_num_workers`: The number of parallel worker processes used for data preprocessing to improve data loading efficiency.
[0084] ## Output Configuration
[0085] - `output_dir`: The save directory for model training outputs, including logs, checkpoints, etc.
[0086] - `logging_steps`: The step interval for logging during training.
[0087] - `save_steps`: The step interval for saving model checkpoints.
[0088] - `plot_loss`: Whether to plot and save the loss curve during training.
[0089] - `overwrite_output_dir`: If `true`, allows the training process to overwrite existing content in the output directory.
[0090] ## Training Parameters
[0091] - `per_device_train_batch_size`: The training batch size per device.
[0092] - `gradient_accumulation_steps`: The number of steps to accumulate gradients, used to simulate larger batch training and reduce memory usage.
[0093] - `learning_rate`: The initial learning rate.
[0094] - `num_train_epochs`: The total number of training epochs.
[0095] - `lr_scheduler_type`: The type of learning rate scheduler, here it is the cosine annealing scheduler.
[0096] - `warmup_steps`: The number of steps in the learning rate warm-up phase, gradually increasing to the initial learning rate.
[0097] - `fp16`: Whether to use mixed-precision training (half-precision floating-point numbers), which can accelerate training and reduce memory consumption.
[0098] ## Evaluation Settings
[0099] - `per_device_eval_batch_size`: The batch size on each device during the evaluation phase.
[0100] - `evaluation_strategy`: The evaluation strategy, where `steps` means evaluating at specified steps.
[0101] - `eval_steps`: The step interval for performing evaluations.
[0102] During the LoRA fine-tuning process, the model will automatically evaluate on the validation set, calculate the validation set loss (i.e., the loss function), and the plot_loss parameter can specify the training loss curve.
[0103] In some embodiments, S3 specifically includes the following steps:
[0104] S31. Divide the dataset into a training set and a validation set according to a preset ratio;
[0105] S32. Set the loss function, perform LoRA fine-tuning training on the large model using the training set, and during each round of training, calculate the loss value based on the label information and the set loss function. Train for multiple rounds until the loss function converges or the number of training times reaches the set number, and then use the validation set for verification. Select the set of weights with the highest accuracy as the weights of the large model to obtain the trained large model;
[0106] S33. Deploy the trained large model to the device side to achieve one-key code generation.
[0107] In some embodiments, S4 specifically includes the following steps:
[0108] S41. Orchestrate the required model using all components, then establish the relationship between the front and rear components using connection lines, and form a complete thought chain based on the relationship between the front and rear components;
[0109] S42. Generate the text description of the complete thought chain for each component within the complete thought chain through their respective mapping rules and sentence structures;
[0110] S43. Input the text description of the complete thought chain obtained in S42 into the trained large model for code generation to obtain the code DM1 for the entire complete process;
[0111] S44. Isolate and split the complete thought chain obtained in S41 according to components into multiple single thought chains; then generate text descriptions of the single thought chains for each component within the single thought chain through their respective mapping rules and sentence structures;
[0112] S45. Input the text descriptions of the single thought chains into the trained large model one by one in the order before and after the single thought chains to obtain the corresponding code;
[0113] S46. Perform data preprocessing or data space alignment processing on the output result of the previous single thought chain, and convert it into a text description and add it to the input text of the next single thought chain to guide the large model to generate input code that matches the data type and scale of the output result of the previous single thought chain, ensuring that the interfaces between the front and back single thought chains are completely consistent;
[0114] S47. Loop S46 multiple times until the codes output by each single thought chain are obtained, and merge the codes output by each single thought chain to form the code DM2 for completing the processing flow;
[0115] S48. Compare code DM1 with code DM2, identify the parts where the code differences reach the set differences, analyze and run the verification on the parts that reach the set differences, and make a choice on the parts that reach the set differences in code DM1 according to the verification results. After the choice is completed, obtain code DM3;
[0116] S49. Demonstrate and run code DM3, display the input and output, and make a final confirmation.
[0117] The code one - key generation method disclosed by the present invention not only improves the accuracy of code generation, but also significantly reduces the time consumption of model calculation, and the code generation efficiency is higher.
[0118] On the other hand, the present invention also provides a code one - key generation system for an orchestration model, including a computer device, and the computer device is programmed or configured to execute the above - mentioned code one - key generation method.
[0119] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Moreover, the technical solutions between various embodiments of the present invention can be combined with each other, but it must be based on the fact that those skilled in the art can implement it. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.
Claims
1. A one-click code generation method for an orchestration model, characterized in that: The steps include: S1. Decouple multiple neural network models to obtain multiple single data processing logics, and encapsulate the multiple single data processing logics into different components; S2. Simulate different orchestration connection modes for all components, collect orchestration cases in multiple different scenarios, generate text descriptions for each orchestration case and each component, and construct a data set using multiple text descriptions; S3. Use the data set to train the large model to obtain the trained large model, and deploy the trained large model to the device end; S4. Use all components to orchestrate the required model, and then use the trained large model to generate code for the orchestrated required model with one click; The S4 specifically includes the following steps: S41. Arrange the required model using all components, then use connecting lines to establish connections between the previous and next components, and form a complete thinking chain based on the connections between the previous and next components; S42, generating a text description of the complete thought chain by using respective mapping rules and sentence structures for each component in the complete thought chain; S43, input the text description of the complete thought chain obtained in S42 into the trained large model for code generation, and obtain the code DM1 of the entire complete process; S44, isolating and splitting the complete thought chain obtained in S41 according to components into multiple single thought chains; Then, each component in a single thought chain is used to generate a text description of the single thought chain through its own mapping rules and sentence structure; S45, according to the sequence of the single thought chain, input the text description of the single thought chain into the trained large model one by one to obtain the corresponding code; S46, perform data preprocessing or data space alignment processing on the output result of the previous single thinking chain, convert it into a text description and add it to the input text of the next single thinking chain to guide the large model to generate input code that matches the data type and scale of the output result of the previous single thinking chain, ensuring that the interface between the previous and next single thinking chains is completely consistent; S47, loop S46 for multiple times until the codes output by each single thought chain are obtained, and the codes output by each single thought chain are merged to form the code DM2 that completes the processing flow; S48, comparing the code DM1 with the code DM2, identifying the part of the code whose difference reaches the set difference, analyzing and running verification on the part that reaches the set difference, selecting and discarding the part of the code DM1 that reaches the set difference according to the verification result, and obtaining the code DM3 after the selection and discarding; S49. Perform a demonstration run on code DM3, display the input and output, and make a final confirmation.
2. The one-click code generation method for an orchestration model according to claim 1, characterized in that: The components include at least data input / output, preprocessing, model selection, hyperparameter setting, deployment framework, deployment platform, and post-processing.
3. The one-click code generation method of the arrangement model according to claim 1 is characterized in that: The multiple components include components with the same component structure but inconsistent weight parameters.
4. The one-click code generation method for an orchestration model according to claim 1, characterized in that: The S2 specifically includes the following steps: S21. Simulate multiple times different orchestration connection modes for all components and collect orchestration cases in multiple different scenarios. S22, establishing a mapping rule from function to text description for each component and each orchestration case; S23. Describe each component and each orchestration case in natural language, and unify the sentence structure of all components and all orchestration cases; S24, according to the mapping rules and sentence structure of each component, and using a program to automatically generate a text description of each component, and then adding the dynamic parameters in each component to the text description of each component; S25, according to the mapping rules and sentence structure of each arrangement, and using a program to automatically generate a text description of each arrangement, and then adding the dynamic parameters of all components in each arrangement to the text description of each arrangement; S26. Build a dataset using the text descriptions generated by each component and each orchestration.
5. The one-click code generation method for an orchestration model according to claim 4, characterized in that: The S2 further comprises the following steps: S27, adding label information to each text description in the data set; S28. Review the label information of each text description in the dataset to improve the accuracy of the samples in the dataset.
6. The one-key code generation method of the arrangement model according to claim 5, characterized in that: The tag information in S27 at least includes one or more of the task type, the used programming language, and the framework version.
7. The one-click code generation method for an orchestration model according to claim 5, characterized in that: The large model is a nine-grid 2B large model.
8. The one-click code generation method for an orchestration model according to claim 5, characterized in that: The S3 specifically includes the following steps: S31, dividing the data set into a training set and a validation set according to a preset ratio; S32, set the loss function, use the training set to perform LoRA fine-tuning training on the large model, and in each round of training, calculate the loss value according to the label information and the set loss function, train multiple rounds until the loss function converges, or the number of training times reaches the set number, and then use the validation set for verification, select the set of weights with the highest accuracy as the weights of the large model, and obtain the trained large model; S33. Deploy the trained large model to the device to achieve one-click code generation.
9. A one-click code generation system for an arrangement model, comprising a computer device, characterized in that: The computer device is programmed or configured to execute the one-key code generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Comprehensive management process design system and method based on AI code interpretation
CN119376711A
Intelligent model arrangement method and device, equipment and storage medium
CN119416862A