Automatic code programming method
By combining rules and deep learning models to analyze natural language requirements and generate high-quality code, the problem of insufficient flexibility and accuracy in the existing technology is solved, and efficient and reliable automatic programming is achieved to adapt to complex business logic and multi-platform development.
Patent Information
- Application Number
- CN202510387539.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing automated programming technologies have poor flexibility and insufficient accuracy, and are unable to effectively deal with complex business logic, resulting in inaccurate and time-consuming code generation.
The semantic analysis model based on rules and deep learning is used to analyze natural language requirements, combine history and context information, generate code templates, and evaluate code quality through performance monitoring and automated testing.
It improves the accuracy and efficiency of code generation, reduces human errors, reduces programming difficulty, supports multiple programming languages and platforms, improves the reliability and consistency of code, and adapts to the needs of rapid iterative development.
Smart Images

Figure CN120335783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software development, and more particularly, to a method for automatic code programming. Background Art
[0002] With the continuous development of software development, the work of programmers has gradually become more complex and onerous. The traditional manual programming method is time-consuming and error-prone. Especially when facing large-scale and complex systems, it becomes increasingly difficult to write high-quality code. To improve development efficiency and reduce human errors, researchers have proposed automated programming techniques, including template-based code generation, rule-driven programming, and intelligent programming assistants. With the rapid development of artificial intelligence and automation technologies, the software development field is facing unprecedented challenges and opportunities. Traditional programming methods rely on manual code writing, which is not only time-consuming and laborious, but also prone to errors due to human factors. In recent years, automated programming techniques have gradually emerged, aiming to reduce manual intervention and improve code generation efficiency and quality through intelligent algorithms and machine learning technologies.
[0003] However, existing automated programming techniques still have many limitations, such as insufficient accuracy of code generation, poor adaptability, and inability to handle complex business logics. There is an urgent need for a more efficient and intelligent solution. Although existing technologies have proposed some automated code generation solutions, these solutions usually have problems such as poor flexibility, limited functionality, or inability to well understand complex requirements. Therefore, developing a system that can accurately generate high-quality code according to natural language requirements has broad application prospects.
[0004] In view of the above problems, the present invention proposes a solution. Summary of the Invention
[0005] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for automatic code programming, by developing a system that can accurately generate high-quality code according to natural language requirements, to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for automatic code programming, comprising the following steps:
[0008] The user inputs a requirement description of the code function in the form of natural language;
[0009] Use a rule-based semantic analysis model and a deep learning semantic analysis model to perform semantic analysis on the user input requirements, and fuse the analysis results of different models to obtain a comprehensive semantic vector;
[0010] Combine the historical requirement records and the context information of the current requirements to correct the fused comprehensive semantic vector, and identify key information based on the corrected comprehensive semantic vector;
[0011] Pre-construct a code template library, search for templates in the template library according to the key information for matching, and generate code based on the replaced template;
[0012] Real-time monitor the performance metrics during the code running process, and calculate the code fitness value according to the monitoring results of the performance metrics;
[0013] Automatically test the code to obtain the code reliability dimension, and evaluate the code automatic programming effect based on fuzzy inference in combination with the code fitness value.
[0014] In a preferred embodiment, the requirements for the code function input in the form of natural language include:
[0015] Describe the basic function requirements and point out the key operations;
[0016] And describe in detail the input data types and formats, as well as the expected outputs;
[0017] And describe the performance requirements, error handling methods or interaction information with other systems.
[0018] In a preferred embodiment, the specific steps for semantic analysis of the requirements input by the rule-based semantic analysis model are as follows:
[0019] Decompose the input user text into individual words or phrases, perform part-of-speech tagging on the words or phrases, and remove stop words;
[0020] Analyze the dependency relationships between the words in the sentence, and construct a syntactic tree according to the grammar rules to represent the structure of the words inside the sentence;
[0021] According to the rules and dictionary mapping, perform reasoning on different words, and obtain a semantic vector based on the rule-based semantic representation.
[0022] In a preferred embodiment, the specific steps for semantic analysis of the requirements input by the deep learning semantic analysis model are as follows:
[0023] Perform text cleaning on the requirements input by the user, including word segmentation and stop word removal;
[0024] Convert each word into a vector with a fixed dimension through a pre-trained word vector model;
[0025] Based on the BERT model, calculate the dynamic vector of each word according to the context of the sentence;
[0026] Train a deep learning model using labeled data and generate semantic vectors based on the output layer of the model.
[0027] In a preferred embodiment, the process of fusing the analysis results of different models to obtain a comprehensive semantic vector is as follows:
[0028] Fuse the analysis results of the rule-based semantic analysis model and the deep learning semantic analysis model, and calculate the comprehensive semantic vector by weighted average. The specific calculation formula is as follows: In the formula, is the comprehensive semantic vector, is the semantic vector obtained from the rule-based semantic analysis model, is the semantic vector of the deep learning semantic analysis model, and α is the weight coefficient, with a value range of [0, 1].
[0029] In a preferred embodiment, the process of correcting the fused comprehensive semantic vector by combining historical demand records and the context information of the current demand is as follows:
[0030] Obtain the context demand record according to the context information of the current demand, and obtain the semantic vector of the context demand according to the context demand record;
[0031] Obtain the semantic vector of the historical demand according to the historical demand record;
[0032] Correct the fused comprehensive semantic vector according to the semantic vectors of the context demand and the historical demand.
[0033] In a preferred embodiment, the process of identifying key information according to the corrected comprehensive semantic vector is as follows:
[0034] Calculate the sum of the dimensions of the corrected comprehensive semantic vector;
[0035] And compare and analyze the sum of the dimensions of the corrected comprehensive semantic vector with a preset dimension threshold;
[0036] If the sum of the dimensions of the corrected comprehensive semantic vector is greater than the preset dimension threshold, it is considered that the information corresponding to the corrected comprehensive semantic vector is key information.
[0037] In a preferred embodiment, the process of generating code according to the replaced template is as follows:
[0038] Pre-construct a template library containing various common code functions. Each template corresponds to a specific code function and contains replaceable placeholders;
[0039] According to the identified key information, search for the most matching code template in the template library;
[0040] If several matching templates are found, replace the placeholders in the template with the highest matching degree to the key information with actual parameters;
[0041] If one matching template is found, replace the placeholders in the template with actual parameters;
[0042] If there is no exactly matching template in the template library, use a machine learning model for code generation.
[0043] In a preferred embodiment, according to the monitoring results of performance metrics, the process of calculating the code fitness value is as follows:
[0044] The performance metrics include execution time and memory occupancy percentage;
[0045] Obtain the time taken for the program to execute from start to end to get the execution time;
[0046] Obtain the ratio of the memory used by the program to the total memory of the computer to get the memory occupancy percentage;
[0047] Calculate the fitness value of the code based on the execution time and memory occupancy percentage in combination with the baseline execution time and baseline memory occupancy percentage.
[0048] In a preferred embodiment, the process of evaluating the automatic programming effect of the code based on fuzzy inference is as follows:
[0049] Obtain the number of times of automated testing of the code, as well as the number of times the code passes the test and the maximum number of consecutive passes of the test;
[0050] Divide the number of times the code passes the test by the number of times of automated testing of the code to get the code automatic programming accuracy coefficient;
[0051] Divide the maximum number of consecutive passes of the test by the number of times of automated testing of the code to get the code automatic programming stability coefficient;
[0052] Obtain the code reliability dimension through weighted synthesis of the code automatic programming accuracy coefficient and stability coefficient;
[0053] Define the code reliability dimension and the code fitness value as input variables and divide them into different fuzzy sets respectively;
[0054] Define the automatic programming effect of the code as the output variable and divide it into a fuzzy set;
[0055] Formulate fuzzy rules to describe the influence of input variables on output variables;
[0056] Perform fuzzy inference according to the fuzzy rules to determine the automatic programming effect of the code.
[0057] Technical effects and advantages of a method for automatic code programming according to the present invention:
[0058] 1. Through automatic code generation, the present invention can significantly reduce the development time, especially for highly repetitive or structured code parts. Programmers can focus more on the core business logic. Since the code generation process is based on an automated system, common errors in traditional manual programming, such as syntax errors and logical errors, can be avoided. The automatically generated code conforms to the best practices of programming languages, is efficient and maintainable, and is easy to expand and optimize. Non-professional users or beginners can also generate high-quality code by inputting requirements in natural language, reducing the programming threshold. The automatic programming method often follows unified code specifications and templates, and the generated code has high consistency and standardization, reducing problems such as syntax errors and logical loopholes caused by human factors, and improving the stability and reliability of the code. The generated code has a clear structure and distinct levels, and the division of each module and function is relatively clear, facilitating subsequent modification, expansion, and maintenance of the code by developers, and reducing the maintenance cost and difficulty. For some fields with high technical thresholds or complex programming tasks, the code automatic programming method enables people without profound professional knowledge to participate in the development, reduces the programming difficulty, and promotes the popularization and application of related technologies.
[0059] 2. By automatically generating a large amount of basic code and even complex logic code, the present invention greatly reduces the time and workload of manual code writing, enabling developers to focus more on the core business logic and innovation, and significantly shortening the project development cycle. In the case of continuous change and iteration of product requirements, it can quickly generate or modify code according to new requirements and specifications, meeting the requirements of rapid iterative development, and enabling enterprises to respond more flexibly to market changes. It can automatically complete some tedious and repetitive code writing tasks, such as data access layer code and interface layout code, avoiding developers from repeatedly writing similar code, and improving the interest and creativity of the work. The best practices in the industry and common programming patterns are solidified in the automatic programming tool or method, realizing the reuse and inheritance of knowledge. New projects can quickly inherit the experience and achievements of previous projects. It reduces the need for a large number of professional programmers. Especially for some regular and standardized development tasks, through code automatic programming, fewer people can complete them, reducing the labor cost of software development. When dealing with a large amount of data and complex business logic, it can ensure the accuracy and consistency of the code, avoiding neglect and inconsistency that may occur in manual writing, and improving the quality and stability of the entire system. Many code automatic programming methods can support multiple operating system platforms and programming languages, facilitating developers to make flexible choices according to project requirements, and improving the portability and generality of the code. Description of the Drawings
[0060] Figure 1 This is a schematic structural diagram of a method for automatic code programming of the present invention. Specific embodiments
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1, Figure 1 A method for automatic code programming of the present invention is given.
[0063] S10, the user inputs a requirement description of the code function in the form of natural language;
[0064] The input of the requirement of the code function in the form of natural language includes:
[0065] Describe the basic function requirements and point out the key operations;
[0066] And describe in detail the data type and format of the input, as well as the expected output;
[0067] And describe the performance requirements, error handling methods or information for interacting with other systems.
[0068] When the user inputs the requirement description of the code function in the form of natural language, it has the following key effects on automatic code programming:
[0069] Reduce communication costs: In the traditional development process, it usually requires multiple communications between developers and users to clarify requirements. Through natural language input, users can directly express their requirements, and developers or automated tools can understand them more easily, reducing misunderstandings and unnecessary communication costs.
[0070] For example, the user may describe: "I need a Python program that can calculate the average value based on a given list." Such a description is more concise than a highly technical requirement (such as: "Please implement a function to calculate the average of a list"), which can help developers quickly understand the requirements.
[0071] Improve development efficiency: After using natural language to describe requirements, code generation tools based on AI or automated systems (such as GPT series models) can quickly parse these requirements and convert them into corresponding code. This not only greatly shortens the development time but also improves development efficiency, especially for simple tasks or common programming problems.
[0072] For example, the user only needs to describe: "I need a sorting algorithm that can sort an array in ascending order." The system can directly generate the corresponding sorting code.
[0073] Help non-professional developers: Not every user understands programming. Many users with non-technical backgrounds also hope to automate some tasks or develop functions. Natural language input enables these non-professional developers to easily describe the functions they want without having to learn a programming language. The AI system can generate the required code based on these requirements, reducing the technical threshold.
[0074] For example, the user might say: "I need a Python script to read my file and count the occurrences of each word in it." This requirement is clear to non-programmers, and the AI can generate code based on this.
[0075] Improve code quality and consistency: Automated code generation tools can follow best practices and common coding standards, reducing human errors and inconsistencies when generating code. In addition, the automatic programming system can optimize performance or select the best algorithm based on requirements analysis, further improving code quality.
[0076] For example, when the user enters the requirement "Please write a function to determine whether a given string is a palindrome", the automated tool will generate a function that conforms to Python programming specifications and ensures it has optimal performance.
[0077] Facilitate rapid iteration and prototyping: Using natural language input can help achieve rapid prototyping or make rapid iterations. Users can change the functionality of the code by modifying the requirement description, and the automated tool can provide a quick response based on the new requirements.
[0078] For example, if an initial requirement is "Create a simple web form", but later it is modified to "Add data validation and display the submission result", the automated tool can quickly modify the generated code according to the changed description.
[0079] Support multiple programming languages and frameworks: Users can describe their requirements in natural language, and the automated tool can select the appropriate programming language, library, and framework based on the requirements to generate code. Users don't need to worry about the underlying implementation or technology stack and can focus on describing the functionality.
[0080] For example, the user only needs to describe: "I want to create a simple user registration API using Flask", and the system can automatically generate the code for the Flask framework.
[0081] Promote code customization: Users can obtain highly customized code by refining natural language descriptions. With the advancement of natural language processing (NLP) technology, automated tools can parse complex requirement descriptions and generate personalized code that meets the requirements.
[0082] For example, when the user describes: "I need a program that receives an Excel file, groups the data according to different conditions, and outputs it to another Excel file", such customized requirements can generate code efficiently through automated tools.
[0083] By inputting requirements in natural language, automatic programming not only makes code generation more convenient and efficient, but also reduces the technical barrier between users and developers, promoting diverse and personalized development experiences. It can help non-technical personnel express their requirements through intuitive language, greatly improving development efficiency, reducing the chance of errors, and promoting rapid innovation and prototyping.
[0084] S20, Use a rule-based semantic analysis model and a deep learning semantic analysis model to perform semantic analysis on the requirements input by the user, and fuse the analysis results of different models to obtain a comprehensive semantic vector;
[0085] The specific steps for the rule-based semantic analysis model to perform semantic analysis on the user input requirements are as follows:
[0086] Decompose the input user text into individual words or phrases, perform part-of-speech tagging on the words or phrases, and remove stop words;
[0087] Analyze the dependency relationships between the words in the sentence, and construct a syntactic tree according to the grammar rules to represent the structure of the internal words in the sentence;
[0088] According to rules and dictionary mapping, reason about different words, and obtain a semantic vector according to the rule-based semantic representation.
[0089] It should be noted that for part-of-speech tagging, the natural language processing tool NLTK is used with the Penn Treebank part-of-speech tagging set to quickly perform part-of-speech tagging on English texts. When performing part-of-speech tagging on Chinese texts, the Harbin Institute of Technology LTP toolkit can be used. This toolkit supports a variety of natural language processing tasks, and with its Python interface, part-of-speech tagging of Chinese texts can be achieved. Dependency parsing uses a neural network-based dependency parsing model, such as the BiLSTM-CRF model in the AllenNLP library, and the training data uses a dataset in CoNLL format. For Chinese dependency parsing, the LTP toolkit can also be used, and its provided dependency syntactic analysis function can analyze the dependency relationships between words in Chinese sentences. According to the results of dependency parsing, a syntactic tree is constructed. Using the Python NetworkX library in combination with the results of dependency parsing, a graphical syntactic tree is constructed. A knowledge graph is introduced to associate words with concepts in the graph, enhancing the reasoning ability. A knowledge graph is built using the Neo4j graph database and operated using the py2neo library. The example steps are as follows: Data import: Organize domain-related knowledge into triple form and import it into the Neo4j database; Lexical reasoning: When reasoning about words in the text, query the knowledge graph through py2neo to obtain relevant concepts and relationships to assist the reasoning process.
[0090] The specific steps for semantic analysis based on the user input requirements using a deep learning semantic analysis model are as follows:
[0091] Perform text cleaning on the user input requirements, including word segmentation and removal of stop words;
[0092] Convert each word into a vector of a fixed dimension through a pre-trained word vector model;
[0093] Based on the BERT model, calculate the dynamic vector of each word according to the context of the sentence;
[0094] Use the labeled data to train the deep learning model and generate a semantic vector according to the output layer of the model.
[0095] The process of fusing the analysis results of different models to obtain a comprehensive semantic vector is as follows:
[0096] Fuse the analysis results of the rule-based semantic analysis model and the deep learning semantic analysis model, and calculate the comprehensive semantic vector using the weighted average method. The specific calculation formula is as follows: In the formula, is the comprehensive semantic vector, is the semantic vector obtained from the rule-based semantic analysis model, is the semantic vector of the deep learning semantic analysis model, and α is the weight coefficient with a value range of [0,1].
[0097] It should be noted that a large number of user requirement texts with annotations and their corresponding accurate semantic vectors are collected. Taking the error between the comprehensive semantic vector after model fusion and the annotated semantic vector as the optimization goal, the gradient descent algorithm is used to train the optimal weight coefficients. The specific process is as follows: First, simulated rule-based semantic vectors, deep learning semantic vectors, and accurate semantic vectors are generated. Then, two weight coefficients, weight_rule and weight_dl, are defined and set as trainable parameters. Next, the mean square error loss function and the stochastic gradient descent optimizer are defined. In the training loop, the comprehensive semantic vector is calculated each time, and then the loss is calculated. The gradients are calculated through backpropagation and the weight coefficients are updated until the specified number of training times is reached. Finally, the optimal weight coefficients obtained through training are output.
[0098] S30. Combine the historical requirement records and the context information of the current requirement to correct the fused comprehensive semantic vector, and identify the key information according to the corrected comprehensive semantic vector;
[0099] Obtain the context requirement records according to the context information of the current requirement, and obtain the semantic vectors of the context requirements according to the context requirement records;
[0100] Obtain the semantic vectors of the historical requirements according to the historical requirement records;
[0101] Correct the fused comprehensive semantic vector according to the semantic vectors of the context requirements and the historical requirements. The specific calculation formula is as follows: In the formula, is the corrected comprehensive semantic vector, is the comprehensive semantic vector, is the semantic vector of the context requirement, is the semantic vector of the historical requirement, β is the context influence coefficient, and γ is the historical requirement influence coefficient.
[0102] The process of identifying the key information according to the corrected comprehensive semantic vector is as follows:
[0103] Calculate the sum of the dimensions of the corrected comprehensive semantic vector;
[0104] And compare and analyze the sum of the dimensions of the corrected comprehensive semantic vector with the preset dimension threshold;
[0105] If the sum of the dimensions of the corrected comprehensive semantic vector is greater than the preset dimension threshold, it is considered that the information corresponding to the corrected comprehensive semantic vector is the key information.
[0106] S40. Pre-construct a code template library, search for templates in the template library for matching according to the key information, and generate code according to the replaced template;
[0107] Pre - construct a template library containing various common code functions. Each template corresponds to a specific code function and contains some replaceable placeholders;
[0108] According to the identified key information, search for the most matching code template in the template library;
[0109] If several matching templates are found, replace the placeholders in the template with the highest matching degree to the key information with actual parameters;
[0110] If one matching template is found, replace the placeholders in the template with actual parameters;
[0111] If there is no completely matching template in the template library, use a machine learning model for code generation.
[0112] S50, Monitor the performance metrics during the code running process in real - time. According to the monitoring results of the performance metrics, calculate the code fitness value;
[0113] The performance metrics include execution time and memory occupancy percentage;
[0114] Obtain the time taken for the program to execute from start to end to get the execution time;
[0115] Obtain the ratio of the memory used by the program to the total memory of the computer to get the memory occupancy percentage;
[0116] Calculate the code fitness value based on the execution time and memory occupancy percentage combined with the baseline execution time and baseline memory occupancy percentage. The specific calculation formula is as follows: In the formula, F is the code fitness value, w1 is the execution - time weight factor, T is the execution time, T b is the baseline execution time, w2 is the memory - occupancy - percentage weight factor, M is the memory occupancy percentage, M b is the baseline memory occupancy percentage.
[0117] S60, Automatically test the code to obtain the code reliability dimension, and evaluate the code automatic programming effect based on fuzzy inference in combination with the code fitness value.
[0118] Obtain the number of times of automatically testing the code, the number of times the code passes the test, and the maximum number of consecutive passes of the test;
[0119] Divide the number of times the code passes the test by the number of times of automatically testing the code to obtain the code automatic programming accuracy coefficient;
[0120] Divide the maximum number of consecutive passes in the test by the number of times the code is automatically tested to obtain the code automatic programming stability coefficient;
[0121] Based on the code automatic programming accuracy coefficient and the stability coefficient, the code reliability dimension is obtained by weighted synthesis.
[0122] The process of evaluating the code automatic programming effect based on fuzzy inference is as follows:
[0123] Step C1, Define the code reliability dimension and the code fitness value as input variables, and divide them into different fuzzy sets respectively.
[0124] For example, "Low", "Medium", "High" for the code reliability dimension, and "Low", "Medium", "High" for the code fitness value.
[0125] Step C2, Define the code automatic programming effect as the output variable, and divide it into a fuzzy set. For example, "Low", "High" for the code automatic programming effect.
[0126] Step C3, Formulate a set of fuzzy rules to describe the influence of different input variables on the output variable. The definition of the rules can be based on professional knowledge or obtained through data analysis and experiments. For example:
[0127] Mark the code reliability dimension as L, the code fitness value as F, and the code automatic programming effect as P, then the following can be defined
[0128] Rule 1: If (L is High) and (F is High), then (P is High)
[0129] Rule 2: If (L is Low) and (F is Low), then (P is Low) ...
[0131] Step C4, Perform fuzzy inference according to the fuzzy rules to determine the code automatic programming effect.
[0132] It should be noted that the division of the fuzzy set can be adjusted according to the actual situation. For example, although three fuzzy sets are used as examples in this embodiment, in fact, the code reliability dimension, the code fitness value, and the code automatic programming effect can be divided into more than three sets to facilitate better and more accurate identification.
[0133] Furthermore, for the judgment of the high or low of the code reliability dimension and the code fitness value, the threshold can be set according to the actual situation for judgment; when the code reliability dimension is higher than 80%, it is marked as "High", and when the code fitness value is higher than 75%, it is marked as "High", etc., which will not be elaborated here.
[0134] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0135] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.
[0136] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0137] In addition, the functional modules in each embodiment of this application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0138] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
[0139] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for automatic code programming, characterized in that, It includes the following steps: The user inputs a requirement description of the code function in the form of natural language; Use a rule-based semantic analysis model and a deep learning semantic analysis model to perform semantic analysis on the requirements input by the user, and fuse the analysis results of different models to obtain a comprehensive semantic vector; Combined with the historical requirement records and the context information of the current requirement, correct the fused comprehensive semantic vector, and identify key information according to the corrected comprehensive semantic vector; Pre-construct a code template library, search for templates in the template library for matching according to the key information, and generate code according to the replaced template; Real-time monitor the performance metrics during the code running process, and calculate the code fitness value according to the monitoring results of the performance metrics; Automatically test the code to obtain the code reliability dimension, and evaluate the code automatic programming effect based on fuzzy inference combined with the code fitness value.
2. The method for automatic code programming according to claim 1, characterized in that, The input of the requirement for the code function in the form of natural language includes: Describe the basic function requirements and point out the key operations; And describe in detail the input data types and formats, as well as the expected outputs; And describe the performance requirements, error handling methods or information for interacting with other systems.
3. A method for automatic code programming according to claim 2, characterized in that, The specific steps for performing semantic analysis on the requirements input by the user by the rule-based semantic analysis model are as follows: Decompose the input user text into individual words or phrases, perform part-of-speech tagging on the words or phrases, and remove stop words; Analyze the dependency relationships between the words in the sentence, and construct a syntactic tree according to the grammar rules to represent the structure of the words inside the sentence; According to the rules and dictionary mappings, perform reasoning on different words, and obtain semantic vectors according to the rule-based semantic representation.
4. A method for automatic code programming according to claim 3, characterized in that, The specific steps for performing semantic analysis on the requirements input by the user by the deep learning semantic analysis model are as follows: Perform text cleaning on the requirements input by the user, including word segmentation and stop word removal; Convert each word into a vector with a fixed dimension through a pre-trained word vector model; Based on the BERT model, calculate the dynamic vector of each word according to the context of the sentence; Use the labeled data to train the deep learning model, and generate semantic vectors according to the output layer of the model.
5. A method for automatic code programming according to claim 4, characterized in that, The specific process of fusing the analysis results of different models to obtain a comprehensive semantic vector is as follows: Fuse the analysis results of the rule-based semantic analysis model and the deep learning semantic analysis model, and calculate the comprehensive semantic vector by weighted average. The specific calculation formula is as follows: In the formula, is the comprehensive semantic vector, is the semantic vector obtained from the rule-based semantic analysis model, is the semantic vector of the deep learning semantic analysis model, and α is the weight coefficient, whose value range is [0, 1].
6. A method for automatic code programming according to claim 5, characterized in that, The process of correcting the fused comprehensive semantic vector combined with the historical requirement records and the context information of the current requirement is as follows: Obtain the context requirement record according to the context information of the current requirement, and obtain the semantic vector of the context requirement according to the context requirement record; Obtain the semantic vector of the historical requirement according to the historical requirement record; The fused comprehensive semantic vector is corrected according to the semantic vectors of the context requirements and historical requirements, and the specific calculation formula is as follows: In the formula, is the corrected comprehensive semantic vector, is the comprehensive semantic vector, is the semantic vector of the context requirements, is the semantic vector of the historical requirements, β is the context influence coefficient, and γ is the historical requirements influence coefficient.
7. A method for automatic code programming according to claim 6, characterized in that The process of identifying key information according to the corrected comprehensive semantic vector is as follows: Calculate the sum of the dimensions of the corrected comprehensive semantic vector; And compare and analyze the sum of the dimensions of the corrected comprehensive semantic vector with a preset dimension threshold; If the sum of the dimensions of the corrected comprehensive semantic vector is greater than the preset dimension threshold, it is considered that the information corresponding to the corrected comprehensive semantic vector is key information.
8. A method for automatic code programming according to claim 7, characterized in that, The process of generating code according to the replaced template is as follows: Pre-construct a template library containing various common code functions, each template corresponding to a specific code function and containing replaceable placeholders; Find the most matching code template in the template library according to the identified key information; If several matching templates are found, replace the placeholders in the template with the highest matching degree to the key information with actual parameters; If one matching template is found, replace the placeholders in the template with actual parameters; If there is no completely matching template in the template library, use a machine learning model for code generation.
9. A method for automatic code programming according to claim 7, characterized in that, The process of calculating the code fitness value according to the performance index monitoring results is as follows: The performance indicators include execution time and memory occupancy percentage; Obtain the time taken for the program to execute from start to end to get the execution time; Obtain the ratio of the memory used by the program to the total memory of the computer to get the memory occupancy percentage; Calculate the fitness value of the code according to the execution time and memory occupancy percentage combined with the benchmark execution time and benchmark memory occupancy percentage. The specific calculation formula is as follows: Where F is the code fitness value, w1 is the execution time weight factor, T is the execution time, and T b is the baseline execution time, w2 is the memory occupancy percentage weight factor, M is the memory occupancy percentage, and M b is the baseline memory occupancy percentage.
10. A method for automatic code programming according to claim 7, characterized in that, The process of evaluating the automatic programming effect of the code based on fuzzy inference is as follows: Obtain the number of times of automated testing of the code, as well as the number of times the code passes the test and the maximum number of consecutive passes of the test; Divide the number of times the code passes the test by the number of times of automated testing of the code to get the code automatic programming accuracy coefficient; Divide the maximum number of consecutive passes of the test by the number of times of automated testing of the code to get the code automatic programming stability coefficient; Obtain the code reliability dimension by weighted synthesis of the code automatic programming accuracy coefficient and stability coefficient; Define the code reliability dimension and the code fitness value as input variables and divide them into different fuzzy sets; Define the code automatic programming effect as the output variable and divide it into a fuzzy set; Formulate fuzzy rules to describe the influence of input variables on output variables; Perform fuzzy inference according to the fuzzy rules to determine the code automatic programming effect.
Citation Information
Patent Citations
Code performance optimization method based on API replacement
CN116400910A
Intelligent auxiliary coding system and method based on large model and machine learning
CN118535142A
Chinese code compiling method and device and storage medium
CN118733043A
Visual intelligent programming method based on large language model
CN119045806A
Railway information system performance test method and system based on cloud platform
CN119336640A
Cited By
Automatic simulation analysis method and system
CN120930288A
Automatic code generation system and method
CN121387267A