Code generation method and system based on template matching and logical reasoning
By constructing a requirement semantic graph and a code template graph, and combining Bayesian optimization algorithms and logical reasoning to generate code, the flexibility and security issues of code generation methods are solved, ensuring that the generated code complies with security specifications and improving the efficiency and security of code generation.
Patent Information
- Application Number
- CN202511214628.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-09
AI Technical Summary
Existing code generation methods lack flexibility and interpretability, have rigid template matching, cannot adapt to complex needs, and generate code with security vulnerabilities that fail to meet security standards.
A code generation system that constructs a requirement semantic graph and combines Bayesian optimization algorithms, template matching, and logical reasoning includes: determining user project requirements and setting edges and nodes to construct a requirement semantic graph; acquiring code template-related data and labeling the data, calculating the corresponding prior probabilities, and constructing a corresponding code template graph; generating matching results based on the code template graph; filtering the matching results using Bayesian optimization algorithms to obtain candidate templates and their security labels; and generating initial code through the requirement semantic graph and logical reasoning.
The final code is generated through security verification and evaluation, combined with Bayesian feedback mechanism and Gaussian process optimization algorithm.
Smart Images

Figure CN121092152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automatic code generation technology, and in particular to a code generation method and system based on template matching and logical reasoning. Background Technology
[0002] With the rapid development of software development, automated code generation technology has gradually become an important tool for improving development efficiency and code quality. Existing automated code generation methods mainly rely on template matching or rule-based reasoning to generate code according to preset templates or rules. However, these traditional methods often suffer from rigid template matching, lack of flexibility, and poor interpretability. Especially when facing complex requirements, the template library is not updated in a timely manner, making it unable to adapt to changing development needs. Furthermore, automatically generated code often fails to fully meet security specifications, easily overlooking potential security vulnerabilities such as permission verification and input validation, thereby increasing system security risks.
[0003] Most existing automatic code generation methods rely on static rule matching or simple reasoning, lacking dynamic optimization capabilities. This results in code defects in terms of security, adaptability, and interpretability. As demands become increasingly complex and development processes require higher levels of automation, traditional methods face growing challenges in practical applications. To address these issues, an automatic code generation method combining Bayesian optimization algorithms, template matching, and logical reasoning has emerged, enabling more flexible, secure, and interpretable automated code generation. Summary of the Invention
[0004] The purpose of this invention is to provide a method for visually and dynamically constructing an intelligence analysis system to improve the aforementioned technical problems.
[0005] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A code generation method based on template matching and logical reasoning includes:
[0007] Determine the user's project requirements and set edges and nodes to construct a requirement semantic graph;
[0008] Acquire code template related data and label the data, calculate the corresponding prior probabilities, and construct the corresponding code template graph; the code template related data includes multiple code templates, triggering conditions, and historical project requirement data; the code template graph includes N templates and their corresponding security tags;
[0009] Matching results are generated based on code template graphs and requirement semantic graphs; candidate templates are obtained by filtering the matching results using a Bayesian optimization algorithm.
[0010] Initial code is generated based on candidate templates and their security tags through requirement semantic graphs and logical reasoning.
[0011] The initial code undergoes security verification and evaluation, and the final code is generated by combining Bayesian feedback mechanism and Gaussian process optimization.
[0012] A code generation system based on template matching and logical reasoning includes:
[0013] The requirement semantic graph construction module is used to determine the user's project requirements and set edges and nodes to construct the requirement semantic graph;
[0014] Code template graph is used to acquire code template related data and label the data, calculate the corresponding prior probabilities, and construct the corresponding code template graph.
[0015] The matching module is used to generate matching results based on the code template graph and the requirement semantic graph; the matching results are filtered by combining the Bayesian optimization algorithm to obtain candidate templates;
[0016] The initial code generation module is used to generate initial code based on candidate templates and their security tags, through requirement semantic graphs and logical reasoning.
[0017] The security verification and evaluation module is used to perform security verification and evaluation on the initial code, and generates the final code by combining Bayesian feedback mechanism and Gaussian process optimization.
[0018] The beneficial effects of this invention are as follows:
[0019] This invention accurately parses project requirements by constructing a requirement semantic graph, achieves multi-dimensional matching by combining it with a code template graph, uses a Bayesian optimization algorithm to select suitable candidate templates, generates initial code through logical reasoning, and obtains the final code through security verification, Bayesian feedback mechanism, and Gaussian process optimization. This not only solves the problems of rigid template selection and lack of flexibility in traditional methods, but also ensures that the generated code complies with security specifications through security label matching and security verification evaluation, reducing potential vulnerabilities. At the same time, it enhances the interpretability of the generation process, supports flexible integration of external code modules, effectively improves the efficiency, accuracy and security of code generation, and can better adapt to complex and ever-changing development needs. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of the method in an embodiment of the present invention;
[0022] Figure 2 This is a system structure diagram in an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0024] Please see Figure 1 This embodiment provides a code generation method based on template matching and logical reasoning, including:
[0025] S1. Determine the user's project requirements and set edges and nodes to construct a requirement semantic graph;
[0026] S1 includes:
[0027] S1-1. Obtain the user's project requirements; where project requirements are the code requirements proposed by the client / owner, which may include natural language text, pseudocode, or functional descriptions.
[0028] S1-2. Use NLP algorithms to segment and identify project requirements to determine operation nodes and concept nodes;
[0029] Specifically, the project requirements are segmented and identified using NLP algorithms and NER (Named Entity Recognition) algorithms. They are then broken down according to part-of-speech tagging or functional descriptions to obtain entities, and their corresponding parts of speech are tagged. Dependency parsing is used to determine the relationships between the entities. Based on the relationships and each part of speech, operation nodes and concept nodes are determined.
[0030] For example, if the project requirement is "verify username", we can analyze it by part of speech. "Verify" is a verb, representing an operation, while "username" is a noun, representing a concept. Therefore, we can segment the project requirement into words, obtaining "verify" and "username", and label them with their corresponding parts of speech. Then, we use dependency parsing to analyze the relationship between "verify" and "username", determining the operation to be performed and the target or object involved. Thus, "verify" and "username" are respectively designated as operation nodes and concept nodes.
[0031] S1-3. Utilize semantic analysis of requirements to obtain the dependencies between operation nodes and concept nodes, and determine the direction and type of edges through logical reasoning. Dependency refers to the sequential or conditional relationship between certain operations, concepts, or functions in the textual description of project requirements; that is, the occurrence of one operation or concept requires another operation or concept as a prerequisite or condition. For example, requiring a username to be entered before verification is a sequential dependency; displaying error messages depends on username or password verification failure, which is a conditional dependency; encrypting passwords requires obtaining the password before execution, which is a data dependency—an operation depends on the result of a previous operation.
[0032] S1-4. Based on operation nodes, concept nodes, and edges, construct an initial requirement semantic graph using a graph structure. The graph structure can be a directed graph or a weighted graph. Add edges between operation nodes and concept nodes according to dependencies.
[0033] S1-5. Perform quality checks on the initial requirement semantic graph to determine whether the logic of each node and edge in the graph is consistent and whether all content in the project requirements has corresponding nodes and dependencies. If the conditions are met, the initial requirement semantic graph is used as the final requirement semantic graph; otherwise, return to S1-2 and re-perform word segmentation and recognition.
[0034] In this embodiment, starting from the user's project requirements, a complete requirement semantic graph that reflects all the dependencies of the requirements is constructed. This helps to understand and analyze the requirements of the code and provides a basic data structure for subsequent operations such as template matching and logical reasoning.
[0035] S2. Obtain code template related data and label the data, calculate the corresponding prior probability, and construct the corresponding code template graph; the code template related data includes multiple code templates and historical project requirement data.
[0036] S2 includes:
[0037] S2-1. Obtain historical project data and the first initial code templates from each open-source code repository; extract the historical initial code templates corresponding to the historical project data; each first initial code template and historical initial code template includes a code snippet and a corresponding triggering condition; the triggering condition refers to the prerequisite for the execution of each code template.
[0038] S2-2. Perform functional identification and annotation on each historical initial code template and each first initial code template, and annotate the operation type and functional module of each code segment in each historical initial code template and each first initial code template; perform security analysis on each functional module to obtain the corresponding security label; the security label includes safe, potential risk, high risk and unsafe.
[0039] Specifically, security analysis refers to determining whether a security module involves sensitive operations (database calls, network requests, etc.), requires special permission authentication (API call permissions or user authentication), and involves resource management (memory release). This can be achieved by analyzing each functional module using industry-standard security criteria or rules, and employing static analysis tools such as SonaQube, Checkmarx, and Fortifu to determine the security characteristics of each module.
[0040] S2-3. Calculate the similarity between each historical initial code template and each first initial code template, and then filter and integrate them to obtain a code template set; the code template set includes historical code templates and first code templates.
[0041] During the code collection process, there may be a large number of similarities or repetitions between historical project templates and open-source code repositories, resulting in template redundancy. This affects subsequent template matching and selection, reduces matching efficiency, and fails to accurately reflect project requirements. Therefore, S2-3 includes:
[0042] Calculate the functional similarity, trigger condition similarity, and security label similarity between each historical initial code template and each first initial code template; the functional similarity, trigger condition similarity, and security label similarity can be calculated using Jaccard similarity or cosine similarity.
[0043] Based on functional similarity, trigger condition similarity, and security label similarity, a weighted sum is used to obtain the similarity score. It is then determined whether each similarity score exceeds a similarity threshold. If so, one of the corresponding historical initial code templates or the first initial code template is removed according to the deduplication selection criteria; otherwise, the corresponding historical initial code template and the first initial code template are retained, resulting in filtered first initial code templates and filtered historical initial code templates. The filtered first initial code templates and filtered historical initial code templates are then merged to obtain the code template set.
[0044] This can be explained by the following: In the deduplication selection criteria, if the execution frequency of the historical initial code template or its success rate in historical projects is higher than the horizontal rate, then the historical template is retained first. If the first initial code template is more in line with the current project's security needs or includes new features, then the first initial code template is retained first.
[0045] In this embodiment, by calculating and filtering the similarity between the historical initial code template and the first initial code module, the complexity of the generated code can be reduced and the matching process optimized, thereby accelerating code generation. In addition, the quality of the generated code can be enhanced. The deduplicated template library can more accurately reflect the project requirements and avoid duplicate or irrelevant templates from affecting the quality of the final code. Especially in terms of template triggering conditions and security tags, deduplication can ensure that the generated code is more in line with actual needs and security standards.
[0046] S2-3. Extract the execution frequency and success frequency of each historical code template in historical projects; use the ratio of success frequency to execution frequency as the prior probability of that historical code template.
[0047] S2-4. Treat each code template as a node and define node attributes. Based on the attributes between each node, use the same method as in S1-3 to determine whether there is a functional dependency or trigger condition dependency between each node. If so, construct the corresponding edge and construct the corresponding code template graph.
[0048] The node attributes include the template ID, functional module, trigger condition, prior probability, and security label of the code template; the prior probability of the first code template can be considered as 0. Functional dependency means that the execution result of one operation depends on the output or result of another operation. This dependency indicates that the input of an operation comes from the result or output of the previous operation. Trigger condition dependency means that the execution of an operation or template will only occur under specific conditions. This condition can be project requirements, user input, or changes in certain system states. Trigger condition dependency typically means that the corresponding operation can only be executed when a certain condition is true.
[0049] S3. Generate matching results based on code template graph and requirement semantic graph; combine Bayesian optimization algorithm to filter matching results and obtain candidate templates;
[0050] S3 includes:
[0051] S3-1. Match the operation nodes in the requirement semantic graph with the functional modules in the code template graph to determine which code templates implement the operations defined in the requirements, and obtain the node-function matching results.
[0052] Specifically, all operation nodes, such as "verify username" and "encrypt password," are extracted from the requirement semantic graph. NLP algorithms combined with semantic analysis techniques are used to ensure that the extracted operation nodes accurately reflect the functional requirements in the requirement. All functional modules (such as "username verification template" and "password encryption template") are extracted from the code template graph. Each functional module contains multiple code snippets and their corresponding triggering conditions. Semantic matching algorithms (such as BERT, GPT, and other deep learning models) are used to calculate the similarity between the operation nodes in the requirement semantic graph and the functional modules in the code template graph. If the similarity is higher than a set threshold (e.g., 0.8), the operation node is considered to match the functional module, generating a node-function matching result to ensure that the template accurately implements the operations in the requirement.
[0053] S3-2. Match the dependencies in the requirement semantic graph with the triggering conditions in the code template graph to ensure that the template is executed only under specific conditions, and obtain the dependency-trigger matching results;
[0054] Specifically, the process involves extracting dependencies between operation nodes, such as sequence dependencies, data dependencies, and conditional dependencies. Dependency parsing techniques are used to identify logical dependencies between nodes. Triggering conditions for each code template (such as specific inputs or system state changes) are extracted. Graph matching algorithms are used to ensure that dependencies in the requirement semantic graph match the triggering conditions of the code templates. By calculating the similarity or matching degree of the triggering conditions, it is ensured that the template executes only when specific conditions are met. Templates with successfully matched dependencies and triggering conditions are marked as meeting the dependency requirements, generating dependency-trigger matching results. This ensures that the generated code templates are correctly triggered when dependencies are met, avoiding erroneous execution; reducing inappropriate triggering conditions; and improving the reliability of templates in practical applications.
[0055] S3-3. Determine the security requirements of the project; match the security requirements with the security tags in the code template diagram to ensure that the security requirements in the requirements are consistent with the security tags in the template, avoid generating code with security vulnerabilities, and obtain the security matching results;
[0056] Specifically, security requirements (such as data encryption, user authentication, and access control) are extracted from project needs. Security tags are extracted for each code template, such as secure, potentially risky, high-risk, and insecure. A security tag matching algorithm is used to ensure that the security tags of the code templates meet the security requirements of the project. If the security tags of a code template do not match the project requirements, that code template is excluded. Through security matching, the selected code templates are ensured to meet the project's security requirements, resulting in security matching results. This avoids generating potentially vulnerable code and reduces security vulnerabilities caused by selecting insecure templates.
[0057] The core of the security tag matching algorithm is to compare the security requirements in the project requirements with the security tags in the code templates. First, for each extracted security requirement, it needs to be labeled to ensure its corresponding security standard (e.g., templates requiring "data encryption" must include encryption operations, or requirements for "preventing SQL injection" must include input validation). These requirements can be categorized into data encryption, SQL injection prevention, access control, and secure operations. Data encryption is used for matching code templates related to encryption; SQL injection prevention is used for matching code templates related to input validation; access control is used for matching code templates related to authentication and authorization; and secure operations are used for matching templates involving sensitive operations.
[0058] Secondly, security labels are matched. If the requirement calls for data encryption, the security label of the code template should be "secure" or "potential risk," and the code template should include encryption-related functionality. If the requirement calls for user authentication, the security label of the code template should be "secure" or "potential risk," and the code template should include an authentication module. If the security label of the code template is "high risk" or "insecure," the code template will be directly excluded during matching unless the user explicitly allows these labels.
[0059] Finally, the code templates are matched against the requirements and security tags. Code templates that meet the criteria (safe or potentially risky) are marked as "safe match", and the security match result is obtained.
[0060] By matching security tags, we ensure that the selected templates meet the requirements without introducing security vulnerabilities, avoiding the use of potentially risky or insecure templates and improving the security of the generated code. We can flexibly select template security tags and standards based on the project's security requirements to meet different security needs.
[0061] S3-4. Based on the node-function matching results, dependency-trigger matching results, and dependency-trigger matching results, the code template set is filtered to obtain a second-filtered code template set, and the matching results are obtained.
[0062] Specifically, the code template set undergoes a comprehensive screening based on node-function matching results, dependency-trigger matching results, and dependency-trigger matching results. This involves determining whether each code template is suitable for the current project requirements based on multiple factors. This includes confirming whether the template implements the operations specified in the requirements; checking whether the template's execution conditions conform to the dependencies in the requirements (ensuring the template executes only under conditions that meet the requirements); and using security tags to ensure the template does not introduce code that does not meet security requirements. This comprehensive screening results in a second-filtered code template set, which has removed templates that do not meet the requirements or security standards, making it more suitable for the actual needs of the project and compliant with security standards.
[0063] For example, if the project requirement is to "implement user login functionality," which involves verifying the username and password (functional operation); displaying a success page if the password is correct, and displaying an error message if incorrect (dependency); and requiring encrypted password storage (security requirement), the code templates include A, B, and C. A verifies the username and password, displays a success page, and encrypts the password (security label "secure"); B verifies the username and password, displays an error page, and encrypts the password (security label "potential risk"); and C verifies the username and password, displays a success page, but does not perform any encryption (security label "high risk"). After matching, the node-functionality matching result shows that templates A, B, and C satisfy the username and password verification function; the dependency-trigger matching result shows that templates A, B, and C satisfy the dependency relationship; and the security matching result shows that template A has a security label of "secure" (meets requirements), template B has a security label of "potential risk" (can be used, but requires additional verification), and template C has a security label of "high risk" (excluded). Based on these node-functionality matching results, dependency-trigger matching results, and dependency-trigger matching results, candidate templates for template A are selected.
[0064] S3-5. Combine the matching results with the Bayesian optimization algorithm, calculate the posterior probability of each code template in the secondary screening code template set, and obtain the candidate templates.
[0065] It should be noted that the posterior probability of each code template is calculated using Bayes' theorem, and the corresponding formula is:
[0066] ;
[0067] in, For the demand semantic graph, For the second-stage filtering code template set A code template, For the first The prior probability of a code template under the current project requirements represents the frequency and success rate of its use in historical projects. For the demand semantic graph and the first The matching likelihood of a code template indicates the degree to which the code template meets the current requirements. For the first The posterior probability of a code template under the current project requirements indicates whether the code template is suitable for the current requirements. They are proportional.
[0068] Sort the posterior probabilities in descending order and select the top k code templates with the highest posterior probabilities as candidate templates.
[0069] S4. Based on candidate templates and their security labels, generate initial code through requirement semantic graphs and logical reasoning;
[0070] S4 includes:
[0071] S4-1. Traverse the dependencies (sequential dependencies / data dependencies / conditional dependencies) in the requirement semantic graph and map the candidate templates to execution units according to the following rules.
[0072] It needs to be explained that, based on the dependencies in the requirement semantic graph, the sequential structure, branching structure, and parallel structure in the requirement semantic graph are determined. Here, the sequential structure refers to the edges in the requirement semantic graph corresponding to data dependencies and sequential dependencies, where the corresponding data flow is linear (e.g., from operation node 1 to operation node 2, and then to operation node 3), strictly executed serially, meaning there is a logical order between operation nodes or concept nodes. That is, the logical order of candidate templates is inferred based on the data flow in the requirement semantic graph. If the output of operation node a is the input of operation node b (the edge corresponding to the data dependency), then the output parameter type of the corresponding candidate template A must match the input parameter type of candidate template B. For example, in the requirement graph "Get Password" → "Encrypt Password", the template execution order must be: Get Password Template → Encrypt Password Template.
[0073] Branching structure refers to the edges corresponding to conditional dependencies in the requirement semantic graph, where the data flow is tree-like, such as if flowing to then or else, with mutually exclusive paths executed, and a "if X then Y" relationship existing between operation nodes or concept nodes. Therefore, the logical order of candidate templates is determined based on the branching structure, according to the edges corresponding to conditional dependencies (e.g., "verification failed → display error").
[0074] Parallel structures refer to closed functional units (independent subgraphs) existing in the requirement semantic graph. This means that the corresponding internal nodes are interconnected, but there is no data or control flow exchange with the outside. For example, one closed functional unit might be used for transaction records, while other nodes are used for user notifications. If multiple independent subgraphs have no cross-dependencies, i.e., no data overlap (sub-... Figure 1 The output is not a child Figure 2 Input), control cross (sub) Figure 1 The execution result does not affect the sub-item. Figure 2 If the triggering conditions are met and resource conflicts occur (such as subgraphs not contributing database connections or file locks), then parallel paths are constructed through thread pools or asynchronous calls.
[0075] Finally, based on the sequential structure, branch structure, and parallel structure, the corresponding logical order is generated, and the execution unit is constructed, that is, the candidate templates are arranged and combined according to logic.
[0076] S4-2. Identify conflicts within the execution unit, optimize the execution unit through logical reasoning, and obtain the optimized execution unit;
[0077] S4-2 includes:
[0078] S4-2-1. Perform conflict detection on the execution unit, identify and mark different types of conflicts. Determine whether there is a resource contention, security priority conflict, or execution efficiency conflict in each candidate template in the execution unit. If any one of these exists, mark the candidate template with the content of "resource contention conflict exists", "security priority conflict exists", or "execution efficiency conflict exists".
[0079] S4-2-2. Based on each marker, perform local optimization on the execution unit to obtain the conflict-mitigated execution unit;
[0080] Specifically, if resource contention exists: for templates sharing resources, a synchronization mechanism (such as locks) is inserted to ensure that threads or tasks do not conflict. For example, mutexes (for CPU-intensive tasks) or semaphores (for I / O-intensive tasks) are used to coordinate resource access. If safety priority conflicts exist: based on the template's safety label, a template with higher safety is selected; if unavailable, the next best option is chosen. If execution efficiency bottleneck conflicts exist: when a computationally intensive task is detected, its posterior probability is checked to see if it exceeds a threshold; if so, a caching mechanism is added to avoid redundant computation.
[0081] S4-2-3. Calculate the priority score of the conflict mitigation execution unit and update the conflict mitigation execution unit;
[0082] It should be noted that priority scores are calculated based on resource contention, execution efficiency, and security requirements of each code template within the conflict mitigation execution unit. The corresponding formula is:
[0083] ;
[0084] in, , , These represent the posterior probability weight, the safety weight, and the matching weight, respectively. Indicates the safety score. Indicates the data stream matching degree. This represents the total posterior probability of the conflict mitigation execution unit.
[0085] Safety rating This is used to quantify the security of each code template in the conflict mitigation execution unit. A "safe" security label indicates that the code template has undergone rigorous review and meets all security standards, with a security score of 1.0. A "potential risk" security label indicates that the code template may contain some known risks, with a security score of 0.7. A "high risk" security label indicates that the code template has obvious security vulnerabilities or has not been adequately verified, with a security score of 0.3. A "safe" security label indicates that the code template may introduce serious security vulnerabilities or potential harm, with a security score of 0.05.
[0086] Each code template in the conflict mitigation execution unit has corresponding input parameters (input data) and output parameters (output results). These parameters play a crucial role in the template execution process; therefore, they need to match the corresponding parts in the project requirements. That is, the input and output of the code template must be identical to the input and output required by the project requirements to ensure that the template can execute correctly and meet the requirements. Data flow matching degree. The formula used to measure the adaptability and compatibility of code templates at the parameter level in conflict mitigation execution units, i.e., the degree of semantic and data type matching between input and output parameters, is as follows:
[0087] ;
[0088] in, , These represent the input weights and the output weights, respectively. , These represent the number of successful matches between the input parameters of the code template and the required input parameters of the project, and the total number of input parameters of the code template, respectively. , These represent the number of successful matches between the output parameters of the code template and the required output parameters of the project, and the total number of output parameters of the code template, respectively.
[0089] Code templates with a sequential relationship are used as preceding and following templates. It is determined whether the priority score of each preceding template is greater than the priority score of the corresponding following template. If so, the sequential order is maintained; otherwise, the sequential order is adjusted, and the code template with the higher priority score is used as the preceding template. The update of the conflict mitigation execution unit is completed.
[0090] S4-3. Perform security hardening and parameter bridging on the execution optimization unit to obtain the final execution unit;
[0091] Even after optimization and sorting, execution units may still present two types of risks: inherent risks and combined risks. Inherent risks refer to defects in the code template itself, such as the lack of input validation; combined risks refer to the potential for new vulnerabilities arising from the combination of multiple code templates, thus requiring security hardening of the execution optimization unit. The process involves determining whether the execution optimization unit contains potential risks or sensitive operations; if potential risks exist, a protective layer is inserted into the execution optimization unit; if sensitive operations exist, permission checking steps are added to the execution optimization unit, resulting in a security-hardened execution unit.
[0092] Even after security hardening of the execution optimization units, interface issues may still exist, such as name mismatches, type incompatibilities, or structural differences. Therefore, introducing parameter bridging can effectively avoid these interface problems. By establishing cross-template variable mapping relationships, interface compatibility issues between execution units are automatically resolved. First, the data flow between nodes in the requirement semantic graph is analyzed to identify input / output parameter pairs between adjacent templates. Then, a parameter mapping table is constructed to handle three typical cases: alias mapping is established when parameter names are different but semantically identical; type conversion functions are automatically inserted when types do not match (e.g., string encoding / decoding); and format reorganization is performed when data structures are inconsistent (e.g., JSON to Protobuf). This process, through a historical conversion rule base and type inference mechanism, ensures that optimized execution units can seamlessly connect, ultimately forming a type-safe, semantically consistent data transmission chain—the final execution unit.
[0093] S4-4. Generate initial code based on the final execution unit.
[0094] S5. Perform security verification and evaluation on the initial code, and generate the final code by combining Bayesian feedback mechanism and Gaussian process optimization.
[0095] S5 includes:
[0096] S5-1. Use the same security analysis method as S2-2 to perform security verification and evaluation on the initial code, and obtain a security verification and evaluation report; determine whether the security verification and evaluation report meets the predetermined security standards. If yes, proceed to S5-2; otherwise, return to S3.
[0097] S5-2. Utilize Bayesian feedback mechanism to optimize the code template selection strategy in reverse, dynamically adjust the prior probability and matching parameters of the code template, and feed them back to the corresponding code template to improve the safety and accuracy of subsequent or next code generation.
[0098] The core of revising the initial code using a Bayesian feedback mechanism is to optimize the code template selection strategy based on the security verification assessment results. Specifically, based on the security assessment results of the initial code (e.g., secure, potentially risky, high-risk), the prior probabilities of corresponding candidate templates are dynamically adjusted—the prior probability of secure code templates is increased, while the prior probability of risky templates is correspondingly decreased. Simultaneously, the likelihood of matching the templates with the requirements is adjusted, taking into account the template's security score (security 1.0, potential risk 0.7, etc.) to adjust the matching degree. Furthermore, for specific risk points identified in the assessment, local repairs are performed on the initial code, such as inserting security repair templates or adding permission verification snippets, to improve code security and the accuracy of template matching.
[0099] S5-3. Use Gaussian process optimization to update the parameter configuration of the initial code to obtain the updated parameter configuration;
[0100] Specifically, Gaussian process optimization is used to fine-tune the parameters of the initial code, aiming to improve code execution efficiency while meeting security constraints. First, the parameter space to be optimized is defined (e.g., the number of iterations of the encryption algorithm, the size of the database connection pool, etc.), and an objective function is constructed with security compliance and execution efficiency as indicators, balancing the two through weights (security weight 0.6~0.8, efficiency weight 0.2~0.4). Then, a Gaussian process model is built based on historical execution data to predict the objective function value corresponding to the parameter configuration. Finally, Bayesian optimization iterations (using the Acquisition Function to select parameter combinations) are used to continuously update the model and find the optimal parameter configuration, ultimately achieving a balance between security and performance in the code.
[0101] S5-4. Apply the updated parameter configuration to the output code to obtain the final code.
[0102] In summary, this invention accurately parses project requirements by constructing a requirement semantic graph, achieves multi-dimensional matching by combining it with a code template graph, uses a Bayesian optimization algorithm to select suitable candidate templates, generates initial code through logical reasoning, and obtains the final code through security verification, a Bayesian feedback mechanism, and Gaussian process optimization. This not only solves the problems of rigid template selection and lack of flexibility in traditional methods, but also ensures that the generated code complies with security specifications and reduces potential vulnerabilities through security label matching and security verification evaluation. At the same time, it enhances the interpretability of the generation process, supports flexible integration of external code modules, and effectively improves the efficiency, accuracy, and security of code generation, making it better able to adapt to complex and ever-changing development needs.
[0103] It should be noted that, Figure 1The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not impose any limitations on this.
[0104] like Figure 2 As shown, a code generation system based on template matching and logical reasoning includes:
[0105] The requirement semantic graph construction module is used to determine the user's project requirements and set edges and nodes to construct the requirement semantic graph;
[0106] Code template graph is used to acquire code template related data and label the data, calculate the corresponding prior probabilities, and construct the corresponding code template graph.
[0107] The matching module is used to generate matching results based on the code template graph and the requirement semantic graph; the matching results are filtered by combining the Bayesian optimization algorithm to obtain candidate templates;
[0108] The initial code generation module is used to generate initial code based on candidate templates and their security tags, through requirement semantic graphs and logical reasoning.
[0109] The security verification and evaluation module is used to perform security verification and evaluation on the initial code, and generates the final code by combining Bayesian feedback mechanism and Gaussian process optimization.
[0110] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the device, and will not be elaborated here.
[0111] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0112] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A code generation method based on template matching and logical reasoning, characterized in that, include: Determine the user's project requirements and set edges and nodes to construct a requirement semantic graph; Acquire code template related data and label the data, calculate the corresponding prior probabilities, and construct the corresponding code template graph; the code template related data includes multiple code templates, triggering conditions, and historical project requirement data. The code template map includes N templates and their corresponding security tags; Generate matching results based on code template graph and requirement semantic graph; By combining the Bayesian optimization algorithm to filter the matching results, candidate templates are obtained; Initial code is generated based on candidate templates and their security tags through requirement semantic graphs and logical reasoning. The initial code undergoes security verification and evaluation, and the final code is generated by combining Bayesian feedback mechanism and Gaussian process optimization.
2. The code generation method based on template matching and logical reasoning according to claim 1, characterized in that, The construction of the requirement semantic graph includes: Obtain user project requirements; use NLP algorithms to segment and identify project requirements, and determine operation nodes and concept nodes; Demand semantic analysis is used to obtain the dependencies between operation nodes and concept nodes, and logical reasoning is used to determine the direction and type of edges; Based on operation nodes, concept nodes, and edges, an initial requirement semantic graph is constructed using a graph structure. The graph structure can be a directed graph or a weighted graph, and edges are added between operation nodes and concept nodes according to dependencies. The initial requirement semantic graph is subjected to quality inspection to determine whether the logic of each node and edge in the graph is consistent and whether all content in the project requirements has corresponding nodes and dependencies. If the conditions are met, the initial requirement semantic graph is used as the final requirement semantic graph; otherwise, word segmentation and recognition are performed again.
3. The code generation method based on template matching and logical reasoning according to claim 1, characterized in that, The construction of the corresponding code template graph includes: Retrieve historical project data and the first initial code template from each open-source code repository; extract the historical initial code template corresponding to the historical project data; each first initial code template and historical initial code template includes a code snippet and a corresponding triggering condition; Functional identification and annotation are performed on each historical initial code template and each first initial code template, and the operation type and functional module of each code segment in each historical initial code template and each first initial code template are annotated; security analysis is performed on each functional module to obtain the corresponding security label; the security label includes safe, potential risk, high risk and unsafe; Calculate the similarity between each historical initial code template and each first initial code template, and then filter and integrate them to obtain a code template set; the code template set includes historical code templates and first code templates. Extract the execution frequency and success frequency of each historical code template in historical projects; use the ratio of success frequency to execution frequency as the prior probability of that historical code template. Each code template is treated as a node and its attributes are defined. Based on the attributes between nodes, it is determined whether there are functional dependencies or triggering condition dependencies between nodes. If so, the corresponding edges are constructed, and the corresponding code template graph is constructed.
4. The code generation method based on template matching and logical reasoning according to claim 3, characterized in that, The process of obtaining the candidate template includes: Match the operation nodes in the requirement semantic graph with the functional modules in the code template graph to obtain the node-function matching results; The dependency relationships in the requirement semantic graph are matched with the triggering conditions in the code template graph to obtain the dependency-trigger matching results; Determine the security requirements of the project; match the security requirements with the security tags in the code template diagram to obtain the security matching results; Based on the node-function matching results, dependency-trigger matching results, and dependency-trigger matching results, the code template set is filtered to obtain a second-filtered code template set, and the matching results are obtained. By combining the matching results filtered by the Bayesian optimization algorithm, the posterior probability of each code template in the secondary filtering code template set is calculated to obtain candidate templates.
5. The code generation method based on template matching and logical reasoning according to claim 1, characterized in that, The process of generating initial code based on candidate templates and their security tags, through demand semantic graphs and logical reasoning, includes: Traverse the dependencies in the requirement semantic graph and map the candidate templates to execution units according to the following rules; Conflicts within the execution unit are identified, and the execution unit is optimized through logical reasoning to obtain the optimized execution unit; The execution optimization unit is then hardened and its parameters are bridged to obtain the final execution unit. Initial code is generated based on the final execution unit.
6. The code generation method based on template matching and logical reasoning according to claim 5, characterized in that, The process of identifying conflicts within an execution unit and optimizing that unit through logical reasoning to obtain an optimized execution unit includes: Conflict detection is performed on the execution unit to identify and mark different types of conflicts. It determines whether each candidate template in the execution unit has resource contention, security priority conflict, or execution efficiency conflict. If any of these exist, the candidate template is marked with the message "Resource contention exists," "Security priority conflict exists," or "Execution efficiency conflict exists." Based on each marker, the execution unit is locally optimized to obtain the conflict-mitigated execution unit; Calculate the priority score of the conflict mitigation execution unit and update the conflict mitigation execution unit.
7. The code generation method based on template matching and logical reasoning according to claim 1, characterized in that, The process of obtaining the final code includes: The initial code is subjected to security verification and evaluation to obtain a security verification and evaluation report; it is then determined whether the security verification and evaluation report does not meet the predetermined security standards. If so, a matching result is regenerated based on the code template graph and the requirement semantic graph. Conversely, the Bayesian feedback mechanism is used to optimize the code template selection strategy, dynamically adjust the prior probability and matching parameters of the code template, and feed them back to the corresponding code template. The updated parameter configuration is applied to the output code to obtain the final code.
8. A code generation system based on template matching and logical reasoning, used to implement the code generation method based on template matching and logical reasoning as described in any one of claims 1 to 7, characterized in that, include: The requirement semantic graph construction module is used to determine the user's project requirements and set edges and nodes to construct the requirement semantic graph; Code template graph is used to acquire code template related data and label the data, calculate the corresponding prior probabilities, and construct the corresponding code template graph. The matching module is used to generate matching results based on the code template graph and the requirement semantic graph; By combining the Bayesian optimization algorithm to filter the matching results, candidate templates are obtained; The initial code generation module is used to generate initial code based on candidate templates and their security tags, through requirement semantic graphs and logical reasoning. The security verification and evaluation module is used to perform security verification and evaluation on the initial code, and generates the final code by combining Bayesian feedback mechanism and Gaussian process optimization.