A code generation method, system, medium and product

CN122816596APending Publication Date: 2026-09-25TONGCHENG NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610704088.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供了一种代码生成方法、系统、介质及产品,用于解决生成的代码虽然在表面逻辑上看似合理,但在实际接入本地工程时往往无法编译或运行的技术问题,提高了代码生成效率

Benefits of technology

1、通过将需求特征向量与包含工具类签名及团队编码规范规则的本地工程知识库进行匹配,并构建包含约束指令的上下文感知提示词,在模型推理阶段明确限定了代码生成边界,克服了现有大模型因缺乏工程上下文而倾向于“重复造轮子”及无视团队规范的问题,从源头引导模型生成高复用性的本地适配代码。同时,将初始代码片段解析为抽象语法树,针对违背规范或未复用工具类签名的违规节点进行迭代校验与大模型修正。这种基于抽象语法树的闭环反馈机制,打破了现有单向开环生成的局限,使得系统能够自发察觉并纠正违规逻辑,从而确保最终输出的目标代码能够严格契合企业级工程标准并顺利编译运行,提高了代码生成效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816596A_ABST
    Figure CN122816596A_ABST
Patent Text Reader

Abstract

A code generation method, system, medium and product, wherein the method comprises: obtaining a document description containing business requirements in a code generation task, performing semantic analysis on the document description to obtain structured requirement data; retrieving target context information related to the business requirements; constructing context-aware prompt words; inputting the context-aware prompt words into a pre-trained code generation large model to generate an initial code snippet; parsing the initial code snippet into an abstract syntax tree and performing iterative verification, and when there is a violation node, modifying it through the code generation large model until a preset stop verification condition is met; if there is no violation node in the abstract syntax tree when the verification is stopped, the corresponding code snippet is determined as the target code; if there is still a violation node in the abstract syntax tree when the verification is stopped, the current initial code snippet with violation prompt information is output. The application improves the code generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of code generation technology, specifically to a code generation method, system, medium, and product. Background Technology

[0002] With the rapid development of artificial intelligence technology, natural language processing and large-scale pre-trained language models are being used more and more widely in the field of software engineering. Currently, using artificial intelligence to assist in code generation has become an important means to improve software development efficiency.

[0003] In existing technologies, developers or systems directly convert business requirement documents into prompts, which are then input into a pre-trained code generation model. Based on general programming paradigms and syntax rules learned from massive open-source code corpora, the model performs a single inference prediction, directly outputting a complete piece of business logic code. To improve generation quality, some related technologies also manually add a small amount of code examples or general programming standards to the prompts, aiming to guide the model to generate code that meets expectations.

[0004] However, because pre-trained models only possess general programming knowledge, the generated code often suffers from a "disconnect from the project context." When implementing specific functions, the model tends to "reinvent the wheel" (i.e., the model tends to write the underlying implementation logic of basic functions from scratch, rather than directly calling mature code assets already existing in the local project) or fabricate general APIs (Application Programming Interfaces), failing to recognize and reuse specific utility classes already encapsulated in the target project. Furthermore, the generated code is unlikely to strictly adhere to the customized coding standards within a specific development team. More seriously, the existing code generation process is usually "one-way and open-loop," meaning the system cannot spontaneously detect when the model output contains illegal calls or non-compliant code. This results in generated code that, while seemingly logically sound on the surface, often fails to compile or run when integrated into a local project, reducing code generation efficiency. Summary of the Invention

[0005] This application provides a code generation method, system, medium, and product to solve the technical problem that although the generated code appears logically reasonable on the surface, it often cannot be compiled or run when actually integrated into a local project, thereby improving code generation efficiency.

[0006] The first aspect of this application provides a code generation method, the method comprising: Obtain the document description containing business requirements from the code generation task, and perform semantic parsing on the document description using a natural language processing model to obtain structured requirement data; The structured requirement data is converted into a requirement feature vector, and the requirement feature vector is matched with a preset local engineering knowledge base to retrieve the target context information associated with the business requirement. The local engineering knowledge base includes the existing tool class signature set and team coding standard rules in the target project. Based on the structured requirement data and the target context information, a context-aware prompt word is constructed. The context-aware prompt word contains constraint instructions. The constraint instructions instruct the pre-trained code generation model to call the utility class signature in the target context information and follow the team coding specification rules. The context-aware prompts are input into the pre-trained code generation model to generate initial code snippets; The initial code snippet is parsed into an abstract syntax tree, and the abstract syntax tree is iteratively verified. When there are violation nodes, the code is used to generate a large model for correction until a preset stopping verification condition is met. The preset stopping verification condition includes that there are no violation nodes in the abstract syntax tree or that the preset maximum number of iterations for correction is met. The violation node is a node that violates the team coding standard rules, or the corresponding parent scope node marked when no call action of the utility class signature in the target context information is detected. If the violation node does not exist in the abstract syntax tree when the verification stops, the corresponding code snippet is identified as the target code; If the violation node still exists in the abstract syntax tree when the verification stops, the current initial code snippet with violation warning information is output.

[0007] Optionally, the requirement feature vector is matched with a preset local engineering knowledge base to retrieve target context information associated with the business requirement, specifically including: Calculate the semantic similarity score between the demand feature vector and each tool signature in the tool signature set, and combine the tool signatures with semantic similarity scores greater than a preset matching threshold into an initial candidate signature set; Determine whether there is a semantic conflict subset in the initial candidate signature set. The semantic conflict subset is at least two of the tool class signatures in the initial candidate signature set whose semantic similarity scores with the requirement feature vector are less than a preset difference threshold, and the at least two tool class signatures belong to different underlying dependency modules in the local engineering knowledge base. If the semantic conflict subset does not exist, then the initial candidate signature set is determined as the target candidate signature set; If the semantic conflict subset exists, the utility class signatures that do not meet the calling constraints are removed from the initial candidate signature set to obtain the target candidate signature set. The target candidate signature set is assembled with the team coding standard rules to obtain target context information related to the business requirements.

[0008] Optionally, before determining the initial candidate signature set as the target candidate signature set, the method further includes: Obtain the target deployment environment identifier corresponding to the code generation task, and parse the metadata annotations of the signatures of each utility class in the local project knowledge base; Based on the metadata annotations, extract the lifecycle status and environment isolation tags of each utility class signature in the initial candidate signature set; Determine whether there is an environment overreach signature in the initial candidate signature set. The environment overreach signature is a signature whose semantic similarity score is greater than or equal to the preset matching threshold and whose environment isolation label conflicts with the target deployment environment identifier. If an unauthorized signature exists in the environment, it is removed from the initial candidate signature set.

[0009] Optionally, utility class signatures that do not meet the calling constraints are removed from the initial candidate signature set to obtain the target candidate signature set, specifically including: Get the target project directory path corresponding to the code generation task; Based on the project configuration information in the local project knowledge base, a module dependency topology diagram is generated; In the module dependency topology graph, the topological reachability and dependency weight of the target module where the target project directory path is located are analyzed to reach the source module where each utility class signature in the semantic conflict subset is located. The utility signatures in the semantic conflict subset whose topological reachability is unreachable or whose dependency weight is less than a preset weight threshold are removed from the initial candidate signature set to obtain the target candidate signature set.

[0010] Optionally, based on the structured requirement data and the target context information, context-aware prompt words are constructed, specifically including: Extract the utility class signature and the team coding standard rules from the target context information, and cross-validate the syntax features of the utility class signature with the team coding standard rules to determine whether there is a standard conflict event. The standard conflict event is that the definition format of the utility class signature itself violates the team coding standard rules. If the aforementioned specification conflict event exists, then extract the conflict signature that triggered the specification conflict event and the corresponding conflict rule; A local exemption statement is generated for the conflict signature, the local exemption statement being used to instruct the code generation model to ignore the conflict rule when invoking the conflict signature; Convert the conflict-free signatures in the utility class signatures into regular constraint declarations along with the team coding standard rules; The local exemption declaration is combined with the regular constraint declaration to generate the constraint instruction used to limit the code generation boundary; The structured requirement data is converted into task description information, and the task description information and the constraint instructions are injected into a preset prompt word template to obtain the context-aware prompt words.

[0011] Optionally, the abstract syntax tree is iteratively validated, and when a violation node is found, it is corrected by generating a large model from the code, until a preset stopping validation condition is met, specifically including: Traverse each node of the abstract syntax tree and, based on the constraint instructions, verify whether there are any illegal nodes in the abstract syntax tree; If the violation node exists in the abstract syntax tree, the source code fragment and violation type corresponding to the violation node are extracted, and a correction prompt word is generated in combination with the constraint instruction. The correction prompt word is then fed back to the pre-trained code generation model to generate candidate code fragments. The candidate code fragment is used as the initial code fragment to regenerate the abstract syntax tree and verify the illegal node until there is no illegal node in the abstract syntax tree or the preset maximum number of iterations for correction is met, at which point the verification stops.

[0012] Optionally, based on the constraint instructions, verifying whether there are any violating nodes in the abstract syntax tree specifically includes: Extract the node attribute information of each node to be verified in the abstract syntax tree; The node attribute information is matched with the regular constraint declaration according to rules. If the node attribute information violates the regular constraint declaration, the corresponding node to be verified is marked as a suspected violation node. Extract the code signature corresponding to the suspected violation node, and determine whether the code signature matches the conflicting signature indicated by the local exemption statement; If a violation is detected, the violation marker on the suspected violation node is removed, and it is determined to be a legitimate node. If no match is found, the suspected violation node is confirmed as the violation node. Extract all actual call signatures from the abstract syntax tree. If the utility class signature required by the constraint instruction is not included in the actual call signature, then it is determined that there is a violation node where the utility class signature is not reused.

[0013] In a second aspect, embodiments of this application provide a code generation system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the code generation system to perform the method described in the first aspect and any possible implementation thereof.

[0014] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a code generation system, cause the code generation system to perform the method described in the first aspect and any possible implementation thereof.

[0015] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a code generation system, cause the code generation system to perform the method described in the first aspect and any possible implementation thereof.

[0016] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: 1. By matching requirement feature vectors with a local engineering knowledge base containing utility class signatures and team coding style rules, and constructing context-aware prompts containing constraint instructions, the code generation boundaries are clearly defined during the model inference phase. This overcomes the problem of existing large models tending to "reinvent the wheel" and ignoring team standards due to a lack of engineering context, guiding the model to generate highly reusable local adapted code from the source. Simultaneously, the initial code snippets are parsed into an abstract syntax tree, and iterative verification and large model correction are performed on violation nodes that violate standards or fail to reuse utility class signatures. This closed-loop feedback mechanism based on the abstract syntax tree breaks the limitations of existing one-way open-loop generation, enabling the system to spontaneously detect and correct violations, thereby ensuring that the final output target code strictly conforms to enterprise-level engineering standards and compiles and runs smoothly, improving code generation efficiency.

[0017] 2. By identifying semantically conflicting subsets belonging to different underlying dependency modules in the initial candidate signature set, and generating a module dependency topology graph based on project configuration information when conflicts exist, the topological reachability and dependency weights of the target module to the source modules of each conflicting signature are analyzed. This eliminates unreachable or low-weight utility class signatures, solving the technical problem in large and complex projects where semantic retrieval alone can easily recall functionally similar but inaccessible utility classes due to module isolation, leading to dependency errors or compilation failures in the generated code of large models. By accurately resolving semantic ambiguities of utility classes at the architectural dependency level, it ensures that the target context information provided to large models is not only highly semantically matched but also legally usable in the engineering physical structure, improving the compilation pass rate and architectural compliance of the generated code.

[0018] 3. By parsing the metadata annotations of utility class signatures, lifecycle states and environment isolation tags are extracted. Combined with the target deployment environment identifier, unauthorized environment signatures are identified and eliminated. This overcomes the problem that conventional semantic retrieval can easily introduce test environment-specific (such as test stubs) or obsolete utility class errors into production code, leading to unauthorized environment, security vulnerabilities, or online operational failures. By implementing strict environment and lifecycle isolation during the context construction phase, it ensures that the utility class signatures provided to the large model fully comply with the permission requirements of the target deployment environment, improving the security, reliability, and environment adaptability of the generated code.

[0019] 4. By cross-validating the syntactic features of utility class signatures with the team's coding style rules, when a specification conflict event is identified where the utility class signature's self-defined format violates the specification, the conflicting signature and conflict rule are extracted, and a local exemption declaration is generated. By combining the local exemption declaration with regular constraint declarations to generate constraint instructions, the technical problem of incompatibility between legacy utility classes and current coding style rules in local projects is solved. This causes the large model to fall into logical paradoxes under strict double constraints (such as forcibly changing the utility class name to conform to the style rules, or abandoning the use of utility classes to comply with the style rules). The local exemption right of specific signatures is accurately granted during the prompt word construction stage, ensuring that the large model can accurately reuse existing utility classes without triggering subsequent style rule verification errors. This achieves the technical effect of balancing the compatibility of historical code and the consistency of overall style rules, and significantly improving the success rate of code generation. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating a code generation method in an embodiment of this application; Figure 2 This is a flowchart illustrating the process of determining target context information in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a code generation system according to an embodiment of this application; Explanation of reference numerals in the attached drawings: 301, Central Processing Unit; 302, Read-Only Memory; 303, Random Access Memory; 304, Bus; 305, Input / Output Interface; 306, Input Section; 307, Output Section; 308, Storage Section; 309, Communication Section; 310, Driver; 311, Removable Media. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0022] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0023] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0024] Figure 1 This is a flowchart illustrating a code generation method in an embodiment of this application.

[0025] Please see Figure 1 This application provides a code generation method, which includes: S101. Obtain the document description containing business requirements in the code generation task, and perform semantic parsing on the document description through a natural language processing model to obtain structured requirement data. In real-world software development scenarios, code generation tasks submitted by product managers or business personnel are typically in the form of unstructured natural language text, such as product requirement documents or work order descriptions in task tracking systems. Natural language text often contains a large amount of background information or descriptive vocabulary unrelated to the core code logic. If this natural language text is directly input into downstream retrieval systems or large code generation models, the redundant information will severely interfere with the accuracy of similarity matching, causing the retrieval system to recall irrelevant utility classes or triggering logical illusions in the large code generation model. Therefore, it is necessary to refine and structurally transform the original document descriptions.

[0026] In practice, the system first retrieves document descriptions containing business requirements from the code generation task on the R&D management platform. Next, the system calls a pre-deployed natural language processing (NLP) model to perform semantic parsing on the document descriptions. The NLP model can be a pre-trained information extraction model based on the Transformer architecture. Internally, the NLP model includes a named entity recognition module and a dependency parsing module. The named entity recognition module uses sequence labeling algorithms to identify words representing specific business entities from the document descriptions, such as table names, field names, or interface names. The dependency parsing module analyzes the grammatical dependencies between words in a sentence to extract the action predicate and its corresponding object. Through parsing by the NLP model, the system transforms the originally unstructured document descriptions into structured requirement data. This structured requirement data is typically stored in key-value pair formats such as JSON or XML, and clearly defines standardized fields such as core actions, operation objects, input parameters, and output results.

[0027] By extracting structured requirement data, the system eliminates useless noise in document descriptions, accurately translating ambiguous business language into machine-readable data that closely aligns with program logic. This structured requirement data provides a clean data foundation for subsequently transforming requirements into high-precision requirement feature vectors, significantly improving recall and accuracy when performing similarity matching with the local engineering knowledge base.

[0028] For example, suppose the document description obtained is: "When a user clicks the checkout button on the front end, the system needs to calculate the total amount of all discounted items in the shopping cart based on the user's membership level. If the total amount is less than zero, an error will be reported." After semantic parsing by a natural language processing model, the resulting structured requirement data can be represented as a JSON object containing: "Core action: calculation", "Operation object: total amount of items in the shopping cart", "Dependency parameter: membership level", and "Constraint condition: total amount is greater than or equal to zero".

[0029] S102. The structured requirement data is converted into a requirement feature vector, and the requirement feature vector is matched with a preset local engineering knowledge base to retrieve the target context information associated with the business requirement. The local engineering knowledge base includes the existing tool class signature set and team coding standard rules in the target project. After acquiring structured requirement data, in order to accurately retrieve reusable code assets from a massive local engineering knowledge base, the system needs to convert the text-based structured requirement data into a mathematical expression that can be directly computed by the computer. Specifically, the system inputs the structured requirement data into a pre-deployed text embedding model. A text embedding model is a vectorization tool based on deep learning. The principle of a text embedding model is to extract deep semantic features from the input text through a multi-layered neural network and map these deep semantic features into a high-dimensional continuous vector space, outputting a one-dimensional array of floating-point numbers, which is the requirement feature vector. For example, if the system inputs structured requirement data containing "calculate the total amount of items in the shopping cart" into the text embedding model, the text embedding model will output a requirement feature vector consisting of hundreds of floating-point numbers, such as [0.15, -0.82, 0.44, ...]. By transforming the data into requirement feature vectors, the system overcomes the limitations of traditional literal keyword matching, enabling subsequent searches to be truly based on the semantics of business intent. After obtaining the requirement feature vectors, the system performs similarity matching with a pre-defined local engineering knowledge base. This local knowledge base stores a set of existing utility class signatures and team coding standards rules from the target project. However, in large enterprise-level projects, there are often utility classes with similar functions but belonging to different underlying dependency modules. Relying solely on simple vector distance retrieval can easily lead to the recall of conflicting utility class signatures. Therefore, to ensure that the retrieved target context information is both semantically highly aligned with business requirements and legally usable within the project architecture, the system executes a specific retrieval process during the similarity matching phase, including conflict detection and elimination mechanisms. Figure 2 This is a flowchart illustrating the process of determining target context information in an embodiment of this application. The following is a summary of the process. Figure 2 Step S102 will be described in detail.

[0030] S201. Calculate the semantic similarity score between the demand feature vector and each tool signature in the tool signature set, and combine the tool signatures with semantic similarity scores greater than a preset matching threshold into an initial candidate signature set; To filter out code assets relevant to business needs from a massive local engineering knowledge base, the system needs to quantify the matching degree between the requirement feature vector and existing code assets. Specifically, the system iterates through the set of utility class signatures in the local engineering knowledge base, using a cosine similarity algorithm to calculate the semantic similarity score between the requirement feature vector and each utility class signature in the set. The principle of the cosine similarity algorithm is to evaluate the semantic closeness of two high-dimensional vectors by measuring the cosine value of the angle between them; the smaller the angle, the closer the semantics. After calculation, the system compares the semantic similarity score with a preset matching threshold. The preset matching threshold is a numerical limit set by developers to filter out noisy data with low relevance. The system extracts utility class signatures with semantic similarity scores greater than the preset matching threshold and combines them into an initial candidate signature set. By calculating semantic similarity scores and setting filtering conditions, the system quickly narrows the search scope, condensing massive code assets into a highly relevant preliminary candidate pool.

[0031] For example, assuming the preset matching threshold is 0.8, the system calculates that the semantic similarity score between the demand feature vector and the "DateFormatter" signature is 0.85, and the semantic similarity score between the demand feature vector and the "StringTrimmer" signature is 0.3. The system will add the "DateFormatter" signature to the initial candidate signature set and discard the "StringTrimmer" signature.

[0032] S202. Determine whether there is a semantic conflict subset in the initial candidate signature set. The semantic conflict subset is at least two of the tool class signatures in the initial candidate signature set whose semantic similarity scores with the requirement feature vector are less than a preset difference threshold, and the at least two tool class signatures belong to different underlying dependency modules in the local engineering knowledge base. After obtaining the initial candidate signature set, the system faces the common problem of name conflicts or functional overlap found in large enterprise-level projects. If all signatures in the initial candidate pool are directly provided to the code generation model, it can easily lead to incorrect package paths or call illusions. Therefore, the system needs to further determine whether there is a semantically conflicting subset in the initial candidate signature set.

[0033] In practice, the system compares each utility signature in the initial candidate signature set pairwise, calculating the difference in semantic similarity scores between the two utility signatures and the requirement feature vector. The system then determines if this difference is less than a preset threshold. This threshold is a very small floating-point number pre-set by developers based on statistical analysis of historical retrieval data, used to define whether two signatures are semantically indistinguishable. When at least two utility signatures are found to have a semantic similarity score difference less than the preset threshold, the system further parses the local engineering knowledge base to examine the physical affiliation of the at least two utility signatures, determining whether they belong to different underlying dependency modules. These underlying dependency modules are derived from the engineering architecture, such as modular directories in build tools, used to achieve physical isolation of code from different business domains. If at least two utility signatures have extremely similar scores and span different underlying dependency modules, the system classifies them as a semantically conflict subset. By investigating semantically conflict subsets, the system accurately identifies potential dependency ambiguity risks in the engineering architecture, providing a clear target for subsequent architecture-level filtering.

[0034] For example, assuming the preset difference threshold is 0.02, the initial candidate signature set contains the "OrderPriceUtil" signature with a score of 0.88 and the "FinancePriceUtil" signature with a score of 0.87. The difference in scores between the two utility signatures is 0.01, and the "OrderPriceUtil" signature belongs to the order module while the "FinancePriceUtil" signature belongs to the finance module. In this case, the system will classify the "OrderPriceUtil" signature and the "FinancePriceUtil" signature together into the semantic conflict subset.

[0035] S203. If the semantic conflict subset does not exist, then the initial candidate signature set is determined as the target candidate signature set; After completing the aforementioned semantic conflict investigation, the system will determine whether a semantic conflict subset exists. If the system determines that the semantic conflict subset does not exist, it means that in the initial candidate signature set, the semantic distinguishability of each utility class signature is high enough, or signatures with similar scores belong to the same underlying dependency module. This indicates that the currently retrieved code assets are clear and unambiguous in terms of engineering architecture and business semantics, and there is no risk of call ambiguity due to module isolation.

[0036] However, clarity in architecture and semantics does not equate to absolute runtime security. Even if a utility class signature is unambiguous in terms of semantics and module affiliation, it may still be a mock interface specifically for testing environments or legacy code that has been marked as obsolete. If the initial set of candidate signatures is provided directly to the large model after confirming there are no semantic conflicts, it can easily lead to the large model generating code that causes unauthorized errors or security vulnerabilities in the production environment. Therefore, before determining the initial candidate signature set as the target candidate signature set, in order to prevent the large code generation model from incorrectly calling test-specific simulation data tools or deprecated old interfaces, causing online operation failures or security vulnerabilities, the system also needs to execute an environment and lifecycle security interception mechanism, which specifically includes the following steps: obtaining the target deployment environment identifier corresponding to the code generation task and parsing the metadata annotations of each tool class signature in the local engineering knowledge base; based on the metadata annotations, extracting the lifecycle status and environment isolation label of each tool class signature in the initial candidate signature set; determining whether there is an environment unauthorized signature in the initial candidate signature set, wherein the environment unauthorized signature is a signature whose semantic similarity score is greater than or equal to the preset matching threshold and whose environment isolation label conflicts with the target deployment environment identifier; if the environment unauthorized signature exists, then the environment unauthorized signature is removed from the initial candidate signature set.

[0037] After completing the aforementioned semantic conflict investigation, to prevent the large code generation model from incorrectly calling test-specific simulation data tools or deprecated interfaces when generating production environment code, thereby causing online operational failures or security vulnerabilities, the system needs to introduce an environment and lifecycle security interception mechanism. Specifically, the system retrieves the target deployment environment identifier corresponding to the code generation task from the continuous integration pipeline configuration or development task work order. The target deployment environment identifier is a string used in the development process to distinguish the final server category for code execution, such as "PROD" for production environment or "TEST" for testing environment. Simultaneously, the system calls the code parser to read and parse the metadata annotations of each utility class signature in the local project knowledge base. Metadata annotations are special marker codes appended to code declarations. Their principle is to provide additional configuration instructions to the compiler or runtime framework through specific syntax structures, declaring special attributes of the code without changing its core logic. By obtaining the environment identifier and parsing the annotations, the system provides the necessary configuration data support for subsequent accurate identification of dangerous code. For example, the system reads the target deployment environment identifier of the current task as "PROD" from the work order, and parses out the metadata annotations "@Environment("TEST")" and "@Deprecated" above the signature of a utility class in the local project knowledge base.

[0038] After obtaining this basic configuration data, the system needs to transform the original annotation text into specific attributes that can be used for logical judgment. In practice, based on the parsed metadata annotations, the system uses regular expressions or attribute mapping algorithms to extract the lifecycle state and environment isolation label of each utility class signature in the initial candidate signature set. The lifecycle state originates from software engineering version iteration management and is used to indicate whether the code is currently active, awaiting obsolescence, or obsolete. This is achieved by issuing status warnings to developers or compilers through specific annotation keywords. The environment isolation label originates from multi-environment deployment architectures and is used to explicitly define the specific server environments in which the code is allowed to be called. This is achieved by controlling the instantiation scope of the code through the dependency injection mechanism at the framework layer. By extracting these two key attributes, the system transforms the originally static code comments into dynamic security verification rules. For example, for the annotation parsed in the previous step, the system extracts the lifecycle state of the utility class signature as "obsolete" and its environment isolation label as "TEST environment only".

[0039] After extracting the specific verification rules, the system immediately conducts a rigorous cross-comparison to screen for potential unauthorized calls. Specifically, the system performs string matching comparison between the extracted environment isolation label and the obtained target deployment environment identifier, while simultaneously checking if the lifecycle status is active. Based on this, the system determines whether an environment-unauthorized signature exists in the initial candidate signature set. An environment-unauthorized signature refers to a signature that, although its semantic similarity score calculated in the previous retrieval stage is greater than or equal to a preset matching threshold, seemingly perfectly matching business requirements, has an environment isolation label that conflicts with the target deployment environment identifier, or a lifecycle status marked as obsolete. By executing this judgment logic, the system can keenly detect dangerous code that disguises itself with high semantic similarity but actually lacks the necessary calling permissions. For example, since the target deployment environment identifier is "PROD," while the environment isolation label of a certain utility class signature is "TEST," the two are clearly inconsistent, and the system determines that this high-relevance signature with a score of 0.9 is an environment-unauthorized signature.

[0040] After identifying these risky signatures, decisive isolation measures must be taken to ensure that the context ultimately input to the large model is absolutely safe and reliable. Specifically, if the system determines that an environment unauthorized signature exists, it will directly perform a memory deletion operation, completely removing the environment unauthorized signature from the initial candidate signature set. Through this removal action, the system cuts off the injection path of dangerous code in the early stages of context construction, effectively overcoming the technical defect of conventional semantic retrieval that easily introduces test stubs or obsolete interface errors into production code. This ensures that the utility classes provided to the large model are not only highly semantically matched but also fully comply with the permission and security requirements of the target deployment environment. For example, the system will permanently delete the aforementioned environment unauthorized signature with the "TEST" tag from the list of initial candidate signatures, ensuring that the large model will never see or call this test-specific utility class when generating the settlement logic for the production environment. It should be noted that although this application embodiment uses annotations with specific syntax symbols (such as @Environment) as an example for illustration, those skilled in the art should understand that any programming language and engineering architecture that adopts a similar metadata tagging mechanism falls within the protection scope of this application.

[0041] After the aforementioned architectural conflict investigation and rigorous environmental and lifecycle security interception, any remaining dangerous or ambiguous code in the initial candidate signature set has been thoroughly cleaned. To provide a stable and clean data source for subsequent context assembly and prevent the introduction of unverified dirty data in later processing, the system needs to solidify the cleaned results. Specifically, the system locks and re-encapsulates the remaining utility class signatures after all pre-filtering conditions in memory. The system moves the remaining signature data from the dynamic temporary cache to the read-only context building area and tags it with the internal state "verified and usable," thus formally instantiating the remaining signature data into the target candidate signature set at the data structure level. The target candidate signature set is a structured data container, derived from a collection of high-quality code assets that have undergone multiple rounds of cleansing, intended as a secure code knowledge base ultimately provided to the large model. By performing state locking and re-encapsulation, the system completes a closed loop from massive fuzzy retrieval to precise and secure matching, ensuring that the utility class signatures referenced when constructing subsequent prompt words are absolutely pure, unambiguous, and compliant. This lays a solid data foundation for generating high-quality code for large models. For example, assuming the initial candidate signature set originally contained five signatures, after confirming there were no semantic conflicts and removing an environment-unauthorized signature with a test label, the system moves the remaining four fully compliant signatures to a read-only memory area, formally establishing them as the target candidate signature set, awaiting final assembly with the team's coding standards and rules.

[0042] S204. If the semantic conflict subset exists, the utility class signatures that do not meet the calling constraints are removed from the initial candidate signature set to obtain the target candidate signature set. If the system determines that a subset of semantic conflicts exists, it means that in the initial candidate signature set, there are utility class signatures with extremely similar semantic scores but whose physical affiliation spans different underlying dependency modules. If signatures with architectural ambiguity are simultaneously provided to the large code generation model, the large model is prone to selection difficulties and may even generate illegal cross-module call code that violates the engineering architecture isolation principle. Therefore, the system must perform strict architecture-level filtering to remove utility class signatures that do not meet the call constraints from the initial candidate signature set, thereby obtaining a unique and legal target candidate signature set. In order to accurately determine which signatures truly meet the call constraints, the system cannot only stay at the semantic level, but also needs to deeply analyze the physical location of the current code generation task in the entire engineering architecture, and combine it with the actual inter-module dependencies within the project. By constructing a global dependency topology, the legality and priority of the call path are quantitatively evaluated, thereby completing the accurate removal operation. The specific implementation steps are as follows.

[0043] Obtain the target project directory path corresponding to the code generation task; generate a module dependency topology graph based on the project configuration information in the local project knowledge base; analyze the topology reachability and dependency weight of the target module where the target project directory path is located in the module dependency topology graph, and respectively reach the source module where each of the utility class signatures in the semantic conflict subset is located; remove the utility class signatures in the semantic conflict subset whose topology reachability is unreachable or whose dependency weight is less than a preset weight threshold from the initial candidate signature set to obtain the target candidate signature set.

[0044] To accurately resolve the architectural semantic conflicts identified in the preceding steps, the system must clearly define the specific location of the code to be generated within the overall physical structure of the project. In practice, the system reads and retrieves the target project directory path corresponding to the code generation task from the context of the integrated development environment or the task configuration of the development management platform. The target project directory path is a string indicating the file system hierarchy, used to pinpoint the final output location of the large code generation model. By clearly defining the target project directory path, the system establishes a unique starting coordinate for subsequently determining the legality of cross-module calls. For example, if the system reads the target project directory path of the current code generation task as " / src / modules / order-service / " from the task configuration, it means that the code to be generated needs to be stored under the order service module.

[0045] After establishing the starting coordinates, the system needs to understand the intricate relationships between the various modules within the entire project to plan the call routes, much like navigation software. In practice, the system invokes a dependency resolution engine to read project configuration information from the local project knowledge base. This configuration information originates from the build tool's configuration files, such as Maven's pom.xml or Node.js's package.json, and declares the compilation and runtime dependencies between various code modules. The system parses the dependency declarations in the project configuration information, abstracting each independent code module as a node in graph theory and the reference relationships between modules as directed edges, thereby generating a module dependency topology graph in memory. A module dependency topology graph is a mathematical model built on a directed acyclic graph data structure, used to transform static text configurations into a dynamic network that a computer can directly execute for path searching. By generating the module dependency topology graph, the system obtains a global architectural map, providing underlying data support for subsequent quantitative evaluation of call legitimacy. For example, after parsing the configuration file, the system generates a graph that clearly shows the "Order Service Module" has a unidirectional dependency on the "Basic Tools Module," while there are no connecting edges between it and the "Financial Service Module."

[0046] With a global architectural map, the system can then conduct specific route surveys for conflicting candidate signatures. In practice, the system first maps the target project directory path to target module nodes in the module dependency topology graph, and simultaneously maps the physical affiliation of each utility class signature in the semantic conflict subset to its corresponding source module node. Subsequently, the system runs graph traversal algorithms, such as breadth-first search, to analyze the topological reachability and dependency weights of the target module to each source module in the module dependency topology graph. Topological reachability, derived from the concept of connectivity in graph theory, is used to determine whether a valid call path exists between two nodes; it works by searching for nodes along directed edges. Dependency weights are quantified scores pre-set by developers for different types of dependencies, used to measure the depth of coupling between modules; for example, direct dependencies have higher scores than indirect transitive dependencies. Through path analysis, the system precisely quantifies the originally vague architectural constraints into comparable mathematical indicators. For example, system analysis revealed that the topological reachability from the "Order Service Module" to the source module containing the "OrderPriceUtil" signature is "reachable," and because it is a direct dependency, the dependency weight is 10; however, the topological reachability to the source module containing the "FinancePriceUtil" signature is "unreachable," because there is no configuration dependency between the two.

[0047] Having obtained the quantified route survey metrics, the system possesses the basis for executing the final architecture-level filtering to completely eliminate contextual ambiguity provided to the large code generation model. In practice, the system iterates through each utility class signature in the semantic conflict subset, checking the topological reachability and dependency weights of the corresponding source module. The system directly identifies utility class signatures with unreachable topological reachability as illegal calls and blocks them; simultaneously, the system compares the dependency weights with a preset weight threshold. The preset weight threshold is a minimum score set by the architect based on the principle of engineering decoupling, intended to prevent the large code generation model from introducing excessively long or weak cross-level dependencies. The system removes utility class signatures with dependency weights below the preset weight threshold, along with those with unreachable topological reachability, from the initial candidate signature set. After this rigorous architecture-level cleaning, the system locks the state of the remaining fully compliant signature data, ultimately obtaining the target candidate signature set. By performing the removal operation, the system perfectly resolves the interference caused by utility classes with the same name or similar functions, ensuring that the large code generation model can only see the code assets that the current module truly has the authority to call. For example, assuming the preset weight threshold is 5, the system will completely remove the topologically unreachable "FinancePriceUtil" signature and only retain the "OrderPriceUtil" signature with a weight of 10, thus establishing the cleaned set containing only the "OrderPriceUtil" signature as the target candidate signature set.

[0048] S205. Assemble the target candidate signature set with the team coding standard rules to obtain target context information related to the business requirements.

[0049] After successfully obtaining a set of target candidate signatures that has undergone multiple rounds of rigorous cleaning and filtering, in order to ensure that the code generation model can not only correctly call existing local security code assets, but also strictly adhere to the company's internal R&D rules, the system needs to deeply integrate the target candidate signature set with the team's coding style guidelines. These team coding style guidelines originate from the company's internal R&D management system and are used to unify the coding styles of different developers and avoid common vulnerabilities. They are a set of constraints in the form of structured text, such as variable naming conventions or exception handling standards.

[0050] In practice, the system reads team coding standard rules from the local engineering knowledge base and launches a text template engine to perform an assembly operation. This assembly operation is not a simple physical string concatenation; instead, it uses preset data exchange format templates, such as JSON or Markdown, to assign clear semantic labels to the target candidate signature set and the team coding standard rules. The system fills the target candidate signature set into the "Available Tool Interfaces" tag in the template, and simultaneously fills the team coding standard rules into the "Mandatory Coding Constraints" tag, thus transforming the two heterogeneous data sets into a well-structured and hierarchical comprehensive text, ultimately obtaining target context information relevant to business requirements. This target context information is a background knowledge base specifically tailored to the cognitive habits of large models, serving as the core material for building prompts in subsequent steps. By performing structured assembly operations, the system provides both the material basis for "what tools can be used" and the institutional boundaries for "how to write code" for large-scale code generation, effectively avoiding the problems of chaotic coding style or reinventing the wheel caused by free rein in large models, greatly improving the engineering usability and compliance of the generated code.

[0051] For example, the system will combine the cleaned and retained "OrderPriceUtil" utility class signature with the team coding convention rule "all local variables must use camelCase" read from the local project knowledge base, and assemble them in Markdown format to generate a target context information containing "### Available API: OrderPriceUtil" and "### Coding Convention: CamelCase", which will be used to build context-aware prompts later.

[0052] S103. Based on the structured requirement data and the target context information, construct context-aware prompt words. The context-aware prompt words contain constraint instructions for limiting the boundaries of code generation. The constraint instructions instruct the code generation model to call the utility class signature in the target context information and follow the team coding standard rules. After successfully acquiring structured requirement data and target context information, the system needs to construct context-aware prompts based on these data. These prompts contain constraint instructions to limit the boundaries of code generation, instructing the code generation model to call utility class signatures from the target context information and adhere to team coding standards. However, in real-world enterprise software development scenarios, historical utility class signatures retrieved from the local engineering knowledge base often carry technical debt. The naming or format of these historical signatures may have already violated the latest team coding standards established by the enterprise. If the system directly throws contradictory calling requirements and specification constraints at the code generation model, it can easily lead to logical confusion during code generation, creating the illusion that it should either forcibly modify historical interfaces or violate the latest standards. Therefore, in order to ensure that the code generation model can receive clear and consistent instructions, the system cannot simply perform text concatenation when constructing context-aware prompts. Instead, it must conduct a deep check on the compatibility of the target context information before assembling the prompts. The final constraint instructions are generated through a sophisticated conflict resolution and exemption declaration mechanism. The specific implementation steps are as follows.

[0053] The tool class signature and the team coding standard rules are extracted from the target context information. The syntax features of the tool class signature and the team coding standard rules are cross-validated to determine whether there is a standard conflict event. The standard conflict event is that the definition format of the tool class signature itself violates the team coding standard rules. If the standard conflict event exists, the conflicting signature that triggered the standard conflict event and the corresponding conflicting rule are extracted. A partial exemption statement is generated for the conflicting signature. The partial exemption statement is used to instruct the code generation model to ignore the conflicting rule when calling the conflicting signature. The non-conflicting signatures in the tool class signature and the team coding standard rules are converted into regular constraint statements. The partial exemption statements and the regular constraint statements are combined to generate the constraint instructions used to limit the boundaries of code generation. The structured requirement data is converted into task description information, and the task description information and the constraint instructions are injected into a preset prompt word template to obtain the context-aware prompt words.

[0054] To ensure the instructions provided to the code generation model are logically consistent, the system first needs to perform an internal inconsistency check on the materials to be issued. Specifically, the system uses a text parsing engine to extract the utility class signatures and team coding standards rules from the target context information. Then, the system calls a static code analysis tool to extract the syntactic features of the utility class signatures, such as method name naming style and parameter list order, and cross-validates these features with the team coding standards rules to determine if any conflicts exist. A conflict refers to a situation where, due to historical technical debt or other reasons, the definition format of the utility class signatures inherited in the local engineering knowledge base violates the company's latest team coding standards rules. By performing cross-validation, the system can identify those "flawed" historical interfaces before prompts are generated, preventing the large model from becoming confused when encountering contradictory materials. For example, the system extracts a utility class signature named "get_user_info" and a team coding convention rule that "all method names must use camelCase". After cross-validation, the system finds that "get_user_info" uses underscores, which clearly violates the camelCase rule, and the system determines that there is a convention conflict.

[0055] After accurately locating the internal conflict points, the system must implement targeted isolation measures to protect the large model from being misled by specification rules when calling historical interfaces. Specifically, if the system determines that a specification conflict event exists, it uses regular expressions or abstract syntax tree node mapping techniques to accurately extract the conflict signature that triggered the conflict event and the corresponding conflict rule. Next, the system uses natural language generation algorithms to generate local exemption statements for the conflict signature. A local exemption statement is a special type of engineering instruction used to create a "rule special zone" for specific legacy code. The principle is to break the absolute constraints of global rules through explicit conditional clauses, instructing the code to generate a large model that ignores conflict rules when calling conflict signatures, but still strictly adheres to the specification when writing other new code. By generating local exemption statements, the system cleverly resolves the contradiction between reusing historical assets and implementing the latest specifications, ensuring that the large model can call old interfaces unchanged without arbitrarily modifying their names, thus preventing compilation errors. For example, in response to the aforementioned naming conflict, the system extracts the "get_user_info" signature and the "lower camelCase" naming rule, and generates a partial exemption declaration: "Note: The lower camelCase naming rule is allowed to be ignored only when calling the get_user_info interface. Please call it as is."

[0056] After handling special cases, the system needs to transform the remaining normal materials into global constraints that the large model can understand. In practice, the system filters out conflict-free signatures from the tool class signatures and converts these conflict-free signatures, along with team coding style guidelines, into regular constraint declarations. These regular constraint declarations are standard text paragraphs in the prompts used to define the boundaries of the large model's normal behavior; their purpose is to clarify the coding principles and available tool pool for the large model in most cases. Subsequently, the system uses text concatenation and hierarchical formatting techniques to logically combine local exemption declarations and regular constraint declarations, generating constraint instructions to limit code generation boundaries. These constraint instructions not only include global mandatory requirements but also nested exception clauses for specific interfaces, forming a rigorous and self-consistent rule system. By combining these two types of declarations, the system constructs an execution framework for the large model that is both principled and flexible, completely eliminating logical dead ends within the prompts. For example, the system converts the normal "OrderPriceUtil" signature and specification rules into a regular constraint declaration, and combines it with the aforementioned local exemption declaration to generate a complete constraint instruction: "Global rule: must use camelCase naming; available tools: OrderPriceUtil, get_user_info; special case: exempt camelCase rule when calling get_user_info."

[0057] After completing the self-consistency processing of all constraints, the system enters the final assembly stage of prompt word construction. Specifically, the system calls a data formatting script to convert the previously acquired structured requirement data into task description information in natural language. This task description information is stripped of the complex data structures required for machine processing, transforming it into a business logic description that is easier for the large model to understand. Next, the system injects the task description information and the previously generated constraint instructions into a preset prompt word template. This preset prompt word template is a text skeleton with placeholders, pre-written by the development team based on the cognitive characteristics of the large model. Its purpose is to standardize the structure of the prompt words, ensuring that modules such as task background, available tools, and constraints are presented to the large model in the optimal order. By injecting dynamic data into a static template, the system finally obtains context-aware prompt words. These context-aware prompt words not only contain the core business logic that the developers want to implement but also perfectly integrate the deeply cleaned and conflict-exempted local project context, enabling the large model to begin writing code with a complete understanding of the current project status and rules. For example, the system takes the task description information of "implementing user login verification" and the constraint instructions containing exemption clauses, and fills them into the "[Task Objective]" and "[Constraint Condition]" placeholders of the preset prompt word template, respectively, to generate a context-aware prompt word with a complete structure and rigorous logic, which is ready to be sent to the pre-trained code generation model to perform the generation task.

[0058] S104. Input the context-aware prompts into the pre-trained code generation model to generate initial code snippets; After completing the refined construction of context-aware prompts, the system needs to leverage a computing engine with powerful natural language understanding and programming language conversion capabilities to truly transform human-readable business requirements and engineering constraints into low-level logic code that can be executed by computers.

[0059] In practice, the system first establishes a communication connection with the pre-trained code generation model. The pre-trained code generation model is a massive neural network built on a deep learning architecture. It is an algorithm model trained over a long period of time on a vast amount of open-source code libraries and programming question-and-answer data. Its purpose is to act as a virtual senior programmer. It learns the syntax rules and logical patterns of code, and predicts and outputs the code character sequence that best fits the probability distribution based on the input text prompts.

[0060] The system initiates a network call request via an application programming interface (API), transmitting and inputting the carefully assembled context-aware prompts from the preceding steps as core parameters into a pre-trained code generation model. To ensure that the code generation model strictly adheres to the constraints contained in the context-aware prompts, rather than over-diverging or creating code illusions, the system simultaneously configures the generation control parameters of the code generation model when initiating the call request. The system lowers the temperature parameter of the code generation model; the temperature parameter is a hyperparameter used to control the randomness of the model's output. Lowering the temperature parameter forces the code generation model to prioritize the deterministic grammatical structure with the highest probability and best conformity to the constraints when generating code. After receiving the context-aware prompts, the code generation model performs complex attention mechanism calculations within the neural network, fusing and reasoning with business requirements and constraints, and finally outputting initial code fragments containing specific business logic line by line. Through the execution of input and generation operations, the system achieves a leap from natural language description to machine programming language. The output code is called the initial code snippet because although the initial code snippet literally responds to the requirements of the context-aware prompt words, it has not yet undergone strict validation at the underlying syntax level of the system. The initial code snippet still needs to be used as the raw material for subsequent iterative validation steps.

[0061] For example, the system inputs a context-aware prompt containing "To implement user login verification, OrderPriceUtil must be called, following camelCase naming" into a pre-trained code generation model via an application programming interface (API), while setting the temperature parameter to an extremely low value. The code generation model, after inference, outputs text containing "public boolean verifyUserLogin() {OrderPriceUtil.calculate();}". The system receives and caches this output text, establishing it as the initial code snippet, ready to be sent to the next stage for abstract syntax tree validation.

[0062] S105. Parse the initial code snippet into an abstract syntax tree, iteratively verify the abstract syntax tree, and correct it by generating a large model through the code when there are violation nodes, until a preset stop verification condition is met. The preset stop verification condition includes that there are no violation nodes in the abstract syntax tree or that the preset maximum number of iterations for correction is met. The violation node is a node that violates the team coding standard rules, or the corresponding parent scope node marked when no call action of the utility class signature in the target context information is detected. The initial code snippet output by the large code generation model is essentially a string of plain text characters based on probability prediction. Due to the inherent uncontrollability and illusion problem of the large model, the initial code snippet is very likely to miss the mandatory requirements in the context-aware prompts, such as forgetting to call the specified utility class signature, or variable naming violating the team's coding style rules. If only simple text regular expression matching is used, the system will find it difficult to accurately identify complex logical errors and scope nesting issues. Therefore, the system needs to call the compiler front-end tool to parse the initial code snippet into an abstract syntax tree. An abstract syntax tree is a tree-like data structure that represents the syntactic structure of the source code in a tree-like form, which can accurately locate the hierarchical relationship of each variable declaration or method call. After the code is transformed into a tree structure, the system has the foundation to perform deep code checking. In order to ensure that the final output code is absolutely compliant and fully reuses local project assets, the system cannot just stop at the stage of finding errors, but must establish an automated closed-loop error correction mechanism. The system needs to delve into the internals of the abstract syntax tree to accurately identify all violation nodes. When a violation node is found, the system guides the code generation model to perform targeted self-correction. Through continuous looping of checks and modifications, the system continues until the code fully meets the standards or triggers a forced stop condition. Specifically, this may include steps S1051-S1053.

[0063] S1051. Traverse each node of the abstract syntax tree and, based on the constraint instructions, verify whether there are any illegal nodes in the abstract syntax tree; After successfully converting the initial code snippet into an abstract syntax tree, the system grasps the hierarchical structure of the code. However, due to the complex constraint instructions introduced during the initial construction of the prompt words, which include both global specifications and special case exemptions, a one-size-fits-all rule matching approach is prone to misjudgments. To ensure that the verification process strictly adheres to the team's coding style guidelines while fully respecting the specific characteristics of legacy code, the system must delve into the internals of the abstract syntax tree to perform refined checks. Specifically, the steps include: extracting node attribute information of each node to be verified in the abstract syntax tree; matching the node attribute information with regular constraint declarations according to rules; if the node attribute information violates the regular constraint declaration, the corresponding node to be verified is marked as a suspected violation node; extracting the code signature corresponding to the suspected violation node and determining whether the code signature matches the conflicting signature indicated by the local exemption declaration; if it matches, removing the violation mark of the suspected violation node and determining it as a legal node; if it does not match, confirming the suspected violation node as the violation node; extracting all actual call signatures in the abstract syntax tree; if the utility class signature required by the constraint instruction is not included in the actual call signature, it is determined that there is a violation node that has not reused the utility class signature.

[0064] To perform a highly accurate "check-up" on the code generated by the large model, the system must delve into the code's skeleton, rather than merely performing surface-level text comparison. Specifically, the system initiates a syntax tree traversal algorithm, visiting each node to be verified in the abstract syntax tree one by one, and extracting node attribute information from each node. Node attribute information is structured metadata generated by the compiler when parsing the code; it is the parsing result of the parser and precisely describes the semantic role of the current code block, mainly including the identifier names of variables or methods and the call relationships between modules. By extracting low-level attributes, the system breaks down the originally continuous code string into independent logical units with clear semantic labels, providing high-precision target data for subsequent rule comparisons. For example, when the system traverses to the line of code "get_user_info()", it extracts the node attribute information from the corresponding node whose identifier name is "get_user_info" and whose call relationship is "external function call".

[0065] After acquiring high-precision target data, the system then uses the team's standard benchmarks to evaluate the compliance of code units. In practice, the system rigorously matches the extracted node attribute information against the standard constraint declarations generated in the preceding steps. These standard constraint declarations represent the company's universally accepted coding bottom line. The system uses regular expressions or a pre-defined logic verification engine to check the naming style and parameter format of identifiers one by one to ensure they meet the requirements. If the system finds that node attribute information violates the terms of the standard constraint declaration during the comparison process, to prevent false positives, the system will not immediately classify the node to be verified as erroneous. Instead, it will first mark the corresponding node with an internal tag, classifying it as a suspected violation node. By performing a global screening, the system can quickly and comprehensively identify all seemingly non-compliant code snippets, greatly narrowing the scope of subsequent in-depth investigations. For example, the system finds that the identifier "get_user_info" extracted earlier uses underscores, while regular constraint declarations require the use of camelCase naming. The system then determines that the two are inconsistent and marks the function call node containing "get_user_info" as a suspected violation node.

[0066] After identifying suspected violation nodes, the system recognizes that within the company's vast historical codebase, there exist some older interfaces that have been granted special permission to continue operating despite their defects due to technical debt. Therefore, a careful secondary review is necessary. Specifically, the system further extracts the code signature corresponding to the suspected violation node. The code signature refers to the complete declaration or call format of the node in the code. The system compares the code signature with the local exemption declaration generated during the warning word construction phase to determine if the code signature matches a conflicting signature indicated by the local exemption declaration. A local exemption declaration acts like a special permit, protecting historical assets that are known to be non-compliant but must be called. By performing this special case verification, the system opens a green channel outside of strict global rules, effectively preventing new rules from unnecessarily affecting older code. For example, if the system extracts the code signature of a suspected violation node as "get_user_info," it then reviews the local exemption declarations to check if "get_user_info" is on the whitelist of allowed conflicting signatures. Specifically, if the extracted code signature is completely identical in character sequence to any conflicting signature in the local exemption declaration whitelist, the system determines it as a hit; otherwise, if no matching character sequence is found after traversing the whitelist, the system determines it as a miss.

[0067] Based on the aforementioned special case verification and comparison, if the system finds a match, it means that the seemingly non-compliant code is actually a legacy interface that the system actively requested the large model to call in the warning message. To protect the reuse of legitimate technical debt, the system will immediately remove the non-compliance mark from the suspected non-compliant node in memory and officially determine the suspected non-compliant node as a legitimate node. By performing the exemption operation, the system effectively prevents the strict specification verification mechanism from mistakenly killing the old core business logic that must be retained. For example, because "get_user_info" does exist in the whitelist of the local exemption declaration, the system removes the non-compliance suspicion that "get_user_info" was burdened with due to its underscore naming, and determines that "get_user_info" is completely legitimate in the current context.

[0068] Conversely, in special case verification and comparison, if the system finds a match, it means that the current code neither conforms to the team's latest established coding standards nor is it on the system's specially approved exemption whitelist. Matches usually occur because the large model has the illusion of free rein when generating code, or has forgotten the constraints in the prompt. In this case, the system will strip the suspected violation node of its suspected status, unhesitatingly confirming it as a genuine violation node. Through rigorous falsification logic, the system ensures that every violation node identified is an undeniable error, providing a reliable target for subsequent precise correction. For example, if the large model creates an underscore variable named "calculate_price," since "calculate_price" is not on the exemption list, the system will definitively classify "calculate_price" as a violation node.

[0069] Besides strictly preventing large-scale models from writing non-compliant code, the system also needs to guard against passive negligence such as cutting corners or reinventing the wheel. In practice, the system shifts its perspective, retracing the abstract syntax tree (AST) at a macro level to extract all actual call signatures. Actual call signatures refer to the set of function or method names actually executed by the large-scale model in the generated code. The system then checks the inclusion relationships between these actual call signatures and the utility class signatures mandated by the constraint instructions. If the system finds that a utility class signature required by the constraint instructions is not included in the extracted actual call signatures, it means the large-scale model has ignored business requirements and failed to reuse existing secure code assets from the local project. The system will then decisively determine that there is a violation node with an unreused utility class signature. By adding this layer of defense, the system not only ensures the surface compliance of the generated code but also enforces deep reuse of existing enterprise project assets at the underlying logic level. For example, the constraint instruction requires that "OrderPriceUtil" must be called to calculate the price. However, after the system extracts all the actual call signatures, it finds that the large model has written a bunch of addition, subtraction, multiplication and division logic itself and has not called "OrderPriceUtil" at all. The system then determines that there is a violation node that has not reused the specified tool and prepares to send the code back for redo.

[0070] It's worth noting that when executing the aforementioned check and correction logic, if the system determines that a utility class signature in the target context information is not reused, it cannot directly mark the missing node because the missing call action does not have a corresponding entity node in the abstract syntax tree. Therefore, the system traces upwards, marking the parent node that should contain the missing logic, such as the method declaration node or the root node of the code block, as a violation node. In the subsequent step of extracting source code snippets for correction, the system extracts the complete code snippet corresponding to the marked parent node and generates a correction prompt based on the violation type, such as "missing specific utility class call," which is then fed back to the code generation model. This prompts the model to perform a global rewrite or local insertion of code within the scope of the parent node. By marking and extracting parent nodes, the system provides the large model with a complete context scope, ensuring that the large model can correctly handle variable dependencies and logical nesting relationships when supplementing missing utility class calls.

[0071] S1052. If the violation node exists in the abstract syntax tree, extract the source code fragment and violation type corresponding to the violation node, combine it with the constraint instruction to generate a correction prompt word, and feed the correction prompt word back to the pre-trained code generation big model to generate candidate code fragments. After identifying the violation node, the system cannot simply discard the entire code segment. Instead, it needs to guide the code generation model to perform targeted self-correction. Specifically, if the system determines that a violation node exists in the abstract syntax tree (API), it uses the node mapping relationship of the API to extract the corresponding source code fragment and the specific violation type. The violation type is an error label assigned by the system during the verification phase, used to clearly indicate the root cause of the code error, such as incorrect naming or a missing utility class. Subsequently, the system concatenates and logically reorganizes the extracted source code fragment and violation type with the constraint instructions built in the previous steps to generate correction prompts. Correction prompts are feedback instructions specifically for error correction, used like a tutor grading homework, clearly telling the large model where the code is wrong and the correct direction for modification. Next, the system feeds the correction prompts back to the pre-trained code generation model via a network interface. Upon receiving the correction prompts, the code generation model adjusts its internal attention weights accordingly, re-infers and generates the erroneous code section, and finally outputs the modified candidate code fragment. By performing targeted feedback and regeneration operations, the system achieves automated repair of code defects, preventing large models from repeating the same mistakes in subsequent generation.

[0072] For example, the system extracts the non-compliant code snippet "calculate_price" and the violation type "violation of camelCase naming convention", and generates a correction prompt based on the constraint instructions: "The code snippet 'calculate_price' violates the camelCase naming convention, please modify it". After receiving the prompt, the large model re-outputs the candidate code snippet "calculatePrice".

[0073] S1053. Using the candidate code fragment as the initial code fragment, regenerate the abstract syntax tree and verify the illegal node until there is no illegal node in the abstract syntax tree or the preset maximum number of iterations for correction is met, then stop the verification.

[0074] After obtaining candidate code snippets after the large model has been repaired, the system understands that modifications to the large model are also uncontrollable. Fixing one error may introduce new syntax errors, therefore a rigorous closed-loop verification mechanism must be established. In practice, the system treats the newly generated candidate code snippets as initial code snippets, calls the compiler front-end tool to parse the code into a completely new abstract syntax tree, and restarts the preliminary fine-grained verification process, checking each node in the abstract syntax tree for any remaining violations. The system places the recurring process of "generation, verification, feedback, and regeneration" into a loop control logic, terminating the loop only when a preset stopping verification condition is met. The preset stopping verification condition includes two cases: first, the system finds no more violations in the latest generated abstract syntax tree, meaning the code has met the standard; second, the loop count reaches the preset maximum number of iterations for correction. The preset maximum number of iterations for correction is a safety threshold set by the developers to prevent the system from falling into an infinite loop. Its purpose is to protect computing resources by promptly triggering a circuit breaker retry mechanism when the large model repeatedly fails to be modified due to limited understanding. By performing closed-loop iterative verification operations, the system not only adds multiple layers of insurance for code quality, but also takes into account the stability and efficiency of project operation.

[0075] For example, the system re-parses and verifies the modified candidate code snippets. If it finds that there is still an error of missing utility classes, the system will initiate a second correction loop. The system will continue until the third loop when there are no non-compliant nodes in the abstract syntax tree, or the set maximum retry limit of five times is reached. At this point, the system will immediately stop the verification process.

[0076] S106. If the violation node does not exist in the abstract syntax tree when the verification stops, the corresponding code segment is determined as the target code; When the iterative verification process terminates normally because the abstract syntax tree no longer contains any non-compliant nodes, it means that the current code snippet has perfectly passed the dual tests of enterprise-level specifications and local project asset reuse. In practice, the system reads the candidate code snippets currently in memory and performs code formatting and final encapsulation. Since the code has undergone a deep scan of the abstract syntax tree, confirming that it neither violates team coding style rules nor completely reuses the utility class signature in the target context information, the system officially designates the corresponding code snippet as the target code. The target code is the final deliverable of the entire code generation task, used for direct integration into the actual business project for compilation and execution. By establishing the compliant code as the target code, the system ensures the engineering usability of the output results and greatly reduces the workload of manual secondary review.

[0077] For example, after the third iteration, the system finds that all nodes in the abstract syntax tree conform to camelCase naming and correctly call OrderPriceUtil. It then marks the corresponding code snippet as target code and pushes it to the developer's editor interface.

[0078] S107. If the violation node still exists in the abstract syntax tree when the verification stops, output the current initial code segment with violation prompt information.

[0079] If, after reaching the preset maximum number of iterations, non-compliant nodes still stubbornly exist in the abstract syntax tree, the system needs to adopt a "output with errors" strategy to prevent the generation task from ending in failure. In practice, the system will stop its error-correction dialogue with the large code generation model and instead call the error aggregation engine to convert the violation type and location information corresponding to the remaining non-compliant nodes into human-readable violation warnings. These warnings are guiding texts designed to assist developers in manual intervention, precisely pointing out logical blind spots that the large model cannot automatically fix. The system will append the violation warnings as comments or sidebar pop-ups to the current initial code snippet and output the entire snippet. The reason for outputting flawed code is that even if the large model cannot be completely fixed, the initial code snippet still contains most of the correct business logic. By outputting the current initial code snippet with violation warnings, the system allows developers to understand the root cause of the error and quickly complete the final "last step" through manual modification, thus ensuring development efficiency even when automation is hindered.

[0080] For example, if the system attempts to correct the code five times but the large model still cannot handle a complex cross-module call correctly, the system will output the current initial code snippet and mark the corresponding line with the violation message "Warning: The specified utility class could not be reused successfully here. Please check manually."

[0081] Please see Figure 3 This is a schematic diagram of the structure of a code generation system in an embodiment of this application.

[0082] It should be noted that, Figure 3 The structure of the code generation system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0083] like Figure 3As shown, a code generation system includes a central processing unit 301, which can perform various appropriate actions and processes based on a program stored in a read-only memory 302 or a program loaded from a storage section 308 into a random access memory 303, such as performing the methods described in the above embodiments. The random access memory 303 also stores various programs and data required for system operation. The central processing unit 301, the read-only memory 302, and the random access memory 303 are interconnected via a bus 304. An input / output interface 305 is also connected to the bus 304.

[0084] The following components are connected to the input / output interface 305: an input section 306 including audio input devices, push-button switches, etc.; an output section 307 including an LCD display, audio output devices, indicator lights, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the input / output interface 305 as needed. A removable medium 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 310 as needed so that computer programs read from it can be installed into the storage section 308 as needed.

[0085] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit 301, it performs the various functions defined in the present invention.

[0086] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0088] Specifically, a code generation system according to this embodiment includes a processor and a memory. The memory stores a computer program, and when the computer program is executed by the processor, it implements a code generation method provided in the above embodiment.

[0089] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in a code generation system described in the above embodiments; or it may exist independently and not assembled into the code generation system. The storage medium carries one or more computer programs, which, when executed by a processor of the code generation system, cause the code generation system to implement the code generation method provided in the above embodiments.

Claims

1. A code generation method, characterized in that, The method includes: Obtain the document description containing business requirements from the code generation task, and perform semantic parsing on the document description using a natural language processing model to obtain structured requirement data; The structured requirement data is converted into a requirement feature vector, and the requirement feature vector is matched with a preset local engineering knowledge base to retrieve the target context information associated with the business requirement. The local engineering knowledge base includes the existing tool class signature set and team coding standard rules in the target project. Based on the structured requirement data and the target context information, a context-aware prompt word is constructed. The context-aware prompt word contains constraint instructions. The constraint instructions instruct the pre-trained code generation model to call the utility class signature in the target context information and follow the team coding specification rules. The context-aware prompts are input into the code generation model to generate initial code snippets; The initial code snippet is parsed into an abstract syntax tree, and the abstract syntax tree is iteratively verified. When there are violation nodes, the code is used to generate a large model for correction until a preset stopping verification condition is met. The preset stopping verification condition includes that there are no violation nodes in the abstract syntax tree or that the preset maximum number of iterations for correction is met. The violation node is a node that violates the team coding standard rules, or the corresponding parent scope node marked when no call action of the utility class signature in the target context information is detected. If the violation node does not exist in the abstract syntax tree when the verification stops, the corresponding code snippet is identified as the target code; If the violation node still exists in the abstract syntax tree when the verification stops, the current initial code snippet with violation warning information is output.

2. The method according to claim 1, characterized in that, The step of performing similarity matching between the requirement feature vector and a preset local engineering knowledge base to retrieve target context information associated with the business requirement specifically includes: Calculate the semantic similarity score between the demand feature vector and each tool signature in the tool signature set, and combine the tool signatures with semantic similarity scores greater than a preset matching threshold into an initial candidate signature set; Determine whether there is a semantic conflict subset in the initial candidate signature set. The semantic conflict subset is at least two tool class signatures in the initial candidate signature set whose semantic similarity scores with the requirement feature vector are less than a preset difference threshold, and the at least two tool class signatures belong to different underlying dependency modules in the local engineering knowledge base. If the semantic conflict subset does not exist, then the initial candidate signature set is determined as the target candidate signature set; If the semantic conflict subset exists, the utility class signatures that do not meet the calling constraints are removed from the initial candidate signature set to obtain the target candidate signature set. The target candidate signature set is assembled with the team coding standard rules to obtain target context information related to the business requirements.

3. The method according to claim 2, characterized in that, Before determining the initial candidate signature set as the target candidate signature set, the method further includes: Obtain the target deployment environment identifier corresponding to the code generation task, and parse the metadata annotations of the signatures of each utility class in the local project knowledge base; Based on the metadata annotations, extract the lifecycle status and environment isolation tags of each utility class signature in the initial candidate signature set; Determine whether there is an environment overreach signature in the initial candidate signature set. The environment overreach signature is a signature whose semantic similarity score is greater than or equal to the preset matching threshold and whose environment isolation label conflicts with the target deployment environment identifier. If an unauthorized signature exists in the environment, it is removed from the initial candidate signature set.

4. The method according to claim 2, characterized in that, The step of removing utility class signatures that do not meet the calling constraints from the initial candidate signature set to obtain the target candidate signature set specifically includes: Get the target project directory path corresponding to the code generation task; Based on the project configuration information in the local project knowledge base, a module dependency topology diagram is generated; In the module dependency topology graph, the topological reachability and dependency weight of the target module where the target project directory path is located are analyzed to reach the source module where each of the utility class signatures in the semantic conflict subset is located. The utility signatures in the semantic conflict subset whose topological reachability is unreachable or whose dependency weight is less than a preset weight threshold are removed from the initial candidate signature set to obtain the target candidate signature set.

5. The method according to claim 1, characterized in that, The construction of context-aware prompts based on the structured requirement data and the target context information specifically includes: Extract the utility class signature and the team coding standard rules from the target context information, and cross-validate the syntax features of the utility class signature with the team coding standard rules to determine whether there is a standard conflict event. The standard conflict event is that the definition format of the utility class signature itself violates the team coding standard rules. If the aforementioned specification conflict event exists, then extract the conflict signature that triggered the specification conflict event and the corresponding conflict rule; A local exemption statement is generated for the conflict signature, the local exemption statement being used to instruct the code generation model to ignore the conflict rule when invoking the conflict signature; Convert the conflict-free signatures in the utility class signatures into regular constraint declarations along with the team coding standard rules; The constraint instruction is generated by combining the local exemption statement with the regular constraint statement. The structured requirement data is converted into task description information, and the task description information and the constraint instructions are injected into a preset prompt word template to obtain the context-aware prompt words.

6. The method according to claim 1, characterized in that, The iterative verification of the abstract syntax tree, and the correction of any violations by generating a large model from the code, until a preset stopping verification condition is met, specifically includes: Traverse each node of the abstract syntax tree and, based on the constraint instructions, verify whether there are any illegal nodes in the abstract syntax tree; If the violation node exists in the abstract syntax tree, the source code fragment and violation type corresponding to the violation node are extracted, and a correction prompt word is generated in combination with the constraint instruction. The correction prompt word is then fed back to the code generation model to generate candidate code fragments. The candidate code fragment is used as the initial code fragment to regenerate the abstract syntax tree and verify the illegal node until there is no illegal node in the abstract syntax tree or the preset maximum number of iterations for correction is met, at which point the verification stops.

7. The method according to claim 6, characterized in that, The step of verifying whether there are any violating nodes in the abstract syntax tree based on the constraint instructions specifically includes: Extract the node attribute information of each node to be verified in the abstract syntax tree; The node attribute information is matched with the regular constraint declaration according to rules. If the node attribute information violates the regular constraint declaration, the corresponding node to be verified is marked as a suspected violation node. Extract the code signature corresponding to the suspected violation node, and determine whether the code signature matches the conflicting signature indicated by the local exemption statement; If a violation is detected, the violation marker on the suspected violation node is removed, and it is determined to be a legitimate node. If no match is found, the suspected violation node is confirmed as the violation node. Extract all actual call signatures from the abstract syntax tree. If the utility class signature required by the constraint instruction is not included in the actual call signature, then it is determined that there is a violation node where the utility class signature is not reused.

8. A code generation system, characterized in that, The code generation system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the code generation system to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are run on the code generation system, the code generation system performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the code generation system, the code generation system performs the method as described in any one of claims 1-7.