Differential symbol execution-oriented large language model code generation result evaluation method and system

CN122817079APending Publication Date: 2026-09-25HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610974400.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]2)路径爆炸问题:随着程序分支数量增加,可行路径数快速增长,导致分析成本显著增加;

Benefits of technology

(1)本发明通过识别参考程序与候选程序之间的语法结构差异,并据此构建差分决策表,在符号执行过程中基于差分决策表引导符号执行,优先探索包含差异代码的执行路径,有效避免传统符号执行在全部路径空间中盲目搜索的问题,使路径探索过程集中于更可能产生功能差异的分支,从而显著提高错误路径被覆盖的概率,并在有限时间内发现更多隐藏于深层路径的错误行为,提升符号执行效率;同时降低与语义等价路径相关的冗余探索开销,缓解路径爆炸问题。因此,本发明能够在不增加测试预算的情况下,提高程序错误检测的深度与完整性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817079A_ABST
    Figure CN122817079A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of software engineering and artificial intelligence, and discloses a large language model code generation result evaluation method and system for differential symbolic execution, which comprises the following steps: identifying the structural differences between a reference program and a candidate program by constructing a differential decision table; driving the symbolic execution to preferentially traverse the differential path based on a differential path exploration mechanism; organizing the execution path and reusing the path condition by using a path decision tree, thereby realizing batch analysis of multiple programs; and generating error reproduction test cases based on the equivalence determination of symbolic abstracts. The application combines differential driving and batch path reuse to significantly improve defect detection efficiency while ensuring analysis accuracy, effectively solves the problems of insufficient test coverage and symbolic execution path explosion in existing methods, and realizes efficient and automated evaluation of large-scale candidate programs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of software engineering and artificial intelligence, and more specifically, relates to a method and system for evaluating the code generation results of a large language model oriented towards differential symbolic execution. Background Technology

[0002] With the development of artificial intelligence technology, especially the emergence of large-scale pre-trained models (LLMs), the technology of automatically generating program code based on natural language descriptions has developed rapidly. These models, trained on large-scale code and text corpora, can generate corresponding program implementations based on user-provided function signatures, interface definitions, and functional descriptions, thus finding wide application in fields such as automatic programming, code completion, and software development assistance.

[0003] In existing technologies, the evaluation of large language model code generation capabilities mainly relies on standardized programming benchmarks, with HumanEval and its extended version HumanEval+ becoming the most commonly used evaluation schemes. This type of evaluation method typically employs the following process: for each programming task, the model generates multiple candidate programs, and these programs are validated using pre-built test cases. If a candidate program passes all test cases, it is considered correct for that task, and the overall performance of the model is quantitatively evaluated using the Pass@K metric.

[0004] Besides test case-based methods, existing technologies also include code similarity-based evaluation methods, such as BLEU and CodeBLEU. These methods assess generation quality by comparing the similarity of textual or syntactic structures between candidate and reference programs. However, due to the many-to-many mapping between program syntax and semantics, methods relying solely on similarity are insufficient to accurately determine whether a program's functionality is correct. Therefore, in practical applications, test cases remain the primary evaluation criterion.

[0005] On the other hand, to improve test coverage, existing technologies have also introduced automated test generation methods, such as generating supplementary test cases based on input variations or language models. However, these methods are essentially still static test set extensions and cannot fundamentally solve the problem of insufficient test coverage.

[0006] Furthermore, symbolic execution, as a classic program analysis technique, can replace input with symbolic variables and systematically enumerate possible execution paths of the program, thereby automatically detecting potential errors and generating test cases. Therefore, some research attempts to introduce symbolic execution techniques into code generation evaluation to improve the rigor and coverage of the evaluation.

[0007] While the above methods have promoted the development of code generation and evaluation for large language models to some extent, they still have the following obvious technical shortcomings in practical applications: (1) Existing evaluation methods mainly rely on predefined test case sets, while program execution paths usually grow exponentially. A fixed number of test cases are difficult to cover all execution paths, especially in the presence of complex branches, recursion, or boundary conditions, where some error paths are difficult to trigger. This results in some programs with logical defects still being able to pass the test, thus causing the evaluation results to be too high.

[0008] (2) Although symbolic execution can systematically explore program paths, it faces the following problems in large-scale program evaluation scenarios: 1) The evaluation results are inconsistent with the user's intent: During symbolic execution, paths that do not conform to the user's intent will be explored.

[0009] 2) Path explosion problem: As the number of program branches increases, the number of feasible paths grows rapidly, leading to a significant increase in analysis costs; 3) Low execution efficiency: In typical evaluation tasks, tens of thousands of programs need to be analyzed, and performing symbol analysis one by one will result in unacceptable time overhead. Summary of the Invention

[0010] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a method and system for evaluating the code generation results of large language models oriented to differential symbolic execution, which aims to improve symbolic execution efficiency and alleviate the path explosion problem.

[0011] To achieve the above objectives, this invention provides a method for evaluating the code generation results of a large language model oriented towards differential symbolic execution, comprising: Obtain the reference program corresponding to the programming task, the candidate program set consisting of at least one candidate program generated by the large language model based on the programming task, and the input constraints of the programming task; After inserting the input constraints into the reference program and candidate program respectively, the difference statements between the reference program and the candidate program are identified; it is determined whether the code area covered by the true branch and false branch of each condition statement in the candidate program contains the difference statement, and then a difference label is set for each condition statement in the candidate program. The difference labels of each condition statement in the candidate program constitute a difference decision table; wherein, the difference label is used to indicate whether the true branch or false branch of the condition statement contains the difference statement. A differential path guidance strategy is used to guide the symbolic execution of the reference program and the candidate program, and to obtain a symbolic summary of the reference program and the candidate program under the current feasible execution path. The differential path guidance strategy includes: during symbolic execution, whenever a conditional statement is encountered, the differential tag of the conditional statement is queried in the differential decision table. When the differential tag indicates that the true branch or the false branch contains the differential statement, symbolic execution is preferentially executed along the true branch or the false branch. Determine whether the symbol digests of the candidate program and the reference program are the same under the current feasible execution path. If they are different, the candidate program is identified as having erroneous code, and test cases are generated under the current feasible execution path to reproduce the erroneous behavior.

[0012] Furthermore, before inserting the input constraints into the reference program and the candidate program respectively, the method further includes converting the assertions in the input constraints into equivalent conditional branch statements; Correspondingly, the transformed input constraints are inserted into the reference program and the candidate program, respectively.

[0013] Furthermore, before inserting the input constraints into the reference program and the candidate program respectively, the method further includes removing the invariant part of the complex structure operation in the input constraints from the complex structure operation to simplify the input constraints; wherein, the complex structure operation includes one or more of string splitting operation, array traversal operation and loop structure operation; Correspondingly, the simplified input constraints are inserted into the reference program and the candidate program, respectively.

[0014] Further, determine whether the symbol digests of the candidate program and the reference program are the same under the currently feasible execution path, including: Construct the logical decision formula φ=¬(M==M′); where M and M′ represent the symbol digests of the reference program and the candidate program in the current feasible execution path, respectively, "¬" represents negation, and "==" represents the equality comparison operator; The logic decision formula is solved using the SMT solver; if the solution result is satisfactory, it means that M and M′ are not equal; otherwise, it means that M and M′ are equal.

[0015] Furthermore, a differential path guidance strategy is employed to guide the symbolic execution of the reference program and candidate programs, resulting in symbolic summaries of the reference program and candidate programs under the current feasible execution path, including: The symbolic execution of the reference program is guided by a differential path guidance strategy, and a symbolic digest of the reference program under the current feasible execution path is obtained. Under the current feasible execution path, the symbolic execution of all candidate programs in the candidate program set is guided by the differential path guidance strategy to obtain the symbolic digest of all candidate programs under the current feasible execution path. Then, the symbolic digest of each candidate program is compared with the symbolic digest of the reference program under the current feasible execution path to determine whether the symbolic digests of each candidate program and the reference program are the same under the current feasible execution path. Correspondingly, if the symbol digests of a candidate program and the reference program are different in the current feasible execution path, the candidate program is deemed to have erroneous code, symbolic execution of the candidate program is stopped, and test cases are generated in the current feasible execution path.

[0016] Furthermore, if a candidate program and a reference program have the same symbol digest under the current feasible execution path, the candidate program and the reference program are considered to be the same under the current feasible execution path. If the candidate program and the reference program have the same symbol digest under all feasible execution paths, the candidate program is considered to have no error code.

[0017] Furthermore, the differential path guidance strategy also includes: when the differential marker indicates that both the true branch and the false branch contain the differential statement, or when neither the true branch nor the false branch contains the differential statement, symbolic execution is randomly performed along the true branch or the false branch.

[0018] The present invention also provides a system for evaluating the code generation results of a large language model oriented to differential symbolic execution, including a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read executable instructions stored in the computer-readable storage medium and execute the large language model code generation result evaluation method described above for differential symbolic execution.

[0019] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for evaluating the code generation results of a large language model oriented to differential symbolic execution as described in any of the preceding claims.

[0020] The present invention also provides a computer program product, including a computer program that, when the computer program is run on a computer, causes the computer to execute the method for evaluating the code generation results of a large language model oriented to differential symbolic execution as described above.

[0021] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects: (1) This invention identifies the syntactic differences between the reference program and the candidate program, and constructs a differential decision table accordingly. During symbolic execution, the differential decision table guides the symbolic execution, prioritizing the exploration of execution paths containing the differing code. This effectively avoids the problem of blindly searching the entire path space in traditional symbolic execution, focusing the path exploration process on branches more likely to produce functional differences. This significantly increases the probability of erroneous paths being covered and discovers more erroneous behaviors hidden in deep paths within a limited time, thus improving the efficiency of symbolic execution. At the same time, it reduces the redundant exploration overhead related to semantically equivalent paths, alleviating the path explosion problem. Therefore, this invention can improve the depth and completeness of program error detection without increasing the testing budget.

[0022] (2) Furthermore, in response to the problem that test inputs in existing evaluation methods are usually obtained by fixed test cases or random generation, which cannot effectively constrain the semantic legality of the inputs and cause the evaluation results to deviate from the user's actual needs, this invention introduces constraint transformation to explicitly express and transform the input constraints implicit in the task. Specifically, by transforming the original assertion-type input constraints in the task, they are transformed from assertion logic that interrupts execution into conditional expressions in the symbolic execution process, so that they participate in the exploration as path conditions in the execution process without interrupting the execution flow. This makes the input space strictly limited to the semantic range defined by the user, and thus makes the path of symbolic execution exploration strictly limited to the input space defined by the user. This avoids the generation of invalid test cases from the source, improves the semantic consistency of the evaluation, avoids the misjudgment problem caused by invalid input in traditional methods, and further makes the evaluation process consistent with the user's intention.

[0023] (3) Furthermore, for input constraints containing complex structures such as arrays or string structures, the present invention simplifies such input constraints by separating the invariant part of the complex structure operation from the complex structure operation, so that symbolic execution is concentrated on the variable part of the complex structure operation. Thus, during symbolic execution, there is no need to explore paths and execute solvers for the invariant part, thereby further alleviating path explosion or solution difficulties.

[0024] (4) Furthermore, in the path exploration process, this invention employs symbolic execution technology to construct executable paths and corresponding symbolic semantic relationships (symbolic summaries), enabling program behavior to be expressed in the form of logical formulas, and detecting functional differences between candidate programs and reference programs based on equivalence determination. Compared to traditional test case-based methods, this invention can cover the potential input space rather than a finite test set; it can discover errors that are only triggered under extreme or boundary conditions; and it does not rely on manually designed test cases. Therefore, this invention significantly improves the theoretical completeness of program evaluation, making code correctness judgment more rigorous.

[0025] (5) Furthermore, this invention constructs an equivalence judgment formula (logical judgment formula) during symbolic execution and automatically generates test cases to reproduce erroneous behavior using a constraint solver. This design enables the system not only to detect errors but also to output corresponding trigger inputs (test cases provide reproducible evidence of errors, facilitating subsequent debugging and repair; simultaneously, it improves the interpretability of the evaluation results). Therefore, this invention achieves an improvement in capability from "error detection" to "error reproduction".

[0026] (6) To address the high structural similarity among multiple candidate programs under the same programming task, this invention proposes a multi-program batch analysis mechanism. By constructing a path decision tree dominated by the path structure of the reference program and sharing executable paths and constraint solution results (symbol digests of the reference program), collaborative analysis of multiple candidate programs is achieved. This mechanism significantly reduces the number of path explorations and constraint solutions, greatly improving the overall system execution efficiency. Specifically, the same executable path only needs to undergo constraint solution once in the reference program to obtain the symbol digest for that executable path; this symbol digest is reused among multiple candidate programs (under that executable path, all candidate programs are judged to be consistent with the symbol digest), and multiple erroneous programs can be identified simultaneously for the same executable path. This design effectively reduces the number of calls to the constraint solver (SMT solver) and reduces redundant path analysis, thereby significantly reducing the overall computational complexity; improving the throughput in large-scale program evaluation scenarios; and supporting efficient evaluation of tens of thousands of candidate programs. Therefore, this invention can significantly reduce evaluation time and computational resource consumption while ensuring analysis accuracy.

[0027] In summary, this paper addresses the problems in existing technologies for evaluating large language model code generation, such as insufficient test case coverage, limited program path exploration capabilities, low symbolic execution efficiency, and inconsistencies between evaluation results and user intent. It provides a method and system for evaluating large language model code generation results based on differential symbolic execution. By constructing a path exploration mechanism oriented towards program differences, a batch analysis mechanism oriented towards multiple programs, and an input constraint modeling mechanism oriented towards user semantics, it achieves efficient and accurate evaluation of large-scale candidate programs, significantly improves error detection capabilities, and reduces evaluation overhead. Attached Figure Description

[0028] Figure 1 This is an overall architecture diagram of the large language model code generation result evaluation system based on differential symbolic execution provided in an embodiment of the present invention.

[0029] Figure 2 This is a schematic diagram of input constraint transformation provided for an example of the present invention.

[0030] Figure 3A schematic diagram illustrating the multi-program batch analysis mechanism provided in this invention for processing multiple programs. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0032] In this invention, the terms "first," "second," etc., used in the invention and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0033] To address the problems existing in the technology: existing large language model code generation and evaluation methods mainly rely on static test cases, which are difficult to cover the complex execution paths of the program, resulting in functional errors not being fully detected; at the same time, traditional symbolic execution methods suffer from severe path explosion, low execution efficiency, and inconsistencies between input constraints and user semantics when applied to the evaluation of large-scale candidate programs, making it difficult to meet the accuracy and efficiency requirements of large-scale code evaluation scenarios.

[0034] This invention proposes a method and system for evaluating code generation in large language models based on differential symbolic execution. By introducing a differential-driven path exploration mechanism, a multi-program batch analysis mechanism, and an input constraint transformation mechanism, this invention enables the symbolic execution process to focus on paths with differences in program behavior and to share analysis results among multiple programs. This significantly improves error detection capabilities and evaluation efficiency while ensuring semantic consistency. Specifically, the differential-driven path exploration mechanism identifies differences between programs and guides symbolic execution to focus on critical paths, thereby improving error detection efficiency. The multi-program batch analysis mechanism significantly reduces the cost of large-scale program evaluation through path reuse and result sharing. The input constraint transformation mechanism ensures that test input aligns with user intent. Figure 1 This will improve the reliability of the evaluation results.

[0035] Figure 1 This is an overall architecture diagram of the large language model code generation result evaluation system based on differential symbolic execution provided in an embodiment of the present invention. Figure 1 This also corresponds to the entire system workflow. For each programming task, the system first obtains a reference program (ground-truth) and several candidate programs generated by a large language model. Figure 1The standard program serves as the reference program, while the candidate programs (#1~N) are the generated programs. Subsequently, the input constraints in the reference program are transformed, and these transformed input constraints are uniformly inserted into both the reference program and each candidate program, ensuring that all programs satisfy consistent input condition constraints. Afterward, the system simultaneously performs differential analysis and symbolic execution analysis, continuously updating the path exploration status until all detectable errors are found or the time limit is reached.

[0036] The entire process can be represented by the following structure: Input: A reference program with transformed input constraints and multiple candidate programs.

[0037] Intermediate processes: Difference decision table generation; Symbolic execution path exploration; Decision tree maintenance; Batch scheduling.

[0038] Output: A set of error programs; test cases triggered by differences.

[0039] Example 1 This invention provides a method for evaluating code generation from a large language model based on differential symbolic execution. This method is applied to an evaluation system for evaluating program code generated by a large language model, and the specific method is as follows.

[0040] (I) System Input and Overall Approach

[0041] The evaluation method is based on the following inputs: 1. A collection of programming tasks (e.g., a HumanEval class benchmark, where each task includes a function signature and a functional description). 2. Reference program (ground-truth) for each programming task; 3. Input each programming task into multiple candidate programs generated by the large language model; where each candidate program is the code to be evaluated.

[0042] The goal of this method is to perform path-level semantic comparisons between candidate programs and the correct program (reference program) using symbolic execution technology, automatically detect erroneous programs in each candidate program, and generate test cases that can reproduce the errors, thereby achieving a rigorous evaluation of the code generation capability of a large language model. Unlike traditional evaluation methods based on static test cases, this method improves the ability to detect potential errors by dynamically exploring the program path space.

[0043] (ii) Input constraint transformation.

[0044] To ensure that the symbol execution result matches the user's intention Figure 1First, the input constraints in the reference program are transformed. Specifically, programming benchmarks typically use assertions to restrict input conditions. If such assertions fail during execution, program execution is terminated, disrupting the continuity of the symbolic execution path. Therefore, this embodiment transforms these assertions in the reference program into equivalent conditional branch statements, enabling them to express input validity in a control flow manner. For example, for constraints such as input type or set length, the execution-interrupting assertion mechanism is no longer used; instead, conditional judgment statements are used for detection, and a default result is returned if the condition is not met. In this way, the symbolic execution process can continue to expand the program path while ensuring the constraint semantics remain unchanged, thereby achieving a systematic exploration of the path.

[0045] Meanwhile, for input constraints involving complex structural operations, such as string splitting, array traversal, and loop structures, which can lead to path explosion or solution difficulties during symbolic execution, this embodiment further simplifies such constraints. Specifically, the invariant parts of the complex structural operations are separated from the complex structural operations themselves, simplifying the input constraints and allowing symbolic execution to focus on the variable parts. Thus, during symbolic execution, path exploration and solver execution are unnecessary for the invariant parts, thereby alleviating path explosion or solution difficulties. For example, as... Figure 2 As shown, for an input constraint containing complex string splitting operations, the input constraint is simplified by finding the invariant part of the string and separating it from the string operations, so that subsequent symbolic execution focuses on the variable part of the string, thereby significantly reducing the complexity of symbolic solution and accelerating the path exploration process.

[0046] This invention avoids the SMT solver from generating invalid test cases that violate task semantics by transforming constraints; ensures that test cases conform to user-defined input constraints; and improves the accuracy and reliability of evaluation results.

[0047] The transformed input constraints are uniformly inserted into the reference program and each candidate program. That is, the transformed input constraints are copied into the reference program and each candidate program to obtain the reference program and each candidate program with inserted constraints.

[0048] (III) Exploration of differential paths.

[0049] When faced with a large number of candidate programs, traditional symbolic execution requires blindly exploring all execution paths, leading to a severe path explosion problem. To address this, this invention proposes a path guidance mechanism based on differential analysis. The reference program and candidate programs are treated as two versions with identical functionality but different implementations. By comparing their abstract syntax trees, the differences in their conditional statements are identified. Based on this difference information, a differential decision table is constructed to indicate which path branches might produce behavioral differences during symbolic execution, thereby avoiding a large amount of redundant computation related to semantically equivalent paths.

[0050] Specifically, after completing the input constraint transformation, this embodiment performs a difference analysis on the reference program and each candidate program after the constraint insertion to guide symbolic execution in prioritizing the exploration of potential error paths. This difference analysis process is based on a comparison at the level of the program's abstract syntax tree. For the current candidate program, the difference statements between the reference program and the current candidate program are identified, and all lines of code with differences are recorded. Subsequently, for each conditional statement in the current candidate program, the code regions covered by its true and false branches are analyzed, and it is determined whether the code regions covered by its true and false branches contain the difference statement. If the difference statement is only contained in the code region covered by the true branch of the conditional statement, then the conditional statement is assigned a first tag value; if the difference statement is only contained in the code region covered by the false branch of the conditional statement, then the conditional statement is assigned a second tag value; if the code regions covered by both the true and false branches of the conditional statement contain the difference statement, or if the code regions covered by both the true and false branches of the conditional statement do not contain the difference statement, then the conditional statement is assigned a third tag value. Thus, a difference table is obtained, which consists of the tag values ​​of each conditional statement in the current candidate program. This table assigns a tag value (differential tag, i.e., first tag value, second tag value, and third tag value) to each conditional statement to represent the branch in the set of difference statements.

[0051] Specifically, the table records whether each conditional statement contains a difference code in different branches. In this embodiment of the invention, the first flag value is 1, the second flag value is 2, and the third flag value is 0; that is, if the difference statement exists only in the true branch of a conditional statement, it is marked as 1; if the difference statement exists only in the false branch of a conditional statement, it is marked as -1; if the difference statement exists in both the true and false branches of a conditional statement or does not exist in either branch, it is marked as 0. In other embodiments, the first flag value, the second flag value, and the third flag value can also be other values.

[0052] Symbolic execution is employed to evaluate the current candidate program, guided by a differential decision table. Based on this differential path guidance, the symbolic execution module expands the program's paths. During symbolic execution, the input variables of the current candidate program are abstracted as symbolic variables, and path conditions are constructed during execution. These path conditions describe the set of constraints that make the current path valid. Each conditional statement encountered during execution is branched based on the path conditions, forming a path decision tree. Each path corresponds to a path condition and an output expression, which together constitute a symbolic summary, describing the relationship between the program's output and input variables. For the entire program, its symbolic semantics are represented as the disjunctive form of all path summaries. In this way, symbolic execution can theoretically cover all feasible execution paths of the current candidate program, and the symbolic semantics of all paths combine to form the overall semantic model of the program.

[0053] A differential path guidance strategy is used to guide symbolic execution, resulting in a symbol digest for a feasible execution path of the current candidate program. The differential path guidance strategy includes: during symbolic execution, whenever a conditional statement is encountered, the flag value of the conditional statement is queried in the differential decision table. If it is the first flag value, symbolic execution first follows the true branch of the conditional statement to obtain a symbol digest for an executable path. If it is the second flag value, symbolic execution first follows the false branch of the conditional statement. If it is the third flag value, symbolic execution randomly selects one of the true or false branches to execute.

[0054] By using a differential decision table to guide symbolic execution, symbolic execution can prioritize exploring execution paths containing differential codes (paths that may contain defects), thereby avoiding the exploration of irrelevant paths.

[0055] (iv) Difference detection based on symbolic execution.

[0056] The system determines whether there is erroneous code in the current executable path based on the symbol digest. If so, the candidate program is considered to have an error, symbolic execution stops, and test cases are generated for the current executable path. If not, the current executable path is considered to be free of erroneous code, symbolic execution continues, and the symbol digest for the next executable path is generated. The evaluation continues until all executable paths for the current candidate program are free of erroneous code, at which point the candidate program is considered correct, and the evaluation is complete. This process is repeated for all candidate programs to evaluate the code generation results of the large language model.

[0057] To determine whether a candidate program and a reference program are functionally equivalent, this embodiment performs an equivalence check on each exploration path. Specifically, under the same path conditions, the symbol digests of the reference program and the current candidate program are calculated respectively, and a decision formula φ=¬(M==M′) is constructed, where M represents the symbol digest of the reference program under the current executable path, M′ represents the symbol digest of the candidate program under the current executable path, and "¬" indicates negation. This formula is solved by the SMT solver. When the solution result is Satisfiable (SAT), that is, M and M′ are not equal, it indicates that there is at least one set of inputs that can cause the two programs to produce different outputs, thus proving that the current candidate program has erroneous code on that executable path; conversely, when the result is Unsatisfiable (UNSAT), that is, M and M′ are equal, it indicates that the two programs are equivalent under the path conditions, and the current candidate program does not have erroneous code on that executable path. In the case of detecting inequivalence (the current candidate program has erroneous code), the SMT solver also provides a test case corresponding to that executable path to reproduce the erroneous behavior.

[0058] like Figure 1 As shown, the symbolic execution process unfolds within a path decision tree, where each node corresponds to a conditional judgment, and the left and right branches represent different path selections. Under the differential path guidance strategy, symbolic execution prioritizes exploring branches containing differing code. For example, when a path results in inconsistent return values, the path condition is recorded, and corresponding test cases are generated.

[0059] (v) Batch analysis of multiple programs.

[0060] To further improve the efficiency of analyzing a large number of candidate programs, this embodiment introduces a multi-program batch analysis mechanism. In practical observation, programs generated by large language models for the same programming task often exhibit high structural similarity, resulting in common path conditions and even common error patterns among multiple candidate programs. Based on this characteristic, this embodiment does not construct an independent path exploration process for each candidate program. Instead, it constructs a path decision tree primarily based on the path structure of the reference program, and analyzes multiple candidate programs on this tree. First, symbolic execution is performed on the reference program to determine the current feasible execution path and its corresponding symbol digest. Then, under the current feasible execution path, symbolic execution is performed on all candidate programs to obtain their symbol digests, which are then compared for equivalence with the symbol digests of the reference program under the same feasible execution path. When a candidate program is determined to be inequivalent on this path, test cases are immediately generated, and the candidate program is removed from the set of candidate programs to be analyzed.

[0061] This approach not only allows for the reuse of SMT solution results but also enables the simultaneous detection of the same errors in multiple candidate programs during a single path exploration. That is, when an error is detected on a particular path, the corresponding test cases can be used to detect errors in all candidate programs simultaneously, thus achieving batch error detection across multiple programs and significantly reducing overall computational overhead. For example, Figure 3 Comparison of the two candidate programs Figure 1 Both the candidate program and the reference program share a common error. When the input is (x=-1, y=0, z=True), the candidate program outputs True, while the reference program outputs False. The batch exploration mechanism of this invention allows it to detect at least two erroneous candidate programs simultaneously.

[0062] (vi) Overall system operation process.

[0063] Based on the above mechanisms, the system execution flow in this embodiment is as follows: 1. Obtain the input for the current programming task, including the reference program, the set of candidate programs to be analyzed, and input constraints; 2. Perform input constraint transformation to generate symbol-friendly input constraints; 3. Perform differential analysis on the program to generate path guidance information (differential decision table); 4. Construct a path decision tree and perform symbolic execution under path guidance; reuse existing path information and share solution results during path exploration; 5. Perform equivalence checks on each path; If discrepancies exist, generate test cases and eliminate candidate programs with errors. Repeat the path exploration until: all paths have been traversed; or the preset time limit has been reached.

[0064] Final output: a collection of error programs and their corresponding test cases.

[0065] This invention can significantly improve evaluation efficiency and reduce evaluation costs while ensuring evaluation accuracy. Specifically, it can achieve the following technical effects: (1) It can dynamically generate test inputs and provide more comprehensive coverage of program execution paths to improve error detection capabilities; (2) It can refer to the differences between the program and candidate programs and prioritize the exploration of potentially defective paths to improve evaluation efficiency; (3) It can reuse existing analysis results when analyzing multiple program sets to avoid redundant calculations; (4) It can effectively model input constraints to ensure that the evaluation process is consistent with user intent; (5) It can run efficiently on large-scale program sets (such as tens of thousands) to meet engineering deployment requirements.

[0066] Example 2 This invention provides a large language model code generation and evaluation system based on differential symbolic execution, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the large language model code generation and evaluation method based on differential symbolic execution in Embodiment 1 above, so as to perform functional correctness analysis on candidate programs generated by the large language model, and can efficiently identify erroneous programs and generate corresponding test cases on a large-scale program set.

[0067] This system is built around a differential-driven symbolic execution process. Its overall structure includes multiple functional components such as input management, input constraint transformation, differential analysis, symbolic execution, path scheduling, constraint solving, and error detection, all coordinated through a unified data flow and control process. By combining differential analysis, symbolic execution, and multi-program batch analysis, the system achieves efficient evaluation of large-scale candidate program sets. Differential analysis is used to narrow down the path exploration scope, symbolic execution is used to build program behavior models, and multi-program batch analysis is used to reuse analysis results, thereby significantly reducing computational overhead while ensuring evaluation accuracy. The system can be deployed on local servers or cloud evaluation platforms to support automated quality evaluation of large-scale language model code generation results.

[0068] The relevant technical solutions are the same as above, and will not be repeated here.

[0069] Example 3 This invention provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the large language model code generation and evaluation method based on differential symbol execution in Embodiment 1 above.

[0070] Specifically, the memory may include high-speed random access memory, as well as non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0071] The relevant technical solutions are the same as above, and will not be repeated here.

[0072] Example 4 This invention provides a computer program product, including a computer program that, when run on a computer, causes the computer to perform the steps of the large language model code generation and evaluation method based on differential symbol execution in Embodiment 1 above.

[0073] The relevant technical solutions are the same as above, and will not be repeated here.

[0074] Those skilled in the art will readily understand that the above embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for evaluating code generation results of a large language model oriented towards differential symbolic execution, characterized in that, include: Obtain the reference program corresponding to the programming task, the candidate program set consisting of at least one candidate program generated by the large language model based on the programming task, and the input constraints of the programming task; After inserting the input constraints into the reference program and candidate program respectively, the difference statements between the reference program and the candidate program are identified; it is determined whether the code area covered by the true branch and false branch of each condition statement in the candidate program contains the difference statement, and then a difference label is set for each condition statement in the candidate program. The difference labels of each condition statement in the candidate program constitute a difference decision table; wherein, the difference label is used to indicate whether the true branch or false branch of the condition statement contains the difference statement. A differential path guidance strategy is used to guide the symbolic execution of the reference program and the candidate program, and to obtain a symbolic summary of the reference program and the candidate program under the current feasible execution path. The differential path guidance strategy includes: during symbolic execution, whenever a conditional statement is encountered, the differential tag of the conditional statement is queried in the differential decision table. When the differential tag indicates that the true branch or the false branch contains the differential statement, symbolic execution is preferentially executed along the true branch or the false branch. Determine whether the symbol digests of the candidate program and the reference program are the same under the current feasible execution path. If they are different, the candidate program is identified as having erroneous code, and test cases are generated under the current feasible execution path to reproduce the erroneous behavior.

2. The method for evaluating the code generation results of a large language model according to claim 1, characterized in that, Before inserting the input constraints into the reference program and candidate program respectively, the method further includes converting the assertions in the input constraints into equivalent conditional branch statements. Correspondingly, the transformed input constraints are inserted into the reference program and the candidate program, respectively.

3. The method for evaluating the code generation results of a large language model according to claim 1, characterized in that, Before inserting the input constraints into the reference program and candidate program respectively, the method further includes removing the invariant part of the complex structure operation in the input constraints from the complex structure operation to simplify the input constraints; wherein, the complex structure operation includes one or more of string splitting operation, array traversal operation and loop structure operation; Correspondingly, the simplified input constraints are inserted into the reference program and the candidate program, respectively.

4. The method for evaluating the code generation results of a large language model according to any one of claims 1-3, characterized in that, Determine whether the symbol digests of the candidate program and the reference program are the same in the currently feasible execution path, including: Construct the logical decision formula φ=¬(M==M′); where M and M′ represent the symbol digests of the reference program and the candidate program in the current feasible execution path, respectively, "¬" represents negation, and "==" represents the equality comparison operator; The logic decision formula is solved using the SMT solver; if the solution result is satisfactory, it means that M and M′ are not equal; otherwise, it means that M and M′ are equal.

5. The method for evaluating the code generation results of a large language model according to claim 4, characterized in that, A differential path guidance strategy is used to guide the symbolic execution of the reference program and candidate programs, resulting in symbolic summaries of the reference program and candidate programs under the current feasible execution path, including: The symbolic execution of the reference program is guided by a differential path guidance strategy, and a symbolic digest of the reference program under the current feasible execution path is obtained. Under the current feasible execution path, the symbolic execution of all candidate programs in the candidate program set is guided by the differential path guidance strategy to obtain the symbolic digest of all candidate programs under the current feasible execution path. Then, the symbolic digest of each candidate program is compared with the symbolic digest of the reference program under the current feasible execution path to determine whether the symbolic digests of each candidate program and the reference program are the same under the current feasible execution path. Correspondingly, if the symbol digests of a candidate program and the reference program are different in the current feasible execution path, the candidate program is deemed to have erroneous code, symbolic execution of the candidate program is stopped, and test cases are generated in the current feasible execution path.

6. The method for evaluating the code generation results of a large language model according to claim 5, characterized in that, If a candidate program and a reference program have the same symbol digest under the current feasible execution path, the candidate program and the reference program are considered to be the same under the current feasible execution path. If the candidate program and the reference program have the same symbol digest under all feasible execution paths, the candidate program is considered to have no error code.

7. The method for evaluating the code generation results of a large language model according to any one of claims 1-3, characterized in that, The differential path guidance strategy further includes: when the differential marker indicates that both the true branch and the false branch contain the differential statement, or when neither the true branch nor the false branch contains the differential statement, symbolic execution is randomly performed along the true branch or the false branch.

8. A system for evaluating the code generation results of a large language model oriented towards differential symbolic execution, characterized in that, Includes computer-readable storage media and processors; The computer-readable storage medium is used to store executable instructions; The processor is used to read executable instructions stored in the computer-readable storage medium and execute the method for evaluating the code generation results of a large language model oriented to differential symbolic execution as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method for evaluating the code generation results of a large language model oriented to differential symbolic execution as described in any one of claims 1-7.

10. A computer program product, characterized in that, The method includes a computer program that, when run on a computer, causes the computer to perform the method for evaluating the code generation results of a large language model oriented to differential symbolic execution as described in any one of claims 1-7.