Database logic error discovery-oriented large model enhanced query rewriting synthesis method
By combining rule tree search and Monte Carlo tree search enhanced by large language model, rewrite rules are dynamically generated and semantic consistency verification is carried out, and problems of insufficient flexibility and insufficient root cause analysis in the existing technology are solved, and efficient database logical error detection and repair are achieved.
Patent Information
- Application Number
- CN202510653093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing rules-based database logic error detection methods have insufficient flexibility, difficulty in rewriting search space management, and lack of effective root cause analysis, resulting in low detection efficiency and limited coverage.
Combining rule tree search and Monte Carlo tree search technology enhanced by large language model, we can dynamically generate rewrite rules, conduct semantic consistency verification and root cause analysis, and optimize the query rewrite process by building rule tree and historical error pattern repository.
It improves the efficiency and coverage of logical error detection, realizes more accurate root cause analysis and repair, reduces redundant testing, and improves the correctness and reliability of the system.
Smart Images

Figure CN120492487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database system testing, and in particular to a large-model enhanced query rewriting synthesis method for discovering database logic errors. Background Art
[0002] With the growth of diverse database systems, such as cloud and HTAP databases, the complexity of query optimization and system management has increased significantly. Database diversity complicates optimization and management, increasing the risk of errors. Ensuring the correctness and reliability of these systems is crucial, as errors can lead to serious consequences, including incorrect query results and system failures.
[0003] Errors in database systems can be broadly categorized as crash errors and logic errors. Crash errors render the system unresponsive, often leading to downtime and data loss, while logic errors are more subtle. They produce incorrect query results without causing a system crash, silently undermining data integrity and decision-making.
[0004] In recent years, various detection methods have been developed to identify logical errors in database systems, with differential testing emerging as a particularly effective approach. Differential testing involves executing the same query across multiple database systems or configurations and comparing the results to detect inconsistencies. This technique systematically generates query variants through rule-based transformations or heuristic searches and compares the outputs to identify differences. By analyzing these differences, differential testing can reveal logical errors that might otherwise remain undetected.
[0005] Transforming queries based on specific rules to identify inconsistencies has been widely studied. For example, the TLP framework decomposes a query into partitioned subqueries that compute results on different tuples, while NoRec compares the results of randomly generated queries with their unoptimized versions. Similarly, DQP generates multiple query plans for the same query to ensure consistent results.
[0006] However, existing rule-based systems have the following limitations:
[0007] These systems rely on static, manually defined rules that are inflexible and difficult to extend, severely limiting coverage and adaptability;
[0008] As query complexity increases, rule-based systems struggle to manage the expanded rewrite search space, leading to slower testing and higher computational costs;
[0009] Traditional methods often lack effective root cause analysis mechanisms, making it difficult to determine the exact reasons behind logic errors. Summary of the Invention
[0010] This invention addresses the shortcomings of existing technologies by providing a large-model-enhanced query rewrite synthesis method for detecting logical errors in database engines. This method systematically explores the query rewrite space by combining rule tree search with Monte Carlo tree search enhanced by a large language model, improving the efficiency and coverage of logical error detection.
[0011] The object of the present invention is achieved through the following technical solution: a large model enhanced query rewriting synthesis method for database logic error discovery, comprising:
[0012] Build a rule tree whose root node represents the original query, non-root nodes represent rewritten queries derived by applying the corresponding rewrite rules to their parent query, and edges between parent and child nodes represent the application of rewrite rules;
[0013] Combining Monte Carlo Tree Search with a large language model improves query rewriting by balancing exploration and exploitation. Monte Carlo Tree Search consists of a selection phase, an expansion phase, a simulation phase, and a backpropagation phase. The large language model is integrated into the expansion and simulation phases of MCTS.
[0014] During the expansion phase, the large language model generates new rewriting rules when predefined rules are insufficient or inapplicable; and analyzes the query structure and semantics to generate new rewrites that maintain the original query intent;
[0015] During the simulation phase, the historical logical errors stored in the historical error pattern repository are used to guide the large language model to prioritize error-prone query patterns.
[0016] Furthermore, it also includes root cause analysis, specifically:
[0017] For queries that generate logical errors, the historical error pattern repository is used to locate the abstract syntax tree node that causes the logical error. After locating the root cause of the error, the query that generates the logical error, its corresponding rules, and the root cause of the error are written back into the historical error pattern repository as a triple for subsequent LLM generation and priority rewriting. The abstract syntax tree is obtained by performing abstract syntax parsing on the input query (i.e., the original query), and each node on the abstract syntax tree is a query operator or parameter.
[0018] Furthermore, it also includes semantic consistency verification for each query.
[0019] Furthermore, semantic consistency verification is performed on each query, including:
[0020] Ensure that each rewritten query maintains logical correctness and verify that the transformation does not introduce unexpected semantic deviations;
[0021] Use semantic equivalence checking to verify that the rewritten query is semantically equivalent to the original query;
[0022] If semantic inconsistency is detected, the rewrite is rejected and the large language model is asked to generate a new rewrite.
[0023] Furthermore, each node in the rule tree corresponds to a query derived through a series of rule applications and is associated with metrics that evaluate its potential to increase the diversity of generated test cases. These metrics help prioritize nodes based on their contribution to generating unique queries.
[0024] Furthermore, we define node efficiency as an indicator for determining node priority. Node efficiency is based on diversity efficiency, which is quantified by the following components:
[0025] Structural difference: measures the degree of structural change introduced by applying the rules;
[0026] Operator changes: Evaluates the changes in operators used before and after applying the rule;
[0027] Data coverage diversity: Evaluate how well rewritten queries explore new and diverse data distributions;
[0028] Historical rule usage: applying penalties to frequently used rules to encourage exploration of less frequently used rules;
[0029] Prior diversity benefit: evaluates the diversity introduced by the query represented by a node relative to its ancestor nodes;
[0030] Subsequent diversity benefit: assessing the diversity potential of a node’s offspring;
[0031] Node comprehensive benefit: Determine the node comprehensive benefit based on previous diversity benefit and subsequent diversity benefit;
[0032] Based on the comprehensive benefits of the nodes, the UCB formula is used in the MCTS selection stage to determine the UCB value.
[0033] Furthermore, structural differences measure the degree of structural changes introduced by the application rules. Specifically, the Zhang-Shasha algorithm is used to calculate the minimum edit distance between two ASTs, that is, to calculate the minimum number of operations required to convert the original AST into the rewritten AST. The calculation formula is:
[0034] d1(v i , v0)TreeEditDistance(AST(Q),AST(Q′))
[0035] AST converts two different queries Q and Q′ into corresponding abstract syntax trees;
[0036] Operator change: Evaluates the change in operators used before and after applying the rule; specifically, the Jaccard similarity index is used to calculate the similarity of the operator set in the query:
[0037]
[0038] Data Coverage Diversity: Evaluates how well the rewritten query explores new and diverse data distributions; specifically, it highlights the contribution of previously unvisited data segments, weighted by their importance:
[0039]
[0040] NewSe is a set representing data segments accessed by Q but not by Q′. Variability(s) represents the diversity and non-uniformity of a data segment s; CoverageWeight(s) represents the relative importance of data segment s;
[0041] Historical rule usage: Penalties are applied to frequently used rules to encourage exploration of less frequently used rules; the calculation formula is:
[0042]
[0043] r i Represents a rewrite rule that has been used, which belongs to the overall rewrite set R; r i Represents a rewrite rule;
[0044] Previous diversity benefit: evaluates the diversity introduced by the query represented by a node relative to its ancestor nodes; the calculation formula is:
[0045] D(v i )=w1·d1(v i ,v0)+w2·d2(v i ,v0)+w3·d3(v i )+w4·d4(v i )
[0046] w1, w2, w3, and w4 all represent weights, and their sum is 1;
[0047] Subsequent diversity benefit: Evaluates the diversity potential of a node's descendants; the calculation formula is:
[0048]
[0049] Among them, Desc(v i ) represents node v i The child node of D(v i ) represents the previous diversity benefit value of node i, D(v j) represents the child node v of node i j 's previous diversity benefit value;
[0050] Node comprehensive benefit: The node comprehensive benefit is determined based on the previous diversity benefit and the subsequent diversity benefit; the calculation formula is:
[0051] B(v i )=D(v i )+D + (v i )
[0052] Based on the comprehensive benefits of the nodes, the UCB formula is used in the MCTS selection stage to determine the UCB value; the calculation formula is:
[0053]
[0054] Among them, F(v i )、F(v j ) are accessed by v i With v j The number of visits to the node, C is the exploration coefficient.
[0055] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned large-model enhanced query rewriting synthesis method for database logic error discovery.
[0056] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned large-model enhanced query rewriting synthesis method for discovering database logic errors.
[0057] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned large-model enhanced query rewriting synthesis method for discovering database logic errors.
[0058] The beneficial effects of the present invention are: 1. Systematically traverse all applicable rewrite rules, minimizing redundant testing by avoiding re-evaluation of previously identified errors, thereby improving the efficiency and coverage of logic error detection;
[0059] 2. Implementing Large Language Model-Enhanced Monte Carlo Tree Search, which combines Monte Carlo Tree Search with a large language model to improve query rewriting by balancing exploration and exploitation. Monte Carlo Tree Search dynamically explores query transformations, while the large language model generates new rewrites when predefined rules are insufficient. A repository of historical error patterns is leveraged to guide the large language model to prioritize error-prone transformations.
[0060] 3. Utilize the historical error pattern repository to record error-prone query templates and their corresponding rule lists, making it possible to pinpoint the exact abstract syntax tree nodes that cause logic errors, enabling more accurate root cause analysis and providing more effective repairs. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a query rewrite example that shows the step-by-step transformation of a query execution plan through rule-based rewrite (R4, R1, and R2). Each step analyzes potential database logic errors and highlights where problems occur during the rewrite process.
[0062] Figure 2 This is an LRS architecture diagram showing how LRS combines Rule Tree Search (RTS) and Rule Tree Search (LMS) based on Monte Carlo Tree Search enhanced by a large language model. RTS builds the search space, while LMS uses Monte Carlo Tree Search with a large language model for dynamic query rewriting and verification.
[0063] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0064] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0065] The present invention is a rule tree query rewriting synthesis (LRS) method based on large language model enhancement, see Figure 1 , used to detect database logic errors, including:
[0066] Rule Tree Search
[0067] Specifically, we systematically traverse the rules in a prefabricated rule base to explore potential transformations of the original query. Given a large number of available rewrite rules, we can sequentially construct an abstract rule tree based on the prefabricated rule base, establishing a structured framework for rule selection. By organizing the search into a hierarchical structure, rule tree search ensures extensive coverage of the rewrite space while maintaining a manageable exploration scope.
[0068] A rule tree is a hierarchical structure where:
[0069] The root node represents the original query;
[0070] Each non-root node represents a rewritten query derived by applying a specific rewrite rule to its parent query;
[0071] The edge between a parent node and a child node represents the application of a rewrite rule, transforming the parent query into a child query;
[0072] Leaf nodes represent queries that cannot be further rewritten by any rules.
[0073] Each node in the rule tree corresponds to a query derived through a series of rule applications and is associated with a metric that evaluates its potential to increase the diversity of generated test cases. These metrics help prioritize nodes based on their contribution to generating unique queries, which is crucial for discovering logic bugs.
[0074] To guide the search of the rule tree, we define node effectiveness as a key metric for determining node prioritization. Node effectiveness is based on diversity effectiveness, which assesses the extent to which a node introduces new and unique logical variations in the test case. This diversity is crucial to ensure that a wide range of query behaviors is explored to detect a variety of logical errors. Diversity effectiveness is quantified by the following components:
[0075] Structural Difference: This measures the degree of structural change introduced by applying rules, such as changes to the Abstract Syntax Tree (AST). The Zhang-Shasha algorithm is used to calculate the minimum edit distance between two ASTs, that is, the minimum number of operations (insertions, deletions, and substitutions) required to transform the original AST into the rewritten AST.
[0076] d1(v i , v0)=TreeEditDistance(AST(Q),AST(Q′))
[0077] AST converts two different queries Q and Q′ into corresponding abstract syntax trees; i Represents a non-root node, v o Represents the root node.
[0078] Operator Change: Evaluates the change in operators (e.g., joins, filters, aggregations) used before and after the rule is applied. The Jaccard similarity index is used to calculate the similarity of the operator sets in the query. The more operators are added or removed, the closer d2 is to 1.
[0079]
[0080] Data Coverage Diversity: Evaluates how well rewritten queries explore new and diverse data distributions. It highlights the contributions of previously unvisited data segments, weighted by their importance.
[0081] Data segmentation: divide the target data table into S segments according to the primary key range or physical partition number;
[0082] Variation Var(s): standard deviation σ, entropy H or Gini coefficient can be selected;
[0083] Weight w(s): The ratio of the sampling segment size to the entire table, denoted as |s| / |D|
[0084]
[0085] NewSe is a set representing data segments accessed by Q but not by Q′. Variability(s) represents the diversity and non-uniformity of a data segment s. CoverageWeight(s) represents the relative importance of data segment s, which is determined by the size and uniqueness of the data segment.
[0086] Historical rule usage: Penalties are applied to frequently used rules to encourage the exploration of less frequently used rules. Frequently used rules have low penalty coefficients, which helps explore less popular rules.
[0087]
[0088] r i Represents a rewrite rule that has been used, which belongs to the overall rewrite set R; r i Represents a rewrite rule.
[0089] Prior Diversity Benefit: Evaluates the diversity introduced by the query represented by a node relative to its ancestor nodes. The weights are adjustable, with w = (0.35, 0.25, 0.25, 0.15) as the initial values.
[0090] D(v i )=w1·d1(v i ,v0)+w2·d2(v i ,v0)+w3·d3(v i )+w4·d4(v i )
[0091] Subsequent diversity benefit: Evaluates the diversity potential of a node’s descendants. It encourages exploration of paths that may produce greater diversity in test cases. i is a leaf node, then D + (v i )=D(v i ). This definition encourages the searcher to extend to potentially high-yield subtrees. j It is v i According to the same algorithm, we can calculate D(v j ) value, and adding them together gives D + (v i ) value, used to reflect the subsequent diversity benefits. i The different subsequent diversity benefit values created by different child nodes can be obtained, which is a key step in calculating the comprehensive benefit.
[0092]
[0093] Node comprehensive benefits B(vi):
[0094] B(v i )=D(v i )+D + (v i )
[0095] The MCTS selection stage uses the UCB formula:
[0096]
[0097] Among them, C is the exploration coefficient (taken as 0.7 in the experiment). And F(v i ) and F(v j ) are accessed by v i With v j The number of times the node was visited.
[0098] Rule Tree Search Based on Monte Carlo Tree Search Enhanced by Large Language Model
[0099] For details, see Figure 2 Based on the rule tree search framework, Monte Carlo Tree Search (MCTS) is used as the core strategy for exploring diverse query rewriting. This step includes three key components: rule tree search based on Monte Carlo Tree Search, adaptive query rewriting driven by a large language model, and semantic consistency verification.
[0100] Rule tree search based on Monte Carlo tree search:
[0101] An adaptive search mechanism is used to explore the structured query space, prioritizing transformations that are most likely to reveal logical errors.
[0102] Iterate through four phases: select, expand, simulate, and post back.
[0103] The selection phase uses the upper confidence bound (UCB) formula to balance exploration and exploitation.
[0104] The expansion phase creates new nodes that represent the results of applying specific rewrite rules.
[0105] The simulation phase evaluates the potential value of new nodes by randomly applying sequences of rules and checking for any logical errors.
[0106] The feedback phase updates the node's statistics to guide future search decisions.
[0107] Adaptive query rewriting driven by large language models:
[0108] Integrate large language models into the expansion and simulation phases of MCTS to achieve dynamic, context-aware query rewriting. Specifically:
[0109] When the predefined rules are insufficient or inapplicable, the large language model generates new rewrite rules.
[0110] A large language model analyzes query structure and semantics to generate new rewrites that maintain the original query intent.
[0111] By storing historical logical errors in the historical error pattern repository, the large language model is guided to prioritize error-prone query patterns.
[0112] Semantic consistency verification:
[0113] Ensure that each rewritten query maintains logical correctness and verify that the transformation does not introduce unexpected semantic deviations.
[0114] Use semantic equivalence checking to verify that the rewritten query is semantically equivalent to the original query.
[0115] If semantic inconsistency is detected, the rewrite is rejected and the large language model is asked to generate a new rewrite.
[0116] Root Cause Analysis
[0117] Specifically, by leveraging a historical error pattern repository that records error-prone query templates and their corresponding rule lists, our approach can more easily pinpoint the exact AST node that causes the logic error. This enables more precise root cause analysis and more effective remediation.
[0118] Figure 3 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 3 The electronic device provided in this embodiment includes: a memory and a processor, wherein the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, a large model enhanced query rewriting synthesis method for database logic error discovery of the present invention is implemented.
[0119] It should be noted that, in addition to Figure 3 In addition to the memory and processor shown, the electronic device may also include other hardware according to its actual functions, which will not be described in detail.
[0120] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned large-model enhanced query rewriting synthesis method for discovering database logic errors.
[0121] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned large-model enhanced query rewriting synthesis method for discovering database logic errors.
[0122] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0123] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0124] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0126] The above embodiments are intended only to illustrate the design concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design concepts disclosed in the present invention are within the scope of protection of the present invention.
Claims
1. A large-model enhanced query rewriting synthesis method for database logic error detection, characterized in that: include: Build a rule tree whose root node represents the original query, non-root nodes represent rewritten queries derived by applying the corresponding rewrite rules to their parent query, and edges between parent and child nodes represent the application of rewrite rules; Combining Monte Carlo Tree Search with a large language model improves query rewriting by balancing exploration and exploitation. Monte Carlo Tree Search consists of a selection phase, an expansion phase, a simulation phase, and a backpropagation phase. The large language model is integrated into the expansion and simulation phases of MCTS. During the expansion phase, the large language model generates new rewriting rules when predefined rules are insufficient or inapplicable; and analyzes the query structure and semantics to generate new rewrites that maintain the original query intent; During the simulation phase, the historical logical errors stored in the historical error pattern repository are used to guide the large language model to prioritize error-prone query patterns.
2. The method according to claim 1, characterized in that It also includes conducting a root cause analysis, specifically: For queries that generate logical errors, the historical error pattern repository is used to locate the abstract syntax tree node that causes the logical error. After locating the root cause of the error, the query that generates the logical error, its corresponding rule, and the root cause of the error are written back into the historical error pattern repository as a triplet for subsequent LLM generation and priority rewriting. The abstract syntax tree is obtained by performing abstract syntax parsing on the input query, and each node on the abstract syntax tree is a query operator or parameter.
3. The method according to claim 1, characterized in that It also includes semantic consistency verification for each query.
4. The method according to claim 3, characterized in that Perform semantic consistency verification on each query, including: Ensure that each rewritten query maintains logical correctness and verify that the transformation does not introduce unexpected semantic deviations; Use semantic equivalence checking to verify that the rewritten query is semantically equivalent to the original query; If semantic inconsistency is detected, the rewrite is rejected and the large language model is asked to generate a new rewrite.
5. The method according to claim 1, characterized in that Each node in the rule tree corresponds to a query derived through a series of rule applications and is associated with metrics that evaluate its potential to increase the diversity of generated test cases. These metrics help prioritize nodes based on their contribution to generating unique queries.
6. The method according to claim 5, characterized in that Node efficiency is defined as an indicator for determining node priority. Node efficiency is based on diversity efficiency, which is quantified by the following components: Structural difference: measures the degree of structural change introduced by applying the rules; Operator changes: Evaluates the changes in operators used before and after applying the rule; Data coverage diversity: Evaluate how well rewritten queries explore new and diverse data distributions; Historical rule usage: applying penalties to frequently used rules to encourage exploration of less frequently used rules; Prior diversity benefit: evaluates the diversity introduced by the query represented by a node relative to its ancestor nodes; Subsequent diversity benefit: assessing the diversity potential of a node’s offspring; Node comprehensive benefit: Determine the node comprehensive benefit based on previous diversity benefit and subsequent diversity benefit; Based on the comprehensive benefits of the nodes, the UCB formula is used in the MCTS selection stage to determine the UCB value.
7. The method according to claim 6, characterized in that Structural Difference: measures the degree of structural change introduced by the application of rules. Specifically, the Zhang-Shasha algorithm is used to calculate the minimum edit distance between two ASTs, that is, to calculate the minimum number of operations required to transform the original AST into the rewritten AST. The calculation formula is: d1(v i ,v0)TreeEditDistance(AST(Q),AST(Q′)) AST converts two different queries Q and Q′ into corresponding abstract syntax trees; Operator change: Evaluates the change in operators used before and after applying the rule; specifically, the Jaccard similarity index is used to calculate the similarity of the operator set in the query: Data Coverage Diversity: Evaluates how well the rewritten query explores new and diverse data distributions; specifically, it highlights the contribution of previously unvisited data segments, weighted by their importance: NewSe is a set representing data segments accessed by Q but not by Q′; Variability(s) represents the diversity and non-uniformity of a data segment s; CoverageWeight(s) represents the relative importance of data segment s; Historical rule usage: Penalties are applied to frequently used rules to encourage exploration of less frequently used rules; the calculation formula is: r i Represents a rewrite rule that has been used, which belongs to the overall rewrite set R; r i Represents a rewrite rule; Prior diversity benefit: evaluates the diversity introduced by the query represented by a node relative to its ancestor nodes; The calculation formula is: D(v i )=w1·d1(v i ,v0)+w2·d2(v i ,v0)+w3·d3(v i )+w4·d4(v i ) w1, w2, w3, and w4 all represent weights, and their sum is 1; Subsequent diversity benefit: Evaluates the diversity potential of a node's descendants; the calculation formula is: Among them, Desc(v i ) represents node v i The child node of D(v i ) represents the previous diversity benefit value of node i, D(v j ) represents the child node v of node i j 's previous diversity benefit value; Node comprehensive benefit: The node comprehensive benefit is determined based on the previous diversity benefit and the subsequent diversity benefit; the calculation formula is: B(v i )=D(v i )+D+(v i ) Based on the comprehensive benefits of the nodes, the UCB formula is used in the MCTS selection stage to determine the UCB value; the calculation formula is: Among them, F(v i )、F(v j ) are accessed by v i With v j The number of visits to the node, C is the exploration coefficient.
8. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the large model enhanced query rewriting synthesis method for database logic error discovery as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a large-model enhanced query rewriting synthesis method for discovering database logic errors as described in any one of claims 1 to 7 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the large-model enhanced query rewriting synthesis method for discovering logical errors in databases as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge graph question and answer retrieval method based on large language model and MCTS algorithm
CN118296114A
Enhanced large language model processing method, system and platform based on Monte Carlo tree search and storage medium
CN120011510A