Reinforced learning SAT solving method and system for chip design verification
By dynamically selecting the rephase heuristic in the SAT solver and optimizing the SAT solution process using a multi-arm game algorithm, the problem of inefficiency in the existing technology is solved, and faster solution speed and higher resource utilization efficiency are achieved.
Patent Information
- Application Number
- CN202510263837.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-01
AI Technical Summary
Existing SAT solvers are inefficient when dealing with large and complex circuits, easily fall into local optimization, and have high resource consumption, making it difficult to complete verification within a reasonable time.
The multi-arm game algorithm is used to dynamically select the re-phase heuristic, combine the best and wandering heuristics to generate multiple re-phase strategies, calculate the reward function value through statistical conflicts and decision times, and use the upper confidence boundary algorithm to select candidate strategies to optimize the SAT solution process.
The SAT solution process is accelerated, the number of solveable instances is increased, the solution time is reduced, and the resource utilization efficiency is improved.
Smart Images

Figure CN120235093A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of chip design and verification, and particularly relates to a reinforcement learning SAT solving method and system for chip design verification. Background Art
[0002] Boolean Satisfiability Problem (SAT) solvers play an important role in chip verification, especially in the field of formal verification. A SAT solver is a tool used to determine whether there exists an assignment that makes a given Boolean expression true. In the process of chip design and verification, SAT solvers can be used to solve a series of problems related to circuit verification, such as equivalence checking, property checking, fault simulation, formal verification, optimization design, etc. Equivalence checking is used to ensure that two logical descriptions (such as RTL-level design and gate-level netlist) are functionally equivalent, which is crucial for verifying whether the synthesis process is correct. Through converting assertions or safety properties into Boolean formulas, property checking can be carried out by SAT solvers to verify whether these properties always hold within the entire state space. To evaluate the impact of potential manufacturing defects on the system, various types of fault modes (such as short circuits, open circuits, etc.) can be simulated by SAT solvers to generate test vectors and predict their impacts. In model checking, SAT solvers can help prove whether a given circuit conforms to certain formal specifications, such as the behavior described by LTL (Linear Temporal Logic) or CTL (Computation Tree Logic) formulas. In some cases, SAT solvers can also assist in logical optimization, such as finding a more efficient implementation method to reduce resource usage or improve performance.
[0003] Although SAT solvers have demonstrated powerful functions and extensive applications in chip verification, they also face some inherent challenges and problems. First of all, for large and complex circuits, SAT problems are often NP-complete, meaning that as the input scale increases, the time required to solve the problem may grow exponentially. For particularly large circuit designs, even the most advanced SAT solvers may have difficulty completing verification within a reasonable time. With the growth of chip design complexity and scale, formal verification tools encounter performance bottlenecks when dealing with large designs. Although technology is constantly advancing, there is still a gap compared with the growth rate of the complexity of SoC systems. Secondly, when dealing with large-scale circuits, SAT solvers may consume a large amount of memory resources. This high memory requirement not only limits the maximum circuit size that can be processed on a single machine, but also increases the hardware cost. In addition, commercial-grade SAT solvers usually have more advanced functions and better performance optimizations, but they are expensive; while open-source tools are more accessible, but may lack in performance and features. This makes small and medium-sized enterprises or academic research institutions face dilemmas when choosing suitable tools.
[0004] In the past two decades, modern SAT solvers based on the Conflict Driven Clause Learning (CDCL) algorithm have shown remarkable efficiency in handling complex formulas, even being able to cope with cases containing millions of variables and clauses. The key step of the CDCL algorithm is the decision step, in which an unassigned variable is heuristically selected and its phase (0 or 1) is determined. The Phase Saving technique records the assignments of variables during propagation or backtracking, and these records are used in subsequent operations to set the phases of variables, helping the SAT solver to return to a similar search space more quickly. The saved phase values can be adjusted as needed, which does not affect the correctness of the CDCL algorithm. During literal propagation, Kissat saves the phases of currently assigned variables and marks the operations that change these saved phases as rephasing heuristics. These heuristics not only broaden the exploration scope of the search space but also increase the diversity of learned clauses.
[0005] The excellent SAT solver Kissat designed six rephasing heuristics and combined them into a rephasing strategy. Kissat's default rephasing strategy is {OI(BWOBWIBW#BWF) ω}. Based on this, the Kissat-MAB-rephasing solver introduced a dynamic rephasing strategy based on Multi-Armed Bandits (MAB). This solver selects the most suitable heuristic in each rephasing step through the MAB algorithm, thus accelerating the instance solving process. For simple instances, the Kissat-MAB-rephasing solver is faster than other top solvers. However, when dealing with complex instances, this solver is prone to falling into local optima, resulting in solving failures. This limitation makes the number of complex instances that Kissat-MAB-rephasing can solve less than that of other advanced solvers.
[0006] Based on Kissat_MAB, Kissat_MAB_Conflict+ introduced a novel rephasing heuristic aimed at helping the solver detect more unsatisfiable cores and learn more clauses. This solver combines this new heuristic with Kissat's original heuristics to form a new rephasing strategy. Due to the introduction of the new heuristic, Kissat_MAB_Conflict+ can generate conflicts faster and more efficiently, so the number of difficult instances it solves exceeds that of other excellent solvers. However, when dealing with simple instances, Kissat_MAB_Conflict+ is relatively slow.
[0007] No single rephasing heuristic can be applied to all instances across different domains. Therefore, when solving different instances, multiple rephasing heuristics need to be combined. However, a fixed offline combination strategy also cannot adapt to all instances. With this consideration, reinforcement learning, especially the multi-armed bandit algorithm, may become an effective method to improve the efficiency of SAT solvers. The multi-armed bandit algorithm (MAB) is a classic decision optimization problem originating from probability theory and decision theory. It describes a scenario where a player faces multiple slot machines (i.e., "multi-armed bandits"), and the return rate (reward distribution) of each slot machine is unknown. The player needs to decide how to allocate limited resources (such as the number of coin tosses) among these slot machines to maximize the total return. Currently, SAT solvers using the multi-armed bandit algorithm have not fully considered the characteristics of SAT solving in the selection of arms and the setting of reward functions, so the potential of the multi-armed bandit algorithm in the SAT solving process has not been fully exploited. Summary of the Invention
[0008] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a reinforcement learning SAT solving method and system for chip design verification are provided. The present invention aims to use the multi-armed bandit algorithm to dynamically select rephasing heuristics during the process of the SAT solver solving instances, accelerate the SAT solving process, increase the number of solved instances, and reduce the instance solving time.
[0009] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A reinforcement learning SAT solving method for chip design verification includes the following steps: converting the chip verification problem to be solved into a Boolean expression and then into a CNF formula. The first line of the CNF formula represents the number of variables and the number of clauses, and each subsequent line represents a clause. Each clause is a disjunction of a group of variables, and the entire clause set is a conjunction of these clauses; solving the CNF formula using a SAT solver based on the conflict-driven clause learning algorithm. When the SAT solver based on the conflict-driven clause learning algorithm heuristically selects an unassigned variable from the clause set in the CNF formula and determines its phase based on the saved phase value, the saved phase value is obtained by assigning values to variables based on the rephasing strategy of the multi-armed bandit algorithm. And obtaining the saved phase value by assigning values to variables based on the rephasing strategy of the multi-armed bandit algorithm includes: S101, select rephasing heuristics including the best heuristic and the walk heuristic, and divide the selected rephasing heuristics into two categories: the first category of rephasing heuristics is the heuristic to prevent the algorithm from falling into local optimality, and the second category of rephasing heuristics is the heuristic focusing on optimizing and refining the search space; S102. Combine the best heuristic, the walk heuristic, and the first type of rephasing heuristic to generate multiple rephasing strategies and use them as multiple arms in the multi-armed bandit algorithm; S103. During the solving process of the SAT solver, count the number of conflicts and the number of decisions to calculate the reward function value of each arm. Use the upper confidence bound algorithm to combine the reward function values of each arm to select candidate rephasing strategies from multiple arms, and set the basic conflict count interval for rephasing heuristic switching and the basic conflict count interval for using the multi-armed bandit algorithm to select candidate rephasing strategies according to the number of variables included in the instance of the CNF formula.
[0010] Optionally, the rephasing heuristics including the best heuristic and the walk heuristic selected in step S101 respectively include: Original heuristic: Set all saved phases to 1; Inversion heuristic: Set all saved phases to 0; Best heuristic: Change all saved phases to the best assignment, which is obtained from the current assignment. If the current assignment does not encounter a conflict and the track length of the current assignment is greater than the track length of the best assignment, save the current assignment as the best assignment and immediately reset the best assignment after applying the best phase heuristic; Walk heuristic: Modify the saved phases according to the results of local search; Flip heuristic: Flip the saved phases; Conflict heuristic: When a conflict occurs, update the saved phases of the variables involved in the conflict clause to the current phases of the variables; When the selected rephasing heuristics are divided into two types, the first type of rephasing heuristic includes the original heuristic, the inversion heuristic, the flip heuristic, and the conflict heuristic, and the second type of rephasing heuristic includes the best heuristic and the walk heuristic.
[0011] Optionally, when combining the best heuristic, the walk heuristic, and the first type of rephasing heuristic to generate multiple rephasing strategies and use them as arms in the multi-armed bandit algorithm in step S102, the functional expression for combining the best heuristic, the walk heuristic, and the first type of rephasing heuristic to generate multiple rephasing strategies is: , where, is the set of multiple rephasing strategies, are respectively the four combined rephasing strategies, where is the best heuristic, is the walk heuristic, is the original heuristic, is the inversion heuristic, is the conflict heuristic, is the flip heuristic.
[0012] Optionally, the function expression used to calculate the reward function value of each arm in step S103 is: , in, For arm The reward function value of the tth solution, is the number of decisions made after the most recent selection of the re-phase strategy using the multi-arm game algorithm, is the number of conflicts that have occurred since the last time a re-phasing strategy was selected using the multi-arm game algorithm.
[0013] Optionally, in step S103, the function expression for selecting a candidate rephasing strategy from multiple arms by using the upper confidence bound algorithm combined with the reward function value of each arm is: , in, For arm The priority of For the front In this operation, the slave arm The average value of the reward function obtained, For the front Select arm in the run The number of times, the final choice The arm with the highest value is used as the rephasing strategy for the solver in that run.
[0014] Optionally, when the number of variables contained in the example of the CNF formula is used to set the basic conflict number interval of the re-phasing heuristic switching and the basic conflict number interval of the candidate re-phasing strategy is selected using the multi-arm game algorithm in step S103, the basic conflict number interval of the re-phasing heuristic switching is set to , the basic conflict interval for selecting the re-phase strategy using the multi-arm game algorithm is set to ,in is the number of variables in the CNF formula.
[0015] Optionally, solving the CNF formula using a SAT solver based on a conflict-driven clause learning algorithm includes: S201, for the clause set in the CNF formula, use variable state-independent decay and heuristics to select unassigned variables. If the selection is successful, jump to step S202; otherwise, determine step S206; S202, assigning values to the selected unassigned variables according to the phase value saved for each variable according to the re-phase strategy based on the multi-arm game algorithm; S203, perform unit propagation. If there is only one unassigned variable in a clause, the variable must be assigned to make the clause true. After each unit propagation, check whether all clauses have been satisfied. If so, a solution has been found. If not, continue searching. S204, perform conflict detection. If the Boolean value of a clause under the current assignment is false, it is determined that a conflict is found and jump to step S205; otherwise, jump to step S201; S205, performing conflict analysis and clause learning, including: deriving a new clause summarizing the causes of the current conflict through analysis of the conflict path, and then adding the new clause to the atomic sentence set to avoid similar situations in the future, and then backtracking to undo the most recent decision and return to an earlier state to try different assignment combinations; jumping to step S203; S206, if all clauses are satisfied during the search process of the clause set, a "satisfiable" solution result is output, and a set of variable assignments that make all clauses true is provided; if an assignment that satisfies all clauses is not found after exhausting all possibilities, an "unsatisfiable" solution result is output, indicating that no such assignment combination exists.
[0016] In addition, the present invention also provides a reinforcement learning SAT solving system for chip design verification, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification.
[0017] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification through a processor.
[0018] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the reinforcement learning SAT solving method for chip design verification through a processor.
[0019] Compared with the prior art, the present invention mainly has the following advantages: the reinforcement learning SAT solving method for chip design verification of the present invention includes selecting a rephasing heuristic, combining the selected rephasing heuristics into a plurality of rephasing strategies as a plurality of arms in a multi-arm game algorithm, counting the number of conflicts and the number of decisions in the solving process to calculate the reward function value of each arm, selecting a candidate rephasing strategy from a plurality of arms using an upper confidence bound algorithm, setting a basic conflict number interval for switching the rephasing heuristic using the number of variables contained in an instance, and using a multi-arm game algorithm to select a basic conflict number interval for candidate rephasing strategies. The present invention uses a multi-arm game algorithm to dynamically select a rephasing heuristic in the process of a SAT solver solving an instance, thereby accelerating the SAT solving process, increasing the number of solved instances, and reducing the instance solving time. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the basic flow of the multi-arm game algorithm in an embodiment of the present invention.
[0021] Figure 2 Schematic diagram of the process of solving the SAT problem by the SAT solver in the embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0023] This embodiment provides a reinforcement learning SAT solving method for chip design verification, including the following steps: converting the chip verification problem to be solved into a Boolean expression, and then converting it into a CNF (Conjunctive Normal Form) formula. The conversion to CNF can be achieved through a standard algorithm such as Tseitin transformation. The first line of the CNF formula is the number of variables and the number of clauses, and each subsequent line represents a clause. Each clause is a disjunction of a set of variables (ie, OR), and the entire clause set is the conjunction of these clauses (ie, AND). Positive integers represent corresponding Boolean variables, and negative integers represent the negation of the variable. The output of the SAT solver is the value of each Boolean variable, that is, 0 or 1. If the Boolean expression contains 1 million variables, the SAT solver will output the values of these variables in sequence after the solution is completed; the CNF formula is solved using the SAT solver based on the conflict-driven clause learning algorithm, and during the solution process, the SAT solver based on the conflict-driven clause learning algorithm heuristically selects an unassigned variable for the clause set in the CNF formula and determines its phase based on the saved phase value. The saved phase value is obtained by assigning values to the variable based on the re-phase strategy based on the multi-arm game algorithm (MAB), and if Figure 1As shown, the phase values saved by assigning variables based on the re-phase strategy of the multi-arm game algorithm include: S101, selecting a rephasing heuristic including an optimal heuristic and a walking heuristic, and dividing the selected rephasing heuristic into two categories: the first type of rephasing heuristic is a heuristic that prevents the algorithm from falling into a local optimum, and the second type of rephasing heuristic is a heuristic that focuses on optimizing and refining the search space; S102, combining the best heuristic and the walking heuristic with the first type of re-phase heuristic to generate multiple re-phase strategies and use them as multiple arms in the multi-arm game algorithm; S103, during the solution process of the SAT solver, the number of conflicts and the number of decisions are counted to calculate the reward function value of each arm, and a candidate rephasing strategy is selected from multiple arms using an upper confidence bound (UCB) algorithm in combination with the reward function value of each arm, and a basic conflict number interval for rephasing heuristic switching is set using the number of variables contained in the instance of the CNF formula, as well as a basic conflict number interval for selecting a candidate rephasing strategy using a multi-arm game algorithm.
[0024] The re-phasing heuristics including the best heuristic and the walking heuristic selected in step S101 of this embodiment respectively include: Original heuristic (O): set all saved phases to 1; Inverted heuristic (Inverted, I): set all saved phases to 0; Best heuristic (Best, B): Set all saved phases to change to the best assignment, which is obtained from the current assignment. If the current assignment does not encounter conflicts and the trajectory length of the current assignment is greater than the trajectory length of the best assignment, then save the current assignment as the best assignment and reset the best assignment immediately after applying the best phase heuristic; Walk heuristic (W): Modify the saved phase according to the results of local search; Flipped heuristic (Flipped, F): flip the saved phase; Conflict heuristic (C): When a conflict occurs, the phase saved by the variables involved in the conflict clause is updated to the current phase of the variable; When the selected re-phasing heuristics are divided into two categories, the first category of re-phasing heuristics includes original heuristics, inversion heuristics, flipping heuristics and conflict heuristics, which enable the conflict-driven clause learning algorithm to avoid local optimality and expand the search space; the second category of re-phasing heuristics includes optimal heuristics and walking heuristics, which give priority to the path that is most likely to lead to the global optimality, thereby accelerating the convergence of the conflict-driven clause learning algorithm. In this embodiment, by combining the above two types of re-phasing heuristics, the conflict-driven clause learning algorithm can explore the solution space more effectively, avoid falling into the local optimality, and enhance the robustness and adaptability in solving complex problems.
[0025] In step S102 of this embodiment, when the best heuristic, the walking heuristic and the first type of re-phasing heuristic are combined to generate multiple re-phasing strategies and used as arms in the multi-arm game algorithm, the function expression for combining the best heuristic, the walking heuristic and the first type of re-phasing heuristic to generate multiple re-phasing strategies is: , in, is a collection of multiple re-phasing strategies, They are four combined re-phasing strategies, among which For the best heuristic, For the walk heuristic, For the original heuristic, For the reversal heuristic, Conflict heuristics, That is, the best heuristic, the walking heuristic and the first type of re-phase heuristic are combined to generate four re-phase strategies, which are used as the four arms in the multi-arm game algorithm.
[0026] The function expression used to calculate the reward function value of each arm in step S103 of this embodiment is: , in, For arm The reward function value of the tth solution, is the number of decisions made after the most recent selection of the re-phase strategy using the multi-arm game algorithm, The number of conflicts that occurred since the last time the rephasing strategy was selected using the multi-arm game algorithm. This reward function is used to evaluate the efficiency of the rephasing strategy, which is generally applicable to all SAT solvers based on multi-arm games and easy to implement. The initial reward value of each arm of the multi-arm game algorithm is set to: use each arm in turn during the solver solution process, count the number of decisions and conflicts made after the last selection of the rephasing strategy using the multi-arm game algorithm, and calculate the initial reward value of each arm according to the reward function.
[0027] In step S103 of this embodiment, the function expression for selecting a candidate rephasing strategy from multiple arms by using the upper confidence bound algorithm combined with the reward function value of each arm is: , in, For arm The priority of For the front In this operation, the slave arm The average value of the reward function obtained, For the front Select arm in the run The number of times, and finally choose The arm with the highest value is used as the rephasing strategy used by the solver in the current run, that is, by maximizing A candidate rephasing strategy is selected from a plurality of arms. For the number of The upper confidence bound algorithm is used to calculate the natural logarithmic function of each arm. When selecting a candidate re-phase strategy using the multi-arm game algorithm, the value of each arm is calculated according to the formula Value, determine which arm The arm with the largest value is selected as the rephasing strategy used by the SAT solver next.
[0028] In step S103 of this embodiment, when the number of variables contained in the example of the CNF formula is used to set the basic conflict number interval of the re-phasing heuristic switching and the basic conflict number interval of the candidate re-phasing strategy is selected using the multi-arm game algorithm, the basic conflict number interval of the re-phasing heuristic switching is set to The basic conflict interval for selecting the re-phase strategy using the multi-arm game algorithm is set to ,in is the number of variables in the CNF formula. In this embodiment, the number of conflicts that occurred after the most recent selection of the re-phasing strategy using the multi-arm game algorithm is used as the criterion for switching the re-phasing heuristic. When the number of conflicts is greater than or equal to the number of variables, the re-phasing heuristic needs to be switched. When the solver just starts to solve the instance, the initial re-phasing heuristic is the original heuristic (O), followed by the inversion heuristic (I), and then each arm in the multi-arm game algorithm is used in turn as the re-phasing heuristic strategy, each arm is used once, and the initial reward value of each arm is obtained. Then select the candidate re-phasing strategy according to the upper confidence bound algorithm. When the number of conflicts that occur after the re-phasing strategy is selected using the multi-arm game algorithm is greater than or equal to three times the number of variables, a re-phasing strategy is selected from the four candidate arms according to the upper confidence bound algorithm. That is: first use the OIBWOBWIBWCBWF re-phasing heuristics in turn, and then select the candidate re-phasing strategy from the four arms according to the upper confidence bound algorithm.
[0029] like Figure 2 As shown, in this embodiment, solving the CNF formula using a SAT solver based on a conflict-driven clause learning algorithm includes: S201, using the variable state independent decaying sum heuristic to select unassigned variables for the clause set in the CNF formula, if the selection is successful, jump to step S202; otherwise, determine step S206; S202, assigning values to the selected unassigned variables according to the phase value saved for each variable according to the re-phase strategy based on the multi-arm game algorithm; S203, perform unit propagation. If there is only one unassigned variable in a clause, the variable must be assigned to make the clause true. After each unit propagation, check whether all clauses have been satisfied. If so, a solution has been found. If not, continue searching. Unit propagation is automatically performed after each new assignment until there are no more unit clauses. S204, perform conflict detection. If the Boolean value of a clause under the current assignment is false, it is determined that a conflict is found and jump to step S205; otherwise, jump to step S201; S205, performing conflict analysis and clause learning, including: deriving a new clause summarizing the causes of the current conflict through analysis of the conflict path, and then adding the new clause to the atomic sentence set to avoid similar situations in the future, and then backtracking to undo the most recent decision and return to an earlier state (i.e., the decision layer) to try different assignment combinations; jumping to step S203; S206, if all clauses are satisfied during the search process of the clause set, a solution result of "Satisfiable" is output, and a set of variable assignments that make all clauses true is provided; if an assignment that satisfies all clauses is not found after exhausting all possibilities, a solution result of "Unsatisfiable" is output, indicating that no such assignment combination exists.
[0030] In summary, the method of this embodiment discloses a conflict-driven clause learning method based on a multi-arm game algorithm. The present invention includes selecting six re-phasing heuristics, combining the six re-phasing heuristics into four re-phasing strategies as four arms in the multi-arm game algorithm, counting the number of conflicts and the number of decisions during the solution process to calculate the reward function value of each arm, selecting candidate re-phasing strategies from the four arms using the upper confidence bound algorithm, setting the basic conflict number interval for switching the re-phasing heuristic using the number of variables contained in the instance, and using the multi-arm game algorithm to select the basic conflict number interval for the candidate re-phasing strategy. The method of this embodiment uses the multi-arm game algorithm to dynamically select the re-phasing heuristic in the process of solving the instance by the SAT solver, speeding up the SAT solution process, increasing the number of solved instances, and reducing the instance solution time.
[0031] In addition, this embodiment also provides a reinforcement learning SAT solving system for chip design verification, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification.
[0032] In addition, this embodiment also provides a computer-readable storage medium, which stores a computer program or instructions, and the computer program or instructions are programmed or configured to execute the reinforcement learning SAT solving method for chip design verification through a processor.
[0033] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the reinforcement learning SAT solving method for chip design verification through a processor.
[0034] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention may be in the form of methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that instructions executed by the processor of a computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0035] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A reinforcement learning SAT solving method for chip design verification, characterized in that: The method comprises the following steps: converting the chip verification problem to be solved into a Boolean expression, and then converting it into a CNF formula, wherein the first line of the CNF formula is the number of variables and the number of clauses, and each subsequent line represents a clause, each clause is a disjunction of a set of variables, and the entire clause set is the conjunction of these clauses; The CNF formula is solved using a SAT solver based on a conflict-driven clause learning algorithm, and during the solving process, the SAT solver based on the conflict-driven clause learning algorithm heuristically selects an unassigned variable for a clause set in the CNF formula and determines its phase based on a saved phase value, the saved phase value is obtained by assigning a value to the variable using a re-phase strategy based on a multi-arm game algorithm, and the saved phase value obtained by assigning a value to the variable using a re-phase strategy based on the multi-arm game algorithm includes: S101, selecting a rephasing heuristic including an optimal heuristic and a walking heuristic, and dividing the selected rephasing heuristic into two categories: the first type of rephasing heuristic is a heuristic that prevents the algorithm from falling into a local optimum, and the second type of rephasing heuristic is a heuristic that focuses on optimizing and refining the search space; S102, combining the best heuristic and the walking heuristic with the first type of re-phase heuristic to generate multiple re-phase strategies and use them as multiple arms in the multi-arm game algorithm; S103, during the solution process of the SAT solver, the number of conflicts and the number of decisions are counted to calculate the reward function value of each arm, and a candidate rephasing strategy is selected from multiple arms using an upper confidence bound algorithm in combination with the reward function value of each arm, and the basic conflict number interval for rephasing heuristic switching is set using the number of variables contained in the instance of the CNF formula, as well as the basic conflict number interval for selecting the candidate rephasing strategy using a multi-arm game algorithm.
2. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: The re-phasing heuristics including the best heuristic and the walking heuristic selected in step S101 include: Original heuristic: set all saved phases to 1; Reversal heuristic: set all saved phases to 0; Best heuristic: Set all saved phases to change to the best assignment, which is obtained from the current assignment. If the current assignment does not encounter conflicts and the trajectory length of the current assignment is greater than the trajectory length of the best assignment, then save the current assignment as the best assignment and reset the best assignment immediately after applying the best phase heuristic; Walking heuristic: modify the saved phase according to the results of local search; Flip heuristic: flip the saved phase; Conflict heuristic: When a conflict occurs, the phase saved by the variables involved in the conflict clause is updated to the current phase of the variable; When the selected rephasing heuristics are divided into two categories, the first category of rephasing heuristics includes original heuristics, inversion heuristics, flipping heuristics and conflict heuristics, and the second category of rephasing heuristics includes optimal heuristics and walking heuristics.
3. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: When the best heuristic, the walk heuristic and the first type of re-phase heuristic are combined to generate multiple re-phase strategies in step S102 and used as arms in the multi-arm game algorithm, the function expression for combining the best heuristic, the walk heuristic and the first type of re-phase heuristic to generate multiple re-phase strategies is: , in, is a collection of multiple re-phasing strategies, They are four combined re-phasing strategies, among which For the best heuristic, For the walk heuristic, For the original heuristic, For the reversal heuristic, Conflict heuristics, is the flip heuristic.
4. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: The function expression used to calculate the reward function value of each arm in step S103 is: , in, For arm The reward function value of the tth solution, is the number of decisions made after the most recent selection of the re-phase strategy using the multi-arm game algorithm, is the number of conflicts that have occurred since the last time a re-phasing strategy was selected using the multi-arm game algorithm.
5. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: In step S103, the function expression of selecting a candidate re-phasing strategy from multiple arms by using the upper confidence bound algorithm combined with the reward function value of each arm is: , in, For arm The priority of For the front In this operation, the slave arm The average value of the reward function obtained, For the front Select arm in the run The number of times, the final choice The arm with the highest value is used as the rephasing strategy for the solver in that run.
6. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: When the number of variables contained in the example of the CNF formula is used in step S103 to set the basic conflict number interval of the re-phasing heuristic switching and the basic conflict number interval of the candidate re-phasing strategy is selected using the multi-arm game algorithm, the basic conflict number interval of the re-phasing heuristic switching is set to , the basic conflict interval for selecting the re-phase strategy using the multi-arm game algorithm is set to ,in is the number of variables in the CNF formula.
7. The reinforcement learning SAT solving method for chip design verification according to claim 1, characterized in that: Solving the CNF formula using a SAT solver based on a conflict-driven clause learning algorithm includes: S201, for the clause set in the CNF formula, use variable state-independent decay and heuristics to select unassigned variables. If the selection is successful, jump to step S202; otherwise, determine step S206; S202, assigning values to the selected unassigned variables according to the phase value saved for each variable according to the re-phase strategy based on the multi-arm game algorithm; S203, perform unit propagation. If there is only one unassigned variable in a clause, the variable must be assigned to make the clause true. After each unit propagation, check whether all clauses have been satisfied. If so, a solution has been found. If not, continue searching. S204, perform conflict detection. If the Boolean value of a clause under the current assignment is false, it is determined that a conflict is found and jump to step S205; otherwise, jump to step S201; S205, performing conflict analysis and clause learning, including: deriving a new clause summarizing the causes of the current conflict through analysis of the conflict path, and then adding the new clause to the atomic sentence set to avoid similar situations in the future, and then backtracking to undo the most recent decision and return to an earlier state to try different assignment combinations; jumping to step S203; S206, if all clauses are satisfied during the search process of the clause set, a "satisfiable" solution result is output, and a set of variable assignments that make all clauses true is provided; if an assignment that satisfies all clauses is not found after exhausting all possibilities, an "unsatisfiable" solution result is output, indicating that no such assignment combination exists.
8. A reinforcement learning SAT solving system for chip design verification, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification as described in any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the reinforcement learning SAT solving method for chip design verification as described in any one of claims 1 to 7 through a processor.