Interactive theorem proving system and method based on large language model

By introducing policy suggestions, proof search and premise selection modules into the interactive theorem proof system, and dynamically generating inference strategies using large language models, the existing system's poor performance on complex theorems and cross-domain problems is solved, and an efficient and automated theorem proof process is realized.

CN120163213APending Publication Date: 2025-06-17JIANGSU IND INNOVATION CENT OF INTELLIGENT EQUIP CO LTD

Patent Information

Application Number
CN202510255953.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing interactive theorem proof system based on large language models has problems such as lack of independence and interactivity in the inference process, low degree of automation, and difficulty in seamless integration with mainstream theorem proofs, resulting in poor performance in the face of complex theorem or cross-domain problems.

Method used

An interactive theorem proof system based on a large language model is designed, including a strategy proposal module, a proof search module and a prerequisite selection module. The reasoning strategy is dynamically generated through LLM, and combined with rule base and user input, a multi-step strategy combination proof of complex theorems is realized, and the prerequisite selection process is optimized.

Benefits of technology

It significantly reduces manual intervention, improves the level of automation, improves the proof success rate and efficiency of complex theorems, reduces the computing resource requirements, and achieves seamless integration with mainstream theorem proofs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163213A_ABST
    Figure CN120163213A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive theorem proving system and method based on a large language model, and the system comprises a strategy suggestion module which is used for analyzing a plurality of candidate strategies of a to-be-proven target based on LLM; the proving search module is used for performing deep reasoning based on a rule base and a plurality of candidate strategies when a to-be-proved target is a complex theorem needing multi-step strategy combination proving, and determining a complete proving path; the premise selection module is used for determining a premise sequence corresponding to the to-be-proved target based on similarity calculation when the premise needs to be cited in the proving process of the to-be-proved target, and performing theorem proving based on the premise selected by the user in the premise sequence; a reasoning strategy can be dynamically generated by utilizing a large language model, the proving process of a complex theorem is intelligently guided, manual intervention is remarkably reduced, and the automation level is improved; the reasoning path is dynamically adjusted according to the theorem target context, mathematical special situations and cross-domain proving tasks are flexibly coped, and the proving success rate and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of theorem proving, and particularly to an interactive theorem proving system and method based on a large language model. Background Art

[0002] Interactive theorem proving (ITP), as a core tool in the fields of mathematics and computer science, plays a crucial role, especially in theorem verification and program correctness checking. Traditional ITP tools highly rely on manual input from users, requiring users to gradually and precisely construct the proof process. However, with the rapid development of artificial intelligence technology, especially the rise of large language models (LLMs), new possibilities have been opened up for automated reasoning-assisted theorem proving.

[0003] Although the proof systems based on LLMs have shown great potential, there are still several key problems to be solved. Firstly, the reasoning processes of these systems often lack independence and interactivity, making it difficult for users to intuitively participate in the proof process, which affects the proof efficiency and user experience. Secondly, the low degree of automation is another major challenge, which makes the system feel powerless when facing complex theorems or cross-domain problems. Thirdly, most of the existing proof tools based on LLMs adopt independent reasoning environments and fail to achieve seamless integration with mainstream interactive theorem provers (such as Lean, Coq, etc.), which limits their extensiveness and practicality in actual applications.

[0004] Specifically, when facing complex theorems, these independent reasoning environments may not be able to provide logically consistent reasoning strategies or even complete the proof. At the same time, the high overhead of computing resources and slow reasoning speed have also become bottlenecks restricting the development of these systems, especially in proof scenarios that require real-time interaction, where these problems are particularly prominent.

[0005] In response to the above problems, Chinese Patent CN118709756A discloses a formal theorem proving generation method and system based on a large model. By using a rule base and an inference algorithm, combined with the analysis of theorem objectives, this patent can automatically generate inference steps and conduct automatic theorem proving. By setting certain inference rules to guide the proof process, this patent aims to reduce manual intervention and improve the proof efficiency.

[0006] However, this patent also has certain limitations: First, it relies on a static rule base and pre-set inference strategies, which may perform well when dealing with standard theorems. However, when faced with complex theorems or targets with special structures, the rule base may not be able to provide sufficiently flexible inference paths, thus limiting its application scope. Second, this patent also lacks seamless integration with existing theorem proving environments during the inference process. Users need to switch between different platforms, increasing the complexity of operations and the usage threshold. Third, in terms of computing resource requirements, this patent still requires a large amount of computing resources when dealing with complex theorems, which poses a major challenge for users or environments with limited resources.

[0007] In summary, although the formal theorem proving method and system based on large models have promoted the development of automatic theorem proving solutions to a certain extent, they still need to be further optimized and improved in terms of inference flexibility, limitations, and computing resource usage costs. Summary of the Invention

[0008] The purpose of the present invention is to provide an interactive theorem proving system and method based on a large language model for the above problems in the prior art, thereby solving all or one of the above problems existing in the prior art.

[0009] To solve the above technical problems, the specific technical solutions of the present invention are as follows: On the one hand, the present invention provides an interactive theorem proving system based on a large language model, including: A strategy recommendation module, configured to: obtain a to-be-proven target input by a user, and analyze several candidate strategies for the to-be-proven target based on the LLM; A proof search module, configured to: when the to-be-proven target is a complex theorem that requires a multi-step strategy combination for proof, perform in-depth reasoning based on a rule base and several of the candidate strategies to determine a complete proof path; A premise selection module, configured to: when the proof process of the to-be-proven target requires the citation of premises, determine a premise sequence corresponding to the to-be-proven target based on a similarity calculation operation, and perform theorem proof based on the premises selected by the user in the premise sequence.

[0010] As an improved solution, the strategy recommendation module includes: a proof target determination unit, a candidate strategy generation unit, a strategy verification unit, and a strategy classification unit; The proof target determination unit is configured to: obtain the to-be-proven target, generate a preliminary proof target according to the to-be-proven target, and input the preliminary proof target into the LLM inference engine; The candidate strategy generation unit is configured to: call the LLM inference engine to generate the next inference strategy as the candidate strategy according to the preliminary proof target; The policy verification unit is used to: apply the candidate policy to the target to be proved, and judge the policy attribute of the candidate policy; The policy classification unit is used to: divide the candidate policy into a successful policy, an intermediate policy, and a failed policy according to the policy attribute.

[0011] As an improved solution, the proof search module includes: a target acquisition unit, a rule library call unit, an inference policy generation unit, a dynamic proof unit, and a result analysis unit; The target acquisition unit is used to: acquire the complex theorem; The rule library call unit is used to: load the aesop rule library; The policy introduction unit is used to: introduce a number of the candidate policies; The dynamic proof unit is used to: call the aesop rule library to perform iterative search for the proof path of the complex theorem based on the best-first search algorithm and the tree structure expansion algorithm until a complete proof path is found or a preset condition is reached; during the iterative search process, combine a number of the candidate policies and the context information of the complex theorem to dynamically generate an inference policy for supporting extended analysis; The result analysis unit is used to: display the result of the complete proof path on the system interface.

[0012] As an improved solution, the premise selection module includes: a target vectorization unit, a premise matching unit, a candidate premise processing unit, a most relevant premise screening unit, and an interaction support unit; The target vectorization unit is used to: call the LLM inference engine to parse the semantics of the target to be proved, and represent the target to be proved in the form of an embedding vector; The premise matching unit is used to: load a premise embedding matrix, and use an efficient mathematical method to calculate the similarity score between the target to be proved represented by the embedding vector and each premise in the premise embedding matrix; The candidate premise processing unit is used to: judge the loading situation of the candidate premise after calculating the similarity score in the current proof environment, determine the loaded premise and the unloaded premise according to the loading situation, explain the information of the loaded premise, and supplement the explanation of the unloaded premise; The most relevant premise screening unit is used to: set a K value, extract the most relevant premise according to the similarity score and the K value, and display the most relevant premise to the user; The interaction support unit is used to: display the premise sequence of the most relevant premise through the interface.

[0013] As an improved solution, the candidate strategy includes: an existing proof step, or a new strategy generated by the LLM inference engine through inference.

[0014] As an improved solution, the strategy classification unit is further configured to: in response to the strategy attribute indicating that the strategy successfully solves the target to be proved, determine that the strategy is the successful strategy; The strategy classification unit is further configured to: in response to the strategy attribute indicating that the strategy can change the target to be proved, determine that the strategy is the intermediate strategy; The strategy classification unit is further configured to: in response to the strategy attribute indicating that the strategy leads to a proof failure and an error occurs, determine that the strategy is the invalid strategy.

[0015] As an improved solution, a preliminary rule set is configured in the aesop rule library; The preliminary rule set includes: apply inference rule, simp inference rule, and cases inference rule.

[0016] As an improved solution, in each search step of the iterative search, several candidate strategies of the aesop rule library and the LLM inference engine act together on the expansion of the target to be proved; The tree structure is a search tree, and each node of the search tree supports expansion using several strategies, and each node of the search tree represents each sub-goal of the target to be proved.

[0017] As an improved solution, the result analysis unit is further configured to: display in the system interface a proof step list including multiple proof steps, a remaining goal set including unsolved proof goals, and multi-path information including multiple possible proof paths.

[0018] On the other hand, the present invention also provides an interactive theorem proving method based on a large language model, including the following steps: Strategy recommendation step: Obtain the target to be proved input by the user, and analyze several candidate strategies for the target to be proved based on the LLM; Proof search step: When the target to be proved is a complex theorem that requires a multi-step strategy combination proof, perform in-depth reasoning based on the rule library and several candidate strategies to determine the complete proof path; Premise selection step: When the proof of the target to be proved requires the citation of premises, determine the premise sequence corresponding to the target to be proved based on the similarity calculation operation, and perform theorem proving based on the premises selected by the user in the premise sequence.

[0019] The beneficial effects of the technical solution of the present invention are: 1. The interactive theorem proving system based on large language models according to the present invention can utilize large language models to dynamically generate inference strategies, intelligently guide the proof process of complex theorems, significantly reduce manual intervention, and improve the automation level; dynamically adjust the inference path according to the theorem target context, flexibly handle mathematical special cases and cross-domain proof tasks, and improve the proof success rate and efficiency; by optimizing the inference process and using CTranslate2 for acceleration, the present invention can run efficiently on conventional hardware, reduce the computing resource requirements, and enhance the popularity and practicality of the system.

[0020] 2. The interactive theorem proving method based on large language models according to the present invention can orderly call system modules, and thus implement the system logic of the interactive theorem proving system based on large language models according to the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 is a schematic diagram of the overall framework of the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention; Figure 2 is a schematic diagram of the definition of the add_abc theorem in the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention; Figure 3 is a schematic diagram of an example of the VSCode window when the strategy suggestion module in the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention runs; Figure 4 is a schematic diagram of an example of the VSCode window when the proof search module in the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention runs; Figure 5 is a schematic diagram of an example of the VSCode window when the premise selection module in the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention runs; Figure 6 is a schematic diagram of the architecture of the interactive theorem proving system based on large language models according to Embodiment 1 of the present invention; Figure 7 is a schematic diagram of the process of the interactive theorem proving method based on large language models according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will elaborate on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.

[0024] In the description of the present invention, it should be noted that the embodiments described in the present invention are part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0025] The terms "first", "second", etc. in the specification and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or equipment. Embodiment 1

[0026] This embodiment provides an interactive theorem proving system based on a large language model, as Figures 1 to 6 shown, including: This system includes the west genesis framework, which enables the interaction between Lean and LLMs through an external function interface; Lean is implemented in C++, and FFI efficiently calls C++ to achieve the native operation of LLMs in the system; in a preferred implementation, the ByT5 pre-trained ReProver model is used by default, and user customization is also supported. Users can simply configure their own models to integrate with the west genesis framework, and then perform different reasoning tasks in Lean; specifically, the framework of this system includes the following three main modules, which further assist users in quantitative proof, as Figure 1 shown, including: a strategy suggestion module, a proof search module, and a premise selection module.

[0027] The functions of the above modules will be described in detail below: (1) The strategy suggestion module is mainly used for: when it is difficult to determine the next strategy during the theorem proving process, automatically generating the next proof strategy according to the current proof goal; among them, based on the powerful reasoning ability of the LLM, multiple strategy options are provided for users, and finally help users select the most effective reasoning steps; Among them, the policy recommendation module mainly includes: a proof goal determination unit, a candidate policy generation unit, a policy verification unit, and a policy classification unit, which are specifically as follows: (1.1) The proof goal determination unit is used to: obtain the theorem to be proved input by the user in the Lean IDE, automatically generate preliminary proof goals according to the obtained theorem to be proved, display these goals in the form of logical formulas, and transfer these goals as inputs to the LLM inference engine; among them, the preliminary proof goals represent several sub-goals that need to be gradually solved in the theorem proof process, that is, the sub-tasks that need to be completed in the proof.

[0028] (1.2) The candidate policy generation unit is used to: when the goals are input to the LLM, automatically call the LLM to complete the policy recommendation analysis, that is, generate the next inference strategy according to these goals; in this process, conduct policy Q&A with the LLM, that is, according to the current goals, ask the LLM about possible next strategies, and the LLM generates several candidate strategies according to the previous training model; the candidate strategies are usually existing proof steps in the system (such as common strategies like "apply", "simp", and "cases"), or may also be new strategies generated by the LLM through inference.

[0029] (1.3) The policy verification unit is used to: apply each candidate policy to the current goal and check whether it will cause errors or unable to continue reasoning, and then judge the policy attribute; among them, in response to a certain policy successfully solving the current proof goal and leaving no unfinished sub-goals, at this time, judge that this policy is a successful policy; in response to a certain policy causing the proof to fail and generating irrecoverable errors (such as syntax errors and logical conflicts, etc.), at this time, judge that this policy is an invalid policy, mark and eliminate this policy.

[0030] (1.4) The policy classification unit is used to: combine the policy attributes verified by the policy verification unit and classify the finally generated policies into successful policies, intermediate policies, and failed policies; Specifically, the successful policy is: a policy that can directly complete the proof goal and solve all remaining sub-goals. Such policies are marked as "green" in the interface (this color can be adjusted according to specific requirements) for the convenience of users to directly use.

[0031] Specifically, the intermediate policy is: a policy that can change the current goal but fails to complete the proof and may introduce new sub-goals; these policies do not completely cause the proof to fail, and they may complete the proof through further steps. Therefore, these policies are marked as "blue" in the interface and display the remaining goals they cause.

[0032] Specifically, the failure strategy is: a strategy that causes an error and cannot continue the proof; these strategies will be automatically excluded and not displayed in the interface, ensuring that the user does not see invalid options.

[0033] In a preferred embodiment, as Figure 3 shown, an operating example of the strategy suggestion module is as follows: (i) Taking the proof of the theorem "add_abc" (defined as "theorem add_abc (abc:Nat):a+b+c=a+c+b") as an example, after the user inputs the theorem definition and manually executes some initial strategies (such as "cases n" and "dsimp"), then the suggest_tactics tool (i.e., the strategy suggestion module) is called; (ii) The suggest_tactics tool sends the current remaining goal ("a+b+c=a+c+b") to the LLM, and the LLM may return candidate strategies such as "apply Nat.add_right_comm" and "rw [Nat.add_assoc]".

[0034] (iii) Verify the candidate strategies: It is found that "apply Nat.add_right_comm" can successfully complete the proof. As a successful strategy, it will be marked green and displayed in the information view later; It is found that strategies such as "rw [Nat.add_assoc]" cannot directly complete the proof but have no errors. As intermediate strategies, they are marked blue and the remaining goals after their application are shown.

[0035] Strategies that cannot complete the proof are excluded and not shown.

[0036] (iiii) Support the user to choose to directly apply "apply Nat.add_right_comm" to complete the proof, or further explore intermediate steps according to the blue strategy tips.

[0037] It should be noted that, as Figure 2 shown, the definition of the simple theorem add_abc is: when no proof steps are applied, the strategy state after the theorem definition is the definition itself.

[0038] (2) The proof search module is mainly used for: dealing with complex theorems that require a combination of multi-step strategies for proof. At this time, it automatically searches and generates a complete and multi-step proof path. This operation conducts in-depth reasoning by combining the candidate strategies generated by the LLM and the existing rule library (such as aesop) to determine the complete proof path; Among them, the proof search module mainly includes: a target acquisition unit, a rule library unit, a strategy introduction unit, a dynamic expansion proof (tree) unit, and a result analysis unit, which are specifically as follows: (2.1) The target acquisition unit is used to: when the current target may involve complex reasoning steps and requires multiple intermediate steps to complete the proof, acquire the current target from the Lean environment.

[0039] (2.2) The rule library unit is used to: carry the aesop rule library; among them, a preliminary rule set is configured in the aesop rule library, and the rule set includes but is not limited to common reasoning rules (such as apply, simp, and cases, etc.), and these rules can be used to guide the expansion of the search tree.

[0040] (2.3) The strategy introduction unit is used to: introduce the candidate strategies generated by the LLM. (2.4) The dynamic expansion proof (tree) unit is used to: call the aesop rule library to continuously iteratively search for the proof path of the complex theorem by expanding different strategies and steps (i.e., the tree structure expansion analysis algorithm) in the tree structure based on the best-first search algorithm; during this continuous search process, combine the candidate strategies generated by the LLM and according to the context information of the current target, dynamically generate inference strategies, and provide diverse candidate steps; until a complete proof path is found or a preset condition (such as timeout) is reached, obtain the most likely successful strategy for expansion; finally, feedback the complete strategy sequence at the time of successful proof or relevant information at the time of failure to the user; among them, the above-mentioned inference strategies can not only adapt to the current target, but also support providing new expansion paths for the search tree according to the target structure and context information; among them, in each search step, the candidate strategies generated by aesop and the LLM act together on the expansion of the proof target; each node of the search tree can be expanded by applying different strategies, that is, each node represents the current sub-goal.

[0041] (2.5) The result analysis unit is used to: when determining the complete proof path, display the result on the system interface, and the specific display content is as follows: 1. The proof step list, which contains multiple proof steps, and each proof step is presented in the form of a strategy, supporting the user to view the detailed information of each step and judge whether it meets the reasoning intention required by the user.

[0042] 2. The remaining target set, which is used to display the unresolved targets when the proof is not completely successful, supporting the user to further select strategies or adjust the proof steps.

[0043] 3. Multiple-path display. This display is used to show multiple possible proof paths in the interface when the LLM generates multiple candidate strategies, support users to select different paths for verification, and help users choose the best proof method.

[0044] In a preferred embodiment, as Figure 4 shown, the running example of the proof search module is as follows: (i) Taking the "add_abc" theorem as an example, after the user inputs the "add_abc" theorem, directly call the search_proofs tool (i.e., the proof search module).

[0045] (ii) Use the initial rule set of aesop to expand the search. During the search process, call the large language model to generate corresponding strategies according to the current goals (such as the initial goal and the intermediate goals generated subsequently). For example, strategies such as "simp[Nat.add_comm,Nat.add_left_comm]" may be generated to expand the rule set.

[0046] (iii) Continuously iterate the search until a complete proof path is found (for example, through the application of a series of strategies, the theorem is finally successfully proved). At this time, display the proof process in the information view, and at the same time update the strategy status to "no goal" to indicate successful proof.

[0047] (3) Premise selection module, mainly used for: according to the user's current proof goal, calculate the corresponding premise queue for the user to select based on similarity; Among them, the premise selection module mainly includes: a goal vectorization unit, a premise matching unit, a candidate premise processing unit, a most relevant premise screening unit, and an interaction support unit, specifically as follows: (3.1) Goal vectorization unit, used for: receiving the user's current proof goal, calling the LLM to parse the semantics of the current proof goal, and finally representing the proof goal in the form of an embedding vector.

[0048] (3.2) Premise matching unit, used for: loading the premise embedding matrix, and after obtaining the embedding vector of the current proof goal, using efficient mathematical methods (such as cosine similarity or dot product, etc.) to measure the correlation between the current proof goal and each premise in the premise embedding matrix, and finally assigning a similarity score to each premise, and arranging the premises in descending order according to the scores; among them, the premise refers to known mathematical facts or theorems, usually configured in an external library (such as Mathlib); the premise embedding matrix stores the vector representations of all premises in the external library (obtained by encoding each premise by the LLM), and each row of the premise embedding matrix represents the embedding vector of a premise.

[0049] The (3.3) candidate premise processing unit is used to: determine whether the candidate premise after similarity calculation has been loaded into the current proof environment, directly provide its type, usage, and detailed description for the loaded premise; for the unloaded premise, prompt the module or library that needs to be imported, and provide relevant code and documentation.

[0050] The (3.4) most relevant premise screening unit is used to: set the K value (this K value supports adjustment by the user according to system suggestions for optimizing the selection range), extract the top K most relevant premises according to the similarity calculation result and this K value, and display them to the user in the order of relevance, helping the user quickly focus on the key premises.

[0051] The (3.5) interaction support unit is used to: display the aforementioned premise sequence through an intuitive interface, support the user to select and apply the most relevant premises recommended through the intuitive interface; among them, the loaded premises support the user to directly call, and the unloaded premises provide detailed loading instructions to facilitate the user to import and perform subsequent operations; in addition, this unit provides real-time feedback on the user's selection, enabling the user to always maintain control over the premise selection result.

[0052] In a preferred embodiment, as Figure 5 shown, the running example of the premise selection module is as follows: (i) When theorems or lemmas in certain mathematical libraries need to be cited during the proof process, it indicates that the proof involves the situation of needing to select relevant premises. At this time, call the select_premises tool (i.e., the premise selection module) at an appropriate position.

[0053] (ii) Encode the current proof goal and calculate it with the pre-computed premise embeddings to match the premises with higher relevance from the mathematical library (such as Mathlib) and other relevant libraries; for example, for the theorem proof related to the addition of natural numbers, premises such as "Nat.add_left_comm" and "Nat.add_right_comm" may be selected.

[0054] (iii) Mark the selected premises, such as displaying type information such as "Nat.add_left_comm:∀(nmk:Nat),n+m+k=n+(m+k)"; if there is a docstring, it will be displayed together; for the premises for which the relevant package has not been imported, prompt the required package and the complete definition code of the premise; based on this, the user can judge whether the premise meets the proof requirements according to these marked information, and the user can import the corresponding package as needed and use the premise to continue the proof.

[0055] In a preferred embodiment, the installation operation of the west genesis framework is as follows: Obtain the westgenesis code and install it as a Lean package according to the installation guide, without complex additional configuration during the installation process.

[0056] In a preferred embodiment, the configuration operation of the large language model in the framework is as follows: Use the default ReProver model or custom-introduce other pre-trained models; When customizing the model, package it according to the interface specifications stipulated by the west genesis framework to ensure that it can interact with the framework normally (for example, correctly process the input proof target and generate appropriate output strategy suggestions or vector encodings, etc.).

[0057] It should be noted that the above examples are only for explaining the present invention and should not limit the protection scope of the present invention. Embodiment 2

[0058] This embodiment is based on the same inventive concept as the interactive theorem proving system based on a large language model described in Embodiment 1, and provides an interactive theorem proving method based on a large language model, as Figure 7 shown, including the following steps: S100, Strategy Suggestion Step: Obtain the proof target to be proved input by the user, and analyze several candidate strategies for the proof target based on the LLM; S200, Proof Search Step: When the proof target is a complex theorem that requires a multi-step strategy combination for proof, perform in-depth reasoning based on the rule base and several of the candidate strategies to determine the complete proof path; S300, Premise Selection Step: When the proof process of the proof target requires the citation of premises, determine the premise sequence corresponding to the proof target based on the similarity calculation operation, and perform theorem proving based on the premises selected by the user in the premise sequence.

[0059] Different from the prior art, by using the interactive theorem proving system and method based on a large language model of the present application, it is possible to dynamically generate inference strategies using the large language model, intelligently guide the proof process of complex theorems, significantly reduce manual intervention, and improve the automation level; Dynamically adjust the inference path according to the theorem target context, flexibly handle mathematical special cases and cross-domain proof tasks, and improve the proof success rate and efficiency; By optimizing the inference process and using CTranslate2 for acceleration, the present invention can run efficiently on conventional hardware, reduce the computing resource requirements, and enhance the popularity and practicality of the system.

[0060] It should be understood that in various embodiments herein, the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments herein.

[0061] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0062] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.

[0063] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific logical processes of the above-described methods can refer to the corresponding working processes of the systems, devices, and units in the foregoing method embodiments, and will not be elaborated herein.

[0064] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0065] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments herein.

[0066] In addition, each functional unit in the various embodiments of this document may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0067] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this document. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0068] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An interactive theorem proving system based on a large language model, characterized in that: include: The strategy suggestion module is used to: obtain the target to be proved input by the user, and analyze several candidate strategies for the target to be proved based on LLM; A proof search module is used to: when the target to be proved is a complex theorem that requires a combination of multi-step strategies to prove, perform deep reasoning based on the rule base and several candidate strategies to determine a complete proof path; The premise selection module is used to: when the proof process of the target to be proved needs to refer to the premise, determine the premise sequence corresponding to the target to be proved based on the similarity calculation operation, and prove the theorem based on the premise selected by the user in the premise sequence.

2. The interactive theorem proving system based on a large language model according to claim 1, characterized in that: The strategy suggestion module includes: a proof target determination unit, a candidate strategy generation unit, a strategy verification unit and a strategy classification unit; The proof target determination unit is used to: obtain the target to be proved, generate a preliminary proof target according to the target to be proved, and input the preliminary proof target into the LLM reasoning engine; The candidate strategy generating unit is used to: call the LLM reasoning engine to generate a next reasoning strategy as the candidate strategy according to the preliminary proof goal; The policy verification unit is used to: apply the candidate policy to the target to be proved, and determine the policy attribute of the candidate policy; The policy classification unit is used to classify the candidate policies into successful policies, intermediate policies and failed policies according to the policy attributes.

3. The interactive theorem proving system based on a large language model according to claim 1, characterized in that: The proof search module includes: a target acquisition unit, a rule base calling unit, a reasoning strategy generation unit, a dynamic proof unit and a result analysis unit; The target acquisition unit is used to: acquire the complex theorem; The rule base calling unit is used to: carry the aesop rule base; The strategy introduction unit is used to: introduce a number of candidate strategies; The dynamic proof unit is used to: call the aesop rule library to perform iterative search of the complex theorem proof path based on the best first search algorithm and the tree structure expansion analysis algorithm until a complete proof path is found or a preset condition is met; in the iterative search process, dynamically generate a reasoning strategy that supports extended analysis by combining several candidate strategies and context information of the complex theorem; The result analysis unit is used to display the result of the complete proof path in the system interface.

4. The interactive theorem proving system based on a large language model according to claim 1, characterized in that: The premise selection module includes: a target vectorization unit, a premise matching unit, a candidate premise processing unit, a most relevant premise screening unit and an interaction support unit; The target vectorization unit is used to: call the LLM reasoning engine to parse the semantics of the target to be proved, and represent the target to be proved in the form of an embedded vector; The premise matching unit is used to: carry the premise embedding matrix and use an efficient mathematical method to calculate the similarity score between the target to be proved represented by the embedding vector and each premise in the premise embedding matrix; The candidate premise processing unit is used to: determine the loading status of the candidate premise after similarity score calculation in the current proof environment, determine the loaded premise and the unloaded premise according to the loading status, provide information description for the loaded premise, and provide supplementary description for the unloaded premise; The most relevant premise screening unit is used to: set a K value, extract the most relevant premise according to the similarity score and the K value, and display the most relevant premise to the user; The interaction support unit is used to: display the premise sequence of the most relevant premise through an interface.

5. The interactive theorem proving system based on a large language model according to claim 2, characterized in that: The candidate strategies include: existing proof steps, or new strategies generated by the LLM reasoning engine through reasoning.

6. The interactive theorem proving system based on a large language model according to claim 2, characterized in that: The strategy classification unit is further configured to: in response to the strategy attribute being that the strategy successfully solves the target to be proved, determine that the strategy is the successful strategy; The policy classification unit is further configured to: in response to the policy attribute being that the policy can change the target to be proved, determine that the policy is the intermediate policy; The policy classification unit is further used for: in response to the policy attribute being that the policy causes the certification to fail and an error to occur, determining that the policy is the invalid policy.

7. The interactive theorem proving system based on a large language model according to claim 3, characterized in that: The aesop rule base is configured with a preliminary rule set; The preliminary rule set includes: apply reasoning rules, simp reasoning rules and cases reasoning rules.

8. The interactive theorem proving system based on a large language model according to claim 3, characterized in that: In each search step of the iterative search, the aesop rule base and several candidate strategies of the LLM reasoning engine act together on the expansion of the target to be proved; The tree structure is a search tree, each node of the search tree supports expansion using several strategies, and each node of the search tree represents each sub-goal of the goal to be proved.

9. The interactive theorem proving system based on a large language model according to claim 3, characterized in that: The result analysis unit is further used to: display in the system interface a proof step list including multiple proof steps, a remaining target set including unresolved proof targets, and multi-path information including multiple possible proof paths.

10. An interactive theorem proving method based on a large language model, characterized in that: The following steps are involved: Strategy suggestion step: obtaining the target to be proved input by the user, and analyzing several candidate strategies for the target to be proved based on LLM; Proof search step: when the target to be proved is a complex theorem that requires a combination of multi-step strategies to prove, deep reasoning is performed based on the rule base and several candidate strategies to determine a complete proof path; Premise selection step: when the proof process of the target to be proved needs to refer to the premise, the premise sequence corresponding to the target to be proved is determined based on the similarity calculation operation, and the theorem is proved based on the premise selected by the user in the premise sequence.

Citation Information

Patent Citations

  • Formal theorem proof generation method and system based on large model

    CN118709756A

Cited By

  • Formalized automatic theorem proving method based on large language model

    CN121436106A