Proposition automatic formalization conversion method and device, equipment and storage medium

By using the dependency search model and a pre-constructed mathematical library, the target formal objects of non-formal statements are determined and inputted into the automatic formal model, the illusion problem during non-formal statement conversion in the prior art is solved, and the accuracy and reliability of automatic formalization are improved.

CN119990066APending Publication Date: 2025-05-13SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062444.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing automatic formalization methods often ignore the definitions and theorems in the formal mathematical library when converting non-formalized statements, resulting in hallucination of generated formal statements, especially in the case of out-of-distribution, resulting in failure of formal verification.

Method used

By obtaining the non-formal statement to be converted, the corresponding embedding vector is determined using the pre-trained dependency search model, the candidate embedding vector is determined as the target formal object based on the pre-constructed mathematical library and embedding vector, and input it into the automatic formal model to generate formal statements.

Benefits of technology

Reduces hallucination phenomena and improves the accuracy and reliability of automatic formalization, especially in out-of-distribution situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990066A_ABST
    Figure CN119990066A_ABST
Patent Text Reader

Abstract

The invention discloses a proposition automatic formalization conversion method and device, equipment and a storage medium. The method comprises the following steps: acquiring a first non-formalized statement to be converted; determining a first embedding vector corresponding to the first non-formalized statement based on a dependency retrieval model; determining a candidate embedding vector according to a math library and the first embedding vector, and taking the candidate embedding vector as a target formalized object on which the first non-formalized statement depends; and inputting the target formalized object and the first non-formalized statement into the automatic formalized model to obtain a corresponding formalized statement. According to the embodiment of the invention, the first embedding vector corresponding to the first non-formalized statement is determined through the dependency retrieval model, so that the target formalized object on which the first non-formalized statement depends is determined according to the math library and the first embedding vector, and the formalized statement corresponding to the first non-formalized statement is obtained through the automatic formalized model; the illusion phenomenon can be reduced, and the accuracy and reliability of automatic formalization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for automatic formal transformation of propositions. Background Art

[0002] Auto formalization aims to convert informal statements (mathematical statements expressed in natural language) into formal statements (statements expressed in formal language, which can be further formally verified by theorem provers). In the process of this conversion, existing methods directly convert given informal statements into formal statements, ignoring the prior knowledge about definitions and theorems in the formal mathematical library. Therefore, the generated formal statements often have hallucinations, generating identifiers and grammatical structures that do not exist in the mathematical library. This hallucination is particularly serious in the case of out-of-distribution (OOD), resulting in the failure of formal verification. Summary of the invention

[0003] In view of this, the present invention provides a proposition automatic formalization conversion method, device, equipment and storage medium, which can reduce the hallucination phenomenon and improve the accuracy and reliability of automatic formalization.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a method for automatic formal transformation of propositions, the method comprising:

[0005] Obtaining a first informal statement to be converted;

[0006] Determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model;

[0007] Determine a candidate embedding vector according to the pre-built mathematical library and the first embedding vector, and use the candidate embedding vector as the target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors with the highest cosine similarity to the first embedding vector; wherein N is a preset value;

[0008] The target formalized object and the first informal sentence are input into a pre-trained automatic formalization model to obtain a formalized sentence corresponding to the first informal sentence.

[0009] According to another aspect of the present invention, an embodiment of the present invention further provides a proposition automatic formal conversion device, the device comprising:

[0010] An acquisition module, used for acquiring a first informal statement to be converted;

[0011] An embedding module, configured to determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model;

[0012] A screening module, configured to determine a candidate embedding vector according to a pre-built mathematical library and the first embedding vector, and use the candidate embedding vector as a target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors with the highest cosine similarity to the first embedding vector; wherein N is a preset value;

[0013] The conversion module is used to input the target formalized object and the first informal sentence into a pre-trained automatic formalization model to obtain a formalized sentence corresponding to the first informal sentence.

[0014] According to another aspect of the present invention, an embodiment of the present invention further provides an electronic device, the electronic device comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the proposition automatic formal conversion method described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the proposition automatic formal conversion method described in any embodiment of the present invention when executed.

[0019] According to another aspect of the present invention, an embodiment of the present invention further provides a computer program product, characterized in that the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the proposition automatic formal conversion method described in any embodiment of the present invention.

[0020] The technical effect of the present invention is that a first embedding vector corresponding to a first informal statement is determined by relying on a retrieval model, thereby determining a target formalized object on which the first informal statement depends based on a mathematical library and the first embedding vector, and then inputting the target formalized object and the first informal statement into an automatic formalization model to obtain a formalized statement corresponding to the first informal statement, thereby reducing hallucinations and improving the accuracy and reliability of automatic formalization.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A flowchart of a method for automatic formal transformation of propositions provided by an embodiment of the present invention;

[0024] Figure 2 A flowchart of another method for automatic formal transformation of propositions provided by an embodiment of the present invention;

[0025] Figure 3 A schematic diagram of the overall process of a method for automatic formal transformation of propositions provided by an embodiment of the present invention;

[0026] Figure 4 A structural block diagram of a proposition automatic formal conversion device provided by an embodiment of the present invention;

[0027] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] In one embodiment, Figure 1 A flowchart of a method for automatic formal conversion of propositions provided in one embodiment of the present invention. This embodiment can be applied to the situation when informal statements are automatically converted into formal statements. The method can be executed by an automatic formal conversion device for propositions. The automatic formal conversion device for propositions can be implemented in the form of hardware and / or software. The automatic formal conversion device for propositions can be configured in an electronic device.

[0031] like Figure 1 As shown, the proposition automatic formal conversion method in this embodiment specifically includes the following steps:

[0032] S110: Obtain a first informal statement to be converted.

[0033] The first informal statement refers to an informal statement that needs to be converted into a formal statement. It can be understood that a mathematical statement expressed in a natural statement is a statement used in a natural language that does not follow strict grammatical rules or logical structures.

[0034] In this embodiment, an informal statement that needs to be formalized is given so as to perform subsequent processing to obtain a corresponding formal statement.

[0035] S120: Determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model.

[0036] The pre-trained dependency retrieval model can also be called an encoder, which can convert the input data into a fixed-dimensional dense representation (i.e., an embedding vector), which can usually capture the semantic information of the input data. The first embedding vector can be understood as the first informal statement in vector form, which is a representation method that maps the first informal statement to a fixed-dimensional vector space, so as to facilitate the subsequent calculation of the similarity between vectors.

[0037] In this embodiment, the first informal sentence is input into a pre-trained dependency retrieval model for embedding, and a first embedding vector corresponding to the first informal sentence can be obtained. It should be noted that the embedding method may include but is not limited to dense embedding, sparse embedding and other embedding methods. Specifically, the first informal sentence is input into the dependency retrieval model, the first informal sentence is segmented and encoded using a tokenizer, the encoded input is passed to the model, an embedding vector is generated, and the generated embedding vector is output.

[0038] In one embodiment, the dependency retrieval model is a trained optimal dependency retrieval model, and the training process of the dependency retrieval model includes: obtaining a first training set; wherein the first training set includes second informal statements corresponding to each formal object in the mathematical library, formal statements of dependent objects on which each formal object depends, or formal statements of dependent objects and second informal statements;

[0039] The first training set is input into the dependency retrieval model, and the dependency retrieval model is trained by contrastive learning until the loss function corresponding to the dependency retrieval model is minimized to obtain a trained dependency retrieval model; wherein the loss function is characterized as the cosine similarity between formalized objects with dependency relationships becoming closer and closer, and the cosine similarity between formalized objects without dependency relationships becoming farther and farther.

[0040] The second informal statement may be understood as the informal statements corresponding to each formal object included in the database.

[0041] In this embodiment, contrastive learning can be understood as a supervised learning method that aims to capture the representation of data by learning the similarities and differences between data. The core idea of ​​contrastive learning is to learn meaningful representations by maximizing the similarity between positive sample pairs (similar samples) and minimizing the similarity between negative sample pairs (dissimilar samples).

[0042] In this embodiment, when training the dependency retrieval model, it is necessary to first obtain a first training set, which includes the second informal statements corresponding to each formalized object in the mathematical library, the formal statements of the dependent objects on which each formalized object depends, or the formal statements of the dependent objects and the second informal statements. Then, the first training set is input into the dependency retrieval model to map the first training set to a low-dimensional dense vector space (embedding space). In the embedding space, the cosine similarity between the formalized objects with dependency relationships and the cosine similarity between the formalized objects with dependency relationships without dependency relationships are calculated until the cosine similarity between the formalized objects with dependency relationships becomes more and more similar (that is, the distance becomes closer and closer), and the cosine similarity between the formalized objects with dependency relationships without dependency relationships becomes more and more dissimilar (that is, the distance becomes farther and farther), that is, the loss function is minimized and a trained dependency retrieval model is obtained.

[0043] S130, determining a candidate embedding vector based on a pre-built mathematical library and a first embedding vector, and using the candidate embedding vector as a target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors having the highest cosine similarity with the first embedding vector; wherein N is a preset value.

[0044] The mathematical library constructed first includes at least two formalized objects, each of which corresponds to corresponding content information; the content information at least includes: object name, definition location, type, declaration, code, comments and dependency. Exemplarily, the formal code of the formalized object Group in the mathematical library is class Group (G: Type u) extends DivInvMonoid...; the corresponding informal statement corresponding to the formalization is expressed as: Class Group represents...; Group depends on DivInvMonoid.zpow / Monoid.npow / and other formalized objects.

[0045] In this embodiment, candidate embedding vectors can be selected from a pre-constructed mathematical library. The candidate embedding vectors are deformations of each formalized object in the mathematical library. It can be understood that each formalized object in the mathematical library first passes through an informal model to obtain a corresponding informal statement. The informal statement is the second informal statement in the above embodiment. Then, the second informal statement is embedded in a pre-trained dependency retrieval model to obtain an embedding vector corresponding to the second informal statement. Furthermore, the N embedding vectors with the highest cosine similarity to the first embedding vector can be found in the embedding vector corresponding to the second informal statement through similarity as candidate embedding vectors.

[0046] In one embodiment, the mathematical library includes at least two formalized objects, each of which corresponds to corresponding content information; the content information at least includes: object name, definition location, type, declaration, code, comment and dependency relationship; the at least two formalized objects in the mathematical library form a data set; the data set includes: the formalized object, the content information corresponding to each formalized object, the second informal statement and the dependent object with dependency relationship;

[0047] Accordingly, the formation of the data set includes:

[0048] Determine the content information corresponding to each of the formalized objects contained in the mathematical library;

[0049] Perform topological sorting based on the dependency relationships in the content information to construct a dependency graph corresponding to each formalized object;

[0050] For each formalized object in the dependency graph, according to the dependency relationship between the formalized objects and the context prompt function of the preset informal large model, each formalized object in the dependency graph is gradually converted into a second informal statement starting from the formalized object of the bottom node; wherein the preset informal large model is a large model trained by a small number of samples or zero-sample training;

[0051] Each formalized object, the content information corresponding to each formalized object, the second informal statement, and the dependent objects with dependency relationships form a data set.

[0052] Among them, the dependency graph can also be called a directed acyclic graph. The preset informal large model is a large model trained by a small number of samples, or zero-sample training, which is used to convert the formal objects contained in the mathematical library into informal statements, and the informal statement is the second formal statement in the above embodiment. In this embodiment, zero-sample learning: the model completes the task through contextual prompts without having seen a specific task. Few-shot learning: the model completes a new task with the prompts of a small number of examples (few-shot demonstrations). For example, given a task description and several examples given to the preset informal large model, the model can infer the task rules and generate the correct answer.

[0053] In this embodiment, the content information corresponding to each of the formalized objects included is obtained from the mathematical library, and topological sorting is performed according to the dependency relationship in the content information to construct a dependency graph corresponding to each formalized object. For each formalized object in the dependency graph, according to the dependency relationship between each formalized object and the context prompt function of the preset informal large model, each formalized object in the dependency graph is gradually converted into a second informal statement starting from the formalized object of the bottom node, so that each formalized object, the content information corresponding to each formalized object, the second informal statement, and the dependent objects with dependency relationships form a data set. It can be understood that the dependency relationship of the formalized objects is formed into a directed acyclic graph, and it is topologically sorted. Then, along the topological order on the dependency graph, for each formalized object, the context learning ability of the large model is used, given the declaration, code, comments of the object, and the informal statements of the objects it depends on, its informal statements are generated, thereby forming a corresponding data set.

[0054] In one embodiment, according to the dependency relationship between the formalized objects and the context prompt function of the preset informal large model, each formalized object in the dependency graph is gradually converted into a second informal statement starting from the formalized object of the bottom node, including:

[0055] Determine the formalized object of the bottom-level node and its corresponding content information;

[0056] The formalized object of the bottom node and its corresponding content information are used as the current bottom node, and the current bottom node is input into the preset informal large model to obtain the second informal statement corresponding to the formalized object of the current bottom node;

[0057] Determine other formalized objects that have dependency relationships with the formalized object of the current bottom-level node, take the other formalized objects that have dependency relationships as dependent objects, and take the content information corresponding to the dependent objects and the second informal statement as the next bottom-level node, and take the next bottom-level node as the current bottom-level node, and return to the step of inputting the current bottom-level node into the preset informal large model until each formalized object in the dependency graph is converted into the second informal statement.

[0058] The bottom-most node can be understood as the most basic mathematical definition, which does not depend on any object. The next bottom-most node can include the first next bottom-most node, the second next bottom-most node, ..., corresponding to the current bottom-most node, until the last next bottom-most node at the end of the loop.

[0059] In this embodiment, the formalization without dependency is first performed, and then the formalization is performed upwards, layer by layer, and the formalization is performed in order of dependency. It can be understood that, according to the formed topological informal dependency graph, each formalized object is gradually informalized. Exemplarily, all formalized objects are topologically sorted first, and then, in the sorting graph, the formalized object (One, Mul) is not dependent on any formalized object, so the two formalized objects (One, Mul) are first informalized. At this time, the content information corresponding to the two formalized objects (One, Mul) is given to the preset informal large model to generate the corresponding informal statements, and then, the content information corresponding to this formalized object and the informal statements of (One, Mul) on which they depend are given to the preset informal large model to generate the informal statements of (Semigroup), and then layer by layer, until each formalized object in the dependency graph is converted into a second informal statement.

[0060] S140: Input the target formalization object and the first informal sentence into a pre-trained automatic formalization model to obtain a formal sentence corresponding to the first informal sentence.

[0061] The target formalized object refers to the formalized object corresponding to the N embedding vectors with the highest cosine similarity to the first embedding vector. The pre-trained automatic formalization model is a model for converting informal sentences into formal sentences.

[0062] In this embodiment, the target formalized object and the first informal sentence are input into a pre-trained automatic formalization model to obtain a formalized statement corresponding to the first informal sentence. Specifically, the target formalized object and the first informal sentence are input into the pre-trained automatic formalization model. The automatic formalization model can understand the semantics of the informal sentence and the candidate embedding vector, identify the key information and logical relationships therein, map the semantic information into the formal language, and generate a sentence that conforms to the rules of the formal language.

[0063] In one embodiment, the training of the automatic formal model includes:

[0064] Obtain a second training set; wherein the second training set includes: second informal statements corresponding to each formal object in the mathematical library, and dependency retrieval results obtained by the trained dependency retrieval model for the second informal statement; the dependency retrieval results include: at least two formal objects with the highest N embedding vectors of cosine similarity with the second informal statement;

[0065] The second training set is input into the automatic formalization model, and the automatic formalization model is trained by adopting the autoregressive generation method. The parameters are optimized by minimizing the cross entropy loss function, and the conditional probability distribution of the input sequence is generated to obtain the trained automatic formalization model.

[0066] Among them, the autoregressive generation method can be understood as gradually generating the next element based on the generated partial sequence until the complete sequence is generated. It is usually based on a probability model and is achieved by maximizing the probability of generating the sequence.

[0067] In this embodiment, the second informal statements corresponding to each formalized object in the mathematical library and the dependency retrieval results obtained by the trained dependency retrieval model for the second informal statements are input into the automatic formalization model, an autoregressive generation method is adopted, and the parameters are optimized by minimizing the cross entropy loss function, and the conditional probability distribution of the input sequence is learned and generated to obtain a trained automatic formalization model.

[0068] The technical solution of the embodiment of the present invention determines the first embedding vector corresponding to the first informal statement by relying on a retrieval model, thereby determining the target formalized object on which the first informal statement depends according to a mathematical library and the first embedding vector, and then inputting the target formalized object and the first informal statement into the automatic formalization model to obtain the formalized statement corresponding to the first informal statement, thereby reducing the hallucination phenomenon and improving the accuracy and reliability of automatic formalization.

[0069] In one embodiment, Figure 2 A flowchart of another method for automatic formal conversion of propositions provided in one embodiment of the present invention. Based on the above embodiments, this embodiment further refines the method for determining a first embedding vector corresponding to a first informal statement based on a pre-trained dependency retrieval model; and determining a candidate embedding vector based on a pre-built mathematical library and the first embedding vector, and taking the candidate embedding vector as a target formal object on which the first informal statement depends.

[0070] like Figure 2 As shown, the proposition automatic formal conversion method in this embodiment may specifically include the following steps:

[0071] S210: Obtain a first informal statement to be converted.

[0072] S220: Input the first informal sentence into the trained dependency retrieval model for embedding to obtain a corresponding first embedding vector.

[0073] In this embodiment, the first informal sentence is input into the trained dependency retrieval model to be mapped to a first embedding vector of fixed dimension, wherein the embedding includes one of the following: dense embedding, sparse embedding, multi-vector embedding, and hybrid embedding.

[0074] S230: Obtain a second informal statement corresponding to the formal object in the mathematics library.

[0075] S240: Input the second informal sentence into a pre-trained dependency retrieval model for dense embedding to obtain a corresponding second embedding vector.

[0076] The second embedding vector refers to the dense vector representation of the second informal sentence after being processed by the dependency retrieval model. It can be understood that the vector represented by a fixed dimension is convenient for subsequent semantic similarity calculation.

[0077] In this embodiment, the second informal sentence is input into the pre-trained dependency retrieval model for dense embedding to obtain the corresponding second embedding vector. Wherein, embedding includes one of the following: dense embedding, sparse embedding. In this embodiment, dense embedding can be understood as mapping a natural language sentence into a dense vector of a fixed dimension (usually a high-dimensional vector), and each dimension in the vector represents a certain semantic feature of the sentence. Specifically, the second informal sentence is input into the dependency retrieval model, and the dependency retrieval model encodes the input sentence to extract semantic information, which may include decomposing the sentence into words or subwords, using the model's encoder (such as Transformer) to encode each word or subword, and performing a pooling operation (such as averaging or taking the maximum value) on the encoded vector to obtain the corresponding dense embedding vector.

[0078] S250: Search for N candidate embedding vectors with the highest cosine similarity to the first embedding vector from the second embedding vector, and use the candidate embedding vectors as target formalized objects that the first informal statement depends on.

[0079] There is a one-to-one correspondence between the formalized object in the mathematical library and the second informal statement, a one-to-one correspondence between the second informal statement and the second embedding vector, and a one-to-one correspondence between the formalized object and the second embedding vector.

[0080] In this embodiment, the N embedding vectors with the highest cosine similarity to the first embedding vector are searched from the second embedding vector, and the embedding vector is used as the candidate embedding vector on which the first informal statement depends. It should be noted that the candidate embedding vector corresponds to the corresponding formalized object, and the formalized object found is the N embedding vector with the highest cosine similarity to the first embedding vector. In this embodiment, the formalized objects in the mathematical library correspond to the corresponding informal statements one by one, and the informal statements obtained by each formalized object will be densely embedded to obtain the corresponding embedding vector, thereby obtaining multiple third embedding vectors, and then the N candidate embedding vectors with the highest cosine similarity to the first embedding vector are found from the third embedding vector. It can be seen that the candidate embedding vector has a corresponding formalized object, so the N candidate embedding vectors with the highest cosine similarity to the first embedding vector are the formalized objects with the highest cosine similarity.

[0081] In this embodiment, given an informal statement, the method first uses a dependency retrieval model to embed the statement, and then, based on the embedding, retrieves several formal objects with the highest cosine similarity from the embeddings of all mathematical objects in the pre-calculated library. Finally, the retrieval results and the informal statement are input into the automatic formal model for conditional generation to obtain a formal statement corresponding to the first informal statement.

[0082] S260: Input the target formalization object and the first informal sentence into a pre-trained automatic formalization model to obtain a formal sentence corresponding to the first informal sentence.

[0083] The technical solution of the embodiment of the present invention is to obtain a corresponding first embedding vector by inputting a first informal sentence into a trained dependency retrieval model for embedding, obtain a second informal sentence corresponding to a formalized object in a mathematical library, input the second informal sentence into a pre-trained dependency retrieval model for dense embedding to obtain a corresponding second embedding vector, search for N candidate embedding vectors with the highest cosine similarity to the first embedding vector from the second embedding vector, and use the candidate embedding vectors as the target formalized object on which the first informal sentence depends, input the target formalized object and the first informal sentence into a pre-trained automatic formalization model, and obtain a formalized sentence corresponding to the first informal sentence, thereby further reducing the hallucination phenomenon and improving the accuracy and reliability of automatic formalization.

[0084] For example, to better understand the automatic formal transformation method of propositions, Figure 3 The overall flow chart of a method for automatic formal transformation of propositions provided by an embodiment of the present invention is as follows: Figure 3 As shown, given an informal sentence, this method first uses the dependency retrieval model to embed the sentence, such as Figure 3 Then, based on this embedding, several objects with the highest cosine similarity are retrieved from the embeddings of all mathematical objects in the pre-computed library, such as Figure 3 Finally, the search results and informal statements are input into the automatic formal model for condition generation, such as Figure 3 ④ generation in . It can be understood as that, given an informal sentence, the encoder is first used to convert the informal sentence into a dense embedding to obtain the corresponding first embedding vector; all formal objects in the mathematical library are also densely embedded to obtain the corresponding second embedding vector, and then the formal objects with the highest cosine similarity with the first embedding vector are found from the second embedding vector, which are the formal objects that the given informal sentence is most likely to depend on, for example, Figure 3 The Group and Fintype in the above are the formalized objects with the highest cosine similarity that are screened out. Then, the search result and the given informal sentence are input into the trained automatic formalization model together to obtain the corresponding formal sentence.

[0085] In one embodiment, Figure 4 This is a structural block diagram of a proposition automatic formal conversion device provided in one embodiment of the present invention. The device is suitable for automatically converting informal statements into formal statements. The device can be implemented by hardware / software. It can be configured in an electronic device to implement a proposition automatic formal conversion method in an embodiment of the present invention.

[0086] like Figure 4 As shown, the device includes: an acquisition module 410, an embedding module 420, a screening module 430 and a conversion module 440;

[0087] The acquisition module 410 is used to acquire a first informal statement to be converted;

[0088] An embedding module 420, configured to determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model;

[0089] A screening module 430 is used to determine a candidate embedding vector according to a pre-built mathematical library and the first embedding vector, and use the candidate embedding vector as a target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors with the highest cosine similarity to the first embedding vector; wherein N is a preset value;

[0090] The conversion module 440 is used to input the target formalized object and the first informal sentence into a pre-trained automatic formalization model to obtain a formalized sentence corresponding to the first informal sentence.

[0091] In an embodiment of the present invention, an embedding module determines a first embedding vector corresponding to a first informal statement by relying on a retrieval model, thereby a screening module determines a candidate embedding vector according to a mathematical library and the first embedding vector, thereby a conversion module inputs the candidate embedding vector and the first informal statement into an automatic formalization model to obtain a formalized statement corresponding to the first informal statement, which can reduce hallucination phenomena and improve the accuracy and reliability of automatic formalization.

[0092] In one embodiment, the mathematical library includes at least two formalized objects, each of which corresponds to corresponding content information; the content information at least includes: object name, definition location, type, declaration, code, comments and dependency relationship;

[0093] At least two formalized objects in the mathematical library form a data set; the data set includes: the formalized objects, content information corresponding to each of the formalized objects, a second informal statement, and a dependent object having a dependent relationship;

[0094] Accordingly, the formation of the data set includes:

[0095] Determine the content information corresponding to each of the formalized objects contained in the mathematical library;

[0096] Performing topological sorting according to the dependency relationships in the content information to construct a dependency graph corresponding to each of the formalized objects;

[0097] For each formalized object in the dependency graph, according to the dependency relationship between the formalized objects and the context prompt function of the preset informal large model, each formalized object in the dependency graph is gradually converted into a second informal statement starting from the formalized object of the bottom node; wherein the preset informal large model is a large model trained with a small number of samples or zero sample training;

[0098] The formalized objects, the content information corresponding to the formalized objects, the second informal statements, and the dependent objects having the dependency relationship form a data set.

[0099] In one embodiment, the step of converting each formal object in the dependency graph into a second informal statement starting from the formal object of the bottom node according to the dependency relationship between the formal objects and the context prompt function of the preset informal large model includes:

[0100] Determine the formalized object of the bottom-level node and its corresponding content information;

[0101] Taking the formalized object of the bottom-level node and its corresponding content information as the current bottom-level node, and inputting the current bottom-level node into a preset informal large model, to obtain a second informal statement corresponding to the formalized object of the current bottom-level node;

[0102] Determine other formalized objects that have dependency relationships with the formalized object of the current bottom-level node, take the other formalized objects that have dependency relationships as dependent objects, and take the content information corresponding to the dependent objects and the second informal statement as the next bottom-level node, and take the next bottom-level node as the current bottom-level node, and return to the step of inputting the current bottom-level node into the preset informal large model until each formalized object in the dependency graph is converted into a second informal statement.

[0103] In one embodiment, the training process of the dependency retrieval model includes: obtaining a first training set; wherein the first training set includes the second informal statements corresponding to each of the formalized objects in the mathematical library, the formal statements of the dependent objects on which each of the formalized objects depends, or the formal statements of the dependent objects and the second informal statements;

[0104] The first training set is input into the dependency retrieval model, and the dependency retrieval model is trained by contrastive learning until the loss function corresponding to the dependency retrieval model reaches a minimum, thereby obtaining a trained dependency retrieval model; wherein the loss function is characterized by the cosine similarity between formalized objects with dependency relationships becoming higher and higher, and the cosine similarity between formalized objects without dependency relationships becoming lower and lower.

[0105] In one embodiment, the training of the automatic formal model includes:

[0106] Obtain a second training set; wherein the second training set includes: second informal statements corresponding to each of the formalized objects in the mathematical library, and dependency retrieval results obtained by the trained dependency retrieval model for the second informal statements; the dependency retrieval results include: at least two formalized objects with the highest N embedding vectors of the cosine similarity with the second informal statements; wherein N is a preset value;

[0107] The second training set is input into the automatic formalization model, and the automatic formalization model is trained by adopting the autoregressive generation method. The parameters are optimized by minimizing the cross entropy loss function, and the conditional probability distribution of the input sequence is generated to obtain the trained automatic formalization model.

[0108] In one embodiment, the embedding module 420 includes:

[0109] An embedding unit is used to input the first informal sentence into the trained dependency retrieval model for embedding to obtain a corresponding first embedding vector; wherein the embedding includes one of the following: dense embedding, sparse embedding, multi-vector embedding, and hybrid embedding.

[0110] In one embodiment, the screening module 430 includes:

[0111] An acquisition unit, configured to acquire a second informal statement corresponding to a formal object in the mathematical library;

[0112] An embedding unit, configured to input the second informal sentence into the pre-trained dependency retrieval model for embedding to obtain a corresponding second embedding vector;

[0113] A search unit is used to search for candidate embedding vectors of N embedding vectors with the highest cosine similarity to the first embedding vector from the second embedding vector, and use the candidate embedding vector as a target formalized object on which the first informal statement depends; wherein the formalized object in the mathematical library corresponds one-to-one to the second informal statement, the second informal statement corresponds one-to-one to the second embedding vector, and the formalized object corresponds one-to-one to the second embedding vector.

[0114] The proposition automatic formal conversion device provided in the embodiment of the present invention can execute the proposition automatic formal conversion method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0115] In one embodiment, Figure 5 A schematic diagram of the structure of an electronic device provided for an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0116] like Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0117] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0118] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the proposition automatic formalization conversion method.

[0119] In some embodiments, the proposition automatic formal conversion method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the proposition automatic formal conversion method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the proposition automatic formal conversion method in any other appropriate manner (for example, by means of firmware).

[0120] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable proposition automatic formal conversion device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0122] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0125] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0126] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0127] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for automatic formal transformation of propositions, characterized in that: The method comprises: Obtaining a first informal statement to be converted; Determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model; Determine a candidate embedding vector according to the pre-built mathematical library and the first embedding vector, and use the candidate embedding vector as the target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors with the highest cosine similarity to the first embedding vector; wherein N is a preset value; The target formalized object and the first informal sentence are input into a pre-trained automatic formalization model to obtain a formalized sentence corresponding to the first informal sentence.

2. The method according to claim 1, characterized in that The mathematical library includes at least two formalized objects, each of which corresponds to corresponding content information; the content information at least includes: object name, definition location, type, declaration, code, comments and dependency relationship; At least two formalized objects in the mathematical library form a data set; the data set includes: the formalized objects, second informal statements corresponding to each of the formalized objects, and dependent objects having a dependency relationship; Accordingly, the formation of the data set includes: Determine the content information corresponding to each of the formalized objects contained in the mathematical library; Performing topological sorting according to the dependency relationships in the content information to construct a dependency graph corresponding to each of the formalized objects; For each formalized object in the dependency graph, according to the dependency relationship between the formalized objects and the context prompt function of the preset informal large model, each formalized object in the dependency graph is gradually converted into a second informal statement starting from the formalized object of the bottom node; wherein the preset informal large model is a large model trained with a small number of samples or zero sample training; The formalized objects, the content information corresponding to the formalized objects, the second informal statements, and the dependent objects having the dependency relationship form a data set.

3. The method according to claim 2, characterized in that The step of converting each formal object in the dependency graph into a second informal statement starting from the formal object of the bottom node according to the dependency relationship between the formal objects and the context prompt function of the preset informal large model includes: Determine the formalized object of the bottom-level node and its corresponding content information; Taking the formalized object of the bottom-level node and its corresponding content information as the current bottom-level node, and inputting the current bottom-level node into a preset informal large model, to obtain a second informal statement corresponding to the formalized object of the current bottom-level node; Determine other formalized objects that have dependency relationships with the formalized object of the current bottom-level node, take the other formalized objects that have dependency relationships as dependent objects, and take the content information corresponding to the dependent objects and the second informal statement as the next bottom-level node, and take the next bottom-level node as the current bottom-level node, and return to the step of inputting the current bottom-level node into the preset informal large model until each formalized object in the dependency graph is converted into a second informal statement.

4. The method according to claim 1, characterized in that: The training process of the dependency retrieval model includes: obtaining a first training set; wherein the first training set includes second informal statements corresponding to each of the formalized objects in the mathematical library, and formal statements of dependent objects on which each of the formalized objects depends; The first training set is input into the dependency retrieval model, and the dependency retrieval model is trained by contrastive learning until the loss function corresponding to the dependency retrieval model reaches a minimum, thereby obtaining a trained dependency retrieval model; wherein the loss function is characterized by the cosine similarity between formalized objects with dependency relationships becoming higher and higher, and the cosine similarity between formalized objects without dependency relationships becoming lower and lower.

5. The method according to claim 1, characterized in that The training of the automatic formal model includes: Obtain a second training set; wherein the second training set includes: second informal statements corresponding to each of the formalized objects in the mathematical library, and dependency retrieval results obtained by the trained dependency retrieval model for the second informal statements; the dependency retrieval results include: at least two formalized objects with the highest N embedding vectors of the cosine similarity with the second informal statements; wherein N is a preset value; The second training set is input into the automatic formalization model, and the automatic formalization model is trained by adopting the autoregressive generation method. The parameters are optimized by minimizing the cross entropy loss function, and the conditional probability distribution of the input sequence is generated to obtain the trained automatic formalization model.

6. The method according to claim 1, characterized in that The determining, based on the pre-trained dependency retrieval model, a first embedding vector corresponding to the first informal sentence includes: The first informal sentence is input into the trained dependency retrieval model for embedding to obtain a corresponding first embedding vector; wherein the embedding includes one of the following: dense embedding, sparse embedding, multi-vector embedding, and hybrid embedding.

7. The method according to claim 1, characterized in that The determining of a candidate embedding vector according to a pre-built mathematical library and the first embedding vector, and using the candidate embedding vector as a target formalized object on which the first informal statement depends, includes: Obtaining a second informal statement corresponding to the formal object in the mathematical library; Inputting the second informal sentence into the pre-trained dependency retrieval model for dense embedding to obtain a corresponding second embedding vector; Searching for N candidate embedding vectors with the highest cosine similarity to the first embedding vector from the second embedding vector, and using the candidate embedding vectors as target formalized objects that the first informal statement depends on; wherein N is a preset value; There is a one-to-one correspondence between the formalized objects in the mathematical library and the second informal statements, a one-to-one correspondence between the second informal statements and the second embedding vectors, and a one-to-one correspondence between the formalized objects and the second embedding vectors.

8. A proposition automatic formal conversion device, characterized in that: The device comprises: An acquisition module, used for acquiring a first informal statement to be converted; An embedding module, configured to determine a first embedding vector corresponding to the first informal sentence based on a pre-trained dependency retrieval model; A screening module, configured to determine a candidate embedding vector according to a pre-built mathematical library and the first embedding vector, and use the candidate embedding vector as a target formalized object on which the first informal statement depends; wherein the candidate embedding vector is the N embedding vectors with the highest cosine similarity to the first embedding vector; wherein N is a preset value; The conversion module is used to input the target formalized object and the first informal sentence into a pre-trained automatic formalization model to obtain a formalized sentence corresponding to the first informal sentence.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the proposition automatic formal transformation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the proposition automatic formal conversion method described in any one of claims 1 to 7 when executed.