A non-autoregressive math problem solver based on a multi-way tree structure
By introducing a non-autoregressive target decomposition module into the math problem solver, combining the multi-head self-attention and mutual attention mechanism to deal with disordered multi-branch decomposition, the problem that multi-forktree structures in the existing technology cannot fully utilize information, and more efficient math problem solving performance and generalization ability are achieved.
Patent Information
- Application Number
- CN202310433743.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-04-21
AI Technical Summary
The existing mathematical problem solver based on multi-forktree structure lacks generalization when dealing with disordered multi-branch decomposition, cannot generate paths that do not appear in the training set, and cannot fully utilize the structural information of multi-forktrees, resulting in insufficient performance.
The encoder-decoder structure is adopted, including the question encoding module, the target-driven multi-tree generation module and the non-autoregressive target decomposition module. The disordered multi-branch decomposition is handled through the multi-head self-attention mechanism and the multi-head mutual attention mechanism, and the pointer network is used to select the most relevant candidate characters to generate mathematical problem expressions.
It improves the performance of the math problem solver, can better explore and capture the relationship between numbers, generate more accurate math problem expressions, and improves the generalization ability and solution efficiency of the model.
Smart Images

Figure CN116401624B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mathematical problem solvers, and more specifically, relates to a non-autoregressive mathematical problem solver based on a multi-way tree structure. Background Art
[0002] Automatically solving math word problems (MWP) is an important sub-problem of machine reasoning. For a mathematical problem solver, it is necessary to generate a solvable equation that conforms to the specification based on the given problem description and mathematical prior knowledge. More specifically, given a math problem consisting of a text description that contains numbers q1, q2, …, q n and certain mathematical definitions, the model is required to automatically give the expression for solving the problem and calculate the final answer of the problem through this expression. The task of solving math problems involves core issues in artificial intelligence research such as in-depth understanding of natural language text, machine intelligence reasoning ability, and interpretability. As an important test benchmark for machine intelligence, math problem solving has always attracted the attention of many researchers.
[0003] Since the task of solving math problems was proposed, it has gone through several stages, including early template-based methods, statistical-based methods, and later deep learning-based methods. The statistical and template-based methods rely heavily on manual annotation and statistics, and the defined models lack generalization. With the assistance of high-performance computer devices and the vast amount of data on the Internet, deep learning has made great progress, and methods based on neural networks have also been applied to the task of solving math problems. A large number of models have achieved amazing performance in the task of solving math problems. Since the tree decoder was proposed in 2019, it has been widely used in the task of solving math problems and has become the mainstream method.
[0004] The tree structure naturally has the advantage of representing mathematical expressions. The depth of the tree can correspond to the operation priority in the mathematical expression, so that operations with higher priority are placed at lower levels, and the root node of the tree is the operator with the lowest priority. The tree structure first has the structural information of the mathematical expression. The tree decoder can generate the target expression in a top-down manner and form an expression tree. The top-down method can be understood as continuously splitting the problem of the question into sub-problems and finally obtaining the answer to the question by continuously solving the sub-problems. This is a problem-solving method that conforms to human intuition and has good interpretability. The sub-expressions generated during the splitting process are actually the intermediate variables in the question, so the tree decoder can well represent the structural information of the math problem.
[0005] Most existing math problem solvers follow an encoder-decoder architecture. The solving process of this architecture is to translate a problem description based on natural language into a math description (solving expression) based on symbolic language. Under such an architecture, the improvement of math problem solvers can also be mainly summarized into two categories: the improvement of the encoder and the improvement of the decoder.
[0006] For the encoder, the key is to enhance the neural network model's understanding of the problem text description. In addition, math problems require a large amount of external math prior knowledge and common sense prior knowledge. How to incorporate these prior knowledge into the encoding to help the model better understand the problem is also an important aspect of improving the model encoder.
[0007] For the decoder, the key issue is to extract the mathematical relationships between variables from the text features and search for the solution expression from these relationships. Therefore, how to narrow the search space of the solution expression is also an effective idea for improving performance. In addition, the expression tree also has some of its own drawbacks. For example, the one-to-one relationship between the binary expression tree and the math expression will make it highly dependent on the form of the math expression. For example, a - b + c and c - b + a are expressions with the same semantics using the commutative law, but they will generate two different binary trees. Later, an expression tree with a multi-way tree structure emerged, which places multiple parts with the same precedence in the math expression on the same layer, effectively solving this semantic inconsistency. However, the solver based on the multi-way tree structure solves based on the encoding. Instead of directly solving on the multi-way tree, it is split into individual paths (a path of a multi-way tree refers to a connection from the root node to the leaf node), and all paths in the training set are used as the retrieval space to retrieve all paths in a problem and reconstruct the multi-way tree. This solver only learns at the path level, cannot fully utilize the structural information of the multi-way tree, and the math problem solver lacks generalization ability and cannot generate paths that do not appear in the training set. The performance and effectiveness of the automatic math problem solver need to be improved. Summary of the Invention
[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a non-autoregressive math problem solver based on a multi-way tree structure to explore and capture the relationships between numbers and improve the performance of the math problem solver.
[0009] To achieve the above invention purpose, the non-autoregressive math problem solver based on a multi-way tree structure of the present invention adopts an encoder-decoder structure, including:
[0010] A problem encoding module for encoding a math problem P = {w1, w2,..., w N}, where w n represents the nth word, n = 1, 2,..., N, into a distributed representation E containing its context informationp , E V , where E p is the problem representation of the entire math problem, and E V is the numerical representation of the math problem;
[0011] The goal-driven multi-way tree generation module is a top-down multi-way tree generator that adopts a goal-driven mechanism. It uses the problem representation E p as the root goal of the multi-way tree and recursively generates sub-goals in a top-down manner. The sub-goals are classified as operands or operators. When the sub-goal is an operand, the result of the current sub-goal is directly obtained. When the sub-goal is an operator, the result of the sub-goal cannot be obtained yet, and further decomposition is required until the sub-goal is an operand;
[0012] It is characterized in that it further includes
[0013] The non-autoregressive goal decomposition module is used to handle the unordered multi-branch decomposition work:
[0014] First, for the parent goal E g in the sub-goal decomposition process, first fuse the parent goal E g with I position encodings p i , i = 1, 2,..., I to obtain the fused goal E pos :
[0015] E pos = [E g + p1; E g + p2; …, E g + p I ;
[0016] Then, use the multi-head self-attention mechanism for processing, that is, input the goal E pos into the multi-head self-attention module for processing to obtain the output
[0017]
[0018] where is the trainable parameter matrix in the multi-head self-attention module;
[0019] Then, connect the output of the multi-head self-attention module and the candidate set Ec through a multi-head cross-attention module. Use the output as the Q matrix of the multi-head cross-attention module. After passing the candidate set E c through a feed-forward neural network and multiplying it by the trainable parameter , obtain the K matrix and V matrix of the multi-head cross-attention module, and thus obtain the output
[0020]
[0021] Among them, d k is the dimension of the encoding vector, and E c The candidate set is:
[0022]
[0023] That is, it is composed of numbers, operators, constants, and special characters. Among them, E V is the output of the question encoding module, and the remaining E op , E con , E N are trainable encodings, and N b is a special character used to represent the number of child nodes;
[0024] Output is the context representation, which is used to select the most relevant candidate E c from the candidate set through the pointer network. The probability ω ij of selecting the j-th candidate character at position i is:
[0025]
[0026] Among them, W p and W b are learning parameters, u is the column weight vector, and are respectively the vector representation of the i-th position and the vector representation of the j-th character in the candidate set;
[0027] In this way, the probability distribution Ptr i of all candidate characters at the i-th position is obtained:
[0028] Ptr i = softmax(ω i )
[0029] Among them, ω i = {ω i1 , ω i2 ,..., ω ij}, and J is the number of candidate characters in the candidate set E c ;
[0030] All probability distributions Ptr iConstruct a probability distribution matrix Ptr row by row; during the training process, the Ptr probability distribution matrix is used to calculate the cross-entropy loss with the true math problem expression to train the non-autoregressive math problem solver based on the multi-way tree structure; during the prediction process, the Ptr probability distribution matrix is used to obtain the character with the highest probability at each position, and use it as the predicted character at each position to obtain the predicted math problem expression.
[0031] The object of the present invention is achieved as follows.
[0032] The non-autoregressive math problem solver of the present invention based on the multi-way tree structure adds a non-autoregressive target decomposition module on the basis of the existing problem encoding module and the target-driven multi-way tree generation module to handle the unordered multi-branch decomposition work. First, the parent target E g is fused with I position encodings p i , i = 1, 2,..., I, then processed using the multi-head self-attention mechanism, and then connected to the output of the multi-head self-attention module with a multi-head mutual attention module and the candidate set E c to obtain the context representation Finally, according to the context representation and the candidate set E c , the most relevant candidate in the candidate set E c is selected through the pointer network to obtain the probability distribution matrix Ptr, which is used to calculate the cross-entropy loss during the training process and the predicted math problem expression during the prediction process. The present invention uses a non-autoregressive target decomposition module to handle the unordered multi-branch decomposition work, explores and captures the relationships between numbers, and improves the performance of the math problem solver. Description of the Drawings
[0033] Figure 1 is a specific instance diagram of the MTree tree structure;
[0034] Figure 2 is a schematic diagram of the principle of a specific implementation manner of the non-autoregressive math problem solver of the present invention based on the multi-way tree structure. Specific Embodiments
[0035] The following describes the specific embodiments of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0036] Given a math problem composed of natural language The problem text contains numbers Q = {q1, q2,..., q M}, q mDenote the m-th number, where m = 1, 2, …, M, and require the math problem solver to give a math expression O s ={e1, e2, ..., e K}, where e k ∈{+, -, ×, / } U Q ∪ C, and C is the constant set {1, π, ...}.
[0037] The present invention designs a non-autoregressive math problem solver based on a multi-way tree structure (hereinafter referred to as MTree), which adopts a goal-driven top-down strategy to generate an expression tree, thereby obtaining a problem-solving expression
[0038] 1. MTree Structure
[0039] The MTree structure was introduced into MWP in 2022 to unify the expression tree structure. MTree is such a multi-way tree, whose internal nodes are operators and external nodes are numbers or constants that appear in the problem. The child nodes of the internal nodes in MTree can be swapped with each other. For a math expression O s ={e1, e2, ..., e M}, where e m ∈{+, -, ×, / } ∪ Q ∪ C, m = 1, 2, ... M. In the expression, the operands of {+, ×} can be swapped, while the operands of {-, / } cannot be swapped. To solve this problem, the MTree structure introduces two new operators {×-, + / } to replace. ×- means the opposite of the product of multiple numbers. For example, ×-{2, 3, 4} is equal to -(2 × 3 × 4). + / means the reciprocal of the sum of multiple operands. For example, + / {2, 3, 4} is equal to In addition, the number n located at the leaf node may have multiple forms, including The structure of the MTree structure is as Figure 1 shown.
[0040] 2. Non-autoregressive Math Problem Solver Based on Multi-way Tree Structure
[0041] In this embodiment, as Figure 1 shown, the non-autoregressive math problem solver of the present invention based on the multi-way tree structure adopts an encoder-decoder structure, including a problem encoding module 1, a goal-driven multi-way tree generation module 2, and a non-autoregressive goal decomposition module 3.
[0042] The problem encoding module 1 is used to encode a given math problem P = {w1, w2, …, w N}, where w n represents the n-th word, n = 1, 2, …, N, into a distributed representation E p 、EV , where E p is the problem representation of the entire math problem, and E V is the numerical representation of the math problem.
[0043] The problem encoding module 1 uses a language model to encode a problem described in natural language into a distributed representation containing its context information. In the prior art, there are two commonly used language models: recurrent neural networks (RNNs), such as LSTM or GRU, and pre-trained language models (PLMs), such as BERT and RoBERTa. Inspired by the superior representation ability of pre-trained models, recent work tends to use pre-trained models as problem encoders. In the present invention, the problem encoding module 1 adopts RoBERTa to obtain the problem representation and the numerical representation. More specifically, given a math problem P = {w1, w2, …, w N}, the problem encoding module 1 adds a special character [CLS] and [SEP] before and after the problem respectively, and feeds them into RoBERTa, and takes the output at the [CLS] position as the representation of the entire problem. The problem encoding module 1 can be expressed by the following formula:
[0044] E p , and E V = RoBERTa([CLS]; P; [SEP])#(1)
[0045] where E p is the problem representation of the entire math problem, and E V is the numerical representation of the math problem. RoBERTa is fine-tuned during the training process.
[0046] Expression tree decoders have been well studied in math problem solving. The goal-driven mechanism gradually decomposes the entire problem into sub-problems, which intuitively conforms to the thinking of humans when solving problems. The present invention uses the goal-driven mechanism to implement a top-down MTree generator. Specifically, the goal-driven multi-way tree generation module 2 is a top-down multi-way tree generator that uses the goal-driven mechanism, and uses the problem representation E p as the root goal of the multi-way tree, and recursively generates sub-goals in a top-down manner. The sub-goals are classified as operands or operators. When the sub-goal is an operand, the result of the current sub-goal is directly obtained. When the sub-goal is an operator, the result of the sub-goal cannot be obtained yet, and it needs to be further decomposed until the sub-goal is an operand.
[0047] In the MTree structure, multiple child nodes of the same parent node are unordered. Therefore, the present invention designs a new non-autoregressive goal decomposition module (NAGD) based on the non-autoregressive Transformer to handle the unordered multi-branch decomposition work. Specifically:
[0048] First, for the parent target E in the sub-goal decomposition process g , first, the parent target E g is fused with I positional encodings p i , where i = 1, 2,..., I, to obtain the fused target E pos :
[0049] E pos = E pos = [E g + p1; E g + p2; …, E g + p I .
[0050] In this embodiment, the positional encoding is generated using the sine-cosine positional encoding in Transformer:
[0051]
[0052]
[0053] where i represents the position, l represents the l-th dimension of the encoding p i at position i, and d k is the dimension of the encoding vector.
[0054] Then, the multi-head self-attention mechanism is used for processing to fuse information from different positions, that is, the target E pos is input into the multi-head self-attention module for processing to obtain the output
[0055]
[0056] where is the trainable parameter matrix in the multi-head self-attention module.
[0057] Then, the output of the multi-head self-attention module is connected to the candidate set Ec through a multi-head cross-attention module, and the output is used as the Q matrix of the multi-head cross-attention module. The candidate set E c is multiplied by the trainable parameter after passing through a feed-forward neural network to obtain the K matrix and V matrix of the multi-head cross-attention module, and thus the output
[0058]
[0059] where d k is the dimension of the encoding vector, and the candidate set E c is:
[0060]
[0061] That is, it is composed of numbers, operators, constants, and special characters. Among them, E V is the output of the question coding module, and the remaining E op , E con , E N are trainable encodings, and N b is a special character used to represent the number of child nodes.
[0062] Output is the context representation, which is used to select the most relevant candidate set E c from the candidate characters through the pointer network. The probability ω ij of selecting the j-th candidate character at position i is:
[0063]
[0064] Among them, W p and W b are learning parameters, u is the column weight vector, and are respectively the vector representation of the i-th position and the vector representation of the j-th character in the candidate set.
[0065] In this way, the probability distribution Ptr i of all candidate characters at the i-th position is obtained:
[0066] Ptr i = softmax(ω i )
[0067] Among them, ω i = {ω i1 , ω i2 ,..., ω iJ}, and J is the number of candidate characters in the candidate set E c .
[0068] All the probability distributions Ptr i form the probability distribution matrix Ptr row by row; during the training process, the Ptr probability distribution matrix is used to calculate the cross-entropy loss with the true math problem expression to train the non-autoregressive math problem solver based on the multi-way tree structure; during the prediction process, the Ptr probability distribution matrix is used to obtain the character with the highest probability at each position and use it as the predicted character at each position to obtain the predicted math problem expression.
[0069] To calculate the loss between the predicted nodes and the actual nodes, we need to align the predicted nodes and the true nodes. Therefore, we define a pseudo-order for the true children nodes of each target. The output i.e., the context representation is classified by an operand format classifier to obtain a pseudo-order, where the operator is in the front and the constants and operands in the question are in the back. For Figure 2 the example in the pseudo-order of the children nodes is [×, ×, -40]. After predicting all the children nodes, the math problem solver needs to indicate the types for the predicted operands
[0070] 3. MTree Precision and MTree-based IoU
[0071] As mentioned above, the MTree results can well solve the defect that there are multiple forms for the same mathematical expression. For the problem-solving expressions generated by multiple math problem solvers, if direct character matching is performed, false negative cases may occur, such as a - b + c and c - b + a. Therefore, the present invention proposes MTree Precision and MTree IoU to more accurately evaluate the accuracy of the problem-solving expressions. Specifically, MTree Precision is to compare the problem-solving expressions generated by different solvers with the true expression on the MTree. At the same time, it is noted that this evaluation of the entire expression tree cannot measure the partial correctness of the expression, which is also an important way to evaluate the capabilities of the solver and human beings. In Figure 1 the example of
[0072]
[0073] 13×(10 + 3)+40 and (13×10 + 3 - 40) are two incorrect expressions. The former only uses the incorrect "+" operation, but the other parts are correct. For this reason, the present invention proposes MTree IoU to calculate the accuracy of the paths connecting the root and the leaves to measure the partial correctness of the expression. The calculation of MTree IoU is as follows: p and P g are all the paths in the predicted MTree and the true MTree respectively.
[0074] 4. Experimental Results
[0075] To evaluate the MTree math problem solver designed in the present invention, the present invention conducted a large number of experiments on two commonly used public datasets, Math23K and MAWPS. Math23K is a Chinese dataset containing 23,162 math problems, and MWPS contains 2,300 math problems.
[0076] First, the effectiveness of the MTree solver was evaluated and analyzed, and its performance was compared with the state-of-the-art. The comparison results are shown in Table (1).
[0077]
[0078] Table 1
[0079] Figure 1 are the performances of different mainstream math problem solvers. As can be seen from Table 1, the present invention is superior to all baseline models and reaches a new level on both datasets, which proves the effectiveness and superiority of the present invention.
[0080] The SUMC Solver uses path prediction to reconstruct the expression tree, and its performance is much lower than that of the present invention. The reason may be that path prediction makes nodes independent and breaks the arithmetic relationship between nodes. The solver of the present invention that implements the attention mechanism can explore and capture the relationships between numbers and produce better results. The performance of DeductReasoner, which realizes complex relationship modeling and deductive reasoning, is very close to our work, which may mean that introducing deductive reasoning into the MTree structure will bring some new insights.
[0081] In addition, ablation studies were also conducted in the experiment to investigate the effectiveness of the proposed cross-object attention. The results are shown in Table (2).
[0082]
[0083] Table 2
[0084] Table 2 is the comparison result of the ablation experiment of cross-object attention. As can be seen from Table 2, after adding cross-object attention, our model has been significantly improved. For example, from 83.2 to 84.4 on Math23K. This shows that through this cross-object attention mechanism, information belonging to different objects can be transmitted and aggregated. This cross-object information integration significantly improves the accuracy of single-object decomposition.
[0085] To study the effectiveness of the MTree precision and MTree IoU proposed in the present invention, this study used 5 representative math problem solvers with open-source code and compared them with the present invention on Math23K. The comparison results are shown in Table (3).
[0086]
[0087] Table 3
[0088] Table 3 shows the comparison results of the MTree precision and MTree IoU of mainstream math problem solvers with open-source code. From a common-sense perspective, the expression precision should be consistent with the value precision. However, as can be seen in Table 3, the expression precision is much lower than the value precision. And the MTree precision is slightly lower than the value precision, which is intuitive. It can be seen from Table 3 that the value precision, MTree precision, and MTree IoU of the present invention are all the highest, and the performance of the math problem solver in the present invention has been improved.
[0089] Although the above description of the illustrative specific embodiments of the present invention is provided for those skilled in the art of the present technology to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. A non-autoregressive math problem solver based on a multi-way tree structure, adopting an encoder-decoder structure, comprising: Problem encoding module, which is used to encode a math problem P = {w1, w2, …, w N}, where w n represents the nth word, n = 1, 2, …, N, into a distributed representation E p and E V , where E p is the problem representation of the entire math problem, and E V is the numerical representation of the math problem; The target-driven multi-way tree generation module is a top-down multi-way tree generator that adopts the target-driven mechanism and utilizes the problem representation E p As the root target of the multi-way tree, it recursively generates sub-targets in a top-down manner. The sub-targets are classified as operands or operators. When the sub-target is an operand, the result of the current sub-target is directly obtained. When the sub-target is an operator, the result of the sub-target cannot be obtained yet and needs to be further decomposed until the sub-target is an operand; It is characterized in that it further comprises A non-autoregressive target decomposition module for handling unordered multi-branch decomposition work: First, for the parent target E in the sub-goal decomposition process g , first take the parent target E g and fuse it with I positional encodings p i , where i = 1, 2, …, I, to obtain the fused target E pos : E pos = [E g + p1; E g + p2; …, E g + p I ; Then, it is processed using the multi-head self-attention mechanism, that is, the target E pos is input into the multi-head self-attention module for processing to obtain the output Among them, is a trainable parameter matrix in the multi-head self-attention module; Then, connect the output of the multi-head self-attention module through a multi-head cross-attention module and the candidate set E c , and use the output as the Q matrix of the multi-head cross-attention module. After passing the candidate set E c through a feed-forward neural network and multiplying it by the trainable parameter , obtain the K matrix and V matrix of the multi-head cross-attention module, and thus obtain the output where d k is the dimension of the encoding vector, and E c The candidate set is: That is, it consists of numbers, operators, constants, and special characters. Among them, E V is the output of the question coding module, and the rest of the E op , E con , E N are trainable encodings, and N b is a special character used to represent the number of child nodes; Output For context representation, used to select the most relevant candidate set E through the pointer network c Among the candidate characters, the probability ω of selecting the j-th candidate character at position i ij is as follows Among them, W p and W b are learning parameters, u is the column weight vector, and are respectively the vector representation at the i-th position and the vector representation of the j-th character in the candidate set; In this way, the probability distribution Ptr of all candidate characters at the i-th position is obtained i : Ptr i = softmax(ω i ) Among them, ω i ={ω i1 , ω i2 , …, ω iJ}, and J is the number of candidate characters in the candidate set E c ; All probability distributions Ptr i Form a probability distribution matrix Ptr row by row; during the training process, the Ptr probability distribution matrix is used to calculate the cross-entropy loss with the true mathematical problem expression to train the non-autoregressive mathematical problem solver based on the multi-way tree structure; during the prediction process, the Ptr probability distribution matrix is used to obtain the character with the highest probability at each position, and use it as the predicted character at each position to obtain the predicted mathematical problem expression.
Citation Information
Patent Citations
Mathematical question answering method and device
CN109961146A
Deep neural network model-based address information feature extraction method
WO2021000362A1