Automatic solution method and system for mathematical word problems based on heterogeneous graph attention network
By using heterogeneous graph attention network in automatic solution of mathematical application problems, combined with graph convolution and multi-layer graph attention neural network, the problem of existing models failing to make full use of fine-grained information is solved, and higher automatic solution accuracy and the quality of solution expressions are achieved.
Patent Information
- Application Number
- CN202310794785.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-06-30
AI Technical Summary
The existing neural network model based on Graph2Tree structure fails to make full use of the structural relationship between clause-level information and words in mathematical application questions text, resulting in many fine-grained information being unused.
The deep learning network model based on heterogeneous graph attention network is adopted, and the relationship between numbers and text is learned by processing mathematical application questions texts, and the clause and sample-level text representations are obtained through the multi-layer graph attention neural network. The final input tree decoder generates the solution expression.
It improves the accuracy of automatic solution of mathematical application problems, and can more effectively utilize the fine-grained information in the text to generate solution expressions that comply with specifications.
Smart Images

Figure CN116775841B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and specifically relates to a method and system for automatically solving mathematical word problems based on a heterogeneous graph attention network. Background Art
[0002] Automatically solving math word problems (MWP) is a hot topic in the field of machine solving. The purpose of automatically solving math word problems is to answer questions using mathematical expressions based on text descriptions. There are many types of math word problems with varying degrees of complexity, and it requires computers to understand how humans understand natural language, apply mathematical rules, and use domain knowledge for logical reasoning.
[0003] Automatic solution of mathematical word problems is one of the important research directions in machine solution research. There are four main methods for automatic solution of mathematical word problems: rule-based methods, statistical machine learning-based methods, semantic parsing-based methods, and deep learning-based methods. In recent years, deep learning methods have been widely used due to their advantages of automatic feature learning and versatility. Some studies use machine learning as a machine translation task to directly convert the input mathematical word problem text into a solution expression. Wang et al. proposed a deep neural solver, which is a seq2seq model that directly converts the input problem into an output expression. Wang et al. proposed a T-RNN model that applies a recursive neural network to predict unknown operators in the prediction template. However, direct conversion to a solution expression does not take into account the rules of mathematical formulas, and it is easy to generate some solution expressions that do not conform to the specifications and cannot be calculated. Therefore, some researchers converted the task of generating solution expressions into a tree structure and proposed the Seq2Tree model, which uses a tree decoder to simulate the idea of human beings solving mathematical word problems, and learns the constraint relationship between numbers and operators in arithmetic expressions to avoid generating some arithmetic expressions that do not conform to the specifications and cannot be calculated. Xie et al. proposed a goal-driven tree-structured MWP solver (GTS), which generates expressions through a tree-structured neural network and generates trees in a goal-driven manner. With the development of graph neural networks, some researchers have introduced graph structures into the text encoding process of mathematical word problems, parsing the text into a graph structure and using structures such as syntactic dependency trees to construct graphs. Zhang et al. proposed the Grap2Tree model, which uses a graph-based encoder to capture the relationship and order information between quantities. Wu et al. proposed the edge-enhanced hierarchical Group2Tree model (EEH-G2T) to construct an edge-labeled graph for problem representation and capture common sense information based on an external knowledge base, achieving good results. However, the existing neural network model based on the Graph2Tree structure simply constructs the entire mathematical word problem text, without fully utilizing the clause-level information in the text and the structural relationship between words in the text, which results in many finer-grained information not being utilized. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for automatically solving mathematical word problems based on a heterogeneous graph attention network, which is conducive to improving the accuracy of automatic solving of mathematical word problems.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a method for automatically solving mathematical word problems based on a heterogeneous graph attention network, comprising the following steps:
[0006] Step A: Collect the text, solution expressions and answers of mathematical word problems and construct the mathematical word problem training set DS;
[0007] Step B: Use the training set DS to train a deep learning network model G based on a heterogeneous graph attention network to solve mathematical word problems;
[0008] Step C: Input the math word problem text into the deep learning network model G, and output the corresponding solution expression and answer of the current math word problem.
[0009] Furthermore, the step B specifically includes the following steps:
[0010] Step B1: Extract the numbers from each training sample in the training set DS and replace them with the special token [NUM], then encode them to obtain the initial representation vector E and dependency tree adjacency matrix A of the math word problem text. s and the component tree adjacency matrix list C s , encode the numbers extracted from each training sample separately to obtain the digital encoding representation E num , and then the numbers are composed according to the size relationship and input into the graph convolutional neural network to obtain the hidden representation of the number h that integrates the size relationship num , the digital representation in the initial representation vector E is combined with the digital hidden representation h of the fusion size relationship num Perform double affine fusion to obtain the digital representation vector h new , and then put it back into the initial representation vector E to obtain the representation vector h of the fused digital representation fs ;
[0011] Step B2: Segment each training sample in the training set DS according to punctuation marks, and transform the representation vector h fs According to the clause segmentation, the representation h of the segmented clause is obtained c_s , and similarly construct a dependency tree adjacency matrix A for each clause after the clause c_s and the component tree adjacency matrix list C c_s ;
[0012] Step B3: Substitute the syntactic dependency adjacency matrix A obtained in step B2 c_s and the component tree matrix list C c_s Fusion is performed to obtain the clause-level adjacency matrix list adj c_s , the clause-level adjacency matrix adj c_s and the clause representation vector h c_s Input into the multi-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then the clause-level text is represented as g c_s Reassemble and get the mathematical word problem representation vector h represented by the fusion clause c_g ;
[0013] Step B4: Substitute the syntactic dependency adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s The sample-level adjacency matrix list adj is fused to obtain the mathematical word problem representation vector h represented by the sample-level adjacency matrix list adj and the fusion clause. c_g Input into the multi-layer graph attention neural network to obtain the sample-level math word problem text representation h, and then average pool the sample-level math word problem text representation h to obtain the initial target vector q;
[0014] Step B5: Input the sample-level math word problem text representation h and the initial target vector q obtained in step B4 into the tree decoder, calculate the gradient of each parameter in the deep learning network model using the back propagation method according to the target loss function loss, and update each parameter using the stochastic gradient descent method;
[0015] Step B6: When the iterative change of the loss value generated by the deep learning network model is less than a given threshold or reaches the maximum number of iterations, the training process of the deep learning network model is terminated.
[0016] Furthermore, the step B1 specifically includes the following steps:
[0017] Step B11: Traverse the training set DS, convert the solution expressions into prefix expressions p, convert the numbers in the math word problem text into special tokens [NUM], and record their values v; construct the decoder's target vocabulary V, which includes the operator op, the constant con, and the numbers n that appear in the math word problems p , that is, V = PV op , V con , n p}; Each sample in DS is represented by ds=(s, v, p); where s is the text of the math word problem, v is the number that appears in the math word problem, and p is the solution expression of the math word problem;
[0018] The math word problem text s is represented as:
[0019]
[0020] in, is the i-th word in the math word problem text s, i = 1, 2, 3..., n, n is the number of words in the math word problem text;
[0021] The number v in math word problems is represented by:
[0022] v={v 1 , v 2 , v 3 , …, vk}
[0023]
[0024] in, is the jth character in the ith number that appears in the math word problem text, i = 1, 2, 3…, k, k is the number of numbers that appear in the math word problem text, j = 1, 2, 3…, m, m is the number v i Length;
[0025] The solution expression p of the mathematical word problem is expressed as:
[0026]
[0027] in, is the ith part of the expression for the solution to the math word problem, i=1, 2, 3…, m, m is the length of the expression for the solution to the math word problem, There are two types {op, number}, where op represents an operator {+, -, ×, ÷} and number represents a number;
[0028] Step B12: For the math word problem text obtained in step B11 Use Bert as the encoder to encode and obtain the initial representation vector E;
[0029] Perform syntactic dependency analysis and syntactic component analysis on each sample in the training set to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A. s and the n-order adjacency matrix list C s , A s It is expressed as:
[0030]
[0031] Among them, A ij 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There are no syntactic dependencies;
[0032] C s It is expressed as:
[0033]
[0034] in represents the adjacency matrix of the i-th level component tree, C ij 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies;
[0035] Step B13: Substitute the number v obtained in step B11 into {v 1 , v 2 , v 3 , …, v k}, get the initial representation E of the number v by random initialization n ; Input the initial representation into the forward layer and reverse layer of a bidirectional gated recurrent neural network to obtain the state vector sequence of the forward hidden layer and the reverse hidden state vector sequence in, The forward hidden state vector and the reverse hidden state vector are concatenated and average pooled to obtain the digital code representation E num :
[0036]
[0037]
[0038]
[0039] in, Represents a splicing operation;
[0040] Step B14: Substitute the number v obtained in step B11 into {v 1 , v 2 , v 3 , …, v k}, construct the adjacency matrix adj according to the size relationship of the numbers num :
[0041]
[0042] Then the adjacency matrix adj num and the numerical eigenvector E num Input into a graph convolutional neural network to obtain the digital hidden representation h of the fusion size relationship num :
[0043] h num =GCN(adj num , E num )
[0044] Step B15: Hidden digital representation h of the fused size relationship obtained in step B14 num With the initial characterization vector E s The numbers in E orgPerform double affine fusion to obtain the digital representation vector h of the fused context representation new :
[0045]
[0046] Among them, W 2 and W 3 is a learnable weight matrix; and h new Replace the digital representation in the initial representation vector E to obtain the representation vector h that integrates the digital representation fs .
[0047] Furthermore, the step B2 specifically includes the following steps:
[0048] Step B21: Segment each training sample according to punctuation marks, and merge the representation vector h of the digit and its neighboring nodes fs Cut according to the clauses to get the representation of each clause after the clause is split Among them, k means there are k clauses in the text;
[0049] Step B22: Perform syntactic dependency analysis and syntactic component analysis on each clause obtained in step 21 to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A c_s and the n-order adjacency matrix list C c_s , A c_s It is expressed as:
[0050]
[0051]
[0052] in, 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There are no syntactic dependencies. Represents the dependency tree adjacency matrix of the i-th clause;
[0053] C c_s It is expressed as:
[0054]
[0055]
[0056]
[0057] in, represents the j-th level component tree adjacency matrix of the ith clause, represents the component tree adjacency matrix of the ith clause, 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies.
[0058] Furthermore, the step B3 specifically includes the following steps:
[0059] Step B31: Select the adjacency matrix C of the component tree of the clause obtained in step B2 c_s The last K layers of the clause are obtained by taking the component tree adjacency matrix The adjacency matrix A of the dependency tree of the clause obtained in step B2 c_s Copy K times And merge it with the component tree matrix by taking the union method, that is, when one of the two is 1, the corresponding position of the adjacency matrix is 1, otherwise it is 0, and the clause-level adjacency matrix list adj is obtained c_s ;
[0060] Step B32: Substitute the clause-level adjacency matrix adj obtained in step B31 c_s and the clause representation vector h c_s Input into the K-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then concatenate them to get the mathematical word problem representation vector h represented by the fusion clause c_g ;
[0061] h c_s =g 0
[0062]
[0063]
[0064]
[0065]
[0066] h c_g =||g K
[0067] Among them, W 4 and W 5 is a learnable weight matrix, || represents a concatenation operation, FFC represents a feedforward neural network, Represents the neighbor nodes of node i in layer 1.
[0068] Furthermore, the step B4 specifically includes the following steps:
[0069] Step B41: Substitute the dependency tree adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s According to the fusion method of clause-level adjacency matrix, the sample-level adjacency matrix adj is obtained;
[0070] Step B42: Input the sample-level adjacency matrix adj obtained in step B41 and the clause-level representation vector obtained in step B3 into the K-layer graph attention neural network to obtain the sample-level representation vector h, and then perform average pooling on it to obtain the initial target vector q;
[0071] h c_g =g 0
[0072]
[0073]
[0074]
[0075]
[0076] h=g K
[0077] q=MeanPool(h)
[0078] Among them, W6, W7 and is a learnable weight matrix, d represents the input dimension, and H represents the number of attention heads.
[0079] Furthermore, the step B5 specifically includes the following steps:
[0080] Step B51: h P Represents the sample-level representation vector of sample P. For each token in the target vocabulary V of sample P obtained in step B1, their embedding is defined as:
[0081]
[0082] Among them, M op and M con is a learnable weight matrix, and the initialized token representation is obtained by querying the embedding matrix. represents the representation of the sample-level representation vector of sample P obtained in step B4 at the y position, that is, the representation of the number appearing in sample P;
[0083] Step B52: A context vector c is derived from the target vector q and the sample-level representation vector h obtained in step B4. The calculation formula is as follows:
[0084]
[0085]
[0086]
[0087] Then generate a tokeny from the target vocabulary V, whose unnormalized log probability is:
[0088]
[0089] Among them, W s is a learnable weight matrix; then normalize it to generate the token with the highest probability As predicted value:
[0090]
[0091]
[0092] Step B53: Assuming that the token predicted in step B52 is an operator, the current target vector q will be achieved by a left sub-goal and a right sub-goal; the left sub-goal is calculated by a two-layer feedforward neural network with a gating mechanism:
[0093]
[0094]
[0095] h l =o l ⊙C l
[0096]
[0097] Q le =tanh(W le h l )
[0098] q l =g l ⊙Q le
[0099] Among them, W ol , W cl , W gl and W le is a learnable weight matrix;
[0100] Step B54: The right sub-target of the right sub-tree needs to consider the left sub-tree, and the left sub-tree is encoded from bottom to top as follows:
[0101]
[0102]
[0103]
[0104]
[0105] Among them, W gt and W ct is a learnable weight matrix;
[0106] Then calculate the target vector q of the right child node r ,h r is a hidden state, and its parent node passes it to its right child node from top to bottom:
[0107]
[0108]
[0109] h r =o r ⊙C r
[0110] g r =σ(W gr [h r , t l ])
[0111] Q re =tanh(W re [h r , t l ])
[0112] q r =g r ⊙Q re
[0113] Among them, W or , W cr , W gr and W re is a learnable weight matrix;
[0114] Step B54: For the training data set ds=(s, v, p), set the target loss function to the negative log-likelihood function:
[0115]
[0116] Where m represents the length of the prefix expression.
[0117] The present invention also provides a system for automatically solving mathematical word problems based on a heterogeneous graph attention network, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method steps can be implemented.
[0118] Compared with the prior art, the present invention has the following beneficial effects: the present invention first encodes the numbers in the mathematical word problems, then composes a graph using the size relationship between the numbers, learns the information of the numbers themselves through a graph convolutional neural network, then divides the text into sentences, and fusedly composes the graph using the component tree and dependency tree of the clauses, obtains the text representation at the clause granularity through a multi-layer graph attention neural network, splices it back to the sample level and fusedly composes the dependency tree and component tree in the same way, obtains the sample-level representation through the graph attention neural network, and inputs it into a tree-structured decoder, thereby improving the accuracy of the model in predicting the expression of the solution to the mathematical word problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0119] Figure 1 is a flow chart of a method implementation of an embodiment of the present invention;
[0120] Figure 2 It is a model architecture diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0121] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0122] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0123] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0124] like Figure 1 As shown, this embodiment provides a method for automatically solving mathematical word problems based on a heterogeneous graph attention network, comprising the following steps:
[0125] Step A: Collect the text, solution expressions and answers of mathematical word problems and construct a mathematical word problem training set DS.
[0126] Step B: Use the training set DS to train a deep learning network model G based on a heterogeneous graph attention network for solving mathematical word problems. In this embodiment, the architecture of the deep learning network model G based on a heterogeneous graph attention network is as follows: Figure 2 shown.
[0127] Step C: Input the math word problem text into the deep learning network model G, and output the corresponding solution expression and answer of the current math word problem.
[0128] In this embodiment, step B specifically includes the following steps:
[0129] Step B1: Extract the numbers from each training sample in the training set DS and replace them with the special token [NUM], then encode them to obtain the initial representation vector E and dependency tree adjacency matrix A of the math word problem text. s and the component tree adjacency matrix list C s , encode the numbers extracted from each training sample separately to obtain the digital encoding representation E num , and then the numbers are composed according to the size relationship and input into the graph convolutional neural network to obtain the hidden representation of the number h that integrates the size relationship num , the digital representation in the initial representation vector E is combined with the digital hidden representation h of the fusion size relationship num Perform double affine fusion to obtain the digital representation vector h new , and then put it back into the initial representation vector E to obtain the representation vector h of the fused digital representation fs .
[0130] In this embodiment, step B1 specifically includes the following steps:
[0131] Step B11: Traverse the training set DS, convert the solution expressions into prefix expressions p, convert the numbers in the math word problem text into special tokens [NUM], and record their values v; construct the decoder's target vocabulary V, which includes the operator op, the constant con, and the numbers n that appear in the math word problems p , that is, V = {V op , V con , n p}; Each sample in DS is represented by ds=(s, v, p); where s is the text of the math word problem, v is the number that appears in the math word problem, and p is the solution expression of the math word problem.
[0132] The math word problem text s is represented as:
[0133]
[0134] in, is the i-th word in the math word problem text s, i=1, 2, 3…, n, and n is the number of words in the math word problem text.
[0135] The number v in math word problems is represented by:
[0136] v={v 1 , v 2 , v 3 , …, v k}
[0137]
[0138] in, is the jth character in the ith number that appears in the math word problem text, i = 1, 2, 3…, k, k is the number of numbers that appear in the math word problem text, j = 1, 2, 3…, m, m is the number v i Length.
[0139] The solution expression p of the mathematical word problem is expressed as:
[0140]
[0141] in, is the ith part of the expression for the solution to the math word problem, i=1, 2, 3…, m, m is the length of the expression for the solution to the math word problem, There are two types {op, number}, where op represents an operator {+, -, ×, ÷} and number represents a number.
[0142] Step B12: For the math word problem text obtained in step B11 Use Bert as the encoder to encode and obtain the initial representation vector E.
[0143] Perform syntactic dependency analysis and syntactic component analysis on each sample in the training set to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A. s and the n-order adjacency matrix list C s , A s It is expressed as:
[0144]
[0145] Among them, A ij 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There are no syntactic dependencies.
[0146] C sIt is expressed as:
[0147]
[0148] in represents the adjacency matrix of the i-th level component tree, C ij 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies.
[0149] Step B13: Substitute the number v obtained in step B11 into {v 1 , v 2 , v 3 , …, v k}, get the initial representation E of the number v by random initialization n ; Input the initial representation into the forward layer and reverse layer of a bidirectional gated recurrent neural network to obtain the state vector sequence of the forward hidden layer and the reverse hidden state vector sequence in, The forward hidden state vector and the reverse hidden state vector are concatenated and average pooled to obtain the digital code representation E num :
[0150]
[0151]
[0152]
[0153] in, Represents a concatenation operation.
[0154] Step B14: Substitute the number v obtained in step B11 into {v 1 , v 2 , v 3 , …, v k}, construct the adjacency matrix adj according to the size relationship of the numbers num :
[0155]
[0156] Then the adjacency matrix adj num and the numerical eigenvector E num Input into a graph convolutional neural network to obtain the digital hidden representation h of the fusion size relationship num :
[0157] h num =GCN(adj num , E num )
[0158] Step B15: Hidden digital representation h of the fused size relationship obtained in step B14 num With the initial characterization vector E s The numbers in E org Perform double affine fusion to obtain the digital representation vector h of the fused context representation new :
[0159]
[0160] Among them, W 2 and W 3 is a learnable weight matrix; and h new Replace the digital representation in the initial representation vector E to obtain the representation vector h that integrates the digital representation fs .
[0161] Step B2: Segment each training sample in the training set DS according to punctuation marks, and transform the representation vector h fs According to the clause segmentation, the representation h of the segmented clause is obtained c_s , and similarly construct a dependency tree adjacency matrix A for each clause after the clause c_s and the component tree adjacency matrix list C c_s .
[0162] In this embodiment, step B2 specifically includes the following steps:
[0163] Step B21: Segment each training sample according to punctuation marks, and merge the representation vector h of the digit and its neighboring nodes fs Cut according to the clauses to get the representation of each clause after the clause is split Among them, k means there are k clauses in the text.
[0164] Step B22: Perform syntactic dependency analysis and syntactic component analysis on each clause obtained in step 21 to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A c_s and the n-order adjacency matrix list C c_s , A c_s It is expressed as:
[0165]
[0166]
[0167] in, 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There are no syntactic dependencies. Represents the dependency tree adjacency matrix of the ith clause.
[0168] C c_s It is expressed as:
[0169]
[0170]
[0171]
[0172] in, represents the j-th level component tree adjacency matrix of the ith clause, represents the component tree adjacency matrix of the ith clause, 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies.
[0173] Step B3: Substitute the syntactic dependency adjacency matrix A obtained in step B2 c_s and the component tree matrix list C c_s Fusion is performed to obtain the clause-level adjacency matrix list adj c_s , the clause-level adjacency matrix adj c_s and the clause representation vector h c_s Input into the multi-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then the clause-level text is represented as g c_s Reassemble and get the mathematical word problem representation vector h represented by the fusion clause c_g .
[0174] In this embodiment, step B3 specifically includes the following steps:
[0175] Step B31: Select the adjacency matrix C of the component tree of the clause obtained in step B2 c_s The last K layers of the clause are obtained by taking the component tree adjacency matrix The adjacency matrix A of the dependency tree of the clause obtained in step B2 c_s Copy K times And merge it with the component tree matrix by taking the union method, that is, when one of the two is 1, the corresponding position of the adjacency matrix is 1, otherwise it is 0, and the clause-level adjacency matrix list adj is obtained c_s .
[0176] Step B32: Substitute the clause-level adjacency matrix adj obtained in step B31 c_s and the clause representation vector h c_s Input into the K-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then concatenate them to get the mathematical word problem representation vector h represented by the fusion clause c_g ;
[0177] h c_s =g 0
[0178]
[0179]
[0180]
[0181]
[0182] h c_g =||g K
[0183] Among them, W 4 and W 5 is a learnable weight matrix, || represents a concatenation operation, FFC represents a feedforward neural network, Represents the neighbor nodes of node i in layer 1.
[0184] Step B4: Substitute the syntactic dependency adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s The sample-level adjacency matrix list adj is fused to obtain the mathematical word problem representation vector h represented by the sample-level adjacency matrix list adj and the fusion clause. c_g The sample-level math word problem text representation h is input into the multi-layer graph attention neural network to obtain the sample-level math word problem text representation h, and then the sample-level math word problem text representation h is average pooled to obtain the initial target vector q.
[0185] In this embodiment, step B4 specifically includes the following steps:
[0186] Step B41: Substitute the dependency tree adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s The clause-level adjacency matrix is fused in the same way as the clause-level adjacency matrix to obtain the sample-level adjacency matrix adj.
[0187] Step B42: Input the sample-level adjacency matrix adj obtained in step B41 and the clause-level representation vector obtained in step B3 into the K-layer graph attention neural network to obtain the sample-level representation vector h, and then perform average pooling on it to obtain the initial target vector q;
[0188] h c_g =g 0
[0189]
[0190]
[0191]
[0192]
[0193] h=g K
[0194] q=MeanPool(h)
[0195] Among them, W 6 , W 7 and is a learnable weight matrix, d represents the input dimension, and H represents the number of attention heads.
[0196] Step B5: Input the sample-level math word problem text representation h and the initial target vector q obtained in step B4 into the tree decoder, calculate the gradient of each parameter in the deep learning network model using the back propagation method according to the target loss function, and update each parameter using the stochastic gradient descent method.
[0197] In this embodiment, step B5 specifically includes the following steps:
[0198] Step B51: h P Represents the sample-level representation vector of sample P. For each token in the target vocabulary V of sample P obtained in step B1, their embedding is defined as:
[0199]
[0200] Among them, M op and M con is a learnable weight matrix, and the initialized token representation is obtained by querying the embedding matrix. It represents the representation of the sample-level representation vector of sample P obtained in step B4 at the y position, that is, the representation of the numbers appearing in sample P.
[0201] Step B52: A context vector c is derived from the target vector q and the sample-level representation vector h obtained in step B4. The calculation formula is as follows:
[0202]
[0203]
[0204]
[0205] Then generate a token y from the target vocabulary V, whose unnormalized log probability is:
[0206]
[0207] Among them, W s is a learnable weight matrix; then normalize it to generate the token with the highest probability As predicted value:
[0208]
[0209]
[0210] Step B53: Assuming that the token predicted in step B52 is an operator, the current target vector q will be achieved by a left sub-goal and a right sub-goal; the left sub-goal is calculated by a two-layer feedforward neural network with a gating mechanism:
[0211]
[0212]
[0213] h l =o l ⊙C l
[0214] g l =σ(W gl h l )
[0215] Q le =tanh(W le h l )
[0216] q l =g l ⊙Q le
[0217] Among them, W ol , W cl , W gl and W leis a learnable weight matrix.
[0218] Step B54: The right sub-target of the right sub-tree needs to consider the left sub-tree, and the left sub-tree is encoded from bottom to top as follows:
[0219]
[0220]
[0221]
[0222]
[0223] Among them, W gt and W ct is a learnable weight matrix.
[0224] Then calculate the target vector q of the right child node r ,h r is a hidden state, and its parent node passes it to its right child node from top to bottom:
[0225]
[0226]
[0227] h r =o r ⊙C r
[0228] g r =σ(W gr [h r , t l ])
[0229] Q re =tanh(W re [h r , t l ])
[0230] q r =g r ⊙Q re
[0231] Among them, W or , W cr , W gr and W re is a learnable weight matrix.
[0232] Step B54: For the training data set ds=(s, v, p), set the target loss function to the negative log-likelihood function:
[0233]
[0234] Where m represents the length of the prefix expression.
[0235] Step B6: When the iterative change of the loss value generated by the deep learning network model is less than a given threshold or reaches the maximum number of iterations, the training process of the deep learning network model is terminated.
[0236] This embodiment also provides a system for automatically solving mathematical word problems based on a heterogeneous graph attention network, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method steps can be implemented.
[0237] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0238] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0239] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0240] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0241] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for automatically solving mathematical word problems based on heterogeneous graph attention networks, It is characterized in that The following steps are involved: Step A: Collect the text, solution expressions and answers of mathematical word problems and construct a mathematical word problem training set DS; Step B: Use the training set DS to train a deep learning network model G based on a heterogeneous graph attention network to solve mathematical word problems; Step C: Input the math problem text into the deep learning network model G, and output the corresponding solution expression and answer of the current math problem; The step B specifically comprises the following steps: Step B1: Extract the numbers from each training sample in the training set DS and replace them with the special token [NUM], then encode them to obtain the initial representation vector E and dependency tree adjacency matrix A of the math word problem text. s and the component tree adjacency matrix list C s , encode the numbers extracted from each training sample separately to obtain the digital encoding representation E num , and then the numbers are composed according to the size relationship and input into the graph convolutional neural network to obtain the hidden representation of the number h that integrates the size relationship num , the digital representation in the initial representation vector E is combined with the digital hidden representation h of the fusion size relationship num Perform double affine fusion to obtain the digital representation vector h new , and then put it back into the initial representation vector E to obtain the representation vector h of the fused digital representation fs ; Step B2: Segment each training sample in the training set DS according to punctuation marks, and transform the representation vector h fs According to the clause segmentation, the representation h of the segmented clause is obtained c_s , and similarly construct a dependency tree adjacency matrix A for each clause after the clause c_s and the component tree adjacency matrix list C c_s ; Step B3: Substitute the syntactic dependency adjacency matrix A obtained in step B2 c_s and the component tree matrix list C c_s Fusion is performed to obtain the clause-level adjacency matrix list adj c_s , the clause-level adjacency matrix adj c_s and the clause representation vector h c_s Input into the multi-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then the clause-level text is represented as g c_s Reassemble and get the mathematical word problem representation vector h represented by the fusion clause c_g ; Step B4: Substitute the syntactic dependency adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s The sample-level adjacency matrix list adj is fused to obtain the mathematical word problem representation vector h represented by the sample-level adjacency matrix list adj and the fusion clause. c_g Input into the multi-layer graph attention neural network to obtain the sample-level math word problem text representation h, and then average pool the sample-level math word problem text representation h to obtain the initial target vector q; Step B5: Input the sample-level math word problem text representation h and the initial target vector q obtained in step B4 into the tree decoder, calculate the gradient of each parameter in the deep learning network model using the back propagation method according to the target loss function loss, and update each parameter using the stochastic gradient descent method; Step B6: When the iterative change of the loss value generated by the deep learning network model is less than a given threshold or reaches the maximum number of iterations, the training process of the deep learning network model is terminated; The step B3 specifically comprises the following steps: Step B31: Select the adjacency matrix C of the component tree of the clause obtained in step B2 c_s The last K layers of the clause are obtained by taking the component tree adjacency matrix The adjacency matrix A of the dependency tree of the clause obtained in step B2 c_s Copy K times And merge it with the component tree matrix by taking the union method, that is, when one of the two is 1, the corresponding position of the adjacency matrix is 1, otherwise it is 0, and the clause-level adjacency matrix list adj is obtained c_s ; Step B32: Substitute the clause-level adjacency matrix adj obtained in step B31 into c_s and the clause representation vector h c_s Input into the K-layer graph attention neural network to obtain the sentence-level mathematical word problem text representation g c_s , and then concatenate them to get the mathematical word problem representation vector h represented by the fusion clause c_g ; h c_s =g 0 h c_g =||g K Among them, W 4 and W 5 is a learnable weight matrix, ∥ represents a concatenation operation, FFC represents a feedforward neural network, Represents the neighboring nodes of node i in layer l; The step B4 specifically comprises the following steps: Step B41: Substitute the dependency tree adjacency matrix A obtained in step B1 s and the component tree adjacency matrix list C s According to the fusion method of clause-level adjacency matrix, the sample-level adjacency matrix adj is obtained; Step B42: Input the sample-level adjacency matrix adj obtained in step B41 and the clause-level representation vector obtained in step B3 into the K-layer graph attention neural network to obtain the sample-level representation vector h, and then perform average pooling on it to obtain the initial target vector q; h c_g =g 0 h=g K q=MeanPool(h) Among them, W 6 , W 7 and is a learnable weight matrix, d represents the input dimension, and H represents the number of attention heads.
2. The method for automatically solving mathematical word problems based on heterogeneous graph attention network according to claim 1, It is characterized in that The step B1 specifically comprises the following steps: Step B11: Traverse the training set DS, convert the solution expressions into prefix expressions p, convert the numbers in the math word problem text into special tokens [NUM], and record their values v; construct the decoder's target vocabulary V, which includes the operator op, the constant con, and the numbers n that appear in the math word problems p , that is, V = {V op ,V con ,n p }; Each sample in DS is represented by ds = (s, v, p); where s is the text of the math word problem, v is the number that appears in the math word problem, and p is the solution expression of the math word problem; The math word problem text s is represented as: in, is the i-th word in the math word problem text s, i = 1, 2, 3..., n, n is the number of words in the math word problem text; The number v in math word problems is represented by: v={v 1 ,v 2 ,v 3 ,…,v k } in, is the jth character in the ith number that appears in the math word problem text, i = 1, 2, 3…, k, k is the number of numbers that appear in the math word problem text, j = 1, 2, 3…, m, m is the number v i Length; The solution expression p of the mathematical word problem is expressed as: in, is the ith part of the expression for the solution to the math word problem, i=1, 2, 3…, m, m is the length of the expression for the solution to the math word problem, There are two types {op, number}, where op represents an operator {+, -, ×, ÷} and number represents a number; Step B12: For the math word problem text obtained in step B11 Use Bert as the encoder to encode and obtain the initial representation vector E; Perform syntactic dependency analysis and syntactic component analysis on each sample in the training set to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A. s and the n-order adjacency matrix list C s , A s It is expressed as: Among them, A ij 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There are no syntactic dependencies; C s It is expressed as: in represents the adjacency matrix of the i-th level component tree, C ij 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies; Step B13: Substitute the number v obtained in step B11 into {v 1 ,v 2 ,v 3 ,…,v k }, get the initial representation E of the number v by random initialization n ; Input the initial representation into the forward layer and reverse layer of a bidirectional gated recurrent neural network to obtain the state vector sequence of the forward hidden layer and the reverse hidden state vector sequence in, t=1,2,3…,m, concatenate the forward hidden state vector and the reverse hidden state vector, and perform average pooling to obtain the digital code representation E num : in, Represents a splicing operation; Step B14: Substitute the number v obtained in step B11 into {v 1 ,v 2 ,v 3 ,…,v k }, construct the adjacency matrix adj according to the size relationship of the numbers num : Then the adjacency matrix adj num and the numerical eigenvector E num Input into a graph convolutional neural network to obtain the digital hidden representation h of the fusion size relationship num : h num =GCN(adj num ,E num ) Step B15: Hidden digital representation h of the fused size relationship obtained in step B14 num With the initial characterization vector E s The numbers in E org Perform double affine fusion to obtain the digital representation vector h of the fused context representation new : Among them, W 2 and W 3 is a learnable weight matrix; and h new Replace the digital representation in the initial representation vector E to obtain the representation vector h that integrates the digital representation fs .
3. The method for automatically solving mathematical word problems based on heterogeneous graph attention network according to claim 2, It is characterized in that The step B2 specifically comprises the following steps: Step B21: Segment each training sample according to punctuation marks, and merge the representation vector h of the digit and its neighboring nodes fs Cut according to the clauses to get the representation of each clause after the clause is split Among them, k means there are k clauses in the text; Step B22: Perform syntactic dependency analysis and syntactic component analysis on each clause obtained in step 21 to obtain a syntactic dependency tree and a component tree, and encode them into an n-order adjacency matrix A c_s and the n-order adjacency matrix list C c_s , A c_s It is expressed as: in, 1 represents word and words There is a syntactic dependency relationship between them, and 0 means that the word and words There is no syntactic dependency. Represents the dependency tree adjacency matrix of the i-th clause; C c_s It is expressed as: in, represents the j-th level component tree adjacency matrix of the ith clause, represents the component tree adjacency matrix of the ith clause, 1 represents word and words There is a component dependency relationship between them, and 0 means that the word and words There are no ingredient dependencies.
4. The method for automatically solving mathematical word problems based on heterogeneous graph attention network according to claim 3, It is characterized in that The step B5 specifically comprises the following steps: Step B51: h P Represents the sample-level representation vector of sample P. For each token in the target vocabulary V of sample P obtained in step B1, their embedding is defined as: Among them, M op and M con is a learnable weight matrix, and the initialized token representation is obtained by querying the embedding matrix. represents the representation of the sample-level representation vector of sample P obtained in step B4 at the y position, that is, the representation of the number appearing in sample P; Step B52: A context vector c is derived from the target vector q and the sample-level representation vector h obtained in step B4. The calculation formula is as follows: Then generate a token y from the target vocabulary V, whose unnormalized log probability is: Among them, W s is a learnable weight matrix; then normalize it to generate the token with the highest probability As predicted value: Step B53: Assuming that the token predicted in step B52 is an operator, the current target vector q will be achieved by a left sub-goal and a right sub-goal; the left sub-goal is calculated by a two-layer feedforward neural network with a gating mechanism: h l =o l ⊙C l g l =σ(W gl h l ) Q le =tanh(W le h l ) q l =g l ⊙Q le Among them, W ol , W cl , W gl and W le is a learnable weight matrix; Step B54: The right sub-target of the right sub-tree needs to consider the left sub-tree, and the left sub-tree is encoded from bottom to top as follows: Among them, W gt and W ct is a learnable weight matrix; Then calculate the target vector q of the right child node r ,h r is a hidden state, and its parent node passes it to its right child node from top to bottom: h r =o r ⊙C r g r =σ(W gr [h r ,t l ]) Q re =tanh(W re [h r ,t l ]) q r =g r ⊙Q re Among them, W or , W cr , W gr and W re is a learnable weight matrix; Step B54: For the training data set ds=(s,v,p), set the target loss function to the negative log-likelihood function: Where m represents the length of the prefix expression.
5. An automatic mathematical word problem solving system based on heterogeneous graph attention network, It is characterized in that The method comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps as described in any one of claims 1 to 4 can be implemented.