Processing method and system of mathematical problem solver based on deep learning

By introducing a deep learning-based mathematical problem-solver into the education platform, using geometric and algebraic problem-solving models to intelligently solve problems, the existing education platform cannot handle problems that are not matched in the question bank, and efficient and accurate mathematical problem answers are achieved.

CN115344811BActive Publication Date: 2025-05-13HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210845072.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-05-13
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

When searching for questions, existing education platforms cannot effectively deal with questions that are not matched in the question bank, making it difficult for users to find the answers they need.

Method used

Using a mathematical problem-solver based on deep learning, the problem-solving model in the geometric problem-solver and algebraic problem-solver is used to intelligently solve the search problems entered by the user to generate problem-solving results.

Benefits of technology

Intelligent solutions for various geometric and algebraic problems can be carried out without storing the question bank, which significantly improves the solution rate and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344811B_ABST
    Figure CN115344811B_ABST
Patent Text Reader

Abstract

The present application provides a processing method and system for a mathematical problem solver based on deep learning. Since the problem-solving models in the geometry solver and the algebra solver are obtained by training using neural networks, one of the geometry solver and the algebra solver can be selected as a target solver according to the search question input by the user, and the search question can be input into the problem-solving model of the target solver for problem-solving processing, and the problem-solving result can be output to the user for viewing. This method does not require the storage of a question bank, and can perform intelligent solutions for various geometry and algebra problems, and the solution rate and accuracy are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a processing method and system for a mathematical problem solver based on deep learning. Background Art

[0002] With the development of web 2.0 and cloud computing technologies, the continuous improvement of question banks, and the rapid development and popularization of smart phones, some question search platform software has emerged. Some of these question search platforms allow users to upload question files, some to input questions, and some to take photos to search for questions. These software are connected to educational platforms and combined with the form of MOOC, and can be regarded as a simple extension of offline education on the Internet.

[0003] At present, most educational platforms and software match questions with the question bank in the database of the platform. If the original question is found, the question and solution will be output; if the original question is not found, similar questions and solutions will be output. Therefore, most students may have encountered this situation when using these platforms. When searching for a question, the platform will output a completely unrelated question and solution, which will bring considerable inconvenience. This is because most of the current educational platforms compare questions in their own question bank, rather than analyzing the questions to generate answers. This results in questions that are not in the question bank not getting the required answers no matter how many times they are searched. Summary of the invention

[0004] In view of this, the purpose of this application is to propose a processing method and system for a mathematical problem solver based on deep learning to solve or partially solve the above-mentioned technical problems.

[0005] Based on the above purpose, the first aspect of the present application provides a processing method of a mathematical problem solver based on deep learning, wherein the mathematical problem solver includes a geometric problem solver and an algebraic problem solver;

[0006] The processing method comprises:

[0007] receiving a search question type, and determining a target problem solver from a geometry solver and an algebra solver of the math problem solver according to the question type;

[0008] Receiving a search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, and performing problem solving processing on the question information through a problem solving model in the target problem solver to obtain a problem solving result, wherein the problem solving model is obtained by training a neural network with problem solving samples;

[0009] The problem-solving result is outputted.

[0010] Based on the same inventive concept, the second aspect of the present application proposes a problem-solving processing system, comprising:

[0011] The user layer is configured to receive search question types and display solution results;

[0012] The functional logic layer is configured to determine a target problem solver from the geometry solver and the algebra solver of the math problem solver according to the problem type, receive the search question, and feed back the problem solving result to the user layer;

[0013] The function execution layer includes a geometry problem solver and an algebra problem solver, and is configured to pre-process the search problem to obtain problem information, input the problem information into the target problem solver, perform problem solving processing on the problem information through the problem solving model in the target problem solver, obtain a problem solving result, and feed the problem solving result back to the function execution layer; wherein the problem solving model is obtained by training a neural network with problem solving samples;

[0014] The storage computing layer is configured to store the data sets required by the geometry problem solver and the algebra problem solver, so that the geometry problem solver and the algebra problem solver can retrieve corresponding data from the data sets for problem solving when performing problem solving.

[0015] From the above, it can be seen that the processing method and system of the deep learning-based mathematical problem solver provided by the present application, since the problem-solving models in the geometric problem solver and the algebraic problem solver are obtained by training using neural networks, it is possible to select one of the geometric problem solver and the algebraic problem solver as a target problem solver according to the search problem input by the user, input the search problem into the problem-solving model of the target problem solver for problem-solving processing, and obtain the problem-solving result and output it to the user for viewing. This method does not require the storage of a question bank, and can perform intelligent solutions for various geometric and algebraic problems, and the solution rate and accuracy are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A flowchart of a processing method of a mathematical problem solver based on deep learning according to an embodiment of the present application;

[0018] Figure 2-A A system use case diagram of an embodiment of the present application;

[0019] Figure 2-BSchematic diagram of input and output of the geometry problem solver model of the embodiment of the present application;

[0020] Figure 2-C A schematic diagram of an algebraic problem solver model according to an embodiment of the present application;

[0021] Figure 2-D A schematic diagram illustrating question distribution by the number of words in a sentence according to an embodiment of the present application;

[0022] Figure 2-E Schematic diagram of an element set P and a symbol set S according to an embodiment of the present application;

[0023] Figure 2-F A schematic diagram of the flow chart of the encoder of the embodiment of the present application;

[0024] Figure 2-G An example diagram of an encoder for a decision tree according to an embodiment of the present application;

[0025] Figure 3-A A schematic diagram of the structure of the problem-solving processing system according to an embodiment of the present application;

[0026] Figure 3-B A schematic diagram of the entire business process when a user uses the subsystem according to an embodiment of the present application;

[0027] Figure 3-C Some fault examples of the geometry problem solver according to the embodiment of the present application;

[0028] Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0030] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the usual meanings understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0031] Based on the description of the background technology, it is very meaningful to develop an educational platform that can analyze geometry and algebra problems and generate solution processes.

[0032] This application mainly uses the a priori method to analyze the formal language obtained from the problem, apply relevant geometric and algebraic theorems, and gradually perform symbolic reasoning until the final answer is predicted. This is very similar to the general problem-solving process. According to the learned theorem, the obtained problem information is substituted into the formula to obtain the answer. This is also the reason why the a priori method is used for research.

[0033] The processing method of a deep learning-based math problem solver proposed in an embodiment of the present application is applied to a problem solver subsystem, in which a math problem solver is provided, and the math problem solver includes a geometry problem solver and an algebra problem solver.

[0034] like Figure 1 As shown, the processing method includes:

[0035] Step 101, receiving a search question type, and determining a target problem solver from a geometry solver and an algebra solver of the math problem solver according to the question type.

[0036] In specific implementation, the user can install the software program of the problem solver subsystem in the terminal device, so that the user can input the corresponding search question type through the display screen of the terminal device, for example, geometry or algebra. For geometry question types, the geometry solver is called, and for algebra, the algebra solver is called.

[0037] Step 102, receiving a search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, performing problem-solving processing on the question information through a problem-solving model in the target problem solver, and obtaining a problem-solving result, wherein the problem-solving model is obtained by training a neural network with problem-solving samples.

[0038] Step 103, outputting the problem-solving result.

[0039] In specific implementation, after receiving the solution result, the user can also determine whether the solution result is correct. The system collects the results of user feedback to determine the use effect of the math problem solver.

[0040] Through the above technical solution, since the problem-solving models in the geometry solver and the algebra solver are both obtained by training using neural networks, one of the geometry solver and the algebra solver can be selected as a target solver according to the search question input by the user, and the search question can be input into the problem-solving model of the target solver for problem-solving processing, and the problem-solving result can be output to the user for viewing. This method does not require the storage of a question bank, and can intelligently solve various geometry and algebra problems, and the solving speed and accuracy are effectively improved.

[0041] In some embodiments, in step 102, in response to determining that the target problem solver is a geometry problem solver, the geometry problem solver includes: a geometry problem solving model obtained by training a neural network with geometry problem solving samples;

[0042] The problem-solving process includes:

[0043] Step A1, receiving the search question, extracting the image element information and text element information in the search question, wherein the image element information and text element information are used as the title information.

[0044] Step A2, input the image element information and the text element information, as well as the determined axiom theorem set into the geometry problem solver, perform geometry problem solving through a geometry problem solving model, obtain a problem solving process sequence, and use the problem solving process sequence as the problem solving result.

[0045] Geometry questions will show what type of question it belongs to, including circle, rectangle, square, parallelogram, etc. At the same time, the type of answer required by the question will be determined and displayed, such as the answer type is length, surface diameter, angle, etc. The system will also include the questions queried this time, and the user can directly display them next time he wants to find the question. In addition, the system also has comments, which can provide feedback on whether the system solves the problem correctly.

[0046] When the user chooses to search for a geometry problem, the system extracts information from the geometry problem searched by the user, uses the extracted information and axiom theorem set as the input of the geometry problem-solving model, outputs a predicted theorem sequence, and brings in the corresponding information to obtain the solution. This method can speed up the problem-solving process of geometry problems and obtain the solution results of geometry problems quickly and accurately.

[0047] In some embodiments, before performing the above step 102, the neural network is first trained to obtain a geometric problem-solving model, and the process includes:

[0048] Step a1, obtaining a geometry problem dataset Geometry3K, wherein the geometry problem dataset includes a plurality of geometry problem samples.

[0049] Step a2, using the geometry problem dataset Geometry3K to train the pre-built deep neural network, and constructing an optimal solution theorem sequence corresponding to the geometry problem dataset through a theorem predictor.

[0050] Step a3, using the predictor to predict the next theorem sequence of the current sequence, determining the optimized negative log-likelihood loss according to the predicted theorem sequence, and then adjusting the parameters of the deep neural network according to the negative log-likelihood loss, completing the training to obtain the geometric problem-solving model.

[0051] In recent years, some datasets of geometric problems have been released, such as GEOS (Seo et al., 2015), GEOS++ (Sachan et al., 2017), GeoShader (Alvin et al., 2017), and GEOS-OS (Sachan and Xing, 2017). These datasets are relatively small in size and contain limited types of problems. The GEOS dataset contains only 186 problems, while the GeoShader dataset only uses 102 shading region problems. There are also some datasets that contain more problems and data, but most of them have not yet been made public, such as the GEOS++ dataset and the GEOSOS dataset.

[0052] A new large-scale benchmark dataset Geometry3K was constructed by a joint research team from UCLA, Zhejiang University, and Sun Yat-sen University. The dataset contains 3,002 geometry problems and is densely annotated using a formal language. The data in the dataset comes from two popular textbooks written for high school students in grades 6-12 in North America by two online digital libraries (McGraw-Hill[2], Geometryonline[3]). It is very suitable for training geometry problem-solving models.

[0053] The geometric problem is formally described as L = {l1,…,lm}. The goal of the theorem predictor is to reconstruct the optimal solution theorem sequence T = {t1,…,tn} one by one. The generated theorem sequence is labeled, and the predictor predicts the next theorem ti, given T = {t1,…ti}. As shown in Formula 3-4, the sequence-to-sequence (Seq2Seq) model is trained to optimize the negative log-likelihood loss,

[0054] Through the above scheme, it can be ensured that the trained geometry problem-solving model has higher problem-solving accuracy and better precision.

[0055] In some embodiments, step A2 comprises:

[0056] Step A21, after the search question is divided into tuples, the tuples include: picture element information and the text element information and numerical format.

[0057] Step A22, convert the text element information in the tuple into text through a deep neural network model to obtain the text part of the geometry problem text, identify the image element information in the tuple, and construct relationships based on the recognition results to obtain a relationship set R.

[0058] Step A23, searching the relationship set R using a theorem search strategy to determine an optimal solution sequence, and determining a geometric answer and a geometric problem-solving process as the problem-solving result based on the optimal solution sequence.

[0059] Through the above scheme, the geometry problem input by the user can be processed to construct the corresponding relationship set R, and each sequence in the relationship set R is solved using the theorem search strategy, and then it is determined whether there is a new solution sequence in the relationship set R. If not, it is proved that the problem is solved and the optimal solution sequence is obtained; if so, the new solution sequence is recorded and the relationship set R is solved again using the theorem search strategy until there is no new solution sequence in R, and finally the optimal solution sequence is obtained. Finally, the geometric answer and the geometric problem-solving process determined by the optimal solution sequence are presented to the user as the problem-solving result output.

[0060] In some embodiments, step A22 includes:

[0061] Step A221, converting the words of the text element information in the tuple into predicates and variables through a deep neural network, and forming a text sequence with the predicates and variables.

[0062] Step A222, extracting geometric elements from the image element information in the tuple to obtain an element set P.

[0063] Step A223, extracting the icon symbols and text areas in the geometric elements through a strong object detector, recognizing the text information in the text area through an optical character recognition tool, and obtaining a symbol set S.

[0064] Step A224, determining each symbol S in the symbol set S i and elements in the set P and S i The corresponding geometric element p j The geometric relationship data F, for each s i and p j To associate, the association algorithm formula is:

[0065] Where dist is S i and p j The Euclidean distance between them, i is the symbol S i The order in the symbol set, j is the geometric element p j In the sorting of the element set P, i and j are both positive integers;

[0066] Step A225, convert the associated results into a quasi-transposition formal language according to the rules to obtain the relationship set R.

[0067] Through the above scheme, it can be ensured that the obtained relationship set R is more accurate and convenient for subsequent problem solving through theorem search strategy.

[0068] In some embodiments, step A23 includes:

[0069] Step A231, expanding the relationship set R into a geometric formal language.

[0070] Step A232, obtaining a theorem set KB, wherein each theorem k in the theorem set KB i Defined as a conditional rule with premise a and conclusion q.

[0071] Step A233, determine theorem k at time sequence t i The premise a and the relationship R currently selected from the relationship set R t-1 Matching, according to the theorem k i Conclusion q updates the relation set R t .

[0072] Step A234, based on R t An equation relationship between known values ​​and unknown problem targets is constructed, and the equation is solved to obtain a geometric answer and a geometric problem-solving process as the problem-solving result.

[0073] Through the above scheme, the obtained relationship set R can be sorted out using theorem set KB to sort out the theorem relationships, and then gradually build the equation relationship between known values ​​and unknown problem targets. In this way, the geometric graphic problem can be converted into an equation solving problem, and the equation can be solved according to the theorem relationship to obtain the geometric answer. The problem-solving process of obtaining the geometric answer can be sorted out and integrated with the geometric answer and the problem-solving process as the problem-solving result.

[0074] In some embodiments, the algebra problem solver includes: an algebra problem-solving model obtained by training a neural network with algebra problem-solving samples;

[0075] The step 102 further includes:

[0076] Step B1, in response to determining that the target problem solver is an algebraic problem solver, receiving the search problem, processing the search problem using a bidirectional long short-term memory model to obtain a key variable generation node, and obtaining internal representation parameters of the algebraic sample problem using a graph converter according to a quantity comparison graph and a quantity unit graph.

[0077] Step B2, inputting the key variable generation node and the internal representation parameters into the algebraic problem solver, performing algebraic problem solving processing through an algebraic problem solving model, obtaining an algebraic solution expression, and using the algebraic solution expression as the problem solving result.

[0078] First, the mathematical word problem queried by the user is read in, and the BiLSTM neural network model is used to initialize the nodes for the text. Then the NLP tool stanford corenlp toolkit is used to analyze text dependencies, regional analysis, and establish quantity unit graphs and quantity comparison graphs. The established nodes, quantity unit graphs, and quantity comparison graphs are used as the input of the graph encoder, and converted into a global context representation graph through the graph convolutional network model in the encoder, and used as the input of the tree decoder. The decision tree is recursively established in a pre-order manner. Finally, the nodes are output by the pre-order traversal method to obtain the solution expression, which is fed back to the user. In this way, algebraic problems can be solved smoothly, and the accuracy and speed of the solution are effectively improved.

[0079] In some embodiments, before implementing the above step 102, the neural network is first trained to obtain an algebraic problem-solving model, and the process includes:

[0080] Step b1, obtaining an algebraic problem data set, wherein the algebraic problem data set includes a plurality of algebraic problem samples.

[0081] Step b2, using the algebraic problem data set to perform learning and training on the pre-constructed deep neural network, determining the sum of the negative log-likelihood of the probability of the predicted node during the training process as the algebraic training loss function, adjusting the parameters of the deep neural network based on the algebraic training loss function, completing the training to obtain the algebraic problem-solving model.

[0082] The spanning tree definition of each algebraic problem P is expressed as (P, T), and the prediction node T is defined as The sum of the negative log-likelihood of the probability is the loss function L(T,P). The goal of training is to minimize the following loss function:

[0083] where q t is the target vector, G c is the representation graph of the global context, E is the number of tags in T, and prob is calculated by the distribution calculation function of GTS.

[0084] The parameters of the neural network are adjusted through the loss function obtained above, and then the training process is completed to obtain an algebraic problem-solving model. This method enables the obtained algebraic problem-solving model to accurately solve algebraic problems and ensure the speed and accuracy of the solution.

[0085] In some embodiments, step 102 includes:

[0086] Step B1', receiving an algebraic problem, and processing the algebraic problem using a bidirectional long short-term memory model to obtain an algebraic problem node representation, wherein the algebraic problem node representation includes algebraic words and algebraic variables.

[0087] Step B2', constructing the nodes of the graph according to the algebraic words and all algebraic variables in the algebraic problem node representation.

[0088] Step B3', establishing a quantity unit diagram and a quantity comparison diagram according to the algebraic problem.

[0089] Step B4', using a graph converter with a multi-head structure to evenly distribute the quantity units and the quantity comparison graph to obtain a root node.

[0090] Step B5', generating a decision tree based on the root node, decoding the data in the root node according to the decision tree to obtain a decoding sequence, and traversing the decoding sequence to obtain a solution result.

[0091] The bidirectional long short-term memory model BiLSTM is used to initialize the question text searched by the user to generate nodes from the key variables in the question text, and then the quantity comparison graph (Quantity ComparisonGraph) and quantity cell graph (Quantity Cell Graph) are constructed according to the quantity units in the text. For these two graphs and nodes, the constructed graph converter is used to obtain the internal representation and implement the graph encoding (Graph Encode). The graph converter is actually a deep learning model containing a graph convolutional network. Finally, the decoder (Tree-Based Decode) based on the constructed tree is used to generate the solution expression.

[0092] Finally, the solution result is fed back to the user for review.

[0093] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.

[0094] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0095] Based on the same inventive concept, a specific embodiment of a processing method of a mathematical problem solver based on deep learning is described below. The details are as follows:

[0096] 1. Problem Solver Subsystem Requirements Analysis

[0097] This step mainly discusses the requirements analysis of the problem solver subsystem. First, we start with the business requirements of searching and solving problems in the online education system, determine the functional modules required by the system through analysis, build the overall framework of the system using the use case diagram, and introduce the brief schematic diagram of each functional module in the system framework, as well as analyze and summarize the non-functionality of the system.

[0098] 1.1 Geometry / Algebra Problem Solver System Requirements Analysis

[0099] The problem solver system mainly analyzes the questions searched by the user. The user must first choose a geometry solver or an algebra solver. Then the system calls the deep learning model to analyze the questions and picture information provided by the user, and outputs the reasoning steps for solving the questions. For geometry questions, it will show what type of question the question belongs to, including circle, rectangle, square, parallelogram, etc. At the same time, the type of answer required by the question will be determined and displayed, such as the answer type is length, surface diameter, angle, etc. The system will also include the questions queried this time, and the user can directly display them next time he wants to find the question. In addition, the system also has comments, which can provide feedback on whether the system solves the problem correctly.

[0100] System use case diagram Figure 2-A shown.

[0101] 1.1.1 Geometry Problem Solver (i.e., Geometry Solver)

[0102] When the user chooses to search for a geometry problem, the system extracts information from the problem the user searches for, uses the extracted information and axiom theorem set as the input of the deep learning model, outputs a predicted theorem sequence, and brings in the corresponding information. The input and output diagram of the geometry problem solver model (i.e., geometry problem solving model) is shown in the figure below. Figure 2-B shown.

[0103] 2.1.2 Algebraic Problem Solver (i.e., Algebraic Problem Solving Model)

[0104] When the user chooses to search for algebraic problems, the system initializes the question text that the user searches for and uses BiLSTM to generate nodes from the key variables in the question text. It then constructs a quantity comparison graph (QuantityComparison Graph) and a quantity cell graph (Quantity Cell Graph) based on the quantity units in the text. For these two graphs and nodes, the constructed graph converter is used to obtain the internal representation and implement graph encoding (Graph Encode). The graph converter is actually a deep learning model that includes a graph convolutional network. Finally, the decoder (Tree-Based Decode) based on the constructed tree is used to generate the solution expression. The schematic diagram of the algebraic problem solver model (i.e., algebraic problem solving model) is shown below. Figure 2-C shown.

[0105] 1.2 System functional requirements analysis

[0106] Combined with the functional description in the business requirements analysis in 1.1, it can be seen that the system has two major functional modules:

[0107] Geometry problem solver module: For the graphics of geometry problems, Hough Transform is used to extract the geometric elements in the graphics. Then the target detection model RetinaNet is used to extract the symbols and text areas in the image. These areas are further identified and extracted by the OCR tool MathPix. Use these elements to establish a relationship set to complete the modeling task. Finally, the theorem predictor is called to select the correct theorem sequence with accurate order in the axiom and theorem set based on the relationship set. Since there are many types of graphics in geometry problems, the axioms and theorems involved in different types will be different, and the answers to the questions also include different types. In order to help users deepen their memory, this module also needs to feedback the basic graphics type of the question and the type of answer required to the user.

[0108] Algebra problem solver module: For algebra problems, a Quantity Cell Graph and a Quantity Comparison Graph are established based on the text. After modeling is completed, the encoder of the graph containing the graph convolutional network model is called on the model, the output is passed to the tree decoder, a decision tree is generated, and the solution expression is output.

[0109] 1.3 System Non-functional Requirements Analysis

[0110] From the perspective of system design, the geometric algebra problem solver subsystem needs to have the following functional requirements:

[0111] (1) Accuracy: Users search for questions and hope to get a correct answer. Therefore, the accuracy of the deep learning model is the cornerstone of the successful construction of the geometry and algebra problem solver subsystem. It is of utmost importance and needs to maintain a certain accuracy. The geometry problem solver is required to achieve an expected accuracy of 40%, and the algebra problem solver is required to achieve an expected accuracy of 33%.

[0112] (2) Ease of use: The users of the geometry and algebra problem solver subsystem are mainly students. It needs to be simple, convenient and easy to use. It should not have a high cost of use, and users should be able to use the subsystem clearly.

[0113] (3) Efficiency: The training and problem-solving process of the geometric algebra problem solver subsystem must ensure a certain degree of efficiency. In particular, when a user searches for a problem, the problem-solving process should be completed in a relatively short period of time, and the resulting problem-solving process should be fed back to the user.

[0114] (4) Robustness: A good system must have good robustness and be able to correctly handle and respond to errors and problems. The geometric algebra problem solver subsystem is required to be able to handle unexpected situations and respond to user comments on the correctness of the answers to ensure the correct operation of the subsystem.

[0115] (5) Visual appeal: The front-end interface of the geometric algebra problem solver subsystem adopts a simple design. In addition to ensuring the basic information feedback of the required results, it also adds additional functions to help users deepen their understanding of the questions. The page is in the form of a web page, which is easy for users to use and operate, and ensures the viewing comfort of the page.

[0116] 1.4 Summary

[0117] The first part discusses the requirements analysis of the geometry and algebra problem solver subsystem from the perspective of software design. It conducts a detailed analysis of the functionality and non-functionality of the subsystem, describes the overall business requirements of the subsystem through a use case diagram, and uses a flowchart to explain in detail the flow charts of the two main modules: the geometry problem solver module and the algebra problem solver module.

[0118] 2. Overview of the Problem Solver Subsystem Design

[0119] The main content of the second part is to analyze the design of geometry problem solvers and algebra problem solvers respectively. First, the data sets used for training, testing and verification are introduced, the test questions of the two solvers are analyzed and modeled, and various operators in the questions are customized to achieve the extraction of key information in the user's search questions. The core of the solver model design is the deep learning algorithm. This chapter will analyze the deep learning algorithms of the two solvers, detailing the theorem predictor constructed in this project, the graph convolutional network algorithm model used, and how the encoding of the graph and the decoding of the tree are completed. Then, an improved algorithm is proposed in this chapter, and the effect of the improved solver is tested and verified, and the implementation effect of this article in the problem-solving process of the solver is given.

[0120] 2.1 Geometry Problem Solver

[0121] 2.1.1 Geometry3K Dataset

[0122] Some datasets of geometric problems have been released in recent years, such as GEOS (Seo et al., 2015), GEOS++ (Sachan et al., 2017), GeoShader (Alvin et al., 2017) and GEOS-OS (Sachan and Xing, 2017) datasets. These datasets are relatively small in size and contain limited types of problems. There are only 186 problems in the GEOS dataset, and the GeoShader dataset only uses 102 shading area problems. There are also some datasets that contain more problems and data, but most of them have not yet been made public, such as the GEOS++ dataset and the GEOSOS dataset. This paper uses an open source dataset Geometry3K that contains 3002 geometric problems. The comparison of Geometry3K with other datasets is shown in Table 2-1.

[0123] Table 2-1 Comparison of Geometry3K and other datasets

[0124]

[0125] This paper uses a new large-scale benchmark dataset Geometry3K built by a joint research team from UCLA (University of California, Los Angeles), Zhejiang University, and Sun Yat-sen University. The dataset contains 3002 geometry problems and is densely annotated using a formal language. The data in this dataset comes from two popular textbooks written by two online digital libraries (McGraw-Hill[2], Geometryonline[3]) for high school students in grades 6-12 in North America. It is very suitable for training geometry problem solver models. Therefore, this paper uses this dataset as the dataset for geometry problem solvers. The basic statistical information of the Geometry3K dataset is shown in Table 2-2.

[0126] Table 2-2

[0127]

[0128]

[0129] This paper divides the 3002 questions in the Geometry3K dataset into training, validation and test sets in a ratio of 7:1:2.

[0130] The Geometry3K dataset contains 6293 question texts and 27213 diagrams. The number of words contained in different questions is also different, such as Figure 2-D As shown, the question distribution is illustrated by the number of words in the sentence.

[0131] 2.1.2 Extraction and modeling of problem elements

[0132] 2.1.2.1 Geometric formal language

[0133] Let's start with the definition. First, define the geometry problem P as a tuple (T, d, c), where T refers to the text part of the geometry problem, d refers to the image part of the geometry problem, and c = {c1, c2, c3, c4} refers to the multiple choice candidate set in numerical format. For a given text T and image d, an algorithm is needed to predict the correct answer c. i ∈c. In the geometry problem solver, a set of literals consisting of predicates and parameters is defined to describe the geometry problem Ω. The basic terms defined are as follows:

[0134] Definition 1: A predicate refers to the geometry, geometric relationship, or operator of a problem.

[0135] Definition 2: A literal is the application of a predicate to a set of parameters such as variables or constants. A set of literals constitutes the semantic description of the formal language space problem text and diagrams Ω.

[0136] Definition 3: A primitive is a basic geometric element, such as a point, line segment, circle, or arc segment extracted from a diagram. Tables 2-3 to 2-8 show the 91 predicates in the formal language defined.

[0137] Table 2-3 Geometric shapes

[0138]

[0139] Table 2-4 Unary geometric attributes

[0140]

[0141] Table 2-5 Geometric attributes

[0142]

[0143] Table 2-6 Binary geometric relations

[0144]

[0145]

[0146] Table 2-7 A-isXOf-B geometric relationship

[0147]

[0148] Table 2-8 Numerical attributes and relationships

[0149]

[0150] 2.1.2.2 Element extraction

[0151] For the text description part of the geometry problem, that is, the text T. The text parser converts the word sequence of T into a set of texts LT, that is, a sequence consisting of predicates and variables. The conversion from sequence to sequence (Seq2Seq) is completed using a deep neural network model.

[0152] For the diagram part of the geometric problem, the diagram is first subjected to Hough transform (Shapiro and Stockman, 2001) to extract the geometric elements in the diagram, namely points, lines, arcs and circles. Then, the strong object detector RetinaNet (Lin et al., 2017) is used to extract the diagram symbols and text areas, and the optical character recognition tool MathPix is ​​used to further identify the text content. In this way, the element set P and the symbol set S are obtained. Figure 2-E Shown are the element set P (left) and the symbol set S (right).

[0153] In this way, the key elements and symbols of the text and diagram parts of geometry problems are obtained.

[0154] 2.1.2.3 Element Modeling (Establishing Relationship Sets)

[0155] In the previous section, we obtained the element set P and the symbol set S. Next, we need to model these element sets to obtain a new set.

[0156] This article needs to connect each symbol with its associated element. Here, the connection task is expressed as an optimization problem with geometric relationship constraints. The algorithm here is shown in Formula 2-1.

[0157]

[0158] In the above formula, dist is s i and p j The Euclidean distance between, i is the symbol s i The order in the symbol set, j is the geometric element p j In the ordering of the element set P, i and j are both positive integers, and F defines the geometric or algebraic relationship that constrains the positioning of symbols. Finally, the associated elements and symbols are converted into the final formal language expression through simple rules to construct the relationship set R.

[0159] 2.1.3 Geometry Symbol Parser

[0160] The core part of the geometry problem solver is the symbol parser. The relation set R contains the geometric properties and relations in the geometry problem, which is initialized using the text in the text and diagram parser. The relation set R is further expanded to define the geometric formal language. For example: Triangle(a, B, C), after the expansion, six predicates are appended to R (Point(A), Point(B), Point(C), Line(A, B), Line(B, C), Line(C, A)).

[0161] Theorem set KB is represented as a set of theorems, where each theorem k in KB i are defined as a conditional rule with premise p and conclusion q. For search step t, if k i The premise p and the current relation set R t-1 If the relationship is matched, update the relationship set R according to the conclusion q of the theorem, as shown in Formula 2-2.

[0162] R t ←k i ∧R t-1 , k i ∈KB (2-2)

[0163] After applying multiple theorems, we can establish an equation between a known value and the unknown problem target g. By solving this equation, we can get the solution g, as shown in Formula 2-3.

[0164] g * ←SOLVEEQUATION(R t , g) (2-3)

[0165] 2.1.4 Theorem Predictor

[0166] Theorem predictor is an important part of symbolic parser. In the previous section, we constructed equations between unknown problem target g and known values ​​by applying multiple theorems. Now we can get the desired result by solving the equation to get g.

[0167] Most of the problems in the dataset Geomety3K require the application of multiple theorems to establish equations between known values ​​and the problem objectives. How to select the applied theorems and the order in which the theorems are applied are urgent issues to be solved. Therefore, a theorem predictor is constructed to complete this task, and the applied theorem sequence and theorem order are recorded as the output of the solver and fed back to the user.

[0168] There are many theorem search algorithms. This paper randomly samples from the theorem set multiple times to generate a sequence of applied theorems. If the solver solves the problem after applying the sequence, the generated sequence is considered a solution sequence. A geometric problem may have multiple solutions, so multiple solution sequences of the geometric problem are obtained, and the solution sequence with the minimum length is regarded as the optimal solution sequence.

[0169] The geometric problem is formally described as L = {l1,…,l m}, the goal of the theorem predictor is to reconstruct the optimal solution theorem sequence T = {t1,…,t n The generated theorem sequence is labeled, and the predictor predicts the next theorem ti, given T = {t1,…t i}. As shown in Formula 2-4, the sequence-to-sequence (Seq2Seq) model is trained to optimize the negative log-likelihood loss.

[0170]

[0171] 2.1.5 Pseudocode and Process of Geometry Problem Solver

[0172] PTP is a parameterized conditional distribution in the theorem predictor model.

[0173] The overall process of the geometry problem solver is to first read the geometry problem P(T, d, c) input by the user, then the text parser and the diagram parser extract the elements and symbols respectively, and establish a relationship set R, and expand R using the defined predicates. Find all the solution theorem sequences of R, determine the optimal solution sequence, and output the answer and process.

[0174] 2.2 Algebraic Problem Solver

[0175] The framework diagram of the algebra problem solver is as follows Figure 2-C As shown in Figure 2, this section will analyze the model in detail. The model is divided into two modules: the graph encoder and the tree decoder. The data modeling, construction ideas, and application algorithms of these two modules will be introduced in turn.

[0176] 2.2.1 Graph Encoder

[0177] 2.2.1.1 Initialization of Node Representation

[0178] This paper uses BiLSTM neural network to learn the mathematical word problems input by users, that is, algebraic problems, and obtains the hidden word-level state representation in the text description of the algebraic problem, and obtains H = {h1,…,h N}∈R N×d , N = m + l. Among them, m represents the number of words, l represents the number of variables, and d represents the dimension of the hidden vector.

[0179] This results in a node representation of the algebraic question text, which will serve as the input of the subsequent graph encoder.

[0180] 2.2.1.2 Definition of quantity unit

[0181] The words and all variables in the algebraic problem are taken as nodes of the graph to be constructed. Define the quantity unit as a subset of the set of nodes associated with the quantity in the graph. For an algebraic problem P, it can be converted into multiple quantity units QC = {Q1, Q2, ..., Q m}, the m here is different from the m in the previous section. The m here refers to the number of quantities in problem P. Each quantity unit Q i ∈QC contains a quantity tag {n i} and the attribute corresponding to the quantity unit {v 1i ,…,vq i}, This paper uses an open source NLP tool Stanford CoreNLP Toolkit (Manning et al., 2014) to complete the dependency parsing, constituency parsing and part-of-speech tagging of the algebra problem text to extract and construct quantity units. The final quantity unit is a subgraph of quantity-related information in the algebra problem.

[0182] The quantity unit in the Algebra Problem Solver contains the following five types of properties:

[0183] (1) Quantity: The numerical value of a variable.

[0184] (2) Associated Nouns: Nouns that are linked by the quantity, number, and preparatory relationship of the relationship.

[0185] (3) Associated Adjectives: adjectives that indicate something related to quantity.

[0186] (4) Associated Verbs: For each quantity, detect the associated verbs based on the relationship between nsubj and dobj.

[0187] (5) Units and Rates: Keywords with the meaning of "each", "every", "per", etc., if associated with nouns, are rates.

[0188] If a quantity cell does not contain any attributes, a window centered on the quantity is selected to select adjacent words as the attributes of the quantity.

[0189] 2.2.1.3 Construction of quantity unit diagram and quantity comparison diagram

[0190] We have already obtained the quantity unit. In this section, we will use the quantity unit to construct two important graphs: the Quantity Cell Graph and the Quantity Comparison Graph. The Quantity Cell Graph contains the association between quantity and informative words, and the Quantity Comparison Graph retains the numerical value of the quantity and uses heuristic methods to improve the relationship between quantities.

[0191] The structures of the two graphs are as follows:

[0192] Quantity Unit Chart G qcell :For each quantity unit Q i ={n i}∪{v 1i ,…,v ni}, where ni and each v j ∈{v 1i ,…,v qi The undirected edge e between ij Add to Figure G qcell middle.

[0193] Quantity comparison chart G qcomp :For any two number nodes n i , n j ∈n P , if n i >n j , then from n i Point to n j The directed edge e ij =(n i , n j ) Add to Figure G qcomp Using this heuristic constraint prevents the situation where a negative number is produced by subtracting a larger number from a smaller number.

[0194] These two graphs can be represented by adjacency matrices: the quantity unit graph and the quantity comparison graph. For each graph, first create an initial adjacency matrix A∈R N×N For two endpoints i and j of an existing edge, in the adjacency matrix (i, j, A i.j ) is set to 1. Otherwise, it is set to 0.

[0195] Through the above method, comp Create an adjacency matrix A qcomp , for graph G qcell Created Aqcell.

[0196] 2.2.1.4 Graph Transformer Based on Graph Convolutional Network

[0197] The graph conversion module is a core module of the graph encoder. The input of this module is multiple graphs. A k ∈{A qcomp , A qcell} and the node H initialized in 1.2.1.1, where K is the number of graphs, each A k ∈R N×N is the adjacency matrix of the kth graph. The graph converter module uses K graphs, so that the multi-head structure adopted can make them evenly distributed between quantity unit graphs and quantity comparison graphs.

[0198] The graph transformer module first uses graph convolutional networks (GCNs) (Kipfand Welling, 2017) to learn the features of graph nodes. Since there are multiple graphs, a K-head graph convolution setting is used. This is similar to the transformer model proposed by Vaswani et al. (2017), which uses K independent graph convolutional networks that need to be connected before applying the residual connection.

[0199] A single GCN has parameters Where dk = d / K. Given the adjacency matrix A representing the graph structure k And the feature matrix X (x is initially set to H) representing the input features of all nodes, the learning definition of GCN is as follows:

[0200] GCN(A k , X) = GConv2(A k , GConv1(A k ,X)) (2-5)

[0201] Here, GCN contains 2 different graph convolution operations:

[0202] GConv(A k , X) = relu(A k X T W gk ) (2-6)

[0203] For each graph Execute graph convolutional network learning tasks in parallel to generate dk-dimensional output values. The output values ​​are concatenated and projected to produce the final value:

[0204]

[0205] Here || represents K GCN head connections.

[0206] The graph converter module also uses a feed-forward network structure, i.e. a positive feedback network, layer-norm layer technology and residual connection to expand K graph convolutional networks:

[0207]

[0208]

[0209] Among them, FFN(x) is a two-layer feedforward network with a relu function between each layer:

[0210] FFN(x)=max(0,xW f1 +b f1 )W f2 +b f2 (2-10)

[0211] Generated node representation Represents quantities, entities, and relations. In order to learn a graph representation of the global context, an element-wise min-pooling operation is applied to all learned node representations. Finally, the global features are fed back to a fully connected neural network (FC) to generate a feature map z g :

[0212]

[0213] like Figure 2-F As shown, a flow chart of the encoder of FIG. 1 is shown.

[0214] 2.2.2 Tree Decoder

[0215] A tree decoder is needed to decode the input from the graph encoder to generate the solution expression. This paper sets each quantity as a leaf node, where each operator must have two child nodes. After the tree is built, an equation can be generated by traversing it in order. The tree construction must first generate the most central operator, and recursively generate the correct child nodes according to the pre-order method.

[0216] 2.2.2.1 Tree Initialization

[0217] The tree decoder first initializes the root node vector q according to the global context representation graph zg obtained in the previous section root For the target word list V of question P dec For each token y in , define a specific token e(y|P) as:

[0218]

[0219] The expression tree in the decoder contains three kinds of nodes: operators, constants, and quantities in the problem P. Constants and quantities in np are always set at leaf node positions. Operators will always occupy non-leaf node positions. p The representation of the quantities in depends on algebraic problems. Take out the corresponding The representations of operators and constants are independent, and they are represented by two independent embedding matrices M op and M com Got it.

[0220] 2.2.2.2 Pre-order decision tree generation

[0221] This article uses the pre-order method to generate a decision tree. The specific steps are as follows:

[0222] First, generate only the root node q root The node embedding is done using the attention module in GTS. Encoded into the global graph vector G c middle:

[0223]

[0224] The decoder of this tree applies the left child node generation module to the derivation in a top-down manner, with the parent node q p and the global graph G c Generate a new left child node q for the condition l , when a new node is generated, a prediction is made

[0225]

[0226] If you continue with this step, it works similarly to breaking down the entire objective into multiple stages of reasoning. is a quantity (either a constant or a quantity derived from n P ) and proceed to the next step.

[0227] The decoder of the tree switches to using the right child node generation module and fills the empty right node position. In each decoding step, the left child node q1, the global graph vector G c The subtree embedding t1 is used as the input of the right generation module and generates the right child node q r and the corresponding

[0228]

[0229] Adding subtree embedding works similarly to the merge subtree replication mechanism.

[0230] The tree embedding component is used to compute the additional subtree embedding t1:

[0231]

[0232] if is an operator, then you should return to the previous step. is a quantity that will proceed to the next step.

[0233] The model switches to backtracking to find a new empty right node position. If the model cannot find a new empty right node position, generation is complete. If an empty right node position still exists, it returns to the above generation of a new node and makes a prediction in the corresponding steps.

[0234] An example diagram of a decision tree encoder is shown below: Figure 2-G .

[0235] 2.2.2.3 Model Learning

[0236] The spanning tree definition for each algebraic problem is expressed as (P, T), and the prediction node T is defined as The sum of the negative log-likelihood of the probability is the loss function L(T,P). The goal of training is to minimize the following loss function:

[0237]

[0238] where q t is the target vector, G c is the representation graph of the global context, E is the number of tags in T, and prob is calculated by the distribution calculation function of GTS.

[0239] 2.2.3 Dataset

[0240] Two commonly used algebra problem datasets, MAWPS (Koncel-Kedziorski et al., 2016) and Math23K (Wang et al., 2017), were used. The MAWPS dataset contains 2,373 problems and the Math23K dataset has 23,162 problems.

[0241] 2.2.4 Flowchart of the Algebraic Problem Solver Model

[0242] The specific process of the algebra problem solver is as follows: the algebra problem solver first reads the mathematical word problem queried by the user, and uses the BiLSTM neural network model to obtain the initialized nodes for the text. Then the NLP tool stanford corenlp toolkit is used to analyze text dependencies and regional analysis, and to establish a quantity unit graph and a quantity comparison graph. The established nodes, quantity unit graphs, and quantity comparison graphs are used as the input of the graph encoder, and converted into a global context representation graph through the graph convolutional network model in the encoder, and used as the input of the tree decoder. The decision tree is recursively established in a pre-order manner, and finally the nodes are output through the pre-order traversal method to obtain the solution expression, which is fed back to the user.

[0243] 2.3 Algorithm Improvement

[0244] 2.3.1 Improvements to the Geometry Solver Text Parser

[0245] In the original algorithm, the text parser converts the word sequence of T into a set of text LT, which is a sequence of predicates and variables. The conversion from sequence to sequence (Seq2Seq) is completed using a deep neural network model. However, this method cannot generate very successful text on the Geometry3K dataset.

[0246] The specific reason may be that the limited rules of geometry datasets such as Geometry3K weaken these highly data-driven methods. In addition, neural semantic analysis tends to introduce noise in the generated results, and geometry solvers with symbolic reasoning are very sensitive to such deviations.

[0247] The improved method is to apply the rule-based parsing method together with regular expressions to the parsing of text descriptions of geometric problems, and use BART (Lewis et al., 2020) to complete the construction of the text parser.

[0248] 2.3.2 Improvement of the search strategy of theorem predictor

[0249] The theorem search strategy used in the original theorem predictor is mainly a low-order theorem search strategy, which is a significant improvement over the brute force enumeration method, but there are many other search strategies that can be tried, each with its own advantages and disadvantages.

[0250] The improved search strategies that can be used are: Random: randomly apply theorems in the theorem set; Low-first: in each round of search, low-order theorems are used first; Predict: first apply the predicted theorems, then randomly apply theorems in the theorem set; Final: first apply the predicted theorems, then give priority to low-order theorems.

[0251] 2.4 Summary

[0252] This chapter is the most important chapter in the realization of the system core and algorithm design of the geometry and algebra problem solver subsystem in this paper. Starting from the problem definition, we first think about the current status of related research and select suitable data sets for the two models. The geometry problem solver selects the Geometry3K data set, and the algebra problem solver selects the Math23 and MAWPS data sets. The training, testing and verification results of the two data sets are compared respectively. These results will be presented in the subsequent process.

[0253] Then, according to the analysis of the schematic diagrams of the two solver models, the text parser, symbol parser module, theorem predictor module of the geometry problem solver and the node initialization module, quantity graph construction module, graph encoder module and tree decoder module of the algebra problem solver are explained in turn. The relevant pseudocode and algorithm flow chart are also given.

[0254] Finally, improvement plans and introductions are proposed for some algorithms.

[0255] Based on the same inventive concept, a problem solving system is proposed corresponding to the method of the above embodiment, including:

[0256] The user layer is configured to receive search question types and display solution results;

[0257] The functional logic layer is configured to determine a target problem solver from the geometry solver and the algebra solver of the math problem solver according to the problem type, receive the search question, and feed back the problem solving result to the user layer;

[0258] The function execution layer includes a geometry problem solver and an algebra problem solver, and is configured to pre-process the search problem to obtain problem information, input the problem information into the target problem solver, perform problem solving processing on the problem information through the problem solving model in the target problem solver, obtain a problem solving result, and feed the problem solving result back to the function execution layer; wherein the problem solving model is obtained by training a neural network with problem solving samples;

[0259] The storage computing layer is configured to store the data sets required by the geometry problem solver and the algebra problem solver, so that the geometry problem solver and the algebra problem solver can retrieve corresponding data from the data sets for problem solving when performing problem solving.

[0260] For the convenience of description, the above devices are described in terms of functions and are divided into various modules / layers / units. Of course, when implementing the present application, the functions of each module / layer / unit can be implemented in the same or multiple software and / or hardware.

[0261] The device of the above embodiment is used to implement the corresponding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0262] 3. Overall Design of the Problem Solver System

[0263] According to the requirements analysis mentioned in the first part, the system should have a simple UI interactive interface that is convenient for users to use, and can efficiently respond to user requests and give feedback. At the same time, the expected accuracy must be guaranteed.

[0264] The geometry problem solver based on the graph convolutional network model question analyzer and theorem predictor can achieve 40% accuracy when using the dataset Geometyr3K. The algebra problem solver based on the graph decoder model of the graph convolutional network and the tree decoder model can achieve 33% accuracy for most datasets.

[0265] In addition, this subsystem also implements extended functions such as user comments.

[0266] 3.2 System Implementation

[0267] like Figure 3-A As shown, the functions of each layer of the problem solver subsystem (i.e., problem solving processing system) are as follows:

[0268] User layer: The UI interface of the problem solver subsystem provides users with simple and convenient interaction.

[0269] Functional logic layer: encapsulates the main functions implemented, provides functional interfaces, and the UI interface calls these interfaces to implement the functions the user wants.

[0270] Functional execution layer: contains two core modules, the geometry solver module and the algebra solver module, for the calling business of the functional logic layer.

[0271] Storage and computing layer: contains the dataset and saved user comments, and directly interacts with the function execution layer.

[0272] like Figure 3-B As shown, the entire business process of the user in the process of using the subsystem is described.

[0273] 3.3 Tool Introduction

[0274] Genism: A simple and efficient natural language processing Python library for extracting semantic topics from documents. Gensim's input is raw, unstructured digital text (plain text). Built-in algorithms include Word2Vec, FastText, Latent Semantic Analysis (LSA), Latent Dirichlet Allocation (LDA), etc., which automatically discover the semantic structure of documents by calculating statistical co-occurrence patterns in the training corpus. These algorithms are all unsupervised, which means that no human input is required - only a set of plain text corpus is required. Once these statistical patterns are discovered, any plain text (sentences, phrases, words) can be concisely expressed using semantic representations.

[0275] Torch: An optimized tensor library for deep learning using GPUs and CPUs.

[0276] Sklearn: A machine learning library based on Python language, with: simple and efficient data analysis tools, reusable in multiple environments, built on data science libraries such as Numpy, Scipy and matplotlib, open source and commercially available - based on BSD license 4.4 system operation process and result display.

[0277] NLTK: NLTK is the leading platform for building Python programs to process human language data. It provides an easy-to-use interface to more than 50 corpora and lexical resources (such as WordNet), as well as a suite of text processing libraries for classification, tokenization, stemming, tagging, parsing, and semantic reasoning, wrappers for industrial-strength NLP libraries, and an active discussion forum.

[0278] Since the core algorithm of this system is time-consuming and labor-intensive, the UI interface is relatively simple and the functions implemented are limited. However, the core functions and the necessary functions in the previous section are successfully implemented. The system's implemented functions are demonstrated below.

[0279] The problem is displayed through the display interface, and the text of the question is displayed in a formal language. Then on the right side of the picture, the extended functions are displayed to show the user the additional text geometric elements in the question and the form of text parsing. The geometric elements in the chart include the basic graphic types and variables.

[0280] The system provides feedback on the answers to the questions asked by the user.

[0281] In addition to providing feedback on the answer, the subsystem also provides feedback on the basic graphic type of the question and the type of answer required for geometry problems.

[0282] The subsystem implements the user comment function for whether the answer is correct or not.

[0283] In the above demand analysis, it is mentioned that this system implements the function of viewing historical questions, which can facilitate users to view the questions they searched in the past.

[0284] 3.5 Summary

[0285] This section first reviews the demand analysis of the geometric algebra problem solver subsystem, and draws the design ideas of the subsystem module based on the source of the subject and the current technical status. The implementation phase is mainly implemented in Python, the tools used are introduced, and the relevant functions and functions of the system are introduced in detail through the architecture diagram and system business analysis diagram, and the operation results are shown.

[0286] IV. Experimental and test design and result analysis

[0287] 4.1 Geometry Solver Experiment and Test Design

[0288] 4.1.1 Experimental setup and results of the geometry problem solver

[0289] Datasets and evaluation metrics: The Geometry3K dataset involves 2101 training data, 300 validation data, and 601 test data. In the geometry solver model, if the one of the four choices closest to the solution found happens to be the basic theorem, the solution found is considered correct. For the numerical value that fails to output the problem target in the specified step, he will randomly select one from the four candidates because the predicted answer has the largest confidence among the candidate choices in terms of the compared neural network baseline.

[0290] 4.1.2 Baselines of Geometry Solver Models

[0291] We implement several deep neural network baselines for geometry problem solvers to compare with our proposed model. By default, these baselines formalize the geometry problem solving task as a classification problem, provided by text embeddings for the sequence encoder and diagram representations for the visual encoder. Q-only encodes the question text in natural language via a bidirectional gated recurrent unit (bi-GRU) encoder (Cho et al., 2014). I-only encodes the question graph as input using a ResNet-50 encoder (He et al., 2016). Q+I encodes the text and graph using Bi GRU and ResNet-50, respectively. RelNet (Bansal et al., 2017) is used to embed the question text because it is a powerful method for modeling entities and relations. FiLM (Perez et al., 2018) achieves effective visual reasoning when answering questions about abstract images, so it is compared. FiLM-BERT uses the BERT encoder (Devlin et al., 2018) instead of the GRU encoder, and FiLM BART uses the recently proposed BART encoder (Lewis et al., 2020). The data comparison is shown in Table 4-1.

[0292] Table 4-1 Comparison of different methods on the Geometry3K dataset

[0293]

[0294] Implementation details. The main hyperparameters used in the experiments are listed below. For this paper solver, a set of 17 geometric theorems were collected to form the knowledge base. To generate a sequence of positive theorems, each problem was attempted 100 times with a maximum sequence length of 20. The transformer model used in the theorem predictor has 6 layers, 12 attention heads, and a hidden embedding size of 768. The search step for theorems is set to 100. For the neural parser, the Adam optimizer was selected and the learning rate was set to 0.01 and the maximum epoch was set to 30. To obtain more precise results, each experiment inside GPS was repeated three times.

[0295] 4.1.3 Search strategy for theorem prediction

[0296] As shown in Table 4-2, the overall accuracy and average steps required for the geometry solver model in this paper to solve the problem by applying different search strategies are shown. Prediction refers to the strategy of using the theorem in the theorem predictor, followed by a random theorem sequence. This strategy greatly reduces the average steps to 6.5 steps. The last strategy is to apply the predicted first-order theorem and low-order theorem in the remaining search steps and obtain the best overall accuracy.

[0297] Table 4-2 Comparison of different search strategies

[0298]

[0299] 4.1.4 Text Parser and Text Source

[0300] The accuracy of the rule-based text parser is 97%, while the accuracy of the semantic text parser is only 67%. Table 4-3 shows the performance of the solvers under different text sources. The accuracy of the text generated by the solver is 57.5%. The current text parser performs very well because there is only a small gap between the generated text and the real text in the context. The performance of the solver with annotated graph text is improved by 17.5%, which shows that there is still a lot of room for improvement in graph parsers.

[0301] Table 4-3 Performance comparison

[0302]

[0303] 4.1.5 Search Step Distribution

[0304] The distribution of correctly solved problems is compared by the average number of search steps in different strategies. The final geometry solver adopts the prediction + low priority strategy, where 65.97% of the problems are solved in two steps and 70.06% of the problems are solved in five steps.

[0305] Geometry Problem Solver. Current neural network baselines for geometry solving have failed to achieve satisfactory results on the Geometry3K dataset. This is because these neural methods have limited data samples to learn meaningful semantics from the problem input. In addition, dense implicit representations may not be suitable for logical reasoning tasks such as geometry problem solving. The inputs of question text and diagrams in the Q+I baseline are replaced with ground truth text and visual form annotations, and the results are shown in Table 4-4. If structural representations with rich semantics are learned, the 9.2% improvement shows that neural network models have great potential in solving problems.

[0306] Table 4-4 Geometry solver performance in different representations of problem text and diagrams

[0307] Diagram(visual) Diagram(formal) Text(natural) 26.7 35.3 Text(formal) 34.6 35.9

[0308] 4.1.6 Failure Cases

[0309] The solver model may fail to find a solution due to inaccurate parsing results and incomplete theorem sets. Figure 3-CSome examples of failures with the geometric solver are shown. For example, if the diagram has ambiguous comments or multiple primitives, then graph parsing tends to fail. The textual parser has a hard time dealing with nested expressions and ambiguous references. And the symbolic solver still can't solve complex problems with combined shapes and shaded regions in the diagram.

[0310] 4.2 Algebraic Problem Solver Experiment and Test Design

[0311] 4.2.1 Dataset

[0312] We use two commonly used algebra problem datasets: MAWPS (Koncel-Kedziorski et al., 2016) and Math23K (Wang et al., 2017). The MAWPS dataset contains 2373 problems and the Math23K dataset has 23162 problems.

[0313] 4.2.2 Comparison with baseline

[0314] The algebra solver model is compared with a wide set of baselines and state-of-the-art models: DNS (Wang et al., 2017) uses a vanilla seq2seq model to generate expressions. Math-EN (Wang et al., 2018) benefits from equation normalization to reduce the target space. T-RNN (Wang et al., 2019) applies a recurrent neural network to a predicted tree-structured template. S-Aligned (Chiang and Chen, 2019) designs a decoder with a stack to keep track of the semantics of operands. GROUPATT (Li et al., 2019) borrows the idea of ​​multi-head attention proposed by Transformer (Vaswani et al., 2017). AST Dec (Liu et al., 2019) uses a tree LSTM decoder to create expression trees. GTS (Xie and Sun, 2019) develops a tree-structured neural network to generate expression trees in a goal-driven manner. IRE (Sahu et al., 2019) is another baseline first proposed in relation extraction and shares some commonalities with the referenced methods in this paper.

[0315] 4.2.3 Implementation Details and Evaluation Metrics

[0316] In the Algebra Solver model, a word embedding with 128 units (not pre-trained), a single-layer graph transformer with 4 GCNs, and the hidden state dimension of each GCN is set to 128. The dimensions of the hidden state of all other layers are set to 512. The model is trained for 80 epochs. The mini-batch size and dropout rate are set to 64 and 0.5, respectively. For the operator, Adam is used, the learning rate is set to 0.001, β1=0.94, β2=0.99, and the learning rate is halved every 20 epochs. In addition, the beam size used in the beam search is 5.

[0317] For the Math23K dataset, some methods are evaluated using 5-fold cross validation, denoted by “Math23K*”, and other methods are evaluated using the available test set (denoted by “Math23K”). Graph2Tree is evaluated on both settings. For the MAWPS dataset, models are evaluated using 5-fold cross validation. Following previous work, the accuracy of the solution is used as the evaluation metric.

[0318] 4.2.4 Overall Results

[0319] Table 4-5 Comparison between the problem solver in this paper and other models (Math23K represents the results on the public test set, Math23K*Gem 5-fold cross validation)

[0320]

[0321] Tables 4-5 show the solving accuracy of Graph2Tree and various baselines. It is observed that the model referenced in this paper outperforms all baselines in both MWP datasets. With the availability of GTS code, GTS was implemented and tested on all dataset settings. This paper also statistically tested the improvement over the strongest baseline (i.e., GTS) and found that the improvement was significant at the 0.01 level using a paired t-test. The excellent performance of the algebraic problem solver demonstrates the importance of rich quantity representation when dealing with algebraic problem tasks.

[0322] 4.2.5 Parameter analysis

[0323] 4.2.5.1 Impact of quantity graphs

[0324] This paper studies the impact of quantity unit graphs and quantity comparison graphs on the model. The results are shown in Tables 4-6. It can be found that the model with both quantity unit graphs and quantity comparison graphs performs best. It can also be found that the implementation with quantity unit graphs and quantity comparison graphs is still better than the implementation without any graph (i.e., fully connected graph). In this task, enriching the quantity representation with any graph will also outperform the baseline GTS model, which shows the importance of quantity representation in the task of solving algebraic problems. If the two types of graphs are merged into an integrated graph, the performance will deteriorate. One possible reason for the poor performance may be due to the noise introduced by the integration of multiple graphs.

[0325] Table 4-6 Solution accuracy of various graphics configurations of the model

[0326]

[0327] 4.2.5.2 Impact of the number of graph convolutional networks

[0328] The number of GCNs is an adjustable hyperparameter in the model. Therefore, the effect of the number of GCNs on the model performance was studied. The number of GCNs was changed from 2, 4, and 8. The reason for using an even number is to facilitate the even splitting of GCNs to create the quantity unit graph and quantity comparison graph. Tables 4-7 show the results of the study. It is observed that the 4-GCN version achieves the best performance. A potential reason could be that using 4 GCNs achieves the best capacity for information aggregation on both quantity graphs.

[0329] Table 4-7 Impact of the number of GCNs

[0330]

[0331] It has the following features:

[0332] (1) The results of geometric problems searched by users can be effectively and quickly fed back to users, and the accuracy is guaranteed to a certain extent.

[0333] (2) For the algebraic problems searched by users, we can solve the street expressions and calculate the answers, with a certain degree of accuracy guaranteed.

[0334] (3) The function of querying search history is implemented, which can facilitate users to query the same question again.

[0335] (4) A comment function is implemented, and users can comment on the correctness of the feedback results.

[0336] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in any of the above embodiments when executing the program.

[0337] Figure 4 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0338] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0339] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0340] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0341] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0342] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0343] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0344] The electronic device of the above embodiment is used to implement the corresponding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0345] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in any of the above embodiments.

[0346] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0347] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0348] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0349] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0350] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0351] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.

Claims

1. A processing method for a mathematical problem solver based on deep learning, characterized in that: The mathematical problem solver includes a geometric problem solver and an algebraic problem solver; The processing method comprises: receiving a search question type, and determining a target problem solver from a geometry solver and an algebra solver of the math problem solver according to the question type; Receiving a search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, and performing problem solving processing on the question information through a problem solving model in the target problem solver to obtain a problem solving result, wherein the problem solving model is obtained by training a neural network with problem solving samples; Outputting the problem-solving result; The geometric problem solver includes: a geometric problem-solving model obtained by training a neural network with geometric problem-solving samples; The receiving of the search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, performing problem solving processing on the question information through a problem solving model in the target problem solver, and obtaining a problem solving result, includes: In response to determining that the target problem solver is a geometry problem solver, receiving the search question, extracting image element information and text element information in the search question, wherein the image element information and the text element information are used as the question information; Inputting the image element information and the text element information, as well as the determined axiom theorem set into the geometry problem solver, performing geometry problem solving processing through a geometry problem solving model, obtaining a problem solving process sequence, and using the problem solving process sequence as the problem solving result; The step of inputting the image element information and the text element information, and the determined axiom theorem set into the geometry problem solver, performing geometry problem solving processing through a geometry problem solving model, obtaining a problem solving process sequence, and using the problem solving process sequence as the problem solving result includes: After the search question is divided into tuples, the tuples include: picture element information and the text element information and numerical format; Converting the words of the text element information in the tuple into predicates and variables through a deep neural network, and forming a text sequence with the predicates and variables; Extracting geometric elements from the image element information in the tuple to obtain an element set P; Extracting the icon symbols and text areas in the geometric elements by using a strong object detector, recognizing the text information in the text area by using an optical character recognition tool, and obtaining a symbol set S; Determine each symbol S in the symbol set S i and elements in the set P and S i The corresponding geometric element p j The geometric relationship data F, for each s i and p j To associate, the association algorithm formula is: min∑ s dist(s i ,p j )×π{s i assigns to p j } st(s i , p j )∈Feasibility set F, where dist is s i and p j The Euclidean distance between, i is the symbol s i The order in the symbol set, j is the geometric element p j In the sorting of the element set P, i and j are both positive integers; The result of the association is converted into a quasi-transposition formal language by rules to obtain the relation set R; The relationship set R is searched using a theorem search strategy to determine an optimal solution sequence, and a geometric answer and a geometric problem-solving process are determined based on the optimal solution sequence as the problem-solving result.

2. The method according to claim 1, characterized in that The process of determining the geometric problem-solving model includes: Acquire a geometry problem dataset Geometry3K, wherein the geometry problem dataset includes a plurality of geometry problem samples; Using the geometry problem dataset Geometry3K to train a pre-built deep neural network, and constructing an optimal solution theorem sequence corresponding to the geometry problem dataset through a theorem predictor; The predictor is used to predict the next theorem sequence of the current sequence, and the optimized negative log-likelihood loss is determined according to the predicted theorem sequence. Then, the parameters of the deep neural network are adjusted according to the negative log-likelihood loss, and the training is completed to obtain the geometric problem-solving model.

3. The method according to claim 1, characterized in that The step of searching the relationship set R using a theorem search strategy to determine an optimal solution sequence, and determining a geometric answer and a geometric problem-solving process as the problem-solving result according to the optimal solution sequence includes: Expanding the relation set R into a geometric formal language; Get a theorem set KB, where each theorem k in the theorem set KB i It is defined as a conditional rule with premise a and conclusion q; Determine theorem k at time t i The premise a and the relationship R currently selected from the relationship set R t-1 Matching, according to the theorem k i Conclusion q updates the relation set R t ; Based on R t An equation relationship between known values ​​and unknown problem targets is constructed, and the equation is solved to obtain a geometric answer and a geometric problem-solving process as the problem-solving result.

4. The method according to claim 1, characterized in that: The algebra problem solver comprises: an algebra problem-solving model obtained by training a neural network with algebra problem-solving samples; The receiving of the search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, performing problem solving processing on the question information through a problem solving model in the target problem solver, and obtaining a problem solving result, includes: In response to determining that the target problem solver is an algebra problem solver, receiving the search problem, processing the search problem using a bidirectional long short-term memory model to obtain a key variable generation node, and obtaining internal representation parameters of the algebra sample problem using a graph converter according to a quantity comparison graph and a quantity unit graph; The key variable generation node and the internal representation parameters are input into the algebraic problem solver, and algebraic problem solving is performed through an algebraic problem solving model to obtain an algebraic solution expression, and the algebraic solution expression is used as the problem solving result.

5. The method according to claim 4, characterized in that The determination process of the algebraic problem-solving model includes: Acquire an algebraic problem data set, wherein the algebraic problem data set includes a plurality of algebraic problem samples; The algebraic problem data set is used to perform learning and training on a pre-constructed deep neural network, and the sum of the negative log-likelihoods of the probabilities of the predicted nodes during the training process is determined as an algebraic training loss function. Based on the algebraic training loss function, the parameters of the deep neural network are adjusted to complete the training and obtain the algebraic problem-solving model.

6. The method according to claim 5, characterized in that The receiving of the search question, preprocessing the search question to obtain question information, inputting the question information into the target problem solver, performing problem solving processing on the question information through a problem solving model in the target problem solver, and obtaining a problem solving result, includes: receiving an algebraic problem, and processing the algebraic problem using a bidirectional long short-term memory model to obtain an algebraic problem node representation, wherein the algebraic problem node representation includes algebraic words and algebraic variables; constructing nodes of a graph according to algebraic words and all algebraic variables in the node representation of the algebraic problem; Establishing a quantity unit diagram and a quantity comparison diagram according to the algebraic problem; The graph converter adopts a multi-head structure to evenly distribute the quantity unit graph and the quantity comparison graph to obtain the root node; A decision tree is generated based on the root node, and the data in the root node is decoded according to the decision tree to obtain a decoding sequence, and the decoding sequence is traversed to obtain a solution result.

7. A problem solving system, characterized in that: The method as claimed in any one of claims 1 to 6, comprising: The user layer is configured to receive search question types and display solution results; The functional logic layer is configured to determine a target problem solver from a geometry problem solver and an algebra problem solver of a math problem solver according to the problem type, receive a search question, and feed back a problem solving result to a user layer; The function execution layer includes a geometry problem solver and an algebra problem solver, and is configured to pre-process the search problem to obtain problem information, input the problem information into the target problem solver, perform problem solving processing on the problem information through the problem solving model in the target problem solver, obtain a problem solving result, and feed the problem solving result back to the function execution layer; wherein the problem solving model is obtained by training a neural network with problem solving samples; The storage computing layer is configured to store the data sets required by the geometry problem solver and the algebra problem solver, so that the geometry problem solver and the algebra problem solver can retrieve corresponding data from the data sets for problem solving when performing problem solving.

Citation Information

Patent Citations

  • A neural network-based method for improving the problem solving capability of a student application problem

    CN109948473A

  • Artificial intelligence science literal question solving method and device, equipment and storage medium

    CN112949410A