A retrieval method and device for engineering design calculation formulas
By converting the design calculation formula into an operation tree and fusing the associated text embedding, a formula vector containing text semantics is generated, which solves the problem that the calculation formula search method in the prior art ignores semantic information, and achieves higher retrieval accuracy and efficiency of the design calculation process.
Patent Information
- Application Number
- CN202111598653.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The existing computational formula search methods mainly rely on the structural information of the formula and ignore the semantic information of the formula, which makes it difficult to accurately distinguish formulas with similar structures but different physical meanings in the search for design calculation formulas in the engineering field.
By identifying the design calculation formula, converting it into an operation tree, and combining the embed vector of the associated text, fusing the operation tree embedding and text embedding, a formula vector containing text semantics is generated, and the vector similarity measurement method is used for searching.
The search accuracy of design calculation formulas is improved, and the results of the design calculation process can be more effectively distinguished, thereby improving the efficiency of the design calculation process.
Smart Images

Figure CN114266228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of information retrieval and natural language processing, and in particular to a retrieval method and a device thereof for designing calculation formulas in the engineering field. Background Art
[0002] Therefore, retrieving calculation formulas has become the key to improving the efficiency of the entire design calculation process.
[0003] The current retrieval methods for calculation formulas, such as patents CN109918473A, CN106372073A, and CN110414319A, all express and retrieve the structure of the formula, ignoring the semantic information of the formula. Compared with mathematical formulas in general scientific and technological documents, design calculation formulas are more likely to have two formulas with similar structures but different physical meanings. For example, the calculation formula for the moment of inertia of the motor in elevator design is J q =GD 2 / 4 and the maximum inertia torque calculation formula M g =Jε / η, these two formulas have similar structures but significantly different physical meanings. Therefore, the expression of formulas not only needs to consider the logical structure of the formula itself, but also needs to integrate the semantic information of the variable names associated with the formula, that is, the formula description and other text.
[0004] The paper DOI: 10.19678 / j.issn.1000-3428.0048934 proposed using ontology to establish the connection between mathematical expressions and their concepts in order to realize the use of phrase query formulas. Although this method takes into account the semantics of the formula, it requires the establishment of a specific ontology for calculation formulas in specific fields, making this method lack of flexibility.
[0005] In this regard, the present application proposes a formula retrieval method that integrates the physical meaning of calculation formulas, enriches the features of embedded vectors of design calculation formulas in the engineering field from a semantic level, and thus improves the accuracy of retrieval of design calculation formulas. Summary of the invention
[0006] In response to the above problems, the present application provides a retrieval method and device for design calculation formulas in the engineering field, which can take into account the physical meaning of the design calculation formulas, thereby enriching the semantic features of their embedded representation vectors and improving the retrieval accuracy.
[0007] A first aspect of the embodiments of the present application provides a method for searching for calculation formulas for engineering design, the specific steps of which include:
[0008] Step 1: Identify the design calculation formula in the document, convert it into an operation tree, and embed the operation tree.
[0009] Step 2: Get the associated text of the calculation formula and embed it into an expression.
[0010] Step 3: Fusion calculates the formula’s operation tree embedding and its associated text embedding to obtain a formula vector that contains text semantics.
[0011] Step 4: When searching, use the vector similarity measurement method to measure the similarity between different formulas and return the result with the highest similarity.
[0012] The step 1 specifically includes:
[0013] Step 1.1: For the printed formulas in the document, use the formula recognition tool to convert them into intermediate expressions f.
[0014] Step 1.2: Use context-free grammar to describe the grammatical patterns of operands and operators in the intermediate expression to build a vocabulary, define the calculation priority order of operators, and segment f to obtain a sequence of operands and operators. Finally, according to the calculation priority of the operators and combined with the stack data structure, the sequence is converted into a formula operation tree T.
[0015] Step 1.3: One-hot encode the node information of the operation tree T and Huffman encode the structure information. After splicing the two, the embedding matrix M of the operation tree of the calculation formula is obtained. OpT .
[0016] The step 2 specifically includes:
[0017] Step 2.1: For the printed formula in the document, locate the sentence describing its output parameters according to the relative position relationship, and remove the stop words to obtain the associated text d of the formula.
[0018] Step 2.2: Using the calculation formulas in books and standards as data sources, construct a sentence similarity annotation dataset in a professional field, use the pre-trained model of text embedding to embed the associated text d, and obtain the vector e of the associated text d .
[0019] The step 3 specifically includes:
[0020] Step 3.1: Using the calculation formulas and their description statements in books and standards as data sources, establish a formula similarity annotation dataset containing text descriptions; combine the formula and its associated text as a sample, and annotate whether different sample pairs are similar.
[0021] Step 3.2: Build a neural network model and train and validate it using the labeled dataset described in step 3.1.
[0022] Step 3.3: Operate the embedding matrix M of the tree using the calculation formula OpT and the embedding vector e of the formula associated text d As input, the neural network model in step 3.2 outputs a formula vector e containing text semantics f .
[0023] The step 4 specifically includes:
[0024] Step 4.1: Use vector similarity metric to measure the formula vector e f The similarity between the formula vector and the data set is returned.
[0025] Step 4.2: Rank the similarities. The higher the similarity, the higher the ranking. Return the formula with the highest ranking as the search result.
[0026] Preferably, the design calculation formula described in step 1 is a formula describing the parameter calculation process in engineering design standards, design calculation manuals and design specifications.
[0027] Preferably, the formula recognition tool described in step 1.1 includes Mathpix and InftyReader tools.
[0028] Furthermore, the formula recognition task is completed using the interface provided by the formula recognition tool Mathpix.
[0029] Preferably, the intermediate expression in step 1.1 is in the form of LaTeX expression and MathML expression. LaTeX expression is used as the form of the intermediate expression.
[0030] Preferably, the context-free grammar described in step 1.2 includes: BNF grammar, regular expression, and the BNF grammar is used to describe the grammatical pattern of operation symbols and operation objects.
[0031] Preferably, the relative position relationship described in step 2.1 includes two situations: the relevant text is in the nearest line above the formula, and the relevant text is in the nearest line below the formula text.
[0032] Preferably, the text embedding pre-training model described in step 2.2 includes a text embedding model BERT and a sentence embedding model SBERT. For relevant texts in the form of short texts, SBERT is used to output their embedding vectors.
[0033] Preferably, the neural network model described in step 3.2 includes a long short-term memory network LSTM, a recurrent neural network RNN and a gated recurrent unit GRU, and GRU is used as the basic structural component of the neural network.
[0034] The second aspect of the embodiment of the present application provides a retrieval device for engineering design calculation formulas. The device includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by one or more processors, the one or more processors implement the method as described above.
[0035] The retrieval method and device for design calculation formulas in the engineering field provided in the embodiments of the present application have the following beneficial effects: the features of the embedded vectors of design calculation formulas in the engineering field are enriched from the semantic level, and the accuracy of retrieval of design calculation formulas is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The above and other objects, features and advantages of the present disclosure will become more apparent by describing in detail example embodiments thereof with reference to the attached drawings.
[0037] Figure 1 This is a general flow chart of a retrieval method and device for design calculation formulas in the engineering field proposed in this application.
[0038] Figure 2 A flowchart for embedding printed formulas.
[0039] Figure 3 A flowchart for generating a formula operation tree from a formula's LaTeX expression.
[0040] Figure 4 A flowchart for embedding the calculation formula.
[0041] Figure 5 A flowchart for obtaining the associated text of a calculation formula and embedding it into an expression.
[0042] Figure 6 Flowchart for building a dataset and tuning SBERT.
[0043] Figure 7 Flowchart for fusing the computation of a formula operation tree embedding and its associated text embedding.
[0044] Figure 8 Figure 2 is a diagram of the neural network structure used to achieve fusion. DETAILED DESCRIPTION
[0045] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0046] This embodiment uses the design calculation formula in the book "Elevator Design Calculation and Examples" in the field of elevator design to establish the model training set and test set. It should be understood that the specific embodiment described here is only used to explain this application and is not used to limit this application. The specific steps of the embodiment of this application include:
[0047] Step 1: Parse the calculation formula and convert it into an operation tree, and embed the operation tree, including:
[0048] 1.1) For the printed formula in the document, convert it into an intermediate expression f. The optional intermediate expressions include but are not limited to LaTex.
[0049] The LaTeX expression is based on the LaTeX typesetting system and consists of grammatical identifiers and parameter symbols. For example, the expression "x = \frac{a}{b}" means "x = a / b". The LaTeX expression defines the positional relationship between symbols and is easy to parse.
[0050] The printed formula recognition may be selected from but not limited to Mathpix. Mathpix can convert formulas in pictures and documents into LaTeX expressions.
[0051] 1.2) Generate a calculation formula operation tree T based on f;
[0052] 1.3) Embedding the operation tree T;
[0053] Step 1.2 specifically includes:
[0054] 1.2.1) Establish a LaTeX expression vocabulary, which defines the general syntax pattern of symbols and variables in LaTeX expressions;
[0055] 1.2.2) Segment the LaTeX expression according to the vocabulary to obtain an infix expression sequence IN consisting of symbols and variables;
[0056] The infix expression is a general method of expressing an arithmetic or logical formula, in which the operator is placed in the middle of the operands in an infix form, such as "1+2".
[0057] 1.2.3) Define the priorities between different operation symbols and convert IN into a postfix expression sequence RPN;
[0058] The postfix expression is also called reverse Polish notation. In this representation method, all operators are placed after the operands, such as the postfix expression "24 / " is equivalent to the infix expression "2 / 4".
[0059] 1.2.4) Generate a calculation formula operation tree T by RPN;
[0060] Step 1.3 specifically includes:
[0061] 1.3.1) Pre-order traversal of the operation tree T, one-hot encoding of the sequence obtained by traversal, and obtaining the embedding matrix M of the operation tree node information nodes ;
[0062] 1.3.2) Depth-first traverse the operation tree, perform Huffman coding on the sequence obtained by traversal, and obtain the embedding matrix M of the position relationship of the operation tree nodes positions ;
[0063] 1.3.3) M nodes With M positions Connect and get the embedding matrix M of the formula operation tree OpT , M OpT Each dimensional vector e in i They all correspond to a node in T;
[0064] Step 2: Get the associated text of the calculation formula and embed it into an expression, including:
[0065] 2.1) For the printed formula in the document, locate the sentence describing its output parameters and remove the stop words as the associated text d of the formula;
[0066] 2.2) Construct a sentence similarity annotation dataset in a professional field to tune the pre-trained model SBERT, use the optimized model to embed the associated text d, and obtain the vector e of the associated text d ;
[0067] The SBERT model is a sentence embedding pre-training model based on the natural language embedding representation model BERT, which uses a twin network and a three-level network structure to obtain sentence vectors containing semantics.
[0068] Wherein 2.2) specifically includes the following steps:
[0069] 2.2.1) Extract parameter description statements from design calculation standards and books in professional fields;
[0070] 2.2.2) If the physical meanings of two parameters are the same, then their description sentences are marked as similar; otherwise, they are marked as dissimilar; thus, a sentence similarity dataset in a professional field is constructed;
[0071] 2.2.3) During tuning, use the SBERT model to predict the similarity of all sentence pairs in the dataset, compare the predicted values with the true values, and back-propagate the errors to optimize the model parameters;
[0072] Step 3: Fusion calculates the embedding of the operation tree of the formula and the embedding of its associated text to obtain a formula vector containing text semantics, including:
[0073] 3.1) Establish a formula similarity annotation dataset containing text descriptions;
[0074] 3.2) Build a neural network model and train and validate it using the labeled data set described in step 3.1;
[0075] 3.3) The embedding matrix M of the tree is operated by the calculation formula OpT and the embedding vector e of the formula associated text d As input, the neural network model in step 3.2 outputs a formula vector e containing text semantics f ;
[0076] Among them, 3.1) specifically includes the following steps:
[0077] 3.1.1) Extract the calculation formula x from professional books and standards i and its corresponding associated text y i , forming a sample of formulas containing descriptions (x i ,y i );
[0078] 3.1.2) According to the sample (x) obtained in step 3.1.1 i ,y i ) to expand the training samples; the specific method is: according to the commutative law, associative law and other operation rules, change the formula x i The structure of does not change its meaning, and repeats n1 times to obtain the positive sample set of the formula Substitute the formula x i The operator symbol in , changes its meaning, and repeats n2 times to obtain the negative sample set of the formula For the description statement y i , add stop words to it, replace synonyms, repeat n3 times to obtain a positive sample set of associated text Randomly select n4 description statements of other formulas and add them to the negative sample set of associated texts
[0079] 3.1.3) Match the formula and associated text obtained in step 3.1.2; the specific method is to set With Collection The elements in are arranged and combined to retain the form For a sample pair,
[0080] 3.1.4) Consider the similarity between the associated text and the calculation process at the same time to mark the similarity between formulas; the specific method is to mark the formulas belonging to (x i ,y i ) + The two samples of (x i ,yi)+ and (x i ,y i ) - The two samples are marked as dissimilar. All possible combinations are made;
[0081] The specific characteristics of the neural network model and its training process described in 3.2) are: a bidirectional GRU model is used, and the model input is the embedding matrix M of the formula operation tree OpT and the embedding vector e described by the formula text d , the output is the formula vector e that integrates text semantics f The training adopts the twin network structure, stochastic gradient descent strategy and cosine embedding loss function. The specific training process is: based on the labeled data set constructed in 3.1, the model outputs the vectors s1 and s2 of the sample pairs, and the training loss is calculated using the following formula:
[0082]
[0083] In the formula, y is the label of the sample pair, 1 means the two are similar, -1 means they are not similar, and cos(s1, s2) is the cosine similarity of s1 and s2. At the end of each training cycle, the training loss is back-propagated to optimize the model.
[0084] The GRU model is a variant of the long short-term memory (LSTM) network. It has a simpler structure than the LSTM network and alleviates the gradient vanishing problem while retaining long-term sequence information.
[0085] The cosine similarity is a method for measuring the similarity between vectors, and its specific calculation formula is:
[0086]
[0087] Optionally, a batch gradient descent strategy is used during training;
[0088] Optionally, the loss function for training is expressed using Euclidean distance or cosine distance;
[0089] Step 4: Compare the similarities between different formula vectors during retrieval and return the results with the highest similarity, including:
[0090] 4.1) Use cosine similarity to measure the similarity between vectors;
[0091] Optionally, use Euclidean distance to measure the similarity between vectors;
[0092] 4.2) Rank the similarities and return the highest ranked formula;
[0093] Figure 1 The overall process of a design calculation formula retrieval method for the engineering field is demonstrated, specifically:
[0094] Identify the printed formulas in the design calculation documents, parse and generate the formula operation tree, and generate the embedding of the formula operation tree based on the operation tree; use the corpus in the elevator design field to tune the pre-trained model SBERT, a general flow chart of the design calculation formula retrieval method for the engineering field, and generate the embedding of the associated text by the tuned SBERT model; for the elevator design field, build a formula similarity annotation dataset containing text descriptions, and based on this dataset, train a neural network model to achieve the fusion embedding of the formula operation tree and the associated text. When searching, compare the similarity between the embedding vectors that integrate the text semantics, and return the result with the highest similarity.
[0095] Figure 2 The process of parsing a printed formula and generating an operation tree is shown. The steps include:
[0096] Step S21: calling the interface of the mathematical formula recognition tool Mathpix to convert the printed formula into a LaTeX expression f.
[0097] In this embodiment, a typical example of the design calculation formula is: v 梯 =d1 / 60×n 主 、v 扶 =d2(1-η 轮 )+(n 主 / 60)×(z4 / z5)×π. The corresponding LaTeX expressions are: “i_total=i×\frac{z_2}{z_1}”, “v_ladder=d_1 / 60×n_main”, “v_support=d_2(1-\eta_{wheel})×(n_main / 60)×(z_4 / z_5)×π”.
[0098] Step S22: Generate an operation tree T of the formula from the LaTeX expression f of the formula.
[0099] Step S23: embedding the operation tree T.
[0100] Figure 3 The process of generating a formula operation tree from a formula's LaTeX expression is shown, specifically:
[0101] Step S31: Establish a LaTeX expression vocabulary table, which defines the general grammatical pattern of symbols and variables in LaTeX expressions. The LaTeX expression vocabulary table established in this embodiment includes an operation symbol table, an identification symbol table, and an operation object table, as shown in Tables 1, 2, and 3:
[0102] Table 1: Operation symbol table
[0103]
[0104] Table 2: Identification symbol table
[0105]
[0106] Table 3: Operation object table
[0107]
[0108]
[0109] The operator is a general calculation symbol, which has the same form and meaning in LaTeX expressions and printed formulas; the identifier is a special symbol in LaTeX syntax, which expresses a specific meaning by describing the positional relationship between two objects in a formula. The identifier has a variety of forms and does not conform to the writing habits of infix expressions. This embodiment converts the identifier into an operator according to its meaning. A typical example is: converting "\frac{x}{y}" into "x / y"; the operation object is the object on which the operator and the identifier act. In this embodiment, one operator will act on two operation objects.
[0110] Step S32: Segment the LaTeX expression according to the vocabulary to obtain an infix expression sequence IN consisting of symbols and variables.
[0111] In this embodiment, a typical infix expression sequence example is: [i_total, I, ×, z_2, / , z_1], [v_ladder, d_1, / , 60, ×, n_main], [v_handrail, =, d_2, (, 1, -, \eta_{wheel},), ×, (, n_main, / , 60,), ×, (, z_4, / , z_5,), ×, π]. Typical examples of associated text are: "Calculation of the total transmission ratio of the main drive"; "Calculation of the step running rate"; "The power required to drive two handrails when fully loaded and rising".
[0112] Step S33: According to the priorities between different operation symbols shown in the operation symbol table in Table 1, IN is converted into a postfix expression sequence RPN.
[0113] This embodiment introduces the concept of "stack" when converting IN to RPN. The specific method is: define a symbol stack s t and the suffix expression sequence stack s r , traverse the elements in the infix expression one by one, determine the type of the element according to the vocabulary, and perform different operations. Specifically:
[0114] 1) If the element is an operation object, push it into s r ;
[0115] 2) If the element is a left bracket "(", push it into s t ;
[0116] 3) If the element is a right bracket ")", then s t The top elements of the stack are popped out and pushed into s one by one. r , until the top element of the stack is a left bracket ")", then pop the ")" out;
[0117] 4) If the element is an operator, then if s t is not empty and s t The priority of the top element of the stack is greater than or equal to the priority of the current element. t The top element of the stack is popped and pushed into s r , and then push the current element into s r ; Otherwise, push the current element directly into s r ;
[0118] After traversing all the elements in the infix expression, s r The elements in are popped out from the bottom of the stack one by one to form the final suffix expression sequence RPN.
[0119] Step S34: Generate a calculation formula operation tree T from the suffix expression sequence RPN.
[0120] In this embodiment, the operation tree is a binary tree.
[0121] In this embodiment, the basis for generating the operation tree is that each operation symbol corresponds to two operation objects. Based on this, the RPN is subsequently traversed to assign left and right nodes to each operation symbol, and finally the calculation formula operation tree T is obtained.
[0122] Figure 4 The process of embedding the calculation formula operation tree is described as follows:
[0123] Step S41: Pre-order traversal of the operation tree T, one-hot encoding of the sequence obtained by the traversal, and obtaining the embedding matrix M of the operation tree node information nodes .
[0124] Step S42: Depth-first traversal of the operation tree, Huffman coding of the sequence obtained by the traversal, and obtaining the embedding matrix M of the position relationship of the operation tree nodes positions .
[0125] Step S43: M nodes With M positions Connect and get the embedding matrix M of the formula operation tree OpT , M OpT Each dimensional vector e in i Each corresponds to a node in T.
[0126] In this embodiment, M nodes With M positions The connection method is end to end. If M nodes The length of M is n, positions The length of is m, then the obtained e i The first n bits encode the node information, and the last m bits encode the node's position in T.
[0127] Figure 5 The process of obtaining the associated text of a calculation formula and embedding it into an expression is shown. The specific steps are as follows:
[0128] Step S51: Locate the sentence describing the output parameters of the printed formula, remove the stop words and use it as the associated text d of the formula.
[0129] In this embodiment, the line of statements closest to the top of the formula is used as the description statement of its output parameter.
[0130] In this embodiment, typical examples of associated texts are: "total transmission ratio of the main drive", "circumferential force at the handrail belt drive", and "rotation rate of the main shaft".
[0131] Step S52: Construct a sentence similarity annotation dataset in a professional field to tune the pre-trained model SBERT
[0132] Step S53: Use the optimized model to embed the associated text d and obtain the vector e of the associated text d ;
[0133] The process of constructing a data set and tuning SBERT in step S52 is as follows: Figure 6 As shown, the specific steps are:
[0134] Step S61: extracting parameter description statements from design calculation standards and books in professional fields.
[0135] Step S62: manually determine whether the physical meanings of the two parameters are the same, if yes, mark their description sentences as similar, otherwise mark them as dissimilar; thereby constructing a sentence similarity annotation dataset in a professional field.
[0136] This embodiment expands the samples when constructing the data set. The specific method is: for a description sentence of a parameter, stop words are added or synonyms are replaced to obtain positive samples similar to it. After expansion, the data set contains 6442 sentence pairs that are manually annotated for similarity.
[0137] Step S63: During tuning, use the SBERT model to predict the similarity of all sentence pairs in the data set, compare the predicted values with the true values, and back-propagate the errors to optimize the model parameters.
[0138] Figure 7 The process of integrating the calculation formula operation tree embedding and its associated text embedding is shown. The specific steps include:
[0139] Step S71: Create a formula similarity annotation dataset containing text descriptions.
[0140] The specific implementation method of step S71 of this application is:
[0141] 1) Extract the calculation formula x from professional books and standards i and its corresponding associated text y i , forming a sample of formulas containing descriptions (x i ,y i );
[0142] 2) According to the sample (x) obtained in step 1) i ,y i ) to expand the training samples; the specific method is: according to the commutative law, associative law and other operation rules, change the formula x i The structure of does not change its meaning, and repeats n1 times to obtain the positive sample set of the formula Substitute the formula x i The operator symbol in , changes its meaning, and repeats n2 times to obtain the negative sample set of the formula For the description statement y i , add stop words to it, replace synonyms, repeat n3 times to obtain a positive sample set of associated text Randomly select n4 description statements of other formulas and add them to the negative sample set of associated texts
[0143] 3) Match the formula and associated text obtained in step 2); the specific method is to set With Collection The elements in are arranged and combined to retain the form For a sample pair,
[0144] 4) Consider the similarity between the associated text and the calculation process at the same time to mark the similarity between formulas; the specific method is to mark the formulas belonging to (x i ,y i ) + The two samples of (x i ,y i ) + and (x i ,y i ) - The two samples are marked as dissimilar. All possible combinations are made;
[0145] Step S72: Build Figure 8 The bidirectional GRU model shown has an input (m is the matrix M OpT The width, x i =M OpT [i, m] means taking M OpT The i-th column vector, The output is the formula vector e that integrates the text semantics. f The twin network structure, stochastic gradient descent strategy and cosine embedding loss function are used during training.
[0146] In this embodiment, the specific method of model training is: based on the labeled data set constructed in step S71, the model outputs the vectors s1 and s2 of the sample pair, and the training loss is calculated using the following formula:
[0147]
[0148] In the formula, y is the label of the sample pair, 1 means the two are similar, -1 means they are not similar, and cos(s1, s2) is the cosine similarity of s1 and s2. At the end of each training cycle, the training loss is back-propagated to optimize the model.
[0149] Step S73: Using the calculation formula to operate the embedding matrix M of the tree OpT and the embedding vector e of the formula associated text d The neural network model in step S72 outputs a formula vector e containing text semantics. f .
[0150] In order to verify the effect of this method in the actual retrieval process, this embodiment is based on the design calculation formula in "Elevator Design Calculation and Examples", and 540 calculation formulas containing associated texts are obtained. These formulas form 4919 similarity comparison pairs, of which 4309 pairs are used as training sets and 610 pairs are used as test sets. The final experimental results show that the accuracy of the model without semantic information integration in the formula similarity matching task is 78.70%, while the formula embedding model with semantic information integration proposed in this application has an accuracy of 85.24% in the formula similarity matching task.
[0151] A design calculation formula retrieval device for the engineering field includes functional components including one or more processors; a storage device for storing one or more programs; when one or more programs are executed by one or more processors, one or more processors implement the method as described above. It should be noted that the information interaction, execution process, etc. between the above-mentioned devices or units are based on the same concept as the method embodiment of this application, and their specific functions and technical effects can be specifically referred to in the method embodiment section.
Claims
1. A retrieval method for design calculation formulas in the engineering field, characterized by: The following steps are involved: Step 1: Identify the design calculation formula in the document, convert it into an operation tree, and embed the operation tree, including: Step 1.1: For the printed formula in the document, use the formula recognition tool to convert it into an intermediate expression f; Step 1.2: Use context-free grammar to describe the grammatical patterns of operands and operators in the intermediate expression to build a vocabulary, define the calculation priority order of operators, and perform word segmentation on f to obtain a sequence of operands and operators. Finally, according to the calculation priority of the operators and combined with the stack data structure, the sequence is converted into a formula operation tree T; Step 1.3: One-hot encode the node information of the operation tree T and Huffman encode the structure information. After splicing the two, the embedding matrix M of the operation tree of the calculation formula is obtained. OpT ; Specifically include: 1.3.1) Pre-order traversal of the operation tree T, one-hot encoding of the sequence obtained by traversal, and obtaining the embedding matrix M of the operation tree node information nodes ; 1.3.2) Depth-first traverse the operation tree, perform Huffman coding on the sequence obtained by traversal, and obtain the embedding matrix M of the position relationship of the operation tree nodes positions ; 1.3.3) M nodes With M positions Connect and get the embedding matrix M of the formula operation tree OpT , M OpT Each dimensional vector e in i They all correspond to a node in T; Step 2: Obtain the associated text of the calculation formula and embed it into an expression, including: Step 2.1: For the printed formula in the document, locate the sentence describing its output parameters according to the relative position relationship, remove the stop words and use it as the associated text d of the formula; Step 2.2: Using the calculation formulas in books and standards as data sources, construct a sentence similarity annotation dataset in a professional field, use the pre-trained model of text embedding to embed the associated text d, and obtain the vector e of the associated text d ; The text embedding pre-training model includes a text embedding model BERT and a sentence embedding model SBERT; SBERT is used to output the embedding vector of the associated text in the form of short text; The step 3: fusing the embedding of the operation tree of the calculation formula and the embedding of its associated text to obtain a formula vector containing text semantics, specifically includes: Step 3.1: Using the calculation formulas and their descriptions in books and standards as data sources, establish a formula similarity annotation dataset containing text descriptions; combine the formula and its associated text as a sample, and annotate whether different sample pairs are similar; Step 3.2: Build a neural network model and train and validate it using the labeled dataset described in step 3.1; Step 3.3: Operate the embedding matrix M of the tree using the calculation formula OpT and the embedding vector e of the formula associated text d As input, the neural network model in step 3.2 outputs a formula vector e containing text semantics f ; Step 4: Use the vector similarity measurement method to measure the similarity between different formulas during retrieval, and return the results with the highest similarity, including: Step 4.1: Use vector similarity metric to measure the formula vector e f The similarity with the formula vector in the data set, and returns the similarity result; Step 4.2: Rank the similarities. The higher the similarity, the higher the ranking. Return the formula with the highest ranking as the search result.
2. A method for retrieving design calculation formulas for the engineering field according to claim 1, characterized in that: The design calculation formula described in step 1 is a formula describing the parameter calculation process in the engineering field design standards, design calculation manuals and design specifications.
3. A method for retrieving design calculation formulas for the engineering field according to claim 1, characterized in that: The formula recognition tools described in step 1.1 include Mathpix and InftyReader tools.
4. A method for retrieving engineering design calculation formulas according to claim 3, characterized in that: Use the interface provided by the formula recognition tool Mathpix to complete the formula recognition task.
5. A method for retrieving design calculation formulas in the engineering field according to claim 1, characterized in that: The intermediate expression described in step 1.1 includes a LaTeX expression and a MathML expression; the LaTeX expression is used as the form of the intermediate expression.
6. A method for retrieving design calculation formulas in the engineering field according to claim 1, characterized in that: The context-free grammar described in step 1.1 includes: BNF grammar, regular expression, and the BNF grammar is used to describe the grammatical pattern of operation symbols and operation objects.
7. A method for searching engineering design calculation formulas according to claim 1, characterized in that: The relative position relationship described in step 2.1 includes two situations: the associated text is in the line closest to the top of the formula, and the associated text is in the line closest to the bottom of the formula text.
8. A method for retrieving design calculation formulas for the engineering field according to claim 1, characterized in that: The neural network model described in step 3.2 adopts a bidirectional gated recurrent unit GRU model.
9. A retrieval device for design calculation formulas in the engineering field, characterized by: The device includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by one or more processors, the one or more processors implement a retrieval method for design calculation formulas for the engineering field as described in any one of claims 1-8; when the processor executes the computer program, it implements a retrieval method for design calculation formulas for the engineering field as described in any one of claims 1-8.
Citation Information
Patent Citations
Mathematical formula retrieval method and apparatus
CN106372073A
A mathematical formula similarity measurement method and a measurement system thereof
CN109918473A
Formula similarity calculation method based on effective matching subtree and scientific and technological document retrieval method and device
CN110414319A