Formula equivalence judgment method based on formula structure embedding and graph neural network

Through the method based on formula structure embedding and graph neural network, the accuracy problem of formula equivalent judgment in the prior art is solved, efficient and accurate judgment of mathematical formulas is achieved, and the efficiency and accuracy of information retrieval and theorem proof are improved.

CN120336873APending Publication Date: 2025-07-18BEIJING FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510405209.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

It is difficult for the prior art to efficiently and accurately determine the equivalence of mathematical formulas, especially in the fields of information retrieval and theorem proof. Traditional methods cannot fully capture the semantic and structural characteristics of formulas.

Method used

Using a method based on formula structure embedding and graph neural network, standardized processing of formula structure and multi-dimensional similarity calculation module through formula standardization module, structural feature extraction module, graph embedding module, graph-level similarity comparison module, node-level similarity comparison module and similarity score calculation module, combined with cross attention mechanism and neural tensor network, standardized processing of formula structure and multi-dimensional similarity calculation are realized.

Benefits of technology

It improves the accuracy and efficiency of formula equivalence judgment, can better understand the semantic and structural characteristics of formulas, and improves the accuracy of information retrieval and theorem proof.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336873A_ABST
    Figure CN120336873A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graph neural networks and formula representation, and provides a formula equivalence judgment method based on formula structure embedding and a graph neural network. Comprising a formula standardization module, a formula structure feature extraction module, a graph embedding module, a graph-level similarity comparison module, a node-level similarity comparison module and a similarity score calculation module. The formula standardization module is used for receiving original formula pairs input by the system, unifying formula structures and transmitting the unified formula structures to the formula structure feature extraction module; the formula structure feature extraction module sets a structure information tag for each node of a formula tree and converts the structure information tag into a node feature vector. According to the method, the accuracy of formula equivalence judgment is improved by splicing similarity vectors, processing and calculating a correlation matrix by adopting a cross attention mechanism, adjusting the weight, mapping into a single similarity score through a full connection layer, calculating an error by utilizing a loss function and adjusting model parameters through a back propagation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of graph neural networks and formula representation, and particularly relates to a method for judging formula equivalence based on formula structure embedding and graph neural networks, aiming to solve the problem of judging the semantic equivalence of mathematical formulas in scenarios such as information retrieval and theorem proving. Background Art

[0002] In today's information age, mathematical formulas, as indispensable key information in scientific literature, their equivalence determination is of great significance for deeply understanding scientific literature and improving the efficiency and accuracy of information extraction and mathematical reasoning. Formula equivalence determination has a wide range of applications in many fields such as formal verification of mathematical theorems, automatic theorem proving, and information retrieval.

[0003] Traditional methods for judging formula equivalence mainly rely on rigorous reasoning processes and a large amount of calculations to achieve. With the continuous development of artificial intelligence technology, more and more models have been applied to the field of formula equivalence determination. Some models only rely on symbolic calculation to achieve formula equivalence determination, such as directly encoding the original character sequence of the formula into a low-dimensional semantic vector and using the KNN algorithm to find the top K equivalent formulas; there are also some models that attempt to introduce partial structural information of the formula, such as using the TreeNN network to generate the embedding vectors of the leaf nodes in the formula parse tree, and then combining them to generate the parent node vectors.

[0004] However, these methods all have certain limitations. Methods based on text or structure matching cannot fully capture the inherent semantic and hierarchical characteristics of formulas. Although large language models have made certain progress in formula similarity comparison, the linear processing method is prone to ignoring the structural information of formulas. Although graph neural networks are good at modeling the structural relationships of formulas, they still face many challenges in accurately defining the nodes and edges in the formula tree, especially in distinguishing nodes at different levels and modeling the edge relationships.

[0005] Therefore, we propose a method for judging formula equivalence based on formula structure embedding and graph neural networks. Summary of the Invention

[0006] The present invention proposes a method for judging formula equivalence based on formula structure embedding and graph neural networks, which solves the problem in related technologies of being difficult to efficiently and accurately judge the similarity of two formulas. By integrating multiple modules and mechanisms, it realizes the standardized processing, feature extraction, and multi-dimensional similarity calculation of formula structures, thereby improving the accuracy of formula equivalence determination.

[0007] The technical solution of the present invention is as follows:

[0008] A method for determining formula equivalence based on formula structure embedding and graph neural network, including a formula normalization module, a formula structure feature extraction module, a graph embedding module, a graph-level similarity comparison module, a node-level similarity comparison module, and a similarity score calculation module; the formula normalization module is used to receive the original formula pair input by the system, process the symbols with commutativity and associativity, unify the formula structure and transfer it to the formula structure feature extraction module; the formula structure feature extraction module sets structure information labels for each node of the formula tree and converts it into a node feature vector, updates the representation of the edge and then transfers it to the graph embedding module; the graph embedding module uses the architecture based on the graph convolutional network GCN to convert the formula into a low-dimensional vector representation, and transmits it to the graph-level similarity comparison module and the node-level similarity comparison module respectively; the graph-level similarity comparison module receives the global embedding vectors of two formulas, uses the neural tensor network NTN to calculate the overall similarity between the graph structures and transfers it to the similarity score calculation module; the node-level similarity comparison module calculates the similarity between the corresponding nodes in the two formula graphs by introducing a cross-attention mechanism and transfers it to the similarity score calculation module; the similarity score calculation module synthesizes the information of the graph-level similarity and the node-level similarity, and adjusts the parameters of the entire system model through the backpropagation algorithm. The finally output similarity score is the determination result of the similarity degree of the input formula pair by the system.

[0009] Further, when the formula normalization module processes the symbols with commutativity and associativity, for the operation symbols such as addition and multiplication that satisfy commutativity and associativity, the order of elements in the formula is adjusted through a specific algorithm, so that the formula structures with the same operation essence are unified, thereby reducing the complexity of subsequent processing.

[0010] Further, when the similarity score calculation module processes the spliced vectors using the cross-attention mechanism, an attention matrix is constructed, and the elements in the matrix are calculated according to the feature correlation of the two formulas in different dimensions, so as to highlight the feature dimensions that contribute significantly to the similarity judgment, and then calculate the correlation matrix more accurately and adaptively adjust the weights.

[0011] A method for determining formula equivalence based on formula structure embedding and graph neural network includes the following method steps:

[0012] Step 1: Process the symbols with commutativity and associativity, eliminate the structural differences caused by the symbol arrangement differences, and unify the structural representation of the formula;

[0013] Step 2: Set structure information labels for each node in the formula tree;

[0014] Step 3: Convert the formula after normalization and structure embedding processing into a low-dimensional vector representation;

[0015] Step 4: Introduce a neural tensor network to capture the complex interaction between the graph structures corresponding to the two formulas;

[0016] Step 5: Introduce the cross-attention mechanism to capture the complex associations and semantic information between nodes and provide local information for similarity score calculation;

[0017] Step 6: Combine the information of graph-level similarity and node-level similarity to obtain the final similarity score of the formula pair.

[0018] Furthermore, in step 1, when processing symbols with commutative and associative properties, for symbols with associative properties, operators with associativity are identified during the traversal process. If two identical and associable operators are found to be in a parent-child node relationship, a merge operation is performed, the child node is deleted and the child node of the child node is changed to the child node of the parent node, and the recursion is continued until all qualified associative operators are integrated into the n-ary tree structure; for commutative operators, all nodes of the abstract syntax tree are traversed, the order of operators and operands is defined, the child nodes of the commutative operator nodes are extracted and sorted from small to large, and reinserted under the corresponding commutative operator nodes to form a standardized representation.

[0019] Furthermore, in step 2, a type identifier T is assigned to the node, type information is given, the hierarchical position L of the node in the formula tree is determined, the position N of the node in the hierarchy is clarified, T, L, and N are converted into corresponding vectors and fused through embedding operations to form a node feature vector containing structural information, and the node ID is replaced with the node feature vector based on the edge information in the graph structure, and the edge representation is updated.

[0020] Furthermore, in step 3, the node features and edge representations with structural information labels output by the formula structure embedding module are obtained, and then a convolution operation is performed on the features of each node and its neighbors using the architecture based on the graph convolutional network GCN. With the help of multi-layer GCN processing, the node representation is continuously updated so that the node can absorb the feature information of the neighboring nodes while integrating its own features. The process is expressed by the formula:

[0021]

[0022] Among them, H (l) represents the node feature matrix of the lth layer, is the adjacency matrix with self-loops added, yes The degree matrix of (l) is the weight matrix of the lth layer, and σ(·) is the nonlinear activation function.

[0023] Further, in step 4, by receiving the global embedding vectors of the two formulas output by the formula embedding module and modeling the graph embedding with the neural tensor network (NTN), the non-linear interaction between the two graph embedding vectors is calculated to capture the complex interaction relationships between the graphs. The specific expression formula is as follows:

[0024]

[0025] where h i and h j are the graph-level feature representations of the two formula graphs after the self-attention mechanism. W [1:K] is a tensor parameter used to capture the quadratic interaction between the feature vectors. V and b are the weight matrix and bias term of the NTN, σ(·) is a non-linear activation function, and K is a hyperparameter that controls the number of interaction scores generated by the model for each graph embedding pair.

[0026] Further, in step 5, the node-level similarity comparison module introduces a cross-attention mechanism. For the node sets in the two formula graphs, the query matrix Q, key matrix K, and value matrix V of each node are first calculated;

[0027] The calculation formulas for the query matrix Q, key matrix K, and value matrix V of each node are as follows:

[0028]

[0029] Then, the similarity scores between the query matrix and the key matrix are calculated and normalized by the softmax function to obtain the attention weights of each node relative to other nodes;

[0030] The value matrix is weighted and summed according to the attention weights to obtain the new representation of each node considering the context information. Finally, the attention weights are multiplied by the value matrix V and summed to calculate the similarity scores between the corresponding nodes in the two formula graphs.

[0031] Further, in step 6, the graph-level similarity vector output by the graph-level similarity comparison module and the node-level similarity vector obtained by the node-level similarity comparison module are concatenated;

[0032] The concatenated vector is processed by the cross-attention mechanism to calculate the correlation matrix, adaptively adjust the weights of different levels of similarity information, and comprehensively integrate the information;

[0033] The integrated vector is mapped to a single similarity score through a fully connected layer. The single similarity score is compared with the true equivalent label, and the error is calculated using a loss function such as the mean square error loss. The model parameters are adjusted through the backpropagation algorithm.

[0034] The working principle and beneficial effects of the present invention are as follows:

[0035] 1. In the present invention, in view of the problems of diverse formula structures and complex processing due to differences in symbol arrangements, a formula standardization module is adopted in the solution. In this module, for symbols with commutative and associative laws, such as arithmetic operators like addition and multiplication, the order of formula elements is adjusted through a specific algorithm. For example, when dealing with associative property symbols, traverse and identify associative operators. If two identical and associative operators are in a parent-child node relationship, perform a merging operation, making the child nodes of the child node become the child nodes of the parent node, and recursively integrate them into an n-ary tree structure. For commutative operators, traverse the nodes of the abstract syntax tree, define the order of operators and operands, extract the sorted child nodes and re-insert them below the corresponding operator nodes to unify the formula structure representation, greatly reducing the subsequent processing complexity.

[0036] 2. In terms of insufficient node semantic mining, it is improved through a formula structure feature extraction module. Rich structure information labels are set for each node of the formula tree, including an assignment type identifier T to clarify the node type, determining the hierarchical position L of the node in the formula tree and its relative position N at its level. Then, through an embedding operation, T, L, and N are converted into vectors and fused to form a node feature vector. At the same time, based on the edge information of the graph structure, the node feature vector is used to replace the node ID to update the edge representation, providing more accurate node semantic information for subsequent graph neural network calculations and improving the accuracy of formula semantic understanding.

[0037] 3. Regarding the problem of the accuracy of overall formula similarity determination, it is solved by the collaborative work of multiple modules. The graph embedding module uses an architecture based on the graph convolutional network GCN to convert the formula into a low-dimensional vector representation, providing a suitable form for subsequent similarity comparison. The graph-level similarity comparison module introduces a neural tensor network NTN. By receiving the global embedding vector output by the formula embedding module, it calculates the overall similarity between graph structures. The node-level similarity comparison module introduces a cross-attention mechanism to calculate the similarity between corresponding nodes in two formula graphs. The similarity score calculation module synthesizes the similarity information at these two levels. By splicing the similarity vectors, it uses the cross-attention mechanism to process and calculate the correlation matrix, adjust the weights, maps them through a fully connected layer to a single similarity score, and then compares it with the true equivalence label. The error is calculated using a loss function and the model parameters are adjusted through the backpropagation algorithm. Finally, the determination result accurately reflecting the similarity degree of the formula pair is output, improving the accuracy of formula equivalence determination. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0039] Figure 1 It is a system framework diagram of the present invention;

[0040] Figure 2Flowchart of the method for determining formula equivalence based on formula structure embedding and graph neural network in the present invention;

[0041] Figure 3 Flowchart of step 1 in the present invention;

[0042] Figure 4 Flowchart of step 2 in the present invention;

[0043] Figure 5 Diagram of the execution order and logical relationship among step 3, step 4 and step 5 in the present invention;

[0044] Figure 6 Flowchart of step 6 in the present invention. Detailed implementation manners

[0045] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0046] Embodiment 1

[0047] Referring to Figure 1 , this embodiment proposes a method for determining formula equivalence based on formula structure embedding and graph neural network, including: formula normalization module, formula structure feature extraction module, graph embedding module, graph-level similarity comparison module, node-level similarity comparison module and similarity score calculation module. This system aims to efficiently and accurately determine the similarity of two formulas;

[0048] Specifically, the formula normalization module receives the original formula pair input by the system, unifies the formula structure by processing symbols with commutativity and associativity, provides a more regular formula representation for subsequent modules, and reduces the complexity of subsequent processing. The processed formula is transmitted to the formula structure information extraction module.

[0049] Through the formula structure information extraction module, structure information labels are set for each node of the formula tree, including type identifiers, hierarchical positions, relative positions, etc., and these information are converted into node feature vectors to update the representation of edges. After completion of processing, the node features and edge representations with structure information labels are transmitted to the graph embedding module.

[0050] The graph embedding module obtains the relevant information output by the previous module, and uses the architecture based on the graph convolutional network GCN to convert the formula into a low-dimensional vector representation, so that the formula can participate in subsequent similarity comparisons in a suitable form. The processed low-dimensional vector representations of the formula are respectively transmitted to the graph-level similarity comparison module and the node-level similarity comparison module.

[0051] The graph-level similarity comparison module is used to receive the global embedding vectors of the two formulas output by the graph embedding module, and use the Neural Tensor Network (NTN) to calculate the overall similarity between the graph structures, providing a reference for the global structural features of the final equivalence determination. The calculated graph-level similarity vector is passed to the similarity score calculation module.

[0052] The node-level similarity comparison module also receives the information output by the graph embedding module. By introducing the cross-attention mechanism, it calculates the similarity between the corresponding nodes in the two formula graphs, providing local information for the similarity score calculation. The generated node-level similarity vector is also passed to the similarity score calculation module.

[0053] The similarity score calculation module then synthesizes the information of the graph-level similarity and the node-level similarity. First, it concatenates the similarity vectors obtained from both, then uses the cross-attention mechanism to process the concatenated vector, calculates the correlation matrix and adaptively adjusts the weights. Then, it maps the integrated vector to a single similarity score through a fully connected layer. Finally, it compares the similarity score with the true equivalence label, calculates the error using the loss function, and adjusts the parameters of the entire system model through the backpropagation algorithm to improve the accuracy of formula equivalence determination. The finally output similarity score is the determination result of the system on the similarity degree of the input formula pair.

[0054] Embodiment 2

[0055] Referring to Figures 2 - 6 , this embodiment proposes a formula equivalence determination method based on formula structure embedding and graph neural network, including the following method steps:

[0056] The formula normalization module makes the feature representations of similar formulas closer, reduces the model processing complexity, and improves the retrieval accuracy. The functions of this module are realized by step 1:

[0057] Step 1: Process the symbols with commutativity and associativity, eliminate the structural differences caused by symbol arrangement differences, and unify the structural representation of the formulas. The specific method of step 1 is as follows:

[0058] Step 1.1 For the symbols with associative properties (such as plus sign, multiplication sign), identify the operators carrying associativity. During the traversal, if two identical and associative operators form a parent-child node relationship, perform the merge operation, delete the child node, and directly use the child node of the child node as the child node of the parent node, and continue to recurse until all eligible associative operators are effectively integrated into the n-ary tree structure.

[0059] Step 1.2 For commutative operators (such as addition and multiplication), traverse all nodes of the abstract syntax tree and define the order of each operator and operand. Extract all child nodes of the commutative operator node, sort them in ascending order, and then re-insert the sorted child nodes below the operator node to form a standardized representation.

[0060] The formula structure information extraction module aims to provide a better input for the subsequent graph neural network model and effectively improve the accuracy of formula similarity calculation. The functions of this module are realized in Step 2:

[0061] Step 2: Set structure information labels for each node in the formula tree to enhance the semantic expression ability of the nodes. The specific method of Step 2 is as follows:

[0062] Step 2.1 Assign a type identifier T to each node in the formula tree to clearly distinguish whether the node is an operator, an operand, or a specific mathematical operator, and endow the node with basic type information semantically.

[0063] Step 2.2 Determine the hierarchical position L of the node in the formula tree, which reflects the depth of the node in the overall structure of the formula and helps the model perceive the hierarchical organization of the formula.

[0064] Step 2.3 Define the relative position N of the node in its hierarchy to further accurately describe the structural role of the node in the formula.

[0065] Step 2.4 Convert the "identifier T", "hierarchy L", and "position N" of each node into corresponding vectors through an embedding operation, and then fuse these three embedding vectors to form a node feature vector containing structure information.

[0066] Step 2.5 According to the edge information recorded in the graph structure, replace the node ID with the node feature vector and update the representation of the edge to provide more accurate information for the subsequent calculation of the graph neural network.

[0067] Provide more accurate information for the subsequent calculation of the graph neural network.

[0068] The graph embedding module enables the formula to participate in the similarity comparison in a suitable form during the subsequent processing in the graph neural network. The functions of this module are realized in Step 3:

[0069] Step 3: Convert the formula after normalization and structure embedding processing into a low-dimensional vector representation. The specific method of Step 3 is as follows:

[0070] Step 3.1 Obtain the node features and edge representations with structure information labels output by the formula structure embedding module.

[0071] Step 3.2 Use the architecture based on the graph convolutional network (GCN) to perform convolutional operations on the features of each node and its neighbors. Through the processing of multiple layers of GCN, the node representations are continuously updated, so that the nodes not only contain their own features, but also incorporate the feature information of neighboring nodes. This process is represented by the formula:

[0072]

[0073] where \(H^{l}\) (l) represents the node feature matrix of the \(l\)-th layer, is the adjacency matrix with self-loops added, is the degree matrix of \(A^{l}\), defined as \(D_{ii}^l=\sum_j A_{ij}^l\), \(W^{l}\) (l) is the weight matrix of the \(l\)-th layer, \(\sigma(\cdot)\) is a non-linear activation function, usually ReLU. After passing through multiple layers of GCN, the node embeddings are ready to be input into the next stage.

[0074] The role of the graph-level similarity comparison module is to provide a global structural feature reference for the final equivalence determination. The functions of this module are implemented in Step 4:

[0075] Step 4: The graph-level similarity comparison module is used to measure the overall similarity between the graph structures corresponding to two formulas by introducing a neural tensor network to capture the complex interaction relationships between the graphs. The specific method of Step 4 is as follows:

[0076] Step 4.1 Receive the global embedding vectors of the two formulas output by the formula embedding module.

[0077] Step 4.2 Use the neural tensor network (NTN) to model the graph embeddings and calculate the non-linear interaction between the two graph embedding vectors. Its calculation formula is:

[0078]

[0079] where \(h^{i}\) i and \(h^{j}\) j are the graph-level feature representations of the two formula graphs after the self-attention mechanism, \(W^{k}\) [1:K] is the tensor parameter used to capture the quadratic interaction between the feature vectors, \(V\) and \(b\) are the weight matrix and bias term of the NTN, \(\sigma(\cdot)\) is a non-linear activation function, and \(K\) is a hyperparameter that controls the number of interaction scores generated by the model for each pair of graph embeddings. Through the modeling method of NTN, the high-order interaction information between the graph embeddings can be better captured. Compared with the simple inner product operation, NTN allows more complex similarity modeling, thus improving the accuracy of graph-graph similarity calculation.

[0080] The role of the node-level similarity comparison module is to compare the similarity between the corresponding nodes in the two formula graphs. The functions of this module are implemented in Step 5:

[0081] Step 5: Introduce a cross-attention mechanism to better capture the complex associations and semantic information between nodes and provide local information for the calculation of similarity scores. The specific method of Step 5 is as follows:

[0082] Step 5.1 For the node sets in two formula graphs, calculate the query matrix Q, key matrix K, and value matrix V for each node. The calculation formulas are as follows:

[0083]

[0084] Step 5.2 Calculate the similarity scores between the query matrix and the key matrix, and normalize them using the softmax function to obtain the attention weights of each node relative to other nodes.

[0085] Step 5.3 Perform a weighted sum on the value matrix according to the obtained attention weights to obtain a new representation of each node after considering the context information.

[0086] Step 5.4 Multiply the obtained attention weights by the value matrix V and sum them to obtain a new representation of each node after considering the context information. Based on this, calculate the similarity scores between the corresponding nodes in the two formula graphs.

[0087] Similarity score calculation module, the function of this module is to generate a final similarity score that accurately reflects the equivalence degree of two formulas, providing a clear quantitative result for formula equivalence determination:

[0088] Step 6: Integrate the information of graph-level similarity and node-level similarity to obtain the final similarity score of the formula pair. The specific method of Step 6 is as follows:

[0089] Step 6.1 Concatenate the graph-level similarity vector obtained by the graph-level similarity comparison module and the node-level similarity vector obtained by the node similarity comparison module.

[0090] Step 6.2 Use the cross-attention mechanism to process the concatenated vector, calculate the correlation matrix between the graph-level and node-level similarity vectors, and adaptively adjust the weights of different levels of similarity information to more comprehensively integrate the information.

[0091] Step 6.3 Map the integrated vector to a single similarity score through a fully connected layer;

[0092] Step 6.4 Compare the similarity score with the true equivalence label, calculate the error using a loss function (such as mean squared error loss), and adjust the model parameters through backpropagation algorithm to improve the accuracy of formula equivalence determination.

[0093] The formula equivalence determination system and method based on graph neural network proposed by the present invention effectively overcome many limitations of traditional formula equivalence determination methods by integrating multiple key modules such as formula structure embedding, formula standardization, and attention mechanism. It has broad application prospects in many fields such as formal verification of mathematical theorems, automated theorem proving, and information retrieval. For example, in the knowledge mining of academic literature, it can quickly and accurately identify similar formulas to assist researchers in efficiently conducting literature analysis and knowledge discovery; in the intelligent tutoring system in the education field, it can accurately judge the equivalence between the formulas input by students and the standard formulas, providing more targeted learning feedback and guidance for students.

[0094] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for determining formula equivalence based on formula structure embedding and graph neural network, characterized in that, It includes a formula standardization module, a formula structure feature extraction module, a graph embedding module, a graph-level similarity comparison module, a node-level similarity comparison module, and a similarity score calculation module; the formula standardization module is used to receive the original formula pair input by the system, process the symbols with commutativity and associativity, unify the formula structure, and transfer it to the formula structure feature extraction module; the formula structure feature extraction module sets structure information labels for each node of the formula tree and converts them into node feature vectors. After updating the representation of the edges, it is passed to the graph embedding module; the graph embedding module uses the architecture based on the graph convolutional network (GCN) to convert the formula into a low-dimensional vector representation and transmits it to the graph-level similarity comparison module and the node-level similarity comparison module respectively; the graph-level similarity comparison module receives the global embedding vectors of the two formulas, uses the neural tensor network (NTN) to calculate the overall similarity between the graph structures, and transfers it to the similarity score calculation module; the node-level similarity comparison module calculates the similarity between the corresponding nodes in the two formula graphs by introducing a cross-attention mechanism and transfers it to the similarity score calculation module; the similarity score calculation module synthesizes the information of the graph-level similarity and the node-level similarity, and adjusts the parameters of the entire system model through the backpropagation algorithm. The finally output similarity score is the determination result of the similarity degree of the input formula pair by the system.

2. The formula equivalence determination method based on formula structure embedding and graph neural network according to claim 1, characterized in that When the formula standardization module processes the symbols with commutativity and associativity, for the operation symbols such as addition and multiplication that satisfy commutativity and associativity, it adjusts the order of the elements in the formula through a specific algorithm, so that the formula structures with the same operation essence reach unification, thereby reducing the complexity of subsequent processing.

3. The formula equivalence determination method based on formula structure embedding and graph neural network according to claim 1, characterized in that When the similarity score calculation module processes the concatenated vectors using the cross-attention mechanism, it constructs an attention matrix. The elements in the matrix are calculated based on the feature correlations of the two formulas in different dimensions to highlight the feature dimensions that contribute significantly to the similarity judgment, and then calculate the correlation matrix more accurately and adaptively adjust the weights.

4. A method for determining formula equivalence based on formula structure embedding and graph neural network, characterized in that, It includes the following method steps: Step 1: Process the symbols with commutativity and associativity, eliminate the structural differences caused by the symbol arrangement differences, and unify the structural representation of the formula; Step 2: Set structure information labels for each node in the formula tree; Step 3: Convert the formula after standardization and structure embedding processing into a low-dimensional vector representation; Step 4: Introduce a neural tensor network to capture the complex interaction relationships between the graph structures corresponding to the two formulas; Step 5: Introduce a cross-attention mechanism to capture the complex correlations and semantic information between nodes and provide local information for the similarity score calculation; Step 6: Synthesize the information of the graph-level similarity and the node-level similarity to obtain the final similarity score of the formula pair.

5. A method for determining formula equivalence based on formula structure embedding and graph neural network according to claim 4, characterized in that In step 1, when processing symbols with commutative and associative properties, for symbols with associative properties, operators with associativity are identified during the traversal process. If two identical and associable operators are found to be in a parent-child node relationship, a merge operation is performed to delete the child node and change the child node of the child node to the child node of the parent node, and the recursion is continued until all eligible associative operators are integrated into the n-ary tree structure; For commutative operators, traverse all nodes of the abstract syntax tree, define the order of operators and operands, extract the child nodes of the commutative operator nodes and sort them from small to large, and reinsert them under the corresponding commutative operator nodes to form a standardized representation.

6. The formula equivalence determination method based on formula structure embedding and graph neural network according to claim 4, wherein The step 2 assigns a type identifier T to the node, gives type information, determines the hierarchical position L of the node in the formula tree, clarifies the position N of the node in the hierarchy, converts T, L, and N into corresponding vectors through embedding operations and merges them to form a node feature vector containing structural information, replaces the node ID with the node feature vector based on the edge information in the graph structure, and updates the edge representation.

7. A method for determining formula equivalence based on formula structure embedding and graph neural network according to claim 4, characterized in that In step 3, the node features and edge representations with structural information labels output by the formula structure embedding module are obtained, and then the convolution operation is performed on the features of each node and its neighbors using the architecture based on the graph convolution network GCN. With the help of multi-layer GCN processing, the node representation is continuously updated so that the node can absorb the feature information of the neighboring nodes while integrating its own features. The process is expressed by the formula: Among which H (l) represents the node feature matrix of the l-th layer, is the adjacency matrix with self-loops added, is the degree matrix of, defined as, W (l) is the weight matrix of the l-th layer, and σ(·) is the non-linear activation function.

8. The method for determining formula equivalence based on formula structure embedding and graph neural network according to claim 4, wherein In step 4, the global embedding vectors of the two formulas output by the formula embedding module are received, and the graph embedding is modeled with the help of the neural tensor network NTN, so as to calculate the nonlinear interaction between the two graph embedding vectors and capture the complex interaction relationship between the graphs. The specific expression formula is: where h i and h j are the graph-level feature representations of the two formula graphs after the self-attention mechanism, W [1:K] is a tensor parameter used to capture the quadratic interactions between the feature vectors, V and b are the weight matrix and bias term of the NTN, σ(·) is a non-linear activation function, and K is a hyperparameter that controls the number of interaction scores produced by the model for each graph embedding pair.

9. A method for determining formula equivalence based on formula structure embedding and graph neural network according to claim 4, characterized in that In step 5, the node-level similarity comparison module first calculates the query matrix Q, key matrix K and value matrix V of each node for the node sets in the two formula graphs by introducing a cross attention mechanism; Calculate the query matrix Q, key matrix K and value matrix V of each node. The calculation formula is expressed as: Then the similarity score between the query matrix and the key matrix is calculated and normalized by the softmax function to obtain the attention weight of each node relative to other nodes; The value matrix is weighted and summed according to the attention weight to obtain a new representation of each node after considering the context information; finally, the attention weight is multiplied by the value matrix V and summed, based on which the similarity score between the corresponding nodes in the two formula graphs is calculated.

10. A method for determining formula equivalence based on formula structure embedding and graph neural network according to claim 4, characterized in that The step 6 is to concatenate the graph-level similarity vector output by the graph-level similarity comparison module and the node-level similarity vector obtained by the node-level similarity comparison module; The cross-attention mechanism is used to process the splicing vectors, calculate the correlation matrix, adaptively adjust the weights of similarity information at different levels, and comprehensively integrate the information; The integrated vector is mapped into a single similarity score through a fully connected layer, the single similarity score is compared with the true equivalent label, the error is calculated using loss functions such as mean square error loss, and the model parameters are adjusted through the back propagation algorithm.