Semantic Vector Encoding for Mathematical Expression Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for processing mathematical content in data search technologies are limited, as they often rely on contextual information that is not available in all cases, such as in pure math texts, and fail to capture cross-disciplinary similarities in mathematical equations.

Innovation Solution

A computer-implemented method that generates a saturated e-graph representation of mathematical expressions, applies rewrite rules, and uses an encoder-decoder framework to produce continuous vector representations that capture the semantic meaning of mathematical expressions, enabling effective clustering and retrieval across different contexts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If mathematical expressions are encoded according to textual context, then equations in similar contexts can be found, but cross-disciplinary retrieval is hampered and equations without context cannot be processed

Engineering Contradiction:
Improvecross-disciplinary retrieval capabilityVSAvoidcontext dependency
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent extracts the mathematical expression from its textual context and encodes it independently using e-graphs and rewrite rules. This allows the expression to be processed and retrieved without relying on surrounding text, enabling cross-disciplinary search while maintaining mathematical meaning.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the encoding parameters from context-based textual features to mathematics-based structural features using e-graph representations. This parameter change enables the system to capture semantic meaning through mathematical equivalence rather than contextual similarity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If existing embedding methods are used for mathematical expressions, then simple transformations can be applied, but they fail to capture semantic meaning and mathematical equivalence

Engineering Contradiction:
Improvesemantic meaning captureVSAvoidencoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary transformation of mathematical expressions into e-graph representations before encoding. This preliminary action captures the structural and semantic properties of expressions, enabling more accurate semantic encoding while the complexity is managed through systematic rewrite rules.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces e-graphs as an intermediary representation between the original mathematical expression and the final embedding. This intermediary structure preserves mathematical equivalence and semantic meaning while providing a systematic framework for encoding that manages complexity through formal rewrite rules.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240135148A1Semantic Representations of Mathematical Expressions in a Continuous Vector Space and Generation of Different but Mathematicallly Equivalent Expressions and Applications Thereof
Publication Date: 2024.04.25 THE BOARD OF TRUSTEES OF THE UNIV OF ILLINOIS
  • US20240135148A1 patent drawing
  • US20240135148A1 patent drawing
  • US20240135148A1 patent drawing

AI summary

Methods are provided herein for training and using models to generate semantically representative continuous vectors for input mathematical expressions. These methods result in models that output continuous vectors that are nearby in an embedding space for equations that are mathematically equivalent but differently written. Such continuous vectors can be used to facilitate indexing and searching of databases of mathematical equations, e.g., to facilitate semantically-aware searching of databases of mathematical texts for equations that are mathematically equivalent to, or mathematically similar to, input query expressions. These training methods include training an encoder together with a decoder to predict pairs of mathematically equivalent but different training expressions, with the output of the encoder being the continuous vector that represents the semantic mathematical content of the pair of training expressions. Also provided are methods for efficiently generating such pairs of mathematically equivalent but different training expressions.