RFQ Pricing Control with Reinforcement Learning and Symbolic Regression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing RFQ pricing systems lack the ability to dynamically adjust pricing strategies to minimize tracking error with respect to target metrics in financial markets, particularly in the context of market-making entities, and require a method to optimize pricing using reinforcement learning and symbolic regression for increased interpretability.

Innovation Solution

A method utilizing reinforcement learning and symbolic regression to formulate a Markov Decision Process (MDP) for defining a pseudo-optimal nonlinear state-feedback controller, incorporating a Monte-Carlo simulation to tune the RFQ pricing controller, and applying a predetermined neural network for algebraic approximation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning and symbolic regression are used to create a pseudo-optimal nonlinear state-feedback controller, then tracking error is minimized and pricing strategy adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvepricing strategy adaptabilityVSAvoidcontroller complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the pricing controller into distinct functional modules: a reinforcement learning agent for strategic decision-making, a symbolic regression component for mathematical model extraction, and a neural network for algebraic approximation. This modular architecture allows each component to handle specific aspects of pricing optimization independently, reducing overall system complexity while maintaining high adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary algebraic approximation layer between the complex reinforcement learning controller and the pricing output. This intermediary component simplifies the controller's behavior by finding polynomial representations that capture essential pricing dynamics, making the system more interpretable and easier to implement while preserving optimal pricing strategy adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If a pseudo-optimal nonlinear state-feedback controller is used to minimize tracking error, then pricing precision is improved, but computational requirements and system complexity increase

Engineering Contradiction:
Improvepricing precisionVSAvoidcontroller complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system dynamically changes parameters based on market conditions by using reinforcement learning to adapt controller behavior in real-time. The symbolic regression component extracts mathematical relationships from observed market data, allowing the controller to adjust pricing parameters optimally without requiring complex manual tuning, thus achieving high pricing precision with manageable complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a simplified algebraic approximation copy of the complex reinforcement learning controller's behavior. This approximation captures the essential pricing logic in a more manageable mathematical form, reducing computational requirements while maintaining pricing precision. The approximation serves as a simplified model that can be executed more efficiently than the full reinforcement learning agent.

Inventive Principle:
Principle #26Copying

3Speed

If reinforcement learning algorithms are applied to tune the RFQ pricing controller, then responsiveness to market dynamics is improved, but computational resources and processing time increase

Engineering Contradiction:
Improveresponsiveness to market dynamicsVSAvoidcomputational resources
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by pre-training the reinforcement learning agent in simulation environments before deploying to live markets. The symbolic regression component pre-extracts mathematical models from training data, and the algebraic approximation is pre-computed. This preliminary preparation allows the actual pricing controller to respond quickly to market dynamics without requiring intensive real-time computational resources.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a simplified algebraic approximation copy of the reinforcement learning controller's decision-making process. This approximation preserves the responsive behavior learned during training but executes much more efficiently, reducing real-time computational requirements while maintaining speed of response to market changes.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250252492A1Method and system for dynamic request for quotation pricing using reinforcement learning and symbolic regression
Publication Date: 2025.08.07 JPMORGAN CHASE BANK NA
  • US20250252492A1 patent drawing
  • US20250252492A1 patent drawing
  • US20250252492A1 patent drawing

AI summary

A method and a system for dynamic request for quotation (RFQ) pricing using reinforcement learning and symbolic regression in order to obtain a pseudo-optimal nonlinear state-feedback controller for tracking a metric with increased interpretability are provided. The method includes: receiving bid price information and ask price information that relates to a financial instrument; selecting a metric to be used in conjunction with a determination of an RFQ price with respect to the financial instrument; formulating a Markov Decision Process (MDP) that relates to the determination of the RFQ price; using the MDP to define an RFQ pricing controller with respect to the metric; applying a reinforcement learning algorithm in order to tune a behavior of the RFQ pricing controller with respect to the metric; and using the RFQ pricing controller, the first information, and the second information to determine the RFQ price.