Stochastic Gradient Langevin Boosting for Non-Convex Loss Convergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Gradient Boosting (GB) methods for building decision-tree Machine Learning Algorithms (MLAs) may not always converge to the global minimum, especially when dealing with non-convex loss functions, potentially getting stuck in local minima or saddle points.

Innovation Solution

The implementation of Stochastic Gradient Langevin Boosting (SGLB), which combines GB with stochastic gradient Langevin dynamics, allows for the generation of stochastic estimated gradient values by shrinking predictions and adding noise, enabling the algorithm to 'jump' between different saddles of a non-convex loss function and converge globally.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If Gradient Boosting methods are used to build decision-tree MLAs, then the algorithm can be trained to make predictions, but it may not converge to the global minimum and can get stuck in local minima or saddle points

Engineering Contradiction:
Improveconvergence to global minimumVSAvoidalgorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies Langevin dynamics by introducing temperature parameter T and noise term ζ to the gradient boosting algorithm. The parameter transformation converts deterministic gradient updates into stochastic updates, allowing the algorithm to escape local minima and saddle points while converging to the global minimum of the loss function.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transforms the static gradient boosting algorithm into a dynamic system by incorporating Langevin dynamics. The algorithm now exhibits stochastic behavior through temperature-controlled noise, enabling it to dynamically explore the loss landscape and transition between different states (local minima, saddle points) to reach the global minimum.

Inventive Principle:
Principle #15Dynamics

2Productivity

If standard Gradient Boosting is used, then training can proceed efficiently, but the algorithm lacks exploration capability and cannot jump between saddles of non-convex loss functions

Engineering Contradiction:
Improvetraining efficiencyVSAvoidexploration capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces temperature parameter T that controls the noise level in gradient updates. By adjusting T, the algorithm can balance between exploitation (low T, efficient training) and exploration (high T, ability to jump between saddles), providing adaptability while maintaining training efficiency.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces noise term ζ as an intermediary element that mediates between the deterministic gradient direction and the stochastic exploration need. This noise acts as a bridge, allowing the algorithm to occasionally deviate from the standard gradient path to escape local minima while generally following the efficient gradient descent trajectory.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250148301A1Methods and systems for training a decision-tree based machine learning algorithm (MLA)
Publication Date: 2025.05.08 Y E HUB ARMENIA LLC
  • US20250148301A1 patent drawing
  • US20250148301A1 patent drawing
  • US20250148301A1 patent drawing

AI summary

Methods and servers for of training a decision-tree based Machine Learning Algorithm (MLA) are disclosed. During a given training iteration, the method includes generating prediction values using current generated trees, generating estimated gradient values by applying a non-convex loss function, generating a first plurality of noisy estimated gradient values based on the estimated gradient values, generating a plurality of noisy candidate trees using the first plurality of noisy estimated gradient values, applying a selection metric to select a target tree amongst the plurality of noisy candidate trees, generating a second plurality of noisy estimated gradient values based on the plurality of estimated gradient values, generating an iteration-specific tree based on the target tree and the second plurality of noisy estimated gradient values, and storing, the iteration-specific tree to be used in combination with the current generated trees.