ML Model Training via Symmetry-Equivariant Kernel Evolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models that rely on probability distributions for sensor data in computer-controlled systems face computational inefficiencies and inaccuracies due to the need for expensive sampling methods, which limits model complexity and training dataset size.
Innovation Solution
The proposed solution involves using symmetries inherent in the sensor data to improve sampling efficiency. By making the probability distribution invariant to these symmetries and using an adapted Stein Variational Gradient Descent (SVGD)-like evolution, the method accounts for symmetries as an inductive bias, leading to more efficient and accurate sampling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sampling methods are used to train machine learning models based on probability distributions of sensor data, then the model can accurately represent the probability distribution and handle uncertainty in sensor data, but the computational cost becomes very expensive, limiting model complexity and training dataset size
Solution Approach 1:
The patent changes the approach from traditional sampling methods to a deterministic transformation method. Instead of randomly sampling from the probability distribution, the method transforms training data points through a learned transformation function to generate samples. This parameter change from stochastic sampling to deterministic transformation significantly reduces computational cost while maintaining the ability to represent the probability distribution accurately.
Solution Approach 2:
The patent replaces the mechanical sampling process (which requires repeated random draws and is computationally intensive) with a learned transformation mechanism. The transformation function, learned from training data, directly maps input data to samples from the probability distribution, substituting the expensive sampling mechanism with a more efficient learned mapping.
2Quantity of substance
If traditional sampling methods are used for training, then the model can be trained on larger datasets, but the training process becomes computationally expensive and time-consuming
Solution Approach 1:
The patent performs preliminary learning of the transformation function during training. Once learned, this transformation function can be applied repeatedly and efficiently to generate samples from the probability distribution. This preliminary action of learning the transformation enables fast sample generation without requiring time-consuming sampling operations during each training iteration, thus reducing overall training time while allowing use of larger datasets.
3Measurement precision
If samples are drawn from the probability distribution to approximate expected values, then the model can make accurate inferences, but the sampling process is computationally expensive and limits the complexity of the model
Solution Approach 1:
The patent creates a learned transformation function that copies the essential characteristics of the probability distribution. Instead of using complex sampling algorithms to approximate the distribution, the transformation function learns to directly produce samples that represent the probability distribution. This copying approach simplifies the model while maintaining accuracy in expected value approximation.
Data Source
AI summary
A computer-implemented method of training a machine learnable model for controlling and/or monitoring a computer-controlled system. The machine learnable model is configured to make inferences based on a probability distribution of sensor data of the computer-controlled system. The machine learnable model is configured to account for symmetries in the probability distribution imposed by the system and/or its environment. The training involves sampling multiple samples of the sensor data according to the probability distribution. Initial values are sampled from a source probability distribution invariant to the one or more symmetries. The samples are iteratively evolved according to a kernel function equivariant to the one or more symmetries. The evolution uses an attraction term and a repulsion term that are defined for a selected sample in terms of gradient directions of the probability distribution and of the kernel function for the multiple samples.


