Modular SoC Inference Engine for Dynamic Neural Network Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI/ML models face challenges in efficiently updating and adapting in the field due to high manufacturing costs, power consumption, and the resource-intensive nature of re-training, which hampers widespread deployment in devices with limited computing power.
Innovation Solution
A modular System-on-a-Chip (SoC) inference engine with a hub-and-spoke topology, referred to as CHAMELEON, allows dynamic model updates and adaptations by loading simple data constants, reducing wire count and power consumption, and enabling real-time modifications without re-programming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional AI/ML models are deployed in devices with limited computing power, then model accuracy and performance can be maintained, but manufacturing costs and power consumption increase significantly
Solution Approach 1:
The patent segments the AI/ML model into multiple layers with different computational requirements. By dividing the model into hierarchical layers, the system can selectively execute only the necessary layers based on device capabilities, reducing power consumption while maintaining accuracy for critical functions.
Solution Approach 2:
The patent applies local quality by optimizing different parts of the neural network with different precision levels. Critical layers use higher precision computations while less critical layers use lower precision, reducing overall power consumption while maintaining necessary model accuracy for specific tasks.
2Adaptability or versatility
If AI/ML models are updated in the field with comprehensive changes, then model adaptability and performance improve, but re-training resources and time requirements increase
Solution Approach 1:
The patent performs preliminary actions by pre-computing and storing intermediate results and model parameters in a hierarchical structure. When field updates are needed, the system can quickly apply changes by loading pre-computed parameters rather than performing complete re-training, significantly reducing update time while maintaining adaptability.
Solution Approach 2:
The patent implements dynamic model updating capabilities that allow the system to adaptively select which model layers to update in the field. The hierarchical architecture enables partial updates of specific layers based on changing requirements, providing adaptability without requiring complete model re-training.
3Adaptability or versatility
If comprehensive model changes are supported in the field, then model versatility improves, but hardware complexity and manufacturing costs increase
Solution Approach 1:
The patent implements a universal hierarchical data transfer topology that can accommodate various model architectures and update scenarios using the same hardware structure. The hub-and-spoke architecture provides multi-functionality by supporting different neural network layers, update strategies, and model types without requiring specialized hardware for each case.
Solution Approach 2:
The patent introduces intermediary hierarchical layers and data transfer structures that mediate between the input data and final output. These intermediary structures enable flexible model changes by providing standardized interfaces and buffer zones that simplify the hardware design while supporting comprehensive model adaptability.
4Ease of operation
If dynamic model updates are implemented, then operational flexibility improves, but system complexity and maintenance costs increase
Solution Approach 1:
The patent segments the model update process into manageable hierarchical stages, where each layer can be independently updated and validated. This segmentation simplifies the operational process by allowing incremental updates rather than requiring complete system reconfiguration, reducing operational complexity while maintaining flexibility.
Solution Approach 2:
The patent implements feedback mechanisms that monitor model performance and automatically trigger updates when needed. The hierarchical architecture enables feedback at multiple levels, allowing the system to adapt operationally based on performance metrics without requiring complex manual intervention, thereby improving ease of operation.
Data Source
AI summary
An electronic circuit system implementing and executing machine learning inference engines. While ML inference engines are based on (architectures and parameters defined by) configured, trained and tuned machine learning models, our design has the novel ability to support data driven, on-the-fly-reconfigured model runs. Reconfiguration and tuning operations include dynamic computational graph modifications, define-by-run alterations, changes to network depth (number of layers) and width (neurons per layer), and adjustments to weights, biases, plus activation function parameters. Neural networks supported include Feed-Forward, RNN, CNN, and Hopfield architectures, plus Ensemble, Federated, Cooperating, Adversarial, and Swarm collections. Decision Trees and Forests are also supported, as are more esoteric approaches such as ART and KAN. Our invention is capable of running both standalone and cooperatively, the cooperative processing being local and/or remote/cloud based, interfacing with telemetry applications to feed data, and machine learning software to feed new or updated models.


