Modular SoC Inference Engine for Dynamic Multi-Model Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI/ML models face challenges in efficiently updating and adapting in the field due to high manufacturing costs, power consumption, and re-training requirements, limiting their deployment in diverse devices.
Innovation Solution
A modular System-on-a-Chip (SoC) inference engine with a hub-and-spoke topology (CHAMELEON) allows dynamic model changes and updates, reducing hardware costs and energy consumption by enabling fast parameter loading and adjustments without re-programming, using a Halt/Update Reload/Resume Interface (HURRI) protocol and flexible circuit designs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional AI/ML models are deployed in diverse devices, then model inference capability is provided, but manufacturing costs and power consumption increase
Solution Approach 1:
The patent segments the AI/ML model into multiple instances that can share common infrastructure (hub nodes, weight memory, bias memory) while maintaining independent inference capabilities. This segmentation allows multiple models to coexist on the same chip, reducing per-model power consumption and manufacturing costs while maintaining versatility across diverse devices.
Solution Approach 2:
The patent creates a universal inference engine architecture where a single chip can support multiple AI/ML model instances with different configurations. The hub-and-spoke topology with shared resources enables one device to perform multiple inference functions, reducing the need for device-specific hardware and lowering both manufacturing costs and power consumption per model.
2Adaptability or versatility
If AI/ML models are updated in the field, then model adaptability improves, but re-training requirements and complexity increase
Solution Approach 1:
The patent implements dynamic model updating through the HURRI protocol, allowing weight and bias parameters to be modified at runtime without reprogramming the entire model architecture. The hub nodes can dynamically load new parameter sets from external sources, enabling field updates while maintaining a fixed, simplified hardware structure that reduces overall system complexity.
Solution Approach 2:
The patent introduces hub nodes as intermediary components that mediate between the fixed hardware architecture and the flexible model parameters. These hub nodes handle parameter storage, retrieval, and updating, isolating the complexity of model management from the core inference engine and enabling simple hardware with complex, updatable model behavior.
3Productivity
If multiple ML models share inference-instance data, then resource efficiency improves, but data management complexity increases
Solution Approach 1:
The patent merges multiple model instances onto a single chip, allowing them to share common infrastructure including hub nodes, weight memory, and bias memory. This consolidation improves resource efficiency by eliminating redundant components while the standardized hub-and-spoke architecture manages data flow between shared and instance-specific resources, keeping data management complexity可控.
Solution Approach 2:
The patent uses a template-based approach where a single hub node design serves as a copyable template for multiple instances. Each model instance copies the hub node structure and associated memory resources, enabling efficient resource sharing while maintaining independent data paths. This copying strategy reduces data management complexity by providing a standardized, repeatable pattern for handling multiple models.
Data Source
AI summary
An electronic circuit system implementing and executing machine learning inference engines. While ML inference engines are based on (architectures and parameters defined by) configured, trained and tuned machine learning models, our design has the novel ability to support data driven, on-the-fly-reconfigured model runs. Reconfiguration and tuning operations include dynamic computational graph modifications, define-by-run alterations, changes to network depth (number of layers) and width (neurons per layer), and adjustments to weights, biases, plus activation function parameters. Neural networks supported include Feed-Forward, RNN, CNN, and Hopfield architectures, plus Ensemble, Federated, Cooperating, Adversarial, and Swarm collections. Decision Trees and Forests are also supported, as are more esoteric approaches such as ART and KAN. Our invention is capable of running both standalone and cooperatively, the cooperative processing being local and/or remote/cloud based, interfacing with telemetry applications to feed data, and machine learning software to feed new or updated models.


