Hierarchical Query Classification Cascade for Policy-Aware Agent Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional query classification and routing systems in AI environments with distributed and access-restricted data sources lack adaptability and context-awareness, leading to inefficiencies and compliance issues, particularly in regulatory and IoT deployments.

Innovation Solution

A hierarchical model cascade architecture that dynamically routes queries through a network of AI agents, each accessing authorized data repositories, with each level evaluating queries based on distinct criteria such as semantic content, contextual metadata, and policy constraints, enabling dynamic escalation or bypass based on query complexity and system load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single monolithic language model is used for query classification and routing, then the system structure is simple, but the system lacks adaptability and context-awareness in distributed environments with access-restricted data sources

Engineering Contradiction:
ImproveadaptabilityVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the single monolithic language model into a hierarchical cascade of multiple specialized language models, where each model handles specific classification tasks at different levels. This segmentation enables the system to process queries through multiple specialized models that can adapt to different data sources and regulatory requirements, thereby improving adaptability while managing complexity through structured organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the query classification system, organizing models across multiple levels (first level, second level, third level) that process queries with increasing specificity. This dimensional structure allows the system to handle diverse query types and regulatory contexts at different abstraction levels, enhancing adaptability without requiring a completely complex monolithic architecture.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If queries are processed through multiple levels of classification models, then query classification accuracy and policy compliance improve, but processing time and computational resources increase

Engineering Contradiction:
Improvepolicy complianceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a dynamic bypass mechanism where the system can skip intermediate levels of the hierarchical model cascade based on query characteristics and previous classification results. This dynamic adjustment allows the system to maintain high policy compliance for complex queries while reducing processing time for simpler queries that can be resolved at lower levels, thus balancing reliability and time efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent enables the query classification system to skip intermediate classification levels when confidence thresholds are met or when query characteristics indicate that lower-level models suffice. This skipping mechanism reduces unnecessary computational steps and processing time while maintaining adequate policy compliance, as the bypassed levels would have provided diminishing returns for those particular queries.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Adaptability or versatility

If a hierarchical cascade of language models is implemented for multi-stage query classification, then query routing accuracy and adaptability improve, but system complexity and computational resource usage increase

Engineering Contradiction:
Improvecontext-awarenessVSAvoidcomputational resource usage
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by selectively engaging different levels of the hierarchical model cascade based on query complexity and characteristics. Rather than always processing queries through all three levels, the system uses only the necessary number of models required to achieve adequate classification accuracy and policy compliance, thereby reducing unnecessary computational resource consumption while maintaining context-awareness for complex queries.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250383970A1Hierarchical cascade architecture of language models for multi-stage query classification and agent routing
Publication Date: 2025.12.18 CITIBANK N A
  • US20250383970A1 patent drawing
  • US20250383970A1 patent drawing
  • US20250383970A1 patent drawing

AI summary

The systems and methods disclosed herein orchestrate task execution among autonomous (or semi-autonomous) AI agentic models (“agents”) responsive to a received query by using a hierarchical model cascade to classify queries into agent domains. Queries are processed iteratively by a series of hierarchical levels containing one or more AI models, where each layer is more complex and imposes fewer resource constraints. Each level generates a classification and a confidence score pertaining to the classification. A dynamic bypass mechanism analyzes the classifications and confidence scores at each level to dynamically determine if one or more levels of the hierarchy can be bypassed while resulting in an accurate classification. The final classifications are matched to one or more agents that process the query. Responses from the candidate agents are aggregated into an output that is responsive to the input.