Generative AI Query Routing Under Variable Processing Load

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing systems using generative AI face efficiency and speed issues due to resource shortages under high usage loads, leading to decreased performance in generating response texts.

Innovation Solution

An information processing apparatus and method that adjusts the number of generative AI models used based on load status, employing a control unit to determine the appropriate number of models for generating response texts, thereby optimizing resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple generative AI models are used to generate response texts, then the quality and accuracy of responses improve, but the resource consumption and processing time increase under high load conditions

Engineering Contradiction:
Improveresponse qualityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts the number of generative AI models based on real-time load status. When load is high, fewer models are activated to maintain processing speed; when load is low, more models are activated to improve response quality. This dynamic configuration resolves the contradiction between reliability and productivity by adapting resource allocation to current system conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of model quantity based on load status thresholds. By monitoring system load and adjusting the number of active models accordingly, the system optimizes the balance between response quality and processing speed, preventing resource exhaustion while maintaining service quality.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the number of generative AI models is increased to handle high demand, then service coverage improves, but resource shortage and system overload worsen

Engineering Contradiction:
Improveservice coverageVSAvoidresource availability
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system employs dynamic model activation where the number of generative AI models in service adjusts according to current load status. This prevents permanent resource allocation that would cause shortages, while still providing adequate service coverage when demand is high. The adaptive nature maintains serviceability without exhausting resources.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system automatically monitors its own load status and self-adjusts the number of active models without external intervention. This self-service mechanism ensures that resource allocation optimally balances service coverage and resource availability, preventing both overload and underutilization.

Inventive Principle:
Principle #25Self-service

3Power

If more computational resources are allocated to generative AI models, then response generation capability improves, but system load and processing efficiency deteriorate

Engineering Contradiction:
Improvecomputational capabilityVSAvoidprocessing efficiency
Core Design Contradiction:
PowerVSLoss of energy

Solution Approach 1:

The system changes the operational parameters by adjusting the number of active models based on load status. This prevents constant high-power consumption that would reduce processing efficiency, while still providing sufficient computational capability when needed. The parameter adjustment optimizes the power-efficiency tradeoff.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260057189A1Information processing apparatus and information processing method
Publication Date: 2026.02.26 TOSHIBA TEC KK
  • US20260057189A1 patent drawing
  • US20260057189A1 patent drawing
  • US20260057189A1 patent drawing

AI summary

According to one embodiment, an information processing apparatus includes a storage unit, a communication unit, and a control unit. The control unit is configured to receive a query text via the communication unit, then generate a prompt based on the query text. The control unit also acquires present load status information corresponding to a current workload of the control unit. The number of generative AI models to which the generated prompt is to be input is determined based at least in part on the present load status information. The control unit receives a response text to the prompt from the determined number of generative AI models, and then outputs a query response text via the communication unit. The query response text reflects each received response text. In some examples, the number of times the prompt is input to a generative AI model may be set based on the current workload.