Selective Layer Activation Control System with Dynamic Capacity Allocation
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- TURKCELL TEKNOLOJI ARASTIRMA & GELISTIRME AS
- Filing Date
- 2026-05-18
- Publication Date
- 2026-06-22
Smart Images

Figure 00000015_0000
Abstract
Description
- 1 - TARIFF SELECTIVE LAYER ACTIVATION WITH DYNAMIC CAPACITY DISTRIBUTION CONTROL SYSTEM TECHNICAL FIELD The invention is based on the Large Language Model (LLM) inference. optimization, hardware-aware artificial intelligence system design, Key- Key-value (KV) cache management, neural network layer activation control, 10 machine learning, deep learning, and AI-based resource management in the field of systems, in inference frameworks that work on large language models, by processing query semantics, hardware telemetry, and performance metrics together. Dynamic resource allocation and selective layer activation are performed dynamically. It relates to a selective layer activation control system with capacity allocation. Invention 15 The subject of the system is LLM-based text generation and reasoning services, chatbot. platforms, enterprise information assistant systems and RAG (Retrieval-Augmented Software that develops Generation-based question-and-answer services, cloud computing, and It can be used in telecommunications applications. PREVIOUS TECHNIQUE In most cases, current Large Language Model inference optimization approaches either token-based budget management only or static tier skipping or KV cache relies on prefetch / eviction techniques. This situation, query 25 The inability to adequately predict complexity differences arising from semantics, The inability to proactively manage hardware resources, layers, and attention The inability to selectively mask heads and the multidimensional nature of capacity consumption This leads to technical problems such as the inability to control it properly. Current systems use selective KV 30 for pre-query complexity classification. cache segment loading, Hardware Activation Map (Hardware Activation Map (HAM) generation, multidimensional capacity vector, two thresholds - 2 - predictive capacity leakage control, quality-based early stop and PPO Proximal Policy Optimization-based meta-adaptation is modular and controlled. It does not integrate holistically in this way; it is proactive, guided by inquiry semantics. orchestration, hardware-aware dynamic masking, and closed-loop adaptive It does not include optimization. 5 Document number US20240370699A1, transformer-based large language In their models, data blocks, weight matrices, and KV- during inference a prefetch that aims to move cache data into memory beforehand It describes an acceleration approach. In this solution, layer-based data... While stream optimization and memory access latency reduction are included, 10 Pre-query C0–C4 complexity classification, Hardware Activation Map (HAM) generation, selective KV segment loading, multidimensional capacity vector, two thresholds Leak control, quality-based early stopping, and PPO meta-adaptation elements. It is not explicitly stated. Document number US20250036876A1, KV 15 in major language models. To reduce memory strain caused by the increase in cache size. It describes token eviction mechanisms. The solution in question involves caching. Eviction policies and token importance rating are in place, however, the query Complexity class-driven selective loading, hardware activation with HAM. integration, capacity leakage control and meta-adaptation integration with PPO 20 It is incomplete, and the complete dynamic source orchestration is not specified. LayerSkip and DASH-like documentation, transformer models It describes the mechanisms of layer skipping and premature exit. Layer activation control in solutions with token-level or static decisions. However, the pre-query predictive complexity classifier, hardware 25 Telemetry-based dynamic HAM generation, selective KV segmentation, multidimensional Capacity vector and PPO meta-adaptation elements are missing, and a proactive holistic approach is lacking. The system architecture is not specified. The current technical documentation presents fragmented teachings; pre-interrogation Complexity classification, selective KV loading, dynamic layer / head 30 with HAM masking, multidimensional capacity vector, two-threshold leakage control, quality-based Early stopping and PPO adaptation query semantics and hardware - 3 - presenting its integrated management around awareness in a single architecture. It does not include. Therefore, current solutions offer dramatic resource savings in simple queries. with maintaining response quality, full capacity utilization in complex queries, and Simultaneous user capacity on the same hardware infrastructure in a single integrated 5 It is unable to integrate them in architecture. A BRIEF DESCRIPTION OF THE INVENTION Large Language Model (LLM) inference optimization, 10 Hardware-aware artificial intelligence system design, Key-Value (KV) KV) cache management, neural network layer activation control, machine learning and query semantics in the field of AI-based resource management systems, a multi-layered system that processes hardware telemetry and performance metrics together The invention involves selective 15-bit dynamic capacity allocation for inference optimization. A layer activation control system has been developed. The developed system includes pre-query complexity classification and selective complexity classification. cache loading, Hardware Activation Map HAM production, multidimensional capacity vector generation, constrained LLM extraction, Capacity leakage control, quality-based early stop, result combining, PPO 20 (Proximal Policy Optimization) meta-adaptation and hardware profiling monitoring steps By working through it, we can create a more efficient, higher quality, and more scalable LLM. It provides the inference. The developed system includes a complexity classifier and selective KV loading. unit, HAM producer, capacity vector, two-threshold leak controller and quality 25 by holistically integrating the base-based stopping unit, unique orchestration and dynamically manages the decision-making process and uses hardware telemetry. It improves optimization quality through dependent adaptation. This integrated architecture provides 60–80% GPU / NPU load on simple queries. Maintaining response quality through reduction, full capacity 30 for complex queries. usage, simultaneous user capacity on the same infrastructure, and measurable energy. It provides savings. - 4 - DESCRIPTION OF THE FIGURES Figure 1. Selective Layer Activation with Dynamic Capacity Allocation. Control System System Architecture The corresponding part numbers shown in the figures are given below. 1. Pre-Query Complexity Classification and Prediction Layer 2. Complexity-Driven Selective KV Cache Loading and Segmentation Unit 3. Hardware Activation Map Generator 10 4. Multidimensional Dynamic Capacitance Vector Generator 5. LLM Inference Engine and Resource-Constrained Execution Layer 6. Two-Threshold Predictive Capacitance Leakage Controller 7. Recursive Quality-Based Early Stop Unit 8. Result Combining and Output Generator 15 9. PPO Meta-Customization and Parameter Update Layer 10. Hardware Profile Monitor and Signal Backlay Layer DETAILED DESCRIPTION OF THE INVENTION 20 Large Language Model (LLM) inference optimization, Hardware-aware artificial intelligence system design, Key-Value (KV) (KV) at least one in the field of cache management and neural network layer activation control. processor, memory, data storage units, Graphics Processing Unit (Graphics Processing 25) Unit - GPU), Neural Processing Unit (NPU), high High-capacity Video Random Access Memory (VRAM) one or more containing VRAM) and real-time telemetry sensors The invention is a dynamic device developed to be run on a computer / server. Capacity allocation and selective layer activation control system; proactive query 30 Pre-query complexity classification and prediction layer (1), which analyzes as follows: Complexity-driven selective KV that manages KV segments according to complexity. - 5 - cache loading and segmentation unit (2), depending on the hardware status activation mask generating hardware activation map generator (3), multidimensional Multidimensional dynamic capacity vector generator that defines resource limits (4), LLM inference engine and resource-constrained execution layer that performs constrained inference. (5), two-threshold predictive capacity leakage controller that predicts capacity overflows 5 (6), early stop unit based on recursive quality providing quality assurance (7), the result combining and output generator (8), which forms the final output, system PPO meta-adaptation and parameter updating that adapts parameters Hardware profile monitor and back up hardware resources with layer (9) It includes a downlink signal layer (10). 10 The invention relates to pre-query complexity classification and system analysis. Predictive layer (1); the incoming user query is processed by the main Large Language Model (Large A lightweight pre-checker that analyzes the Language Model (LLM) before it is run. It is a layer. This layer contains the raw query, the context window, and the history. Number of talks, tool references, domain tag (field 15 (tag), expected response format, and user request task type as input. It receives the query and assigns it to one of the C0–C4 complexity classes. The C0 class Factual recall or queries requiring a single unit of information, C1 Queries requiring single-step inference, C2 multi-step Problems requiring multi-step reasoning, using the C3 tool (tool 20) C4 is for tasks requiring deep planning, while C4 is for expert-level research. long-context synthesis or multi-branch reasoning Branch reasoning represents tasks requiring output complexity. ComplexityVector is a data structure. (ComplexityVector); class, confidence, est_token_cost 25 (estimated token cost), est_branch_count (estimated number of branches), est_kv_scope_ratio (estimated KV coverage ratio), task_type (task type), domain_risk_level and expected_tool_need (needs) includes components. If the confidence score falls below the threshold, the system By using the safe-side principle, we are raising the query to a higher complexity class and 30 It creates a proactive execution plan for processor and memory resources. - 6 - The invention describes a complexity-driven selective KV cache within the system. loading and segmentation unit (2); context window dividing into semantic segments and for each segment semantic similarity with the query, the segment within the speech freshness (recency_score), entity_match_density, 5 task-related tool reference (tool_reference_overlap) and previous success It calculates the records (prior_usefulness_score). KV_relevance(seg_i, q) = α · semantic_similarity(embed(q), embed(seg_i)) + β · recency_score(seg_i) + γ · entity_match_density(q, seg_i) + δ · tool_reference_overlap(q, seg_i) + ε · The segments are 10 with the prior_usefulness_score(seg_i) scoring function. They are prioritized, and only the closest and highest safety minimal in class C0. Segments are loading, limited historical context and related information in classes C1 and C2. Asset segments are being retrieved, C3 class vehicle call history, broader context. and task interim outputs are included, and full context can be loaded in class C4. However, eviction prioritization is performed. Output KV Cache 15 The Load Plan (KVCacheLoadPlan) is a data structure. KV Cache Load Plan (KVCacheLoadPlan); segment_ids (segment IDs), priority_scores (priority scores), load_order, eviction_candidates, kv_scope_ratio and memory_estimate It includes. This unit is Video Random Access Memory (Video Random Access 20 It minimizes memory (VRAM) usage in proportion to complexity and Graphics Processing Unit (GPU) / Neural Processing Unit The Neural Processing Unit (NPU) optimizes memory access. The hardware activation map generator included in the system that is the subject of the invention (3); ComplexityVector allows you to determine the instantaneous hardware state (VRAM 25). occupancy rate, NPU utilization, thermal headroom Using the area together, the Hardware Activation Map - RAM) produces. ActivationMask; layer_range (layer range), active_head_set (active attention header set), moe_expert_set (Mixture-of- Experts (expert set), quant_level (quantification level: INT4, INT8, BF16, FP16), 30 kv_precision_level (KV precision level), fallback_applied (fallback C0 includes the components (implemented) and hardware_route. - 7 - active_layer = [L-4, L] for INT4 and active_layer = [L-8, L] for C1 For INT8, C2, use active_layer = [L-16, L] and for INT8 + selective heads, C3. active_layer = [L-32, L] and for BF16, C4 active_layer = [0, L] and FP16 / BF16+ A wide range of MoE expert options are applied. VRAM_used / total > 0.85, Automatic reset if NPU_utilization > 0.90 or thermal_headroom < 5°C. A downgrade is being implemented and the Graphics Processing Unit (GPU) physical execution of Neural Processing Units (NPUs) The paths are dynamically controlled. The invention describes a multidimensional dynamic capacity vector within the system. generator (4); multidimensional 10 that deviates from token budget approaches It provides capacity management. Capacity vector B; reasoning_token_limit (judgment token limit), active_layer_range, branch_count (number of branches), kv_cache_scope (KV cache scope), tool_call_limit (tool call limit) limit), latency_ceiling, memory_bandwidth_cap width limit), max_parallel_decoding_paths (maximum parallel decoding 15 It includes the components (path) and quantization_policy. vector, Hardware Activation Map (HAM) and KV It works in conjunction with the loading decision. The invention concerns an LLM inference engine and resource-constrained system. Execution layer (5); KVCacheLoadPlan, 20 The main inference is performed using the ActivationMask and vector B. Only the KVCacheLoadPlan is included in the prefill phase. The segments inside are being loaded, ActivationMask is in the decode phase. The layer / head / expert sets allowed by the (Activation Mask) are active. is being held and the reasoning branches and the number of tool-calls (tool calls 25) The number of B vectors remains within the limits of this layer. This layer is a standard complete layer. Layered, fully loaded KV and allocated for each query instead of fixed precision inference. It operates within the limitations of established physical and logical resources. The invention describes a system involving two-threshold predictive capacity leakage. controller (6); by deriving the consumption rate during extraction, the budget exhaustion is 30 It predicts in advance. consumption_rate(t) > θ_warn In this case, branch_count is reduced, and the quantization level is lowered. - 8 - (The level of slack) is being reduced and low priority KV segments are being unloaded. (being) cleared; remaining_budget(t) < θ_critical An early stop signal is generated and quality_signal(t) > θ_quality The quality is sufficient when coverage_ratio(t) > θ_coverage. Production is completed early thanks to this observation. This unit is a closed-loop control 5. It operates as a mechanism. The invention relates to early detection based on recursive quality within the system. stop unit (7); quality_score after each N token production window It calculates the (quality score). quality_score; q1·semantic_coherence (semantic consistency) + q2·factual_grounding_score + 10 q3·coverage_ratio (coverage ratio) + q4·instruction_compliance (instruction compliance) - It is defined as q5·repetition_penalty. Medical, legal. or in areas requiring high accuracy, the weight of q2 is increased and θ_stop is further adjusted. It is being made stricter. If quality_score ≥ θ_stop, production is stopped, early In case of pause, the intermediate result with the highest quality score is recovered. 15 The system of invention includes a result combining and output generator (8); Interim results when production is completed or when an early stoppage occurs, It combines the reasoning summary and the final token sequence. The FinalOutput data structure includes: final_response, complexity_class (complexity class), quality_score, consumed_token_count 20 (number of tokens consumed), active_layer_usage (active layer usage), kv_scope_used, latency_ms (latency in milliseconds), and It contains fallback_events. The invention concerns the PPO meta-adaptation and parameters included in the system. Update layer (9); quality, cost and 25 after each inference cycle receiving delay signals and reward = wq·quality_score - wc·compute_cost_normalized - wl·latency_normalized - wm·memory_pressure_penalty - wf·fallback_penalty reward function with PPO Complexity-driven selective optimization using the Proximal Policy Optimization algorithm. θ_kv and α / β / γ / δ / ε 30 in KV cache loading and segmentation unit (2) their weights, fallback thresholds in the hardware activation map generator (3), two threshold predictive capacity leakage controller (6) and recursive quality based - 9 - Thresholds in early stop unit (7), q1–q5 quality weights in Module 7 and initial capacity in multidimensional dynamic capacity vector generator (4) It updates the vector values. The invention concerns a hardware profile monitor and a fallback system. Signal layer (10); Graphics Processing Unit (Graphics Processing Unit - 5 GPU) / Neural Processing Unit (NPU) / Random Access Memory (Random Access Memory - RAM) usage, Video Random Access Memory (Video Random Access Memory - VRAM) occupancy rate, thermal header area (thermal headroom), memory bandwidth and the number of simultaneous inferences are constantly being increased. While monitoring, if VRAM_used / total > 0.85, layer_range should be 10. throttling, consumption_rate (consumption rate) if NPU_utilization > 0.90 The warning indicates that the thermal_headroom < 5°C will trigger a system-wide Class C upgrade. It generates signals to lower its limit to C2. The invention concerns dynamic capacity allocation and selective layer activation. From receiving the user query in the control system to the optimized response, 15 a closed-loop and hardware-aware inference extending to its production It functions as an optimization pipeline. Architecture, pre-query complexity. Classification and prediction layer (1) with pre-query C0–C4 complexity classification and Complexity Vector generation It starts with this vector complexity-driven selective KV cache loading and 20 segmentation unit (2), hardware activation map generator (3) and multidimensional The dynamic capacity vector is distributed simultaneously to the generator (4); complexity-driven selective KV cache loading and segmentation unit (2) KVCacheLoadPlan, hardware activation map. Manufacturer (3) with ActivationMask and multidimensional dynamic 25 Capacity vector B is created with capacity vector generator (4); LLM Inference engine and resource-constrained execution layer (5) with these inputs restricted LLM Inference from Graphics Processing Unit (GPU) / Neural Processing It is being run on the Neural Processing Unit (NPU); two-threshold predictor. Two-threshold predictive leakage control with capacity leakage controller (6), recursive 30 Quality-based early stop unit (7) with quality_score (quality) in each N token (score) based recursive quality assessment and early stop - 10 - is being carried out; result merging and output generator (8) with FinalOutput is being produced and all with PPO meta-adaptation and parameter update layer (9) The system is made adaptive by updating the parameters. Hardware profile monitor and fallback signal layer (10) with continuous hardware telemetry It provides real-time feedback; thus, the system query semantics 5 Guided proactive orchestration, hardware-aware dynamics. layer / head / Mixture-of-Experts / quantization masking, multidimensional capacity management, predictive leak control and quality assurance with early stop, classic It differs from other approaches, using the Graphics Processing Unit (GPU) in simple queries. Processing Unit (GPU) / Neural Processing Unit (NPU) 10 It reduces the load by 0.80% while maintaining response quality and handling complex queries. It allows hardware resources to be used to their full capacity. 20 30
Claims
- 11 - SYSTEMS 1. Optimization of Large Language Model (LLM) inference, Hardware-aware artificial intelligence system design, Key-Value - KV) At least 5 in cache management and neural network layer activation control. a processor, memory, data storage units, Graphics Processing Unit (Graphics Processing Unit (GPU), Neural Processing Unit (NPU), High-capacity Video Random Access Memory (VRAM) one or more systems containing Memory - VRAM) and real-time telemetry sensors. The invention is designed to run on multiple computers / servers. It is a selective layer activation control system with dynamic capacity allocation, This feature proactively analyzes the query pre-query complexity. Classification and prediction layer (1), KV segments according to complexity Managing complexity-driven selective KV cache loading and segmentation unit (2), hardware that generates activation mask according to hardware status 15 activation map generator (3), defining multidimensional resource boundaries 3D dynamic capacity vector generator (4), LLM performing constrained inference inference engine and resource-constrained execution layer (5), capacity overflows Two-threshold predictive capacity leakage controller (6), quality assurance early stop unit (7) which provides recursive quality based on the final output 20 The resulting combination and output generator (8) forms the system parameters. adapting PPO meta-adaptation and parameter update layer (9) Hardware profile monitor that monitors hardware resources and reduces signal feedback. It is characterized by containing layer (10).
2. Selective layer activation control system according to claim 1, its feature is; raw 25 query, context window, history, conversation, tour count, tool Complexity Vector from tool references and task type Pre-query complexity classification and prediction generating (ComplexityVector) It is characterized by containing layer (1).
3. According to claim 1, it is a selective layer activation control system, and its characteristic is; trust 30 If the score falls below the threshold, it elevates the query to a higher complexity class. - 12 - by including prequery complexity classification and prediction layer (1) It is characteristic.
4. Selective layer activation control system according to claim 1, its characteristic is; C0– Query that defines C4 complexity classes and performs class representation. 5 by including the prior complexity classification and prediction layer (1) It is characteristic.
5. According to claim 1, it is a selective layer activation control system, the characteristic of which is; semantic similarity, recency score, and entity_match_density and KV_relevance score Complexity-driven selective KV cache loading using the function and 10 It is characterized by containing a segmentation unit (2).
6. According to claim 1, it is a selective layer activation control system, and its characteristic is; class complexity-driven selective KV loading policy based on (C0-C4) by including selective KV cache loading and segmentation unit (2) It is characterized by: 15 7. Selective layer activation control system according to claim 1, its characteristic is; KV Complexity-driven generation of Cache Loading Plan (KVCacheLoadPlan) by including selective KV cache loading and segmentation unit (2) It is characteristic.
8. Selective layer activation control system according to claim 1, with the characteristic; C0-20 layer_range and quant_level are defined according to C4 classes. ActivationMask (including proving level) policies It is characterized by containing the hardware activation map generator (3).
9. According to claim 1, it is a selective layer activation control system, the characteristic of which is; 25 that implements automatic fallback based on hardware telemetry. It is characterized by containing a hardware activation map generator (3).
10. According to claim 1, it is a selective layer activation control system, the characteristic of which is; reasoning_token_limit (reasoning token limit), active_layer_range (active layer range), branch_count (number of branches) and quantization_policy (quantization policy) Multidimensional dynamic capacity 30 forming capacity vector B including policy) It is characterized by containing the vector generator (4). - 13 - 11. Selective layer activation control system according to claim 1, its characteristic is; KV Cache Loading Plan (KVCacheLoadPlan), ActivationMask (Activation Mask) LLM performs restricted prefill and decode stages with the Mask) and vector B. by including an inference engine and a resource-constrained execution layer (5) It is characterized by: 5 12. According to claim 1, it is a selective layer activation control system, the characteristic of which is; The derivative of consumption_rate(t) and the thresholds θ_warn and θ_critical Two-threshold predictive capacity leakage controller with phased intervention (6) It is characterized by its inclusion.
13. Selective layer activation control system according to claim 1, its characteristic is; 10 early with quality_signal and coverage_ratio Two-threshold predictive capacity leakage controller (6) that makes completion decision It is characterized by its inclusion.
14. It is a selective layer activation control system according to claim 1, and its characteristic is; semantic_coherence, factual_grounding_score quality_score includes (baseline score) and coverage_ratio (coverage ratio). early stop unit (7) which calculates the score based on recursive quality. It is characterized by its inclusion.
15. According to Claim 1, it is a selective layer activation control system, the characteristic of which is; 20 recursive quality-based systems that increase the weight of q2 in medical and legal fields. It is characterized by containing an early stop unit (7).
16. According to Claim 1, it is a selective layer activation control system, the characteristic of which is; final_response, complexity_class, quality_score and consumed_token_count (8) 25 FinalOutput generator that produces the result combine and output generator containing (number) It is characterized by its inclusion.
17. It is a selective layer activation control system according to claim 1, and its characteristic is; The reward function and the PPO (Proximal Policy Optimization) algorithm. PPO meta-adaptation and parameters update using It is characterized by containing an update layer (9). 30 18. It is a selective layer activation control system according to claim 1, and its characteristic is; Graphics Processing Unit (GPU) / Neural Processing Unit - 14 - (Neural Processing Unit - NPU) / Video Random Access Memory (Video) Hardware profile that monitors Random Access Memory (VRAM) telemetry. It is characterized by containing a monitor and a fallback signal layer (10).
19. It is a selective layer activation control system according to Claim 1, and its characteristic is; VRAM occupancy and thermal headroom thresholds are 5. according to the signal generating hardware profile monitor and the signal degrade layer. (10) is characterized by its inclusion.
20. Selective layer activation control system according to claim 1, its characteristic is; C0– Selective KV loading, layer activation, and according to C4 complexity classes. 10 through the holistic integration of modules that perform capacity allocation It is characteristic.
21. According to claim 1, it is a selective layer activation control system, characterized by its simplicity. queries regarding Graphics Processing Unit (GPU) / Neural While reducing the Neural Processing Unit (NPU) load by 60–80% 15 by including an optimization mechanism that preserves response quality It is characteristic. 25