An agent layered information bottleneck processing method
Patent Information
- Application Number
- CN202611161218.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-03
- Publication Date
- 2026-08-28
AI Technical Summary
但现有方法多停留在离线分析或单阶段压缩层面,缺乏面向智能体系统的分层处理结构,且未能结合任务约束与资源预算进行动态优化
[0029]A device for processing hierarchical information bottlenecks in intelligent agents includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described method steps.
Smart Images

Figure CN122655848A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information processing technology, and in particular to a method for handling hierarchical information bottlenecks in intelligent agents. Background Technology
[0002] With the widespread application of large model and intelligent agent technologies in various fields, their capabilities in semantic understanding, multimodal processing, and complex task reasoning are constantly improving, greatly promoting the development of artificial intelligence systems towards automation and intelligence.
[0003] However, in practical applications, intelligent agents still face problems such as severe information redundancy, excessive consumption of computing resources, and unstable inference efficiency when processing large-scale, multimodal information.
[0004] Meanwhile, traditional methods typically rely on a single model to directly process all input information, lacking an information filtering mechanism for specific task objectives and failing to effectively combine computational budgets for resource scheduling, resulting in low information utilization efficiency and difficulty in meeting the dual requirements of real-time performance and accuracy in complex scenarios.
[0005] To address these issues, researchers have begun exploring information compression and representation learning methods based on the information bottleneck theory, which extract key semantic features by compressing input information. However, existing methods mostly remain at the level of offline analysis or single-stage compression, lacking a hierarchical processing structure for intelligent agent systems, and failing to incorporate task constraints and resource budgets for dynamic optimization.
[0006] Based on the above, this invention proposes a method for handling hierarchical information bottlenecks in intelligent agents. Summary of the Invention
[0007] To overcome the shortcomings of existing technologies, this invention provides a simple and efficient method for handling hierarchical information bottlenecks in intelligent agents.
[0008] This invention is achieved through the following technical solution: A method for handling hierarchical information bottlenecks in intelligent agents involves first collecting multimodal raw data related to the task, preprocessing it, and generating feature-encoded data. Custom task constraint models are built according to specific application scenarios, and input information is parsed and its importance is evaluated. Priority identifiers are generated to provide a basis for subsequent information filtering, context construction and reasoning generation. A hierarchical information processing structure is constructed, an information bottleneck mechanism is introduced, and a hierarchical information bottleneck processing module is built to process the input information step by step, so as to achieve a balance between information compression efficiency and semantic expression capability. Based on task constraints and compressed semantic representation output by the hierarchical information bottleneck processing module, a task-aware information selection module is constructed to uniformly model different types of information and dynamically adjust the retention weights of various types of information according to the target task. An attention mechanism and dynamic information selection strategy are introduced to enhance the task relevance of each information unit. A computational budget model is constructed to uniformly model the budget constraints in the information processing process, forming a budget constraint vector, which provides a quantitative basis for the dynamic scheduling of subsequent information processing paths; A mutual information-based information quality assessment mechanism is constructed to quantitatively evaluate the information retention capability between the original input information and the compressed semantic representation, and to evaluate the expressive capability of the compressed representation for task-related information in combination with the target task constraints, so as to provide feedback basis for subsequent compression strategy optimization. The intelligent agent reasoning module performs task parsing and reasoning planning for the target task based on intermediate semantic representation, and performs reasoning step by step according to the preset reasoning strategy to solve complex tasks. A model deployment module is built to deploy the model after intelligent agent inference to actual business application scenarios, providing unified intelligent service capabilities for different types of tasks.
[0009] Includes the following steps: Step S1: Multimodal data acquisition and preprocessing a. Multimodal data acquisition Based on the application scenario of the intelligent agent performing the task, multimodal raw data related to the task is collected. The multimodal raw data includes at least one or more of the following: text, images, audio, video, and structured data. b. Multimodal data preprocessing Preprocessing operations are performed on the acquired multimodal raw data according to their data types to eliminate noise interference and generate feature-coded data: After preprocessing, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation, providing a unified input for subsequent hierarchical information bottleneck processing driven by task constraints and budget. In step S1, when performing preprocessing on the acquired multimodal raw data, the specific method is as follows: For text data, the following steps are performed in sequence: text cleaning, word segmentation, stop word removal, semantic encoding, and vectorization representation. For image data, image denoising, size normalization, target region extraction, and visual feature encoding are performed; For audio data, noise suppression, speech enhancement, speech recognition, and acoustic feature extraction are performed. For structured data, outlier detection, missing value imputation, field standardization, and numerical normalization are performed.
[0010] Step S2: Task Constraint Modeling and Goal Definition a. Task constraint modeling and target configuration Custom task constraint models are built according to specific application scenarios, and task types, execution objectives, constraints, and evaluation indicators are modeled in a unified manner. The task types include question-answering tasks, decision-making tasks, analysis tasks, generation tasks, and reasoning tasks. The corresponding objective functions, output forms, and key output requirements are determined in combination with different task types to form constraint rules adapted to the target tasks, providing task guidance for subsequent information filtering and model reasoning. b. Input information semantic annotation and importance assessment Based on the constructed task constraint model, semantic parsing, semantic annotation, and importance assessment are performed on the input information to identify core semantic units, key entities, task-related information, and contextual relationships in the input content. Priority identifiers are generated according to semantic relevance, task contribution, and information importance. Key content areas that need to be retained and information areas that can be downgraded or filtered are also defined to form an information retention strategy oriented towards the target task, providing a basis for subsequent information filtering, context construction, and reasoning generation. Step S3: Design of the Layered Information Bottleneck Processing Module a. Construction of a hierarchical information processing structure A hierarchical information processing structure is constructed, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, for processing input information step by step. The primary filtering layer is used to initially filter irrelevant, repetitive, and low-value content in the input information; the intermediate compression layer is used to perform semantic aggregation, redundancy merging, and context compression on the information after primary filtering; the high-level semantic abstraction layer is used to extract high-level semantic features, core intentions, and key reasoning bases related to the target task, thereby forming a concise semantic representation suitable for subsequent model reasoning or task execution. b. Information bottleneck constraint compression mechanism Information bottleneck mechanisms are introduced in the primary filtering layer, intermediate compression layer and advanced semantic abstraction layer respectively, and a hierarchical information bottleneck processing module is constructed to perform constrained compression processing on the input information. In step S3, based on task relevance, semantic importance, information redundancy, and context contribution, and based on a predefined information bottleneck mechanism, the input information of each layer is filtered, compressed, and reconstructed. While reducing the amount of irrelevant and redundant information transmitted, effective information related to the target task is retained. By customizing the information retention ratio, compression granularity, and semantic abstraction level of each layer, a balance is achieved between information compression efficiency and semantic expression capability.
[0011] Step S4, Task-Aware Information Selection Mechanism a. Task-aware dynamic weight allocation Based on task constraints and compressed semantic representation output by the hierarchical information bottleneck processing module, a task-aware information selection module is constructed to uniformly model information from different sources, modalities, and semantic levels, and dynamically adjust the retention weights of various types of information according to the target task. The information includes multimodal information such as text information, image information, audio information, and structured data.
[0012] Input information is represented as ,in, This represents the set of input information for the task-aware information selection module. Indicates the first There are 1 information units, where n represents the total number of information units; Let the task constraints be represented as The task relevance score of each information unit is calculated based on task constraints: , in, This represents a function for evaluating task relevance. Representation of information unit With the current task The correlation score between them; To allocate importance among different information units, the Softmax function is used to calculate the dynamic weight of each information unit. , in, Indicates the first The retention weight of each information unit, and satisfying ; Subsequently, the input information is weighted according to the calculated dynamic weights: , in, This represents the information representation after task-aware enhancement. Through the above dynamic weight allocation, the adaptive retention and fusion of information from different modalities are achieved, providing task-oriented information representation for the subsequent reasoning process.
[0013] b. Task-related information enhancement mechanism Based on the dynamic weight allocation results, an attention mechanism and a dynamic information selection strategy are introduced to enhance the task relevance of each information unit; the details are as follows: By combining preset retention thresholds, TopK filtering strategies, or dynamic gating mechanisms, task-related information is enhanced, and redundant, repetitive, and irrelevant information is suppressed to reduce the propagation of invalid information in subsequent reasoning processes. In step S4, when a threshold gating strategy is adopted, the information retention flag is represented as follows: , , Where, m i To preserve the information's tags, The preset threshold; Alternatively, the TopK strategy can be used, retaining only the k information units with the highest weights, and its output is represented as: , in, This represents the TopK filtering operator, where k represents the number of information units to be retained.
[0014] Through the aforementioned task-related information enhancement mechanism, while ensuring the full expression of task-related information, irrelevant and redundant information is effectively suppressed, thereby improving information utilization efficiency, model inference accuracy, and overall task execution performance.
[0015] Step S5: Budget Constraint Calculation and Scheduling Module a. Construction of the budget calculation model A computational budget model is constructed to uniformly model the budget constraints in the information processing process, including computing resources, response latency, and operating costs, and to form a budget constraint vector, providing a quantitative basis for the dynamic scheduling of subsequent information processing paths; The computing resources include processor computing power, GPU computing resources, video memory capacity, and memory usage; the response latency is used to describe the maximum allowable response time for the target task; the operating cost is used to describe the model call cost, interface call cost, and computing resource consumption cost; The budget constraint vector is represented as Where B represents the budget constraint vector, C represents the computing power resource index that supports the call, L represents the response latency constraint index, and R represents the operating cost constraint index; The indicators in the budget constraint vector are normalized, and the comprehensive budget score is calculated based on the importance of different constraint vectors. In step S5, the comprehensive budget score is calculated as follows: , This indicates the overall budget score; , and These represent the normalized computing power resource indicators, response latency constraint indicators, and operating cost constraint indicators, respectively. , and Let represent the weight coefficients of the corresponding indicators, and satisfy the following conditions: , The comprehensive budget score is used to quantify the computing budget available to support the current system and serves as an evaluation basis for the selection of subsequent information processing paths.
[0016] b. Dynamic path scheduling based on budget constraints Based on comprehensive budget scoring, a dynamic path scheduling mechanism is constructed to adaptively select different information processing paths according to budget constraints, so as to achieve a dynamic balance between computing resource consumption, response efficiency and model inference performance. In step S5, the comprehensive budget score is compared with a preset budget threshold, and the target processing path is determined based on the comparison result. The path selection rule is expressed as follows: , in, This indicates a lightweight processing path. Indicates the deep reasoning path, Indicates the threshold for budget decision-making; When the overall budget score is lower than the budget threshold, the scheduling controller selects a lightweight processing path and adopts a lightweight model, few-step inference and shallow information compression strategy to reduce computing resource consumption and response latency. When the overall budget score reaches or exceeds the budget threshold, the scheduling controller selects a deep inference path and adopts a large-scale model, multi-step inference strategy and toolchain collaboration mechanism to improve the inference ability and result accuracy of complex tasks.
[0017] During task execution, the scheduling controller continuously monitors the model's running status, resource utilization, response latency, and inference performance, and dynamically adjusts the comprehensive budget score based on real-time feedback. The update method is as follows: , in, Indicates the first Comprehensive budget score during the next scheduling This represents the budget adjustment amount obtained based on system operation feedback. Indicates the budget update factor, and satisfies ; The scheduling controller re-executes path selection based on the updated comprehensive budget score, thereby enabling dynamic switching between lightweight processing paths and deep inference paths. This allows the system to adaptively adjust its information processing strategy according to the current budget constraints, achieving an optimal balance between computational resource consumption, response efficiency, and inference performance.
[0018] Step S6: Mutual Information Evaluation and Rate-Distortion Optimization a. Information quality assessment based on mutual information A mutual information-based information quality assessment mechanism is constructed to quantitatively evaluate the information retention capability between the original input information and the compressed semantic representation, and to evaluate the expressive capability of the compressed representation for task-related information in combination with the target task constraints, so as to provide feedback basis for subsequent compression strategy optimization. Let the original input information be represented as The compressed representation information after layered information processing is: The task constraint representation information is as follows The mutual information between the input information and the compressed representation is then expressed as: ; Mutual information represents the conditional mutual information between the original input information and the compressed representation information under task constraints. It is used to measure the degree to which the compressed representation information retains task-related information. The larger the mutual information value, the richer the effective task information contained in the compressed representation. The smaller the mutual information value, the greater the information loss during the compression process.
[0019] In step S6, to evaluate the information retention efficiency corresponding to a unit of encoded resources, an information rate index is introduced, which is calculated as follows: , Where A represents the encoding length or number of encoding bits in the compressed representation, and E represents the information retention efficiency per unit encoding length, which is used to comprehensively evaluate compression efficiency and information retention capability; The mutual information and information rate are calculated using Monte Carlo sampling, contrastive learning estimation, or neural mutual information estimation methods to form task-oriented information quality evaluation indicators.
[0020] b. Rate and distortion-based dynamic optimization mechanism Based on the mutual information evaluation results, a rate and distortion optimization mechanism is constructed to jointly optimize the compression ratio, information retention level, and computational resource consumption during the information compression process, so that the compressed representation achieves better information expression capability while meeting the task performance requirements. Let the information distortion generated during the compression process be represented as: , in, Indicates compression distortion. The distortion metric function between the initial input and the compressed representation is described by semantic distance, reconstruction error, or task performance loss. Furthermore, the joint optimization objective of rate-distortion is expressed as: , in, Indicates the overall optimization objective; Indicates information rate; This indicates information distortion; The representation rate versus distortion tradeoff is used to balance compression efficiency and information retention.
[0021] During system operation, the information compression threshold, compression ratio, model size, and information processing path are dynamically adjusted based on the mutual information evaluation results and current task performance feedback. When the mutual information drops below the preset threshold or the task performance declines at a rate exceeding a custom threshold, the system reduces the compression ratio or increases the amount of information retained. When the mutual information meets the task requirements and computing resources are limited, the compression ratio is increased to reduce computing costs and achieve a dynamic optimization balance between task performance, information retention capacity, and computing resource consumption.
[0022] Step S7: Agent Reasoning and Result Generation a. Agent reasoning based on intermediate representations The intermediate semantic representation, after task constraint modeling, hierarchical information bottleneck processing, and task perception information selection, is input into the agent reasoning module as a unified input for the agent to perform task planning, reasoning analysis, and decision generation. The intermediate semantic representation includes target task information, contextual semantic information, key constraints, and high-value task-related information after screening, so as to avoid redundant information from interfering with the reasoning process. The intelligent agent reasoning module performs task understanding, task decomposition and reasoning planning for the target task based on intermediate semantic representation, and generates multiple sub-tasks according to the preset reasoning strategy. Based on the dependencies between subtasks, a reasoning execution sequence is constructed, and complex tasks are solved through step-by-step reasoning, state updates, and intermediate result feedback mechanisms. During the reasoning process, the agent continuously receives intermediate results generated by the reasoning and uses these intermediate results as input for subsequent reasoning steps, thereby achieving dynamic updates and iterative optimization of the reasoning state and ultimately generating reasoning results that meet the requirements of the target task. b. Collaborative Reasoning Mechanism between Large Models and Tools During the agent's reasoning process, the reasoning capabilities of a large language model and an external tool invocation mechanism are introduced to construct a collaborative framework for model reasoning and tool execution. When the agent determines that the current task involves external knowledge acquisition, data retrieval, mathematical calculation, code execution, database query, or interface invocation capabilities, it automatically generates a tool invocation request based on the current reasoning state and selects the appropriate tool to complete the target operation. After the tool completes its execution, the results returned by the tool are fed back to the agent's inference module as observation information. Together with the current inference state, they form new context information, driving the next round of inference and realizing a closed-loop execution mechanism of "inference-tool call-result feedback-continued inference". The agent continuously corrects the inference path based on the continuous feedback until the task termination condition is met, and outputs the final task solution result.
[0023] Furthermore, after the task is completed, the intermediate states generated during the reasoning process, tool call logs, task execution trajectory, and key decision-making basis are output to form a reasoning trajectory that supports explanation, providing support for subsequent result verification, task backtracking, and system optimization.
[0024] Step S8: Application and Feedback Optimization a. Application-scenario-oriented model deployment and service integration A model deployment module is constructed to deploy the model, after task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The specific implementation method is as follows: Based on the target business requirements, the model is encapsulated as an API interface, microservice interface, or containerized service, and a data interaction channel is established with the business system. This enables applications, including intelligent question answering, decision support, and multimodal analysis, to uniformly call the model service. After receiving a user request, the business system inputs the relevant task into the intelligent agent system, which then completes task understanding, information processing, reasoning analysis, and result generation. The system then returns the task execution result to the business system, achieving collaborative integration of model capabilities and business processes.
[0025] Meanwhile, the model deployment module supports multi-instance deployment, elastic resource scheduling, and service expansion mechanisms. It also supports dynamically adjusting the number of model instances and the allocation of computing resources according to business load to improve the system's scalability, stability, and concurrent processing capabilities.
[0026] b. Strategy optimization and closed-loop update based on feedback information A feedback optimization module is constructed to uniformly collect, analyze, and evaluate task execution results, user feedback information, and system operation indicators generated during model operation. Based on the feedback results, the system operation strategy is dynamically optimized to achieve continuous iterative optimization of model capabilities. The specific implementation method is as follows: After the model completes the task, the feedback optimization module synchronously collects operational metrics including task execution accuracy, response latency, operating cost, resource utilization, and user satisfaction. It also combines user feedback information to comprehensively evaluate the current information filtering strategy, budget allocation strategy, and path scheduling strategy. When a decline in task performance, an increase in response latency, a decrease in resource utilization, or a decrease in user satisfaction is detected, the feedback optimization module automatically analyzes the influencing factors and generates corresponding strategy adjustment plans.
[0027] Furthermore, based on the feedback analysis results, the feedback optimization module dynamically updates the information filtering threshold, task-related weights, budget allocation parameters, path scheduling rules, and model calling strategies, and reapplies the updated strategies to the model service system, forming a closed-loop optimization mechanism of "task execution - feedback collection - strategy optimization - parameter update - re-execution". This enables the system to continuously optimize the information processing flow according to the actual operating status, thereby improving task execution accuracy, resource utilization efficiency, and overall system performance.
[0028] A hierarchical information bottleneck processing system for intelligent agents, used to implement the above method, includes: The multimodal data acquisition and preprocessing module is responsible for acquiring multimodal raw data related to the task based on the application scenario of the intelligent agent's task execution, and performing preprocessing operations according to the data type to eliminate noise interference and generate feature-encoded data. Subsequently, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation. The task constraint modeling and goal definition module is responsible for customizing and constructing task constraint models according to specific application scenarios. It performs unified modeling of task types, execution goals, constraints, and evaluation indicators. It performs semantic parsing, semantic annotation, and importance assessment on input information, identifying core semantic units, key entities, task-related information, and contextual relationships in the input content. Based on semantic relevance, task contribution, and information importance, it custom-generates priority identifiers, and custom-determines the key content areas that need to be retained and the information areas that can be downgraded or filtered, forming an information retention strategy oriented towards the target task. The hierarchical information bottleneck processing module is responsible for constructing a hierarchical information processing structure, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, which are used to process the input information step by step; and introduces information bottleneck mechanisms in the primary filtering layer, intermediate compression layer, and high-level semantic abstraction layer respectively to perform constrained compression processing on the input information. The task-aware information selection module is responsible for uniformly modeling information from different sources, modalities, and semantic levels based on task constraints and the compressed semantic representation output by the hierarchical information bottleneck processing module, and dynamically adjusting the retention weights of various types of information according to the target task; based on the dynamic weight allocation results, an attention mechanism and dynamic information selection strategy are introduced to enhance the task relevance of each information unit. The budget constraint computation scheduling module is responsible for constructing a computation budget model. It performs unified modeling of budget constraint factors in the information processing process, including computing resources, response latency, and operating costs, and forms a budget constraint vector. It normalizes the various indicators in the budget constraint vector and calculates a comprehensive budget score based on the importance of different constraint vectors. Based on the comprehensive budget score, it constructs a dynamic path scheduling mechanism to adaptively select different information processing paths according to budget constraints, so as to achieve a dynamic balance between computing resource consumption, response efficiency, and model inference performance. The mutual information evaluation and rate-distortion optimization module is responsible for constructing an information quality evaluation mechanism based on mutual information, quantitatively evaluating the information retention capability between the original input information and the compressed semantic representation, evaluating the expressive capability of the compressed representation for task-related information in combination with the target task constraints, and constructing a rate-distortion optimization mechanism based on the mutual information evaluation results to jointly optimize the compression ratio, information retention degree and computational resource consumption in the information compression process. The agent reasoning module is responsible for understanding, decomposing, and planning the reasoning of the target task based on intermediate semantic representations, and generating multiple sub-tasks according to a preset reasoning strategy. It constructs a reasoning execution sequence based on the dependencies between sub-tasks, and solves complex tasks through step-by-step reasoning, state updates, and intermediate result feedback mechanisms, ultimately generating a reasoning result that meets the requirements of the target task. In the agent reasoning process, it introduces the reasoning capabilities of a large language model and an external tool invocation mechanism to build a collaborative framework for model reasoning and tool execution. It automatically generates tool invocation requests based on the current reasoning state, and feeds back the tool's return results as observation information to the agent reasoning module. Together with the current reasoning state, these form new contextual information to drive the next round of reasoning. The model deployment module is responsible for deploying the model, which has undergone task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The feedback optimization module is responsible for the unified collection, analysis and evaluation of task execution results, user feedback information and system operation indicators generated during model operation, and dynamically optimizes the system operation strategy based on the feedback results to achieve continuous iterative optimization of model capabilities.
[0029] A device for processing hierarchical information bottlenecks in intelligent agents includes a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described method steps.
[0030] A readable storage medium storing a computer program that, when executed by a processor, implements the above-described method steps.
[0031] The beneficial effects of this invention are: the hierarchical information bottleneck processing method of the intelligent agent has a clear structure, is easy to deploy in engineering, has a wide range of applications and strong scalability, effectively reduces redundant information participating in the calculation process, significantly reduces the overall computing overhead and system operating cost, improves resource utilization efficiency, system inference efficiency and response speed, and has strong task adaptability and high result stability, thereby improving the inference accuracy and result stability of the intelligent agent in complex task scenarios. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of the large model graph inference method based on adaptive multimodal mode of the present invention.
[0034] Figure 2 This is a schematic diagram of the multimodal data acquisition and preprocessing method of the present invention.
[0035] Figure 3 This is a schematic diagram of the task constraint modeling method of the present invention.
[0036] Figure 4 This is a schematic diagram of the hierarchical information bottleneck processing method of the present invention.
[0037] Figure 5 This is a schematic diagram of the task-aware information selection method of the present invention.
[0038] Figure 6 This is a schematic diagram of the budget constraint calculation and scheduling method of the present invention.
[0039] Figure 7 This is a schematic diagram of the information quality assessment and optimization method based on mutual information of the present invention.
[0040] Figure 8 This is a schematic diagram of the intelligent agent reasoning and result generation method of the present invention.
[0041] Figure 9This is a schematic diagram illustrating the application and feedback optimization method of the present invention. Detailed Implementation
[0042] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.
[0043] The method for handling hierarchical information bottlenecks in intelligent agents includes the following steps: Step S1: Multimodal data acquisition and preprocessing a. Multimodal data acquisition Based on the application scenario of the intelligent agent performing the task, multimodal raw data related to the task is collected. The multimodal raw data includes at least one or more of the following: text, images, audio, video, and structured data. Text data can originate from user input, business documents, or online text; image and video data can originate from camera equipment, image databases, or internet resources; audio data can originate from voice interaction devices or audio databases; and structured data can originate from relational databases, knowledge graphs, or business systems. A unified data access interface is established for data from different sources, and metadata such as source information, timestamps, and data types for each modality is recorded to ensure the integrity, consistency, and traceability of multimodal data, providing a reliable data foundation for subsequent task constraint analysis and information processing.
[0044] b. Multimodal data preprocessing Preprocessing operations are performed on the acquired multimodal raw data according to their data types to eliminate noise interference and generate feature-coded data: After preprocessing, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation, providing a unified input for subsequent hierarchical information bottleneck processing driven by task constraints and budget. In step S1, when performing preprocessing on the acquired multimodal raw data, the specific method is as follows: For text data, the following steps are performed in sequence: text cleaning, word segmentation, stop word removal, semantic encoding, and vectorization representation. For image data, image denoising, size normalization, target region extraction, and visual feature encoding are performed; For audio data, noise suppression, speech enhancement, speech recognition, and acoustic feature extraction are performed. For structured data, outlier detection, missing value imputation, field standardization, and numerical normalization are performed.
[0045] Step S2: Task Constraint Modeling and Goal Definition a. Task constraint modeling and target configuration Custom task constraint models are built according to specific application scenarios, and task types, execution objectives, constraints, and evaluation indicators are modeled in a unified manner. The task types include question-answering tasks, decision-making tasks, analysis tasks, generation tasks, and reasoning tasks. The corresponding objective functions, output forms, and key output requirements are determined in combination with different task types to form constraint rules adapted to the target tasks, providing task guidance for subsequent information filtering and model reasoning. b. Input information semantic annotation and importance assessment Based on the constructed task constraint model, semantic parsing, semantic annotation, and importance assessment are performed on the input information to identify core semantic units, key entities, task-related information, and contextual relationships in the input content. Priority identifiers are generated according to semantic relevance, task contribution, and information importance. Key content areas that need to be retained and information areas that can be downgraded or filtered are also defined to form an information retention strategy oriented towards the target task, providing a basis for subsequent information filtering, context construction, and reasoning generation. Step S3: Design of the Layered Information Bottleneck Processing Module a. Construction of a hierarchical information processing structure A hierarchical information processing structure is constructed, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, for processing input information step by step. The primary filtering layer is used to initially filter irrelevant, repetitive, and low-value content in the input information; the intermediate compression layer is used to perform semantic aggregation, redundancy merging, and context compression on the information after primary filtering; the high-level semantic abstraction layer is used to extract high-level semantic features, core intentions, and key reasoning bases related to the target task, thereby forming a concise semantic representation suitable for subsequent model reasoning or task execution. b. Information bottleneck constraint compression mechanism Information bottleneck mechanisms are introduced in the primary filtering layer, intermediate compression layer and advanced semantic abstraction layer respectively, and a hierarchical information bottleneck processing module is constructed to perform constrained compression processing on the input information. In step S3, based on task relevance, semantic importance, information redundancy, and context contribution, and based on a predefined information bottleneck mechanism, the input information of each layer is filtered, compressed, and reconstructed. While reducing the amount of irrelevant and redundant information transmitted, effective information related to the target task is retained. By customizing the information retention ratio, compression granularity, and semantic abstraction level of each layer, a balance is achieved between information compression efficiency and semantic expression capability.
[0046] Step S4, Task-Aware Information Selection Mechanism a. Task-aware dynamic weight allocation Based on task constraints and compressed semantic representation output by the hierarchical information bottleneck processing module, a task-aware information selection module is constructed to uniformly model information from different sources, modalities, and semantic levels, and dynamically adjust the retention weights of various types of information according to the target task. The information includes multimodal information such as text information, image information, audio information, and structured data.
[0047] Input information is represented as ,in, This represents the set of input information for the task-aware information selection module. Indicates the first There are 1 information units, where n represents the total number of information units; Let the task constraints be represented as The task relevance score of each information unit is calculated based on task constraints: , in, This represents a function for evaluating task relevance. Representation of information unit With the current task The correlation score between them; To allocate importance among different information units, the Softmax function is used to calculate the dynamic weight of each information unit. , in, Indicates the first The retention weight of each information unit, and satisfying ; Subsequently, the input information is weighted according to the calculated dynamic weights: , in, This represents the information representation after task-aware enhancement. Through the above dynamic weight allocation, the adaptive retention and fusion of information from different modalities are achieved, providing task-oriented information representation for the subsequent reasoning process.
[0048] b. Task-related information enhancement mechanism Based on the dynamic weight allocation results, an attention mechanism and a dynamic information selection strategy are introduced to enhance the task relevance of each information unit; the details are as follows: By combining preset retention thresholds, TopK filtering strategies, or dynamic gating mechanisms, task-related information is enhanced, and redundant, repetitive, and irrelevant information is suppressed to reduce the propagation of invalid information in subsequent reasoning processes. In step S4, when a threshold gating strategy is adopted, the information retention flag is represented as follows: , , Where, m i To preserve the information's tags, The preset threshold; Alternatively, the TopK strategy can be used, retaining only the k information units with the highest weights, and its output is represented as: , in, This represents the TopK filtering operator, where k represents the number of information units to be retained.
[0049] Through the aforementioned task-related information enhancement mechanism, while ensuring the full expression of task-related information, irrelevant and redundant information is effectively suppressed, thereby improving information utilization efficiency, model inference accuracy, and overall task execution performance.
[0050] Step S5: Budget Constraint Calculation and Scheduling Module a. Construction of the budget calculation model A computational budget model is constructed to uniformly model the budget constraints in the information processing process, including computing resources, response latency, and operating costs, and to form a budget constraint vector, providing a quantitative basis for the dynamic scheduling of subsequent information processing paths; The computing resources include processor computing power, GPU computing resources, video memory capacity, and memory usage; the response latency is used to describe the maximum allowable response time for the target task; the operating cost is used to describe the model call cost, interface call cost, and computing resource consumption cost; The budget constraint vector is represented as Where B represents the budget constraint vector, C represents the computing power resource index that supports the call, L represents the response latency constraint index, and R represents the operating cost constraint index; The indicators in the budget constraint vector are normalized, and the comprehensive budget score is calculated based on the importance of different constraint vectors. In step S5, the comprehensive budget score is calculated as follows: , This indicates the overall budget score; , and These represent the normalized computing power resource indicators, response latency constraint indicators, and operating cost constraint indicators, respectively. , and Let represent the weight coefficients of the corresponding indicators, and satisfy the following conditions: , The comprehensive budget score is used to quantify the computing budget available to support the current system and serves as an evaluation basis for the selection of subsequent information processing paths.
[0051] b. Dynamic path scheduling based on budget constraints Based on comprehensive budget scoring, a dynamic path scheduling mechanism is constructed to adaptively select different information processing paths according to budget constraints, so as to achieve a dynamic balance between computing resource consumption, response efficiency and model inference performance. In step S5, the comprehensive budget score is compared with a preset budget threshold, and the target processing path is determined based on the comparison result. The path selection rule is expressed as follows: , in, This indicates a lightweight processing path. Indicates the deep reasoning path, Indicates the threshold for budget decision-making; When the overall budget score is lower than the budget threshold, the scheduling controller selects a lightweight processing path and adopts a lightweight model, few-step inference and shallow information compression strategy to reduce computing resource consumption and response latency. When the overall budget score reaches or exceeds the budget threshold, the scheduling controller selects a deep inference path and adopts a large-scale model, multi-step inference strategy and toolchain collaboration mechanism to improve the inference ability and result accuracy of complex tasks.
[0052] During task execution, the scheduling controller continuously monitors the model's running status, resource utilization, response latency, and inference performance, and dynamically adjusts the comprehensive budget score based on real-time feedback. The update method is as follows: , in, Indicates the first Comprehensive budget score during the next scheduling This represents the budget adjustment amount obtained based on system operation feedback. Indicates the budget update factor, and satisfies ; The scheduling controller re-executes path selection based on the updated comprehensive budget score, thereby enabling dynamic switching between lightweight processing paths and deep inference paths. This allows the system to adaptively adjust its information processing strategy according to the current budget constraints, achieving an optimal balance between computational resource consumption, response efficiency, and inference performance.
[0053] Step S6: Mutual Information Evaluation and Rate-Distortion Optimization a. Information quality assessment based on mutual information A mutual information-based information quality assessment mechanism is constructed to quantitatively evaluate the information retention capability between the original input information and the compressed semantic representation, and to evaluate the expressive capability of the compressed representation for task-related information in combination with the target task constraints, so as to provide feedback basis for subsequent compression strategy optimization. Let the original input information be represented as The compressed representation information after layered information processing is: The task constraint representation information is as follows The mutual information between the input information and the compressed representation is then expressed as: ; Mutual information represents the conditional mutual information between the original input information and the compressed representation information under task constraints. It is used to measure the degree to which the compressed representation information retains task-related information. The larger the mutual information value, the richer the effective task information contained in the compressed representation. The smaller the mutual information value, the greater the information loss during the compression process.
[0054] In step S6, to evaluate the information retention efficiency corresponding to a unit of encoded resources, an information rate index is introduced, which is calculated as follows: , Where A represents the encoding length or number of encoding bits in the compressed representation, and E represents the information retention efficiency per unit encoding length, which is used to comprehensively evaluate compression efficiency and information retention capability; The mutual information and information rate are calculated using Monte Carlo sampling, contrastive learning estimation, or neural mutual information estimation methods to form task-oriented information quality evaluation indicators.
[0055] b. Rate and distortion-based dynamic optimization mechanism Based on the mutual information evaluation results, a rate and distortion optimization mechanism is constructed to jointly optimize the compression ratio, information retention level, and computational resource consumption during the information compression process, so that the compressed representation achieves better information expression capability while meeting the task performance requirements. Let the information distortion generated during the compression process be represented as: , in, Indicates compression distortion. The distortion metric function between the initial input and the compressed representation is described by semantic distance, reconstruction error, or task performance loss. Furthermore, the joint optimization objective of rate-distortion is expressed as: , in, Indicates the overall optimization objective; Indicates information rate; This indicates information distortion; The representation rate versus distortion tradeoff is used to balance compression efficiency and information retention.
[0056] During system operation, the information compression threshold, compression ratio, model size, and information processing path are dynamically adjusted based on the mutual information evaluation results and current task performance feedback. When the mutual information drops below the preset threshold or the task performance declines at a rate exceeding a custom threshold, the system reduces the compression ratio or increases the amount of information retained. When the mutual information meets the task requirements and computing resources are limited, the compression ratio is increased to reduce computing costs and achieve a dynamic optimization balance between task performance, information retention capacity, and computing resource consumption.
[0057] Step S7: Agent Reasoning and Result Generation a. Agent reasoning based on intermediate representations The intermediate semantic representation, after task constraint modeling, hierarchical information bottleneck processing, and task perception information selection, is input into the agent reasoning module as a unified input for the agent to perform task planning, reasoning analysis, and decision generation. The intermediate semantic representation includes target task information, contextual semantic information, key constraints, and high-value task-related information after screening, so as to avoid redundant information from interfering with the reasoning process. The intelligent agent reasoning module performs task understanding, task decomposition and reasoning planning for the target task based on intermediate semantic representation, and generates multiple sub-tasks according to the preset reasoning strategy. Based on the dependencies between subtasks, a reasoning execution sequence is constructed, and complex tasks are solved through step-by-step reasoning, state updates, and intermediate result feedback mechanisms. During the reasoning process, the agent continuously receives intermediate results generated by the reasoning and uses these intermediate results as input for subsequent reasoning steps, thereby achieving dynamic updates and iterative optimization of the reasoning state and ultimately generating reasoning results that meet the requirements of the target task. b. Collaborative Reasoning Mechanism between Large Models and Tools During the agent's reasoning process, the reasoning capabilities of a large language model and an external tool invocation mechanism are introduced to construct a collaborative framework for model reasoning and tool execution. When the agent determines that the current task involves external knowledge acquisition, data retrieval, mathematical calculation, code execution, database query, or interface invocation capabilities, it automatically generates a tool invocation request based on the current reasoning state and selects the appropriate tool to complete the target operation. After the tool completes its execution, the results returned by the tool are fed back to the agent's inference module as observation information. Together with the current inference state, they form new context information, driving the next round of inference and realizing a closed-loop execution mechanism of "inference-tool call-result feedback-continued inference". The agent continuously corrects the inference path based on the continuous feedback until the task termination condition is met, and outputs the final task solution result.
[0058] Furthermore, after the task is completed, the intermediate states generated during the reasoning process, tool call logs, task execution trajectory, and key decision-making basis are output to form a reasoning trajectory that supports explanation, providing support for subsequent result verification, task backtracking, and system optimization.
[0059] Step S8: Application and Feedback Optimization a. Application-scenario-oriented model deployment and service integration A model deployment module is constructed to deploy the model, after task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The specific implementation method is as follows: Based on the target business requirements, the model is encapsulated as an API interface, microservice interface, or containerized service, and a data interaction channel is established with the business system. This enables applications, including intelligent question answering, decision support, and multimodal analysis, to uniformly call the model service. After receiving a user request, the business system inputs the relevant task into the intelligent agent system, which then completes task understanding, information processing, reasoning analysis, and result generation. The system then returns the task execution result to the business system, achieving collaborative integration of model capabilities and business processes.
[0060] Meanwhile, the model deployment module supports multi-instance deployment, elastic resource scheduling, and service expansion mechanisms. It also supports dynamically adjusting the number of model instances and the allocation of computing resources according to business load to improve the system's scalability, stability, and concurrent processing capabilities.
[0061] b. Strategy optimization and closed-loop update based on feedback information A feedback optimization module is constructed to uniformly collect, analyze, and evaluate task execution results, user feedback information, and system operation indicators generated during model operation. Based on the feedback results, the system operation strategy is dynamically optimized to achieve continuous iterative optimization of model capabilities. The specific implementation method is as follows: After the model completes the task, the feedback optimization module synchronously collects operational metrics including task execution accuracy, response latency, operating cost, resource utilization, and user satisfaction. It also combines user feedback information to comprehensively evaluate the current information filtering strategy, budget allocation strategy, and path scheduling strategy. When a decline in task performance, an increase in response latency, a decrease in resource utilization, or a decrease in user satisfaction is detected, the feedback optimization module automatically analyzes the influencing factors and generates corresponding strategy adjustment plans.
[0062] Furthermore, based on the feedback analysis results, the feedback optimization module dynamically updates the information filtering threshold, task-related weights, budget allocation parameters, path scheduling rules, and model calling strategies, and reapplies the updated strategies to the model service system, forming a closed-loop optimization mechanism of "task execution - feedback collection - strategy optimization - parameter update - re-execution". This enables the system to continuously optimize the information processing flow according to the actual operating status, thereby improving task execution accuracy, resource utilization efficiency, and overall system performance.
[0063] This agent-based hierarchical information bottleneck processing system is used to implement the above method, including: The multimodal data acquisition and preprocessing module is responsible for acquiring multimodal raw data related to the task based on the application scenario of the intelligent agent's task execution, and performing preprocessing operations according to the data type to eliminate noise interference and generate feature-encoded data. Subsequently, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation. The task constraint modeling and goal definition module is responsible for customizing and constructing task constraint models according to specific application scenarios. It performs unified modeling of task types, execution goals, constraints, and evaluation indicators. It performs semantic parsing, semantic annotation, and importance assessment on input information, identifying core semantic units, key entities, task-related information, and contextual relationships in the input content. Based on semantic relevance, task contribution, and information importance, it custom-generates priority identifiers, and custom-determines the key content areas that need to be retained and the information areas that can be downgraded or filtered, forming an information retention strategy oriented towards the target task. The hierarchical information bottleneck processing module is responsible for constructing a hierarchical information processing structure, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, which are used to process the input information step by step; and introduces information bottleneck mechanisms in the primary filtering layer, intermediate compression layer, and high-level semantic abstraction layer respectively to perform constrained compression processing on the input information. The task-aware information selection module is responsible for uniformly modeling information from different sources, modalities, and semantic levels based on task constraints and the compressed semantic representation output by the hierarchical information bottleneck processing module, and dynamically adjusting the retention weights of various types of information according to the target task; based on the dynamic weight allocation results, an attention mechanism and dynamic information selection strategy are introduced to enhance the task relevance of each information unit. The budget constraint computation scheduling module is responsible for constructing a computation budget model. It performs unified modeling of budget constraint factors in the information processing process, including computing resources, response latency, and operating costs, and forms a budget constraint vector. It normalizes the various indicators in the budget constraint vector and calculates a comprehensive budget score based on the importance of different constraint vectors. Based on the comprehensive budget score, it constructs a dynamic path scheduling mechanism to adaptively select different information processing paths according to budget constraints, so as to achieve a dynamic balance between computing resource consumption, response efficiency, and model inference performance. The mutual information evaluation and rate-distortion optimization module is responsible for constructing an information quality evaluation mechanism based on mutual information, quantitatively evaluating the information retention capability between the original input information and the compressed semantic representation, evaluating the expressive capability of the compressed representation for task-related information in combination with the target task constraints, and constructing a rate-distortion optimization mechanism based on the mutual information evaluation results to jointly optimize the compression ratio, information retention degree and computational resource consumption in the information compression process. The agent reasoning module is responsible for understanding, decomposing, and planning the reasoning of the target task based on intermediate semantic representations, and generating multiple sub-tasks according to a preset reasoning strategy. It constructs a reasoning execution sequence based on the dependencies between sub-tasks, and solves complex tasks through step-by-step reasoning, state updates, and intermediate result feedback mechanisms, ultimately generating a reasoning result that meets the requirements of the target task. In the agent reasoning process, it introduces the reasoning capabilities of a large language model and an external tool invocation mechanism to build a collaborative framework for model reasoning and tool execution. It automatically generates tool invocation requests based on the current reasoning state, and feeds back the tool's return results as observation information to the agent reasoning module. Together with the current reasoning state, these form new contextual information to drive the next round of reasoning. The model deployment module is responsible for deploying the model, which has undergone task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The feedback optimization module is responsible for the unified collection, analysis and evaluation of task execution results, user feedback information and system operation indicators generated during model operation, and dynamically optimizes the system operation strategy based on the feedback results to achieve continuous iterative optimization of model capabilities.
[0064] The present invention discloses a hierarchical information bottleneck processing device for intelligent agents, comprising a memory and a processor; the memory is used to store a computer program, and the processor is used to execute the computer program to implement the above-described method steps.
[0065] The readable storage medium stores a computer program that, when executed by a processor, implements the above-described method steps.
[0066] Compared with existing technologies, this method for handling hierarchical information bottlenecks in intelligent agents has the following characteristics: a. Significantly improved information processing efficiency: This invention uses a hierarchical information bottleneck structure to filter and compress input information step by step, and dynamically adjusts the processing path in conjunction with computational budget constraints, effectively reducing redundant information in the computation process, significantly reducing overall computational overhead, and improving system inference efficiency and response speed.
[0067] b. Strong task adaptability and high result stability: By introducing a task constraint-driven information selection mechanism, the information retention strategy can be dynamically adjusted according to different task objectives, the expression of key semantic information can be strengthened, and the interference of irrelevant information can be suppressed, thereby improving the reasoning accuracy and result stability of the agent in complex task scenarios.
[0068] c. High resource utilization and low operating cost: This invention uses a budget-driven computing scheduling mechanism to rationally allocate computing resources while ensuring task performance, thereby achieving an optimal balance between information compression and computing cost, effectively reducing system operating costs and improving resource utilization efficiency.
[0069] d. Clear structure and easy engineering deployment: Its adaptive multimodal characteristics enable it to flexibly adapt to different types of multimodal data and diverse application scenarios, demonstrating good performance in various fields and having wide applicability and compatibility.
[0070] e. Wide applicability and strong scalability: This method can support the processing of multimodal data such as text and images, and can be combined with large model inference, intelligent agent systems and tool calling mechanisms. It is suitable for various application scenarios such as intelligent question answering, decision analysis, and multimodal understanding, and has good versatility and scalability.
[0071] The embodiments described above are merely one specific implementation of the present invention. Ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for handling hierarchical information bottlenecks in intelligent agents, characterized in that: Includes the following steps: Step S1: Collect multimodal raw data related to the task, preprocess it, and generate feature-encoded data; Step S2: Customize and construct a task constraint model according to the specific application scenario, parse and evaluate the importance of the input information, and generate a priority identifier to provide a basis for subsequent information filtering, context construction and reasoning generation; Step S3: Construct a hierarchical information processing structure, introduce an information bottleneck mechanism, construct a hierarchical information bottleneck processing module, process the input information step by step, and achieve a balance between information compression efficiency and semantic expression ability. Step S4: Based on task constraints and the compressed semantic representation output by the hierarchical information bottleneck processing module, construct a task-aware information selection module, perform unified modeling of different information, and dynamically adjust the retention weight of various types of information according to the target task; introduce attention mechanism and dynamic information selection strategy to perform task relevance enhancement processing on each information unit. Step S5: Construct a computational budget model to uniformly model the budget constraints in the information processing process, forming a budget constraint vector to provide a quantitative basis for the dynamic scheduling of subsequent information processing paths; Step S6: Construct an information quality assessment mechanism based on mutual information to quantitatively evaluate the information retention capability between the original input information and the compressed semantic representation, and evaluate the ability of the compressed representation to express task-related information in combination with the target task constraints, so as to provide feedback basis for subsequent compression strategy optimization. Step S7: The intelligent agent reasoning module performs task parsing and reasoning planning for the target task based on the intermediate semantic representation, and performs reasoning step by step according to the preset reasoning strategy to complete the solution of complex tasks. Step S8: Build a model deployment module to deploy the model after intelligent agent inference to actual business application scenarios, providing unified intelligent service capabilities for different types of tasks.
2. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S1, based on the application scenario of the intelligent agent performing the task, multimodal raw data related to the task is collected. The multimodal raw data includes at least one or more of text, images, audio, video, and structured data. The acquired multimodal raw data are preprocessed according to their data types to eliminate noise interference and generate feature-coded data. The specific methods are as follows: For text data, the following steps are performed in sequence: text cleaning, word segmentation, stop word removal, semantic encoding, and vectorization representation. For image data, image denoising, size normalization, target region extraction, and visual feature encoding are performed; For audio data, noise suppression, speech enhancement, speech recognition, and acoustic feature extraction are performed. For structured data, outlier detection, missing value imputation, field standardization, and numerical normalization are performed. After preprocessing, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation, providing a unified input for subsequent hierarchical information bottleneck processing driven by task constraints and budget.
3. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S2, task constraint modeling and target definition are performed. Custom task constraint models are built according to specific application scenarios, and task types, execution objectives, constraints, and evaluation indicators are modeled in a unified manner. The task types include question-answering tasks, decision-making tasks, analysis tasks, generation tasks, and reasoning tasks. The corresponding objective functions, output forms, and key output requirements are determined in combination with different task types to form constraint rules adapted to the target tasks, providing task guidance for subsequent information filtering and model reasoning. Based on the constructed task constraint model, semantic parsing, semantic annotation, and importance assessment are performed on the input information to identify core semantic units, key entities, task-related information, and contextual relationships in the input content. Priority identifiers are generated according to semantic relevance, task contribution, and information importance. Key content areas that need to be retained and information areas that can be downgraded or filtered are also defined to form an information retention strategy oriented towards the target task, providing a basis for subsequent information filtering, context construction, and reasoning generation.
4. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S3, a hierarchical information processing structure is constructed, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, for processing the input information step by step. The primary filtering layer is used to initially filter irrelevant, repetitive, and low-value content in the input information. The intermediate compression layer is used to perform semantic aggregation, redundancy merging, and context compression on the information after primary filtering. The high-level semantic abstraction layer is used to extract high-level semantic features, core intentions, and key reasoning bases related to the target task, thereby forming a concise semantic representation suitable for subsequent model reasoning or task execution. Information bottleneck mechanisms are introduced in the primary filtering layer, intermediate compression layer, and advanced semantic abstraction layer, respectively, to construct a hierarchical information bottleneck processing module for constrained compression of input information; specifically as follows: Based on task relevance, semantic importance, information redundancy, and contextual contribution, and using a predefined information bottleneck mechanism, the input information at each layer is filtered, compressed, and reconstructed. By customizing the information retention ratio, compression granularity, and semantic abstraction level of each layer, a balance is achieved between information compression efficiency and semantic expressive power.
5. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S4, based on task constraints and the compressed semantic representation output by the hierarchical information bottleneck processing module, a task-aware information selection module is constructed to uniformly model information from different sources, different modalities and different semantic levels, and dynamically adjust the retention weight of various types of information according to the target task. Input information is represented as ,in, This represents the set of input information for the task-aware information selection module. Indicates the first There are 1 information units, where n represents the total number of information units; Let the task constraints be represented as The task relevance score of each information unit is calculated based on task constraints: , in, This represents the task relevance evaluation function; Representation of information unit With the current task The correlation score between them; To allocate importance among different information units, the Softmax function is used to calculate the dynamic weight of each information unit. , in, Indicates the first The retention weight of each information unit, and satisfying ; Subsequently, the input information is weighted according to the calculated dynamic weights: , in, This represents the information representation after task-aware enhancement. Based on the dynamic weight allocation results, an attention mechanism and a dynamic information selection strategy are introduced to enhance the task relevance of each information unit; the details are as follows: By combining preset retention thresholds, TopK filtering strategies, or dynamic gating mechanisms, task-related information is enhanced, and redundant, repetitive, and irrelevant information is suppressed to reduce the propagation of invalid information in subsequent reasoning processes. When using a threshold gating strategy, the information retention flag is represented as follows: , , Where, m i To preserve the information's tags, The preset threshold; Alternatively, the TopK strategy can be used, retaining only the k information units with the highest weights, and its output is represented as: , in, This represents the TopK filtering operator, where k represents the number of information units to be retained.
6. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S5, a computational budget model is constructed to uniformly model the budget constraints in the information processing process, including computing resources, response latency, and operating costs, and to form a budget constraint vector, providing a quantitative basis for the dynamic scheduling of subsequent information processing paths. The computing resources include processor computing power, GPU computing resources, video memory capacity, and memory usage; the response latency is used to describe the maximum allowable response time for the target task; the operating cost is used to describe the model call cost, interface call cost, and computing resource consumption cost; The budget constraint vector is represented as Where B represents the budget constraint vector, C represents the computing power resource index that supports the call, L represents the response latency constraint index, and R represents the operating cost constraint index; The indicators in the budget constraint vector are normalized, and the comprehensive budget score is calculated based on the importance of different constraint vectors. The calculation method for the comprehensive budget score is as follows: , This indicates the overall budget score; , and These represent the normalized computing power resource indicators, response latency constraint indicators, and operating cost constraint indicators, respectively. , and Let each represent the weight coefficient of the corresponding indicator, and satisfy the following: , The comprehensive budget score is used to quantify the computing budget available to support the current system and serves as an evaluation basis for the selection of subsequent information processing paths; The comprehensive budget score is compared with a preset budget threshold, and the target processing path is determined based on the comparison result. The path selection rule is expressed as follows: , in, This indicates a lightweight processing path. Indicates the deep reasoning path, Indicates the threshold for budget decision-making; When the overall budget score is lower than the budget threshold, the scheduling controller selects a lightweight processing path and adopts a lightweight model, few-step inference and shallow information compression strategy to reduce computing resource consumption and response latency. When the overall budget score reaches or exceeds the budget threshold, the scheduling controller selects a deep inference path and adopts a large-scale model, multi-step inference strategy and toolchain collaboration mechanism to improve the inference ability and result accuracy of complex tasks. During task execution, the scheduling controller continuously monitors the model's running status, resource utilization, response latency, and inference performance, and dynamically adjusts the comprehensive budget score based on real-time feedback. The update method is as follows: , in, Indicates the first Comprehensive budget score during the next scheduling This represents the budget adjustment amount obtained based on system operation feedback. Indicates the budget update factor, and satisfies ; The scheduling controller re-executes path selection based on the updated comprehensive budget score, thereby enabling dynamic switching between lightweight processing paths and deep inference paths. This allows the system to adaptively adjust its information processing strategy according to the current budget constraints, achieving an optimal balance between computational resource consumption, response efficiency, and inference performance.
7. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S6, an information quality assessment mechanism based on mutual information is constructed to quantitatively evaluate the information retention capability between the original input information and the compressed semantic representation, and to evaluate the ability of the compressed representation to express task-related information in combination with the target task constraints, so as to provide feedback basis for subsequent compression strategy optimization. Let the original input information be represented as The compressed representation information after hierarchical information processing is The task constraint representation information is as follows The mutual information between the input information and the compressed representation is then expressed as: ; Mutual information represents the conditional mutual information between the original input information and the compressed representation information under task constraints. It is used to measure the degree to which the compressed representation information retains task-related information. The larger the mutual information value, the richer the effective task information contained in the compressed representation. The smaller the mutual information value, the greater the information loss during the compression process. Based on the mutual information evaluation results, a rate and distortion optimization mechanism is constructed to jointly optimize the compression ratio, information retention degree and computing resource consumption in the information compression process. During system operation, the information compression threshold, compression ratio, model size, and information processing path are dynamically adjusted based on mutual information evaluation results and current task performance feedback. When mutual information drops below a preset threshold or the task performance declines at a rate exceeding a custom threshold, the system reduces the compression ratio or increases the amount of information retained. When mutual information meets task requirements and computing resources are limited, the compression ratio is increased to reduce computing costs and achieve a dynamic optimization balance between task performance, information retention capacity, and computing resource consumption. To evaluate the information retention efficiency of a unit of coded resources, an information rate index is introduced, which is calculated as follows: , Where A represents the encoding length or number of encoded bits in the compressed representation, and E represents the information retention efficiency per unit encoding length, which is used to comprehensively evaluate compression efficiency and information retention capability; The mutual information and information rate are calculated using Monte Carlo sampling, contrastive learning estimation, or neural mutual information estimation methods to form a task-oriented information quality evaluation index. Let the information distortion caused by the compression process be represented as: , in, Indicates compression distortion. The distortion metric function between the initial input and the compressed representation is described by semantic distance, reconstruction error, or task performance loss. The joint optimization objective of establishment rate and distortion is expressed as: , in, Indicates the overall optimization objective; Indicates information rate; This indicates information distortion; The representation rate versus distortion tradeoff is used to balance compression efficiency and information retention.
8. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S7, the intermediate semantic representation after task constraint modeling, hierarchical information bottleneck processing and task perception information selection is input into the agent reasoning module as a unified input for the agent to perform task planning, reasoning analysis and decision generation; the intermediate semantic representation includes target task information, contextual semantic information, key constraints and high-value task-related information after screening. The intelligent agent reasoning module performs task understanding, task decomposition and reasoning planning for the target task based on intermediate semantic representation, and generates multiple sub-tasks according to the preset reasoning strategy. Based on the dependencies between subtasks, a reasoning execution sequence is constructed, and complex tasks are solved through step-by-step reasoning, state updates, and intermediate result feedback mechanisms. During the reasoning process, the agent continuously receives intermediate results generated by the reasoning and uses these intermediate results as input for subsequent reasoning steps, thereby achieving dynamic updates and iterative optimization of the reasoning state and ultimately generating reasoning results that meet the requirements of the target task. During the agent's reasoning process, the reasoning capabilities of a large language model and an external tool invocation mechanism are introduced to construct a collaborative framework for model reasoning and tool execution. When the agent determines that the current task involves external knowledge acquisition, data retrieval, mathematical calculation, code execution, database query, or interface invocation capabilities, it automatically generates a tool invocation request based on the current reasoning state and selects the appropriate tool to complete the target operation. After the tool completes its execution, the results returned by the tool are fed back to the agent's inference module as observation information. Together with the current inference state, they form new context information, driving the next round of inference and realizing a closed-loop execution mechanism. The agent continuously corrects the inference path based on the continuous feedback until the task termination condition is met, and then outputs the final task solution result. After the task is completed, the intermediate states generated during the reasoning process, tool call logs, task execution trajectory, and key decision-making basis are output to form a reasoning trajectory that supports explanation, providing support for subsequent result verification, task backtracking, and system optimization.
9. The method for handling hierarchical information bottlenecks in intelligent agents according to claim 1, characterized in that: In step S8, a model deployment module is constructed to deploy the model, which has undergone task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The specific implementation method is as follows: Based on the target business needs, the model is encapsulated as an API interface, microservice interface, or containerized service, and a data interaction channel is established with the business system so that applications including intelligent question answering, decision support, and multimodal analysis can uniformly call the model service. After receiving a user request, the business system inputs the relevant task into the intelligent agent system, which then completes task understanding, information processing, reasoning analysis, and result generation, and returns the task execution result to the business system, thereby achieving the collaborative integration of model capabilities and business processes. Meanwhile, the model deployment module supports multi-instance deployment, elastic resource scheduling, and service expansion mechanisms. It supports dynamically adjusting the number of model instances and the allocation of computing resources according to business load to improve the system's scalability, stability, and concurrent processing capabilities. A feedback optimization module is constructed to uniformly collect, analyze, and evaluate task execution results, user feedback information, and system operation indicators generated during model operation. Based on the feedback analysis results, information filtering thresholds, task-related weights, budget allocation parameters, path scheduling rules, and model invocation strategies are dynamically updated, and the updated strategies are reapplied to the model service system, forming a closed-loop optimization mechanism to achieve continuous iterative optimization of model capabilities. The specific implementation method is as follows: After the model completes the task, the feedback optimization module synchronously collects operational metrics including task execution accuracy, response latency, operating cost, resource utilization, and user satisfaction. It also combines user feedback information to comprehensively evaluate the current information filtering strategy, budget allocation strategy, and path scheduling strategy. When a decline in task performance, an increase in response latency, a decrease in resource utilization, or a decrease in user satisfaction is detected, the feedback optimization module automatically analyzes the influencing factors and generates corresponding strategy adjustment plans.
10. A hierarchical information bottleneck processing system for intelligent agents, characterized in that: To implement the method according to any one of claims 1 to 9, comprising: The multimodal data acquisition and preprocessing module is responsible for acquiring multimodal raw data related to the task based on the application scenario of the intelligent agent's task execution, and performing preprocessing operations according to the data type to eliminate noise interference and generate feature-encoded data. Subsequently, the features of different modalities are mapped to a unified feature representation space to form a standardized multimodal feature representation. The task constraint modeling and goal definition module is responsible for customizing and constructing task constraint models according to specific application scenarios. It performs unified modeling of task types, execution goals, constraints, and evaluation indicators. It performs semantic parsing, semantic annotation, and importance assessment on input information, identifying core semantic units, key entities, task-related information, and contextual relationships in the input content. Based on semantic relevance, task contribution, and information importance, it custom-generates priority identifiers, and custom-determines the key content areas that need to be retained and the information areas that can be downgraded or filtered, forming an information retention strategy oriented towards the target task. The hierarchical information bottleneck processing module is responsible for constructing a hierarchical information processing structure, including a primary filtering layer, an intermediate compression layer, and a high-level semantic abstraction layer, which are used to process the input information step by step; and introduces information bottleneck mechanisms in the primary filtering layer, intermediate compression layer, and high-level semantic abstraction layer respectively to perform constrained compression processing on the input information. The task-aware information selection module is responsible for uniformly modeling information from different sources, modalities, and semantic levels based on task constraints and the compressed semantic representation output by the hierarchical information bottleneck processing module, and dynamically adjusting the retention weights of various types of information according to the target task; based on the dynamic weight allocation results, an attention mechanism and dynamic information selection strategy are introduced to enhance the task relevance of each information unit. The budget constraint computation scheduling module is responsible for constructing a computation budget model. It performs unified modeling of budget constraint factors in the information processing process, including computing resources, response latency, and operating costs, and forms a budget constraint vector. It normalizes the various indicators in the budget constraint vector and calculates a comprehensive budget score based on the importance of different constraint vectors. Based on the comprehensive budget score, it constructs a dynamic path scheduling mechanism to adaptively select different information processing paths according to budget constraints, so as to achieve a dynamic balance between computing resource consumption, response efficiency, and model inference performance. The mutual information evaluation and rate-distortion optimization module is responsible for constructing an information quality evaluation mechanism based on mutual information, quantitatively evaluating the information retention capability between the original input information and the compressed semantic representation, evaluating the expressive capability of the compressed representation for task-related information in combination with the target task constraints, and constructing a rate-distortion optimization mechanism based on the mutual information evaluation results to jointly optimize the compression ratio, information retention degree and computational resource consumption in the information compression process. The agent reasoning module is responsible for understanding, decomposing, and planning the reasoning of the target task based on intermediate semantic representations, and generating multiple sub-tasks according to a preset reasoning strategy. It constructs a reasoning execution sequence based on the dependencies between sub-tasks, and solves complex tasks through step-by-step reasoning, state updates, and intermediate result feedback mechanisms, ultimately generating a reasoning result that meets the requirements of the target task. In the agent reasoning process, it introduces the reasoning capabilities of a large language model and an external tool invocation mechanism to build a collaborative framework for model reasoning and tool execution. It automatically generates tool invocation requests based on the current reasoning state, and feeds back the tool's return results as observation information to the agent reasoning module. Together with the current reasoning state, these form new contextual information to drive the next round of reasoning. The model deployment module is responsible for deploying the model, which has undergone task constraint modeling, information filtering, hierarchical information processing, budget scheduling, and agent inference, to actual business application scenarios. It also integrates with external business systems through standardized interfaces, providing unified intelligent service capabilities for different types of tasks. The feedback optimization module is responsible for the unified collection, analysis and evaluation of task execution results, user feedback information and system operation indicators generated during model operation, and dynamically optimizes the system operation strategy based on the feedback results to achieve continuous iterative optimization of model capabilities.