Large model agent system based on multilayer iterative optimization feedback mechanism

By employing a multi-layered iterative optimization feedback mechanism for multimodal data fusion, uncertainty quantification, and dynamic knowledge graph construction, the shortcomings of large-scale intelligent agents in multimodal understanding, collaborative optimization, and dynamic environment adaptability are addressed, resulting in higher decision-making accuracy and dynamic response capabilities.

CN120930778APending Publication Date: 2025-11-11ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510946921.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing large-scale intelligent agents have shortcomings in multimodal understanding, collaborative optimization, and dynamic environment adaptation, including insufficient feature coupling, uncontrollable error propagation, cross-modal attention imbalance, vertical information fragmentation, rigid resource allocation, discrete optimization interval, and semantic-physical space decoupling.

Method used

By adopting a multi-layer iterative optimization feedback mechanism, which combines a standardized input processing layer, a core decision-making layer, and an optimization feedback layer, the above problems are solved by achieving multi-modal data fusion, uncertainty quantification, dynamic knowledge graph construction, and cross-level gradient propagation.

Benefits of technology

It effectively solves the problems of semantic understanding bias, collaborative optimization defects and environmental adaptability in traditional single-layer architecture, and achieves higher decision-making accuracy and dynamic response capability, making it suitable for complex decision-making scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930778A_ABST
    Figure CN120930778A_ABST
Patent Text Reader

Abstract

The invention discloses a large model agent system based on a multi-layer iterative optimization feedback mechanism, which comprises a standardized input processing layer used for carrying out feature fusion on input multi-modal data and quantifying uncertainty of different modal data to obtain standardized multi-modal feature data, and a multi-modal feature data processing layer used for carrying out feature fusion on the input multi-modal data. The normalized multi-modal feature data input is used for performing time sequence knowledge graph construction on the normalized multi-modal feature data to obtain decision-making tree feature data and generate a decision-making tree, and is used for a core decision-making layer of decision-making reasoning; decision tree feature data input is subjected to cross-level propagation joint optimization processing, and an optimization feedback layer, with balanced parameters and high dynamic adaptability, of the decision inference model is obtained. According to the method, the problems of accumulation of understanding deviation of an existing large model agent, defects of cross-level collaborative optimization and limitation of real-time response of a dynamic environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and intelligent agent technology, specifically relating to a large model intelligent agent system based on a multi-layer iterative optimization feedback mechanism. Background Technology

[0002] With the continuous development of large model technology, large model agents are increasingly becoming the core of practical applications of large models. However, existing large model agents have problems in areas such as comprehension bias, collaborative optimization, and dynamic scenarios, specifically manifested in the following aspects:

[0003] 1. The problem of multimodal understanding bias accumulation in single-layer decision-making mechanism: Most current mainstream agent architectures adopt single-layer decision-making mechanism. Its processing flow is usually a linear structure of "input → feature extraction → decision output". This design has significant defects in multimodal scenarios: (1) insufficient feature coupling. When processing text, speech and visual input at the same time, the features of each modality are compressed into a single vector in the early fusion stage, resulting in the loss of fine-grained semantic information; (2) uncontrollable error propagation. The single-layer structure lacks intermediate verification links. The recognition error in the initial stage (such as speech-to-text error) will directly pollute the subsequent decision; (3) cross-modal attention imbalance. Existing methods use a cross-modal attention mechanism with fixed weights, which is difficult to dynamically adapt to the needs of different scenarios.

[0004] 2. Defects of collaborative optimization in one-way feedback mechanism: Existing feedback systems mostly adopt a one-way correction mode of "decision layer → execution layer", which has the following limitations: (1) vertical information discontinuity, the feedback flow of each level (perception, reasoning, execution) is isolated from each other; (2) lack of temporal correlation, the traditional method adopts a feedback mechanism with independent timestamps, ignoring the causal relationship of action sequence; (3) rigid resource allocation: Microsoft Azure Cognitive Service adopts a fixed proportion of resource allocation strategy (70% of computing resources are used for perception and 30% for decision-making), which cannot be dynamically adjusted according to the complexity of the task. When dealing with sudden hot events, the bottleneck of decision resources leads to a 400% increase in response delay.

[0005] 3. Theoretical limitations of dynamic environmental adaptability: There are two fundamental problems with the current environmental adaptation mechanism: (1) Discrete optimization interval: The mainstream method adopts a parameter update strategy with a fixed interval (such as updating once every 10 minutes). In the test of high-frequency trading scenario in finance, this mechanism leads to: when the market fluctuation is <5%, the decision accuracy is 91%, and when the fluctuation is >15%, it drops sharply to 32%. (2) Semantic-physical space decoupling: The existing intelligent agent architecture separates and optimizes the environmental understanding (semantic layer) and the action execution (physical layer), resulting in a semantic layer recognition accuracy of 92%, a physical layer execution success rate of 78%, but a joint success rate of only 58%. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, the present invention aims to provide a large model intelligent agent system based on a multi-layer iterative optimization feedback mechanism, which solves the problems of accumulated understanding bias, defects in cross-level collaborative optimization, and limitations in real-time response to dynamic environments in existing large model intelligent agents.

[0007] The technical solution of the present invention is as follows:

[0008] A large-model intelligent agent system based on a multi-layer iterative optimization feedback mechanism includes a normalized input processing layer for fusing features from input multimodal data and quantifying the uncertainty of different modal data to obtain normalized multimodal feature data. The normalized multimodal feature data input is used to construct a temporal knowledge graph from the normalized multimodal feature data to obtain decision tree feature data and generate a decision tree, which is the core decision layer for decision reasoning. The decision tree feature data input undergoes cross-level propagation joint optimization processing to obtain an optimized feedback layer for a decision reasoning model with balanced parameters and strong dynamic adaptability.

[0009] Furthermore, the normalized input processing layer includes a multimodal feature fusion module and an uncertainty quantization module. The multimodal feature fusion module takes multimodal data as input and sequentially passes it through an attention weight dynamic allocation unit, a tensor fusion unit, a gating unit, and a contrastive learning unit to obtain multimodal feature data. The uncertainty quantization module, based on the multimodal feature data processed by the multimodal feature fusion module, uses a Bayesian method to model the uncertainty of the multimodal feature data through probability distribution to obtain normalized multimodal feature data.

[0010] Furthermore, the multimodal feature fusion module aligns the temporal or spatial sequences of different modalities through dynamic temporal warping and cross-modal attention; it achieves feature extraction and representation, cross-modal alignment and interaction through specific encoders for different modalities; it completes missing modalities through generative adversarial networks; and it improves computational efficiency by integrating a lightweight attention mechanism with PCA dimensionality reduction technology.

[0011] Furthermore, the method of using Bayesian approaches to model the uncertainty of multimodal feature data through probability distributions specifically involves:

[0012] 1) Construct a Bayesian neural network, assuming that the weights are probability distributions, and approximate the posterior through variational inference or Markov chain Monte Carlo.

[0013] 2) The Monte Carlo Dropout method is used to calculate uncertainty. During inference, Dropout is kept active, and the prediction results are sampled multiple times to calculate the variance. The calculation formula is as follows:

[0014]

[0015] Where T is the number of samples, θ t Let be the model parameters for the t-th sampling.

[0016] Furthermore, the core decision-making layer includes a dynamic knowledge graph construction module and a decision tree generation module; the dynamic knowledge graph construction module inputs standardized multimodal feature data, and sequentially passes through a streaming data processing unit, an incremental knowledge extraction unit, and a time-aware modeling unit to obtain a temporal knowledge graph; the decision tree generation module inputs the temporal knowledge graph, and sequentially passes through a feature selection unit, a node partitioning unit, and a computational optimization unit to generate a decision tree.

[0017] Furthermore, the streaming data processing unit is used to store dynamic data streams, employs Apache Flink for real-time knowledge extraction, uses a sliding window to continuously update recent knowledge fragments, and uses a session window to process bursty data streams.

[0018] The incremental knowledge extraction unit takes the dynamic data stream as input, incrementally updates the entity dictionary based on the online learning NER model, uses context embedding to link existing entities in the knowledge graph in real time, matches dynamic events based on predefined templates, uses online learning to update the relation extraction model for incremental training, and combines a timestamp graph neural network to capture relation evolution, performs temporal relation modeling, and outputs the knowledge graph.

[0019] The time-aware modeling unit takes the knowledge graph as input, adds a time dimension to the knowledge graph embedding, models the timeliness of the knowledge graph, segments the knowledge graph by time window, supports historical state query, eliminates expired knowledge based on time decay function or event triggering, and uses a graph database to record the evolution process of the knowledge graph, outputs a time-series knowledge graph, and stores it in the graph database.

[0020] Furthermore, the feature selection unit uses squared error for feature selection, selecting feature A and segmentation point s, with the objective function being:

[0021]

[0022] Where D1 and D2 are the partitioned subsets, and c1 and c2 are the mean of the subsets;

[0023] For discrete features, the node partitioning unit directly partitions child nodes according to feature values; for continuous features, a binary search method is used to find the optimal split point; a substitution splitting strategy is used to handle missing values, pre-selecting substitution splitting features for each node, and using the substitution feature for partitioning when the main feature is missing; a cost complexity pruning method is adopted, balancing tree complexity and error through parameter α, with the loss function being:

[0024] C a(T)=C(T)+α|T|

[0025] Where C(T) is the error and |T| is the number of leaf nodes;

[0026] The computational optimization unit sorts continuous feature values ​​and quickly traverses possible split points to achieve pre-sorting; it adopts a parallelization method to distribute the computation of information gain or Gini index of each feature to accelerate feature selection.

[0027] Furthermore, the optimization feedback layer includes a cross-level gradient propagation module and a dual-mode feedback module. The cross-level gradient propagation module is used to achieve parameter update balance between different levels, and the dual-mode feedback module enhances the dynamic adaptability of the model through two complementary feedback mechanisms.

[0028] Furthermore, the cross-level gradient propagation module includes a gradient routing unit, a gradient correction unit, and an adaptive optimization unit;

[0029] The gradient routing unit dynamically allocates gradient traffic at different levels through gating weights; using an attention mechanism, attention weights are generated based on the decision tree feature data to adjust the gradient contributions at different levels, thereby obtaining gradient weights at different levels.

[0030] The gradient correction unit is based on gradient pruning technology to limit the maximum norm of cross-level gradients and prevent gradient explosion; it uses gradient normalization technology to standardize cross-level gradients and balance the update magnitude of each level; and it uses second-order optimization technology to adjust the update direction using the Hessian matrix.

[0031] The adaptive optimization unit adopts a hierarchical learning rate strategy, assigning independent learning rates to different levels; at the same time, for the Transformer model, the parameter matrix is ​​decomposed and optimized in blocks to reduce cross-layer gradient coupling.

[0032] Furthermore, the dual-mode feedback module includes an explicit feedback unit, an implicit feedback unit, and a dual-mode fusion unit;

[0033] The explicit feedback unit performs gradient reweighting on the gradient using a learnable weight matrix W. e The gradient is then recalibrated in terms of both spatial and channel dimensions to obtain the calibrated gradient, as shown in the following formula:

[0034]

[0035] in, For the upstream gradient, F in G represents the current input layer features. e This represents the calibrated gradient;

[0036] The calibrated gradient is injected into the feature using a gating mechanism for feature correction, as shown in the following formula:

[0037] F explicit =F in +α·Conv(G e )

[0038] Where α is an adaptive scalar parameter, dynamically generated through a lightweight MLP, and F explicit Represents the explicit features after feature correction;

[0039] The implicit feedback unit is based on the current layer feature F in For cross-layer features, calculate the long-range dependencies within the features and perform self-attention aggregation, as shown in the formula:

[0040]

[0041] Where Q, K, and V are derived from F in The linear transformation yields d, where d is the feature dimension and A represents the dependency matrix within the feature.

[0042] A dynamic convolution kernel is used to capture local patterns, and dynamic convolution compensation is applied to the input features. The formula is as follows:

[0043] F implicit =F in +Conv(W d A)

[0044] Among them, W d It is a dynamic convolution kernel, and its parameters are generated from feature statistics, F implicit Represents the implicit features after dynamic convolution compensation;

[0045] The dual-mode fusion unit balances the contributions of explicit and implicit features using learnable parameters, as shown in the following formula:

[0046] β=Sigmoid(MLP([F explicit ,F implicit ]))

[0047] F out =β·F explicit +(1-β)·F implicit

[0048] Where β represents the learnable parameter, sigmoid function, MLP represents the multilayer perceptron model, and F out Characteristics representing the output;

[0049] By using residual connections, we ensure that the final output features retain the original input information.

[0050] Ffinal =F out +F in

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] This invention employs a dual mechanism of multimodal data fusion and uncertainty quantification to effectively address the semantic understanding bias problem of traditional single-layer architectures. It utilizes dynamic knowledge graph construction and decision tree mechanisms to effectively solve the problem of collaborative optimization of large models. Furthermore, it employs a joint optimization mechanism of cross-level gradient propagation to effectively address environmental adaptability issues, making it suitable for dynamic optimization and adaptive interaction systems in complex decision-making scenarios. Through systematic innovation, this invention achieves a generational leap in error control, system synergy, and environmental adaptability, laying the technological foundation for building a new generation of adaptive intelligent agents. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of the system structure of the present invention. Detailed Implementation

[0054] The present invention will be further described in detail below with reference to specific embodiments. These descriptions are for explanation purposes only and are not intended to limit the scope of the invention.

[0055] refer to Figure 1 A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism is proposed, employing a three-level progressive architecture. It includes a normalized input processing layer, which fuses features from input multimodal data and quantifies the uncertainty of different modalities to obtain normalized multimodal feature data. This normalized input processing layer includes a multimodal feature fusion module and an uncertainty quantification module. The normalized multimodal feature data is input to a core decision layer, which incrementally constructs a knowledge graph from the normalized multimodal feature data processed by the normalized input processing layer and generates a decision tree for decision reasoning. This core decision layer includes a dynamic knowledge graph construction module and a decision tree generation module. The decision tree feature data is input to an optimization feedback layer, which performs cross-layer propagation joint optimization processing on the neural network model (multiple layers, each with gradients during optimization) based on the decision tree feature data output from the core decision layer, resulting in a parameter-balanced and dynamically adaptable inference model. This optimization feedback layer includes a cross-layer gradient propagation module and a dual-modal feedback module.

[0056] 1. Normalized Input Processing Layer

[0057] 1.1 Multimodal Feature Fusion Module

[0058] The multimodal feature fusion module includes an attention weight dynamic allocation unit, a tensor fusion unit, a gating unit, and a contrastive learning unit, which are performed sequentially: weight dynamic allocation unit → tensor fusion unit → gating unit → contrastive learning unit.

[0059] The multimodal feature fusion module dynamically allocates the importance of different modalities through attention weights to capture cross-modal correlations; it expands the feature vectors of different modalities into high-dimensional tensors through tensor fusion to capture all possible interaction combinations between modalities; based on the gating mechanism, it controls the contribution weights of different modal features through gating units; and it utilizes contrastive learning to bring related modal features closer together and push away irrelevant features through contrastive loss.

[0060] The multimodal feature fusion module addresses the temporal or spatial alignment issues between different modalities through Dynamic Time Warping (DTW) and cross-modal attention; it addresses the feature distribution differences between different modalities through modality-specific encoders (such as CNNs for images and LSTMs for text); it addresses the potential missing modal data by using Generative Adversarial Networks (GANs) to complete missing modalities; and it addresses the high computational cost when fusing high-dimensional features by using PCA dimensionality reduction techniques combined with a lightweight attention mechanism.

[0061] 1.2 Uncertainty Quantification Module

[0062] Based on the multimodal feature data obtained from the multimodal feature fusion module, a Bayesian method is used to model parameter uncertainty through probability distribution, and the posterior distribution is calculated during inference. The specific implementation is as follows:

[0063] (1) Construct a Bayesian neural network, assuming that the weights are probability distributions (generally Gaussian distributions are used), and approximate the posterior through variational inference (VI) or Markov chain Monte Carlo (MCMC).

[0064] (2) The Monte Carlo Dropout method is used to calculate uncertainty. Dropout is kept active during inference, and the prediction results are sampled multiple times to calculate the variance. The calculation formula is as follows:

[0065]

[0066] Where T is the number of samples, θ t Let be the model parameters for the t-th sampling.

[0067] 2. Core Decision Layer (CDL)

[0068] The core area decision-making layer consists of two parts: a dynamic knowledge graph construction module and a decision tree generation module.

[0069] 2.1 Dynamic Knowledge Graph Construction Module

[0070] The dynamic knowledge graph construction module uses units such as streaming data processing, incremental knowledge extraction, and time-aware modeling to continuously extract, fuse, and infer knowledge from the standardized multimodal feature data output from the standardized input processing layer (because it is processed online, it is dynamic).

[0071] (1) Streaming data processing unit

[0072] Apache Flink was chosen for real-time knowledge extraction. A window mechanism was designed: a sliding window was used to continuously update recent knowledge fragments; a session window was used to handle bursty data streams. This unit is a tool for storing streaming data (i.e., normalized multimodal feature data output from the normalized input processing layer; because it is processed online, it is a dynamic data stream), and this data type needs to be stored.

[0073] (2) Incremental knowledge extraction unit

[0074] The input dynamic data stream is processed as follows: ① Incrementally update the entity dictionary based on the online learning NER model; ② Use context embedding to link existing entities in the knowledge graph in real time; ③ Match dynamic events based on predefined templates; ④ Use online learning to update the relation extraction model and perform incremental training; ⑤ Combine a timestamp-based graph neural network to capture relation evolution, perform temporal relation modeling, and output the knowledge graph.

[0075] (3) Time-aware modeling unit

[0076] It comprises three parts: time-aware representation, time slicing, and lifecycle management. The input is a knowledge graph. ① Time-aware representation: A time dimension is added to the knowledge graph embedding to model the timeliness of entity-relationship relationships; ② Time slicing: The knowledge graph is divided into slices according to time windows, supporting historical state queries; ③ Full lifecycle management: Expired knowledge is discarded based on time decay functions or event triggers. Simultaneously, a graph database is used to record the evolution process of the time-series knowledge graph.

[0077] A knowledge graph is an entity-relationship model. The output of an incremental knowledge extraction unit is a knowledge graph. The first step is to realize the knowledge graph, and subsequent incremental expansion and temporal evolution follow.

[0078] 2.2 Decision Tree Generation Module

[0079] Decision trees are generated based on the temporal knowledge graph output by the dynamic knowledge graph module. The knowledge graph model generated by the aforementioned module has high complexity, leading to reduced generalization ability and weakened reasoning capabilities. After a round of feature selection, node partitioning, and computational optimization, a more applicable decision tree can be generated.

[0080] The decision tree generation module recursively selects the optimal features to partition the data, and combines pruning strategies to balance model complexity and generalization ability. It consists of a feature selection unit, a node partitioning unit, and a computation optimization unit.

[0081] (1) Feature selection unit

[0082] Feature selection is performed using squared error, selecting feature A and segmentation point s, with the objective function being:

[0083]

[0084] Where D1 and D2 are the partitioned subsets, and c1 and c2 are the mean of the subsets.

[0085] (2) Node partitioning unit

[0086] For discrete features, child nodes are directly partitioned based on feature values; for continuous features, a binary search method is used to find the optimal split point. A substitution splitting strategy is used to handle missing values, pre-selecting substitution splitting features for each node, and using these features for partitioning when the main feature is missing. A cost-complexity pruning method is employed, balancing the tree complexity and error through parameter α. The loss function is:

[0087] C a (T)=C(T)+α|T|

[0088] Where C(T) is the error and |T| is the number of leaf nodes.

[0089] (3) Calculation optimization unit

[0090] By sorting continuous feature values ​​and quickly traversing possible split points, pre-sorting is achieved. Parallel methods are used to distribute the calculation of information gain or Gini index of each feature to accelerate feature selection.

[0091] 3. Optimization Feedback Layer (OFL)

[0092] The optimized feedback layer consists of two parts: a cross-level gradient propagation module and a dual-mode feedback module.

[0093] 3.1 Cross-level gradient propagation module

[0094] The cross-level gradient propagation module is mainly used to solve problems such as uneven parameter updates between different levels. This module mainly consists of a gradient routing unit, a gradient correction unit, and an adaptive optimization unit.

[0095] (1) Gradient routing unit

[0096] Gradient flows at each level are dynamically allocated through gating weights; attention weights are generated based on the decision tree feature data using an attention mechanism to adjust the gradient contributions at different levels, thereby obtaining gradient weights at different levels.

[0097] (2) Gradient correction unit

[0098] Based on gradient pruning, the maximum norm of cross-level gradients is limited to prevent gradient explosion; gradient normalization is used to standardize cross-level gradients and balance the update magnitude of each level; second-order optimization is used to adjust the update direction using the Hessian matrix.

[0099] (3) Adaptive optimization unit

[0100] A hierarchical learning rate strategy is adopted, which assigns independent learning rates to different layers. At the same time, for the Transformer model, the parameter matrix is ​​decomposed and optimized in blocks to reduce cross-layer gradient coupling.

[0101] 3.2 Dual-mode feedback module

[0102] The dual-mode feedback module aims to enhance the dynamic adaptability of the model through two complementary feedback mechanisms, and mainly consists of an explicit feedback unit, an implicit feedback unit, and a dual-mode fusion unit.

[0103] (1) Explicit feedback unit

[0104] ① Perform gradient reweighting on the input gradient signal (the gradient signal is the gradient of each layer during optimization in the neural network; if not mentioned in the previous steps, the input of this module is the decision tree features output by the aforementioned modules and the neural network model currently used in the agent). This is done using a learnable weight matrix W. e The gradient is recalibrated in terms of spatial and channel dimensions, as shown in the following formula:

[0105]

[0106] in, For the upstream gradient, F in G represents the current input layer features. e This represents the calibrated gradient.

[0107] ② The calibrated gradient is injected into the feature using a gating mechanism for feature correction, as shown in the following formula:

[0108] Fexplicit =F in +α·Conv(G e )

[0109] Where α is an adaptive scalar parameter, dynamically generated through a lightweight MLP, and F explicit This represents the explicit features after feature correction.

[0110] (2) Implicit Feedback Unit

[0111] ①Based on the current layer feature F in For cross-layer features (features present in many layers of the network), calculate the long-range dependencies within the features and perform self-attention aggregation, as shown in the formula:

[0112]

[0113] Where Q, K, and V are derived from F in The linear transformation yields d, where d is the feature dimension and A represents the dependency matrix within the feature.

[0114] ② A dynamic convolution kernel is used to capture local patterns and perform dynamic convolution compensation, as shown in the following formula:

[0115] F implicit =F in +Conv(W d A)

[0116] Among them, W d It is a dynamic convolution kernel, and its parameters are generated from feature statistics, F implicit This represents the implicit features after dynamic convolution compensation.

[0117] (3) Dual-mode fusion unit

[0118] ① The contributions of explicit and implicit feedback are balanced by learnable parameters, as shown in the following formula:

[0119] β=Sigmoid(MLP([F explicit ,F implicit ]))

[0120] F out =β·F explicit +(1-β)·F implicit

[0121] β represents the learnable parameter, sigmoid function, MLP represents the multilayer perceptron model, and F... out The characteristics that represent the output.

[0122] ② By using residual connections, we ensure that the final output retains the original input information:

[0123] Ffinal =F out +F in

[0124] Applications of this invention:

[0125] Scenario 1: Industrial Quality Inspection System

[0126] Hardware configuration:

[0127] Image acquisition: 10x12MP high-speed industrial camera

[0128] Processing Units: NVIDIA Jetson AGX Orin × 2

[0129] Workflow:

[0130] The input layer fuses multi-view images (confidence-weighted);

[0131] The decision-making level compared the 3D product model with the quality standard map;

[0132] The feedback layer adjusts camera parameters and detection thresholds in real time.

[0133] Scenario 2: Application of Intelligent Customer Service System

[0134] System Configuration

[0135] Hardware environment: GPU cluster (4×A100)

[0136] Software framework: PyTorch 2.0 + LangChain

[0137] Workflow

[0138] Input layer processing: Performs multimodal input normalization processing, including text cleaning, speech conversion, and visual analysis;

[0139] Decision-making process: Generating high-quality entity-relationship graph features;

[0140] Optimize feedback implementation: Monitor the perplexity metric in real time, and trigger a deep optimization process when ppl>50 is detected.

Claims

1. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism, characterized in that, The system includes a normalized input processing layer for fusing features from input multimodal data and quantifying the uncertainty of different modal data to obtain normalized multimodal feature data. The normalized multimodal feature data input is used to construct a temporal knowledge graph from the normalized multimodal feature data to obtain decision tree feature data and generate a decision tree, which is the core decision layer for decision reasoning. The decision tree feature data input undergoes cross-level propagation joint optimization processing to obtain an optimization feedback layer for a decision reasoning model with balanced parameters and strong dynamic adaptability.

2. The large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 1, characterized in that, The normalized input processing layer includes a multimodal feature fusion module and an uncertainty quantization module. The multimodal feature fusion module takes multimodal data as input and passes it sequentially through an attention weight dynamic allocation unit, a tensor fusion unit, a gating unit, and a contrastive learning unit to obtain multimodal feature data. The uncertainty quantization module, based on the multimodal feature data processed by the multimodal feature fusion module, uses a Bayesian method to model the uncertainty of the multimodal feature data through probability distribution to obtain normalized multimodal feature data.

3. The large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 2, characterized in that, The multimodal feature fusion module aligns the temporal or spatial order of different modalities through dynamic time warping and cross-modal attention; it achieves feature extraction and representation, cross-modal alignment and interaction through specific encoders for different modalities; it completes missing modalities through generative adversarial networks; and it improves computational efficiency by integrating a lightweight attention mechanism with PCA dimensionality reduction technology.

4. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 2, characterized in that, The method of using Bayesian approaches to model the uncertainty of multimodal feature data through probability distributions specifically involves: 1) Construct a Bayesian neural network, assuming that the weights are probability distributions, and approximate the posterior through variational inference or Markov chain Monte Carlo. 2) The Monte Carlo Dropout method is used to calculate uncertainty. During inference, Dropout is kept active, and the prediction results are sampled multiple times to calculate the variance. The calculation formula is as follows: Where T is the number of samples, θ t Let be the model parameters for the t-th sampling.

5. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 1, characterized in that, The core decision layer includes a dynamic knowledge graph construction module and a decision tree generation module. The dynamic knowledge graph construction module takes normalized multimodal feature data as input and passes it through a streaming data processing unit, an incremental knowledge extraction unit, and a time-aware modeling unit in sequence to obtain a temporal knowledge graph. The decision tree generation module takes the temporal knowledge graph as input and passes it through a feature selection unit, a node partitioning unit, and a computational optimization unit in sequence to generate a decision tree.

6. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 5, characterized in that, The streaming data processing unit is used to store dynamic data streams and uses Apache Flink for real-time knowledge extraction. A sliding window is used to continuously update recent knowledge snippets; a conversation window is used to handle bursts of data. The incremental knowledge extraction unit is input into the dynamic data stream and incrementally updates the entity dictionary based on the online learning NER model. Utilize contextual embedding to link to existing entities in the knowledge graph in real time; Dynamic events are matched based on predefined templates; The relation extraction model is updated using online learning for incremental training. By combining timestamps with graph neural networks to capture relationship evolution, perform temporal relationship modeling, and output a knowledge graph; The time-aware modeling unit takes the knowledge graph as input, adds a time dimension to the knowledge graph embedding, models the timeliness of the knowledge graph, segments the knowledge graph by time window, supports historical state query, eliminates expired knowledge based on time decay function or event triggering, and uses a graph database to record the evolution process of the knowledge graph, outputs a time-series knowledge graph, and stores it in the graph database.

7. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 5, characterized in that, The feature selection unit uses squared error to select features A and segmentation point s. The objective function is: Where D1 and D2 are the partitioned subsets, and c1 and c2 are the mean of the subsets; For discrete features, the node partitioning unit directly partitions child nodes according to feature values; for continuous features, a binary search method is used to find the optimal split point; a substitution splitting strategy is used to handle missing values, pre-selecting substitution splitting features for each node, and using the substitution feature for partitioning when the main feature is missing; a cost complexity pruning method is adopted, balancing tree complexity and error through parameter α, with the loss function being: C a (T)=C(T)+α|T| Where C(T) is the error and |T| is the number of leaf nodes; The computational optimization unit sorts continuous feature values ​​and quickly traverses possible split points to achieve pre-sorting; it adopts a parallelization method to distribute the computation of information gain or Gini index of each feature to accelerate feature selection.

8. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 1, characterized in that, The optimization feedback layer includes a cross-level gradient propagation module and a dual-mode feedback module. The cross-level gradient propagation module is used to achieve parameter update balance between different levels, and the dual-mode feedback module enhances the dynamic adaptability of the model through two complementary feedback mechanisms.

9. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 8, characterized in that, The cross-level gradient propagation module includes a gradient routing unit, a gradient correction unit, and an adaptive optimization unit; The gradient routing unit dynamically allocates gradient traffic at different levels through gating weights; Using an attention mechanism, attention weights are generated based on the decision tree feature data, and gradient contributions at different levels are adjusted to obtain gradient weights at different levels. The gradient correction unit is based on gradient clipping technology to limit the maximum norm of gradients across layers and prevent gradient explosion. Gradient normalization is used to standardize the gradient across layers and balance the update magnitude of each layer; second-order optimization is used to adjust the update direction using the Hessian matrix. The adaptive optimization unit adopts a hierarchical learning rate strategy, assigning independent learning rates to different levels; at the same time, for the Transformer model, the parameter matrix is ​​decomposed and optimized in blocks to reduce cross-layer gradient coupling.

10. A large-scale intelligent agent system based on a multi-layer iterative optimization feedback mechanism according to claim 8, characterized in that, The dual-mode feedback module includes an explicit feedback unit, an implicit feedback unit, and a dual-mode fusion unit; The explicit feedback unit performs gradient reweighting on the gradient using a learnable weight matrix W. e The gradient is then recalibrated in terms of both spatial and channel dimensions to obtain the calibrated gradient, as shown in the following formula: in, For the upstream gradient, F in G represents the current input layer features. e This represents the calibrated gradient; The calibrated gradient is injected into the feature using a gating mechanism for feature correction, as shown in the following formula: F explicit =F in +α·Conv(G e ) Where α is an adaptive scalar parameter, dynamically generated through a lightweight MLP, and F explicit Represents the explicit features after feature correction; The implicit feedback unit is based on the current layer feature F in For cross-layer features, calculate the long-range dependencies within the features and perform self-attention aggregation, as shown in the formula: Where Q, K, and V are derived from F in The linear transformation yields d, where d is the feature dimension and A represents the dependency matrix within the feature. A dynamic convolution kernel is used to capture local patterns, and dynamic convolution compensation is applied to the input features. The formula is as follows: F implicit =F in +Conv(W d ,A) Among them, W d It is a dynamic convolution kernel, and its parameters are generated from feature statistics, F implicit Represents the implicit features after dynamic convolution compensation; The dual-mode fusion unit balances the contributions of explicit and implicit features using learnable parameters, as shown in the following formula: β=Sigmoid(MLP([F explicit ,F implicit ])) F out =β·F explicit +(1-β)·F implicit Where β represents the learnable parameter, sigmoid function, MLP represents the multilayer perceptron model, and F out Characteristics representing the output; By using residual connections, we ensure that the final output features retain the original input information. F final =F out +F in

Citation Information

Cited By

  • Ternary collaborative multi-agent framework and environmental assessment report intelligent auditing method

    CN121959428A

  • Hybrid large model cascade intelligent decision-making method

    CN122021910A