Multi-modal agent RAG-ReAct double-engine cooperative training method
Through the multimodal data high-dimensional vectorization and dynamic distillation technology, the knowledge graph is constructed, combined with distributed architecture and recursive reflection mechanism, the semantic fusion and resource utilization problems of multimodal agents in complex environments is solved, efficient data access and model training closed loops are realized, and inference accuracy and stability are improved.
Patent Information
- Application Number
- CN202510749085.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The traditional multimodal agent RAG-ReAct dual-engine collaborative training method has problems such as insufficient semantic fusion capability, low data access efficiency, poor model robustness and low resource utilization in multimodal data processing, making it difficult to adapt to dynamic and complex environments.
The knowledge graph is constructed using multimodal data high-dimensional vectorization and dynamic distillation technology, a distributed architecture is designed for mixed search and reasoning, a recursive reflection mechanism is introduced for parallel data detection, and load balancing is achieved through node resource scheduling, and a three-level reflection system and cognitive distillation technology are combined to optimize the model training process.
The reasoning accuracy and resource utilization of multimodal agents in complex environments are improved, the adaptability and stability of the model are enhanced, the closed-loop path from semantic modeling to model injection is realized, and the task execution efficiency is improved.
Smart Images

Figure CN120256971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a multi-modal agent RAG-ReAct dual-engine collaborative training method. Background Art
[0002] In the traditional multi-modal agent RAG-ReAct dual-engine collaborative training, during the process of multi-modal data vectorization and knowledge graph construction, it often relies on fixed templates or static feature extraction mechanisms, making it difficult to adapt to dynamic and semantically complex multi-modal environments, resulting in insufficient semantic fusion capabilities; in the hybrid retrieval stage, a static vector index structure is often used, lacking the ability to perceive dynamic changes in data distribution, making it difficult to achieve efficient data access and retrieval strategy updates, affecting the accuracy and efficiency of downstream inference tasks; the recursive reflection mechanism lacks fine-grained perception and correction strategies for abnormal inference paths or conflicting information in traditional implementations, resulting in the reflection process being difficult to effectively feedback to the training framework, affecting the robustness and generalization ability of the model; in the data parallel detection and node resource scheduling link, the task dependency characteristics of the inference path and the differences in node communication overhead are not fully considered, resulting in low resource utilization and scheduling bottlenecks; in the model fine-tuning process, a single loss function and fixed injection strategy are generally used, without distinguishing the different requirements for explicit and implicit parameter injection in different task stages, resulting in the training process being difficult to form a stable and effective knowledge internalization path. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a multi-modal agent RAG-ReAct dual-engine collaborative training method to solve at least one of the above technical problems.
[0004] To achieve the above object, a multi-modal agent RAG-ReAct dual-engine collaborative training method includes the following steps: Step S1: Obtain multi-modal data, and convert the multi-modal data into high-dimensional vectors; construct a knowledge graph based on the high-dimensional vectors; use dynamic distillation technology to encode the semantic relationships of the knowledge graph into the gradient direction of model fine-tuning; Step S2: Design a distributed architecture based on the high-dimensional vectors; perform hybrid retrieval based on the distributed architecture to obtain hybrid retrieval data; perform distributed inference based on a preset recursive reflection mechanism and the hybrid retrieval data to generate distributed inference data; perform data parallel detection based on the distributed inference data to obtain data parallel parameters; Step S3: Perform node resource scheduling based on the data parallel parameters to obtain node load balancing data; construct a three-level reflection system based on the node load balancing data; perform cognitive three-level reflection based on the three-level reflection system to generate three-level reflection data; perform cognitive distillation on the three-level reflection data to obtain cognitive distillation data; Step S4: Visualize the cognitive distillation data to obtain visualized thinking data; train the reasoning engine based on the visualized thinking data; and inject model parameters into the reasoning engine in a directed manner according to the model fine-tuning gradient direction to obtain the intelligent agent reasoning model.
[0005] The present invention introduces high-dimensional vectorization of multimodal data and knowledge graph construction based on dynamic distillation mechanism, so that the intelligent agent has stronger semantic modeling ability and the ability to adapt to dynamic context, thereby effectively overcoming the defect of weak expression ability of traditional methods when dealing with heterogeneous semantic relations. Through the collaborative design of distributed architecture and hybrid retrieval mechanism, the system can dynamically update the index structure and optimize the data access path according to the data distribution, significantly improving the retrieval efficiency and reasoning accuracy. In the reasoning stage, the recursive reflection mechanism and data parallel detection module are introduced to enhance the model's adaptive adjustment ability to complex task dependencies and abnormal reasoning paths. At the same time, the task decomposition and node load balancing are optimized and scheduled through data parallel parameters, effectively solving the problems of low resource utilization and communication bottlenecks in traditional parallel reasoning. The three-level reflection system further enhances the model's self-correction ability for abnormal paths in reasoning, forming a feedback loop at the cognitive level. Through cognitive distillation and thinking display technology, the system not only realizes the explicit output of cognitive layer knowledge, but also transforms deep semantic information into thinking representation with intuitive expression ability, providing transparent and controllable feedback basis for the model training process. Finally, through the targeted injection technology combined with the fine-tuning gradients obtained during the training phase, the differentiated internalization and dynamic adaptation of the inference engine parameters were achieved, opening up a closed-loop path from semantic modeling, distributed reasoning to model injection, greatly enhancing the reasoning stability, semantic generalization ability and task execution efficiency of the intelligent agent in complex, changeable and multimodal environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments thereof made with reference to the following drawings: Figure 1 A schematic diagram of the steps of a multimodal agent RAG-ReAct dual-engine collaborative training method of the present invention; Figure 2 Detailed step flow diagram of step S1 in the present invention; The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0007] The technical method of the present invention patent will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0008] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0009] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0010] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a multi-modal agent RAG-ReAct dual-engine collaborative training method, and the method includes the following steps: Step S1: Obtain multi-modal data and convert the multi-modal data into high-dimensional vectors; construct a knowledge graph based on the high-dimensional vectors; use dynamic distillation technology to encode the semantic relationships of the knowledge graph into the model fine-tuning gradient direction; In this embodiment, by setting a multi-modal data acquisition scheme, three types of data, namely images, text, and speech, are obtained. Image data is collected by an industrial camera with a resolution of 4000×3000 and a frame rate of 30fps; text data is exported through a logging system with a unified character encoding of UTF-8; speech data is encoded in PCM with a sampling rate of 44.1kHz and a bit depth of 16bit. The above multi-modal data is respectively vectorized using specific embedding mechanisms: image data is input to the layer before the fully connected layer of the ResNet-152 network, and the output is a 2048-dimensional feature vector; text data is extracted with a 768-dimensional embedding through a BERT-Base model with a fixed vocabulary size of 50,000; speech data uses a CNN network based on log-Mel spectrogram to extract 512-dimensional features. After normalizing the different modal vectors (mean is 0, variance is 1), Canonical Correlation Analysis (CCA) is used to map them into a shared semantic space, and finally a unified 4096-dimensional high-dimensional vector is constructed. Based on this high-dimensional vector, a knowledge graph is constructed by combining rule-driven and clustering: First, the number of clusters is set to K = 100, and KMeans is used to cluster and partition the vector space, and each cluster center represents a concept node; Subsequently, according to the artificial relationship templates set by domain knowledge (such as three edge types of "contains", "causes", "belongs to"), edge relationships are established between semantically similar vector instances to form relationship triples. The constructed knowledge graph is stored and accessed in the RDF (Resource Description Framework) format. In the process of converting semantic relationships into the gradient direction of model fine-tuning, a graph attention mechanism (Graph Attention Network, GAT) based on backpropagation is adopted, and the attention weights between nodes in the graph are used as edge gradient contribution factors, and the weight upper limit of each edge is set to 1.0, and edge connections with weights lower than 0.05 are discarded. By introducing a dynamic distillation mechanism, the gradient direction of the target model is dynamically adjusted during training: The teacher model is set as a structure-preserving graph embedding model (such as DGI), the student model is an inference engine, the KL divergence is used as the loss function, the distillation temperature τ is set to 2.0, and the distillation ratio weight α is set to 0.7, and the target model is dynamically adjusted to update parameters along the semantic gradient direction in adjacent iteration rounds.
[0011] Step S2: Design a distributed architecture based on the high-dimensional vector; perform hybrid retrieval based on the distributed architecture to obtain hybrid retrieval data; perform distributed inference based on a preset recursive reflection mechanism and the hybrid retrieval data to generate distributed inference data; perform data parallelism detection based on the distributed inference data to obtain data parallelism parameters; In this embodiment, for the unified 4096-dimensional high-dimensional vector, a distributed vector storage and computing architecture is designed, and the Faiss distributed index system is used to process the vector data. First, an IVF+PQ index structure is constructed, with the number of clustering centers set to nlist = 1000, the number of quantization segments set to M = 32, and the single-segment bit width set to 8 bits. To support high-concurrency access and data partitioning, the vector data is sharded to 32 nodes, and each node can accommodate up to 1 million vector records at most. The hybrid retrieval module is based on a dual-engine structure and consists of a vector retrieval engine and a rule-based relationship matching engine. For the high-dimensional vector queried by the user, the vector retrieval part performs a top-k = 50 nearest neighbor retrieval using the L2 distance metric. The relationship matching part matches all nodes with "preceding cause and effect" or "superordinate and subordinate" relationships according to the knowledge graph structure corresponding to the query. After fusing the two results, a correlation sorting mechanism is adopted. With the retrieval confidence threshold θ = 0.6 as the standard, only the data with a confidence higher than the threshold is retained as the hybrid retrieval data. When performing distributed inference, it is set that the inference process includes three stages: initial matching, logical synthesis, and rule abstraction. In the initial matching stage, known causal rules are called for a single matching judgment; in the logical synthesis stage, a structured rule template (such as a premise-conclusion template) is used to merge adjacent retrieval items; in the rule abstraction stage, a graph convolutional neural network (GCN) is used to perform inductive propagation in the graph structure to transmit deep semantics. Each stage is computed in parallel on GPU nodes, and finally the merged results are generated as distributed inference data. After the distributed inference data is recorded with timestamps, its logical flow is detected, and content differences are marked, data parallel detection is performed based on the Spark framework. The data conflict detection threshold is set to a conflict frequency > 0.3 / second, and parameter information such as conflict partitions, redundant segments, and synchronization status is output to form the final data parallel parameter set.
[0012] Step S3: Based on the data parallel parameters, perform node resource scheduling to obtain node load balancing data; construct a three-level reflection system based on the node load balancing data; perform cognitive three-level reflection based on the three-level reflection system to generate three-level reflection data; perform cognitive distillation on the three-level reflection data to obtain cognitive distillation data; In this embodiment, data parallel parameters are used to perform load balancing resource scheduling for distributed nodes. The scheduling mechanism is implemented based on the Kubernetes custom scheduler. The scheduler takes the load intensity, data synchronization latency, and historical conflict frequency in the parallel parameters as scheduling metrics and adopts a weighted summation scoring function, where the load intensity weight is 0.5, the latency weight is 0.3, and the conflict frequency weight is 0.2. Nodes with a total score higher than the threshold of 0.75 are preferentially allocated. The scheduling period is set to once every 5 seconds. The scheduling unit takes Pod as the smallest unit, and when the node CPU usage exceeds 85%, the Pod migration mechanism is automatically triggered. After obtaining the node load balancing data, a three-level reflection system is constructed based on its running state and interaction path. This system includes a task behavior layer, a knowledge backtracking layer, and a model structure layer, corresponding to real-time task feedback, historical knowledge path tracking, and model parameter update respectively. Each layer is equipped with an independent reflection trigger mechanism: the task behavior layer is based on the average task execution duration fluctuating by more than 20%; the knowledge backtracking layer is based on the confidence of the previous round of reasoning being lower than 0.4; the model structure layer is based on the loss reduction rate <1e-4 after multiple trainings as the trigger basis. When performing three-level cognitive reflection, each layer of reflection unit will combine the node state and historical reasoning trajectory in the load balancing data to perform time series modeling on the reflection records and use an LSTM network to generate reflection outputs. The three-level reflection data is input into the cognitive distillation module for information compression and knowledge abstraction processing. The distillation process adopts a teacher-student structure, where the teacher model is a structured attention network and the student model is a three-layer perceptron structure. The input is the time series encoded vector of the reflection data, and the output is the knowledge unit of the compressed representation, and the vector length is compressed to 1 / 4 of the original. The distillation loss is composed of the MSE loss plus the entropy regularization term, and the total loss function is set as: L = MSE_loss + λ * Entropy, where λ is the regularization factor and is set to 0.02. The finally generated cognitive distillation data is output in JSON format, recording the knowledge abstraction nodes and structure labels corresponding to each reflection task.
[0013] Step S4: Visualize the cognitive distillation data to obtain visualized thinking data; train the inference engine according to the visualized thinking data; perform directional injection of model parameters into the inference engine according to the model fine-tuning gradient direction to obtain the intelligent agent inference model.
[0014] In this embodiment, the cognitive distillation data is processed for thought visualization. The specific operation is to use the graph visualization engine Graphviz for graph structure rendering, and construct a graph display by combining the hierarchical structure and relationship directions of thought nodes. The nodes adopt shape encoding (circles represent behavioral information, squares represent knowledge concepts, and hexagons represent model structures); the edges use colors to distinguish different relationships (red represents causality, blue represents reference, and green represents parallel relationships); the node labels are determined based on the knowledge summary fields in the cognitive distillation data. Each graph contains no more than 100 nodes, and the graph output format is SVG for easy front-end interaction. The inference engine is trained based on the visualized thought data. The inference engine is based on a logical induction program, and the training data consists of triples (entity 1 - relationship - entity 2) extracted from the visualized graph. A structured rule learner (such as ILP) is used to extract induction rules. The number of training rounds is set to 500, the learning rate is 0.001, and the Adam optimizer is used. During the training process, a consistency check module is used to eliminate rule groups with semantic contradictions (the conflict rate threshold is >5%), and only the valid rule set is retained for training update. Combining the model fine-tuning gradient direction obtained in step S1, a directional injection of model parameters is performed on the inference engine. This operation is achieved by multiplying the fine-tuning gradient direction vector (after unit vectorization) by a specific coefficient γ = 0.8 and then weighting it into the target parameter tensor, realizing a parameter shift in a specific semantic direction. This injection operation only takes effect on the second hidden layer and the output layer in the inference engine, and the injection frequency is executed once every 20 iterations, and the parameter update step size is set to 1e-4. The inference engine obtained in this way is the final trained intelligent agent inference model.
[0015] Preferably, step S1 is specifically as follows: Step S11: Obtain multimodal data, and extract text data and image data; In this embodiment, data from multi-source heterogeneous data systems (such as medical imaging devices, industrial monitoring systems, text entry systems, etc.) is collected and sorted out uniformly. The image data format is limited to JPEG or PNG, the resolution is not less than 512×512 pixels, and the source of the image data needs to be accompanied by metadata such as shooting time, device number, and acquisition conditions. The text data is read in UTF-8 encoding format, the maximum single-line length shall not exceed 4096 characters, and the text data needs to be segmented by full stops, semicolons, or line breaks. Regular expressions (such as \w+ to match word tokens) are used to perform preliminary cleaning on the text to remove special symbols, HTML tags, and non-standard character encodings. The image data needs to be size-normalized by bilinear interpolation (the normalized target size is 224×224 pixels), and converted into a three-dimensional matrix of RGB channels, and the values are normalized to the [0,1] interval. All processed data needs to be uniformly stored in a directory structure named after the data hash value for subsequent retrieval and association operations.
[0016] Step S12: Identify key entities based on the text data; extract entity relationships according to the key entities; In this embodiment, a named entity recognition tool based on part-of-speech tagging rules is used to perform entity recognition operations on the text data in step S11. The entity recognition adopts the CRF (Conditional Random Field) tagging method and relies on a part-of-speech tagging tool for preprocessing. The BIO coding tagging rule is used (i.e., B - entity, I - entity, O - non - entity). The dictionary for entity recognition covers three categories of terms: medicine, industry, and geography. The dictionary is constructed from public corpora (such as UMLS, GeoNames, and mechanical product vocabulary sets) and low - frequency terms (with a frequency of less than 5 occurrences) are removed. After entity recognition, the subject - verb - object relationship is extracted based on dependency syntactic analysis (using a grammar analysis tool such as the dependency parser of spaCy) as the way of entity relationship extraction. The triple format is defined as <entity 1, relationship, entity 2>, where the relationship words are limited to 30 categories, including system - defined keywords such as "belong to", "be located in", "control", "detect".
[0017] Step S13: Construct a semantic graph according to the key entities and entity relationships; In this embodiment, the process of constructing the semantic graph is based on the graph database architecture, and Neo4j is used for graph data storage. Each key entity is defined as a node, and its node label type is determined by the category in the entity dictionary (such as "device", "symptom", "location"). Each entity relationship is defined as a directed edge, and the edge type is the relationship vocabulary extracted in step S12. All graph nodes are added with attribute fields, including: entity name, original text position index, entity category, confidence score (based on the CRF prediction probability, reserved to three decimal places). During the graph structure construction process, if a homonymous entity is found, its context similarity (Jaccard similarity, with a threshold set to 0.75) is calculated. If the similarity is higher than the threshold, they are merged into the same node, and the merging rule is mainly to retain the entity with a higher confidence. After the semantic graph is completed, it is exported in JSON format through the Cypher query language for subsequent processing.
[0018] Step S14: Identify image targets based on the image data, map the image targets to the semantic graph, and convert the semantic graph into a high - dimensional vector using a preset vector database; In this embodiment, image target recognition adopts the target recognition method based on YOLOv5, and the input of the model is the normalized image data processed in step S11. The recognition categories are limited to 20 categories, such as "heart", "brain tissue", "transmission line", "pipe leakage", etc. The training set uses the COCO format and is pre-trained with mAP>0.85 as the evaluation standard. The target detection results include the coordinates of the bounding box (BoundingBox), class label, and confidence. The coordinates of the bounding box are normalized to [0,1]. The recognized image targets are mapped to the semantic graph. By calculating the edit distance (Levenshtein distance) between the class label of the image target and the node label of the semantic graph, the maximum edit distance is set to 2. If the mapping condition is met, an "image correspondence" edge relationship from the image target node to the semantic graph node is created. The process of converting the semantic graph into a high-dimensional vector adopts graph embedding technology, using the Node2Vec algorithm, with the dimension set to 128, walk length = 40, number of walks = 10, context size = 5, and window size = 10. The generated embedding vectors are stored in a vector database (such as FAISS) for subsequent knowledge graph construction.
[0019] Step S15: Construct a knowledge graph based on the high-dimensional vectors; In this embodiment, the knowledge graph construction is based on the embedded semantic graph vectors in step S14. The graph clustering method (such as K-Means clustering, with the number of clustering centers k = 50) is used to divide the embedded vectors into semantic topics. Each cluster represents a semantic subgraph, and the edge weights between the nodes within the subgraph are calculated by cosine similarity. The edges with weights less than 0.3 will be removed. The cross-subgraph connection relationships between the subgraphs are established through co-occurring entities, and the connection threshold is set to the co-occurrence times ≥ 3. Finally, all the subgraphs are merged to form a complete knowledge graph with a semantic clustering structure. The knowledge graph is exported in the RDF triple format and uses a ttl file for semantic description. The constructed ontology relationships include 15 types of system-defined attributes such as "is-a", "part-of", "caused-by", etc.
[0020] Step S16: Use dynamic distillation technology to encode the semantic relationships of the knowledge graph into the gradient direction for model fine-tuning.
[0021] In this embodiment, the dynamic distillation technology is adopted to convert the semantic relationships in the knowledge graph into the gradient direction for model fine-tuning. The specific process is as follows: First, each relationship path in the knowledge graph is encoded, and the path length is controlled within 3. The Positional Path Embedding method is used to embed each path into a vector. The path embedding dimension is 128, and the attention mechanism (based on multi-head attention, head = 8) is used to extract the relationship weights. The path embedding is aligned with the current model gradient space, and the correlation between the path embedding and the model gradient vector is calculated by the Mutual Information Maximization method. In dynamic distillation, in each backpropagation operation, the graph embedding is added to guide the gradient change direction, and its control weight λ is set to 0.1 and added to the loss function as the graph guidance term. The gradient direction adjustment method is to project the gradient vector in the direction of the graph embedding and perform weighted summation to generate the fine-tuning direction vector, which is finally used for subsequent model updates.
[0022] Preferably, step S2 is specifically as follows: Step S21: Perform vector partitioning based on the high-dimensional vectors to obtain the partition vector space; In this embodiment, the high-dimensional vectors are derived from the knowledge graph embedding vectors constructed in step S15, and the vector dimension is fixed at 128. First, all vectors are normalized by the Euclidean norm to ensure that the vectors are distributed on the unit sphere. The vector partitioning adopts a clustering-based partitioning strategy, using the K-Means clustering algorithm, where the number of cluster centers k depends on the overall vector scale, and the specific value rule is k = ⌊√n⌋, where n is the total number of vectors. For example, if the total number of vectors is 100,000, then k is set to 316. The K-Means++ strategy is used to initialize the cluster centers, the maximum number of iterations is set to 300, and the convergence threshold is set to 1e-4. After clustering, each cluster is defined as a subspace, and all vectors belonging to the cluster form the partition vector space. To avoid the problem of blurred boundaries, the Soft Assignment method is used to calculate the distance from each vector to all cluster centers and save the first 3 nearest clusters as the soft membership labels. Repeated membership is allowed when the distance threshold is within 0.2. Finally, each partition space is numbered, and its center vector, member index, and distance matrix are recorded for constructing the subsequent index structure.
[0023] Step S22: Design a distributed vector index structure according to the partition vector space; design a distributed architecture based on the distributed vector index structure; In this embodiment, based on the partition vector space obtained in step S21, a distributed vector index structure is designed, adopting a hybrid method that combines inverted index and quantization coding. An independent index node is established for each partition vector space. Inside the index node, an IVF (Inverted File Index Structure) is used, and each IVF sub-structure is compressed and stored using PQ (Product Quantization). The PQ coding dimension is divided into 8 groups, each group has a dimension of 16, the quantization codebook size is 256, and the quantization error is controlled within 1e-3. Each partition index node is registered with the master coordinator through the gRPC protocol. The master coordinator maintains a mapping table of all partition numbers and their corresponding index nodes, and records the number of CPU cores, memory capacity, and network bandwidth status of each node. The distributed architecture design adopts the master-slave mode (Master-Worker architecture). The master coordinator schedules all query distributions and node load balancing. The load strategy combines the round-robin method with real-time node load monitoring (queue length, response latency). When the average response time of a certain node exceeds 200 ms or the CPU utilization rate continuously exceeds 85%, the task allocation ratio of this node will be automatically adjusted to not exceed 5% of the total. The entire distributed system is deployed in a Docker container cluster, and Kubernetes is used to manage the scaling of container nodes. The communication bandwidth requirement between nodes is not less than 1 Gbps, and the latency is controlled within 10 ms.
[0024] Step S23: Perform hybrid retrieval based on the distributed architecture, including vector retrieval, full-text retrieval, and re-ranking retrieval, to obtain hybrid retrieval data; In this embodiment, the query input by the user is first encoded into a 128-dimensional vector in the hybrid retrieval process. This vector is sent to the index nodes in the corresponding partition through the master coordinator in step S22, triggering the vector retrieval process. The HNSW algorithm of FAISS is used for vector retrieval, with the adjacency graph parameter M set to 32, efConstruction set to 200 during construction, and efSearch set to 100 during retrieval. In parallel, the query is simultaneously parsed into keyword phrases and enters the full-text retrieval module. The full-text retrieval module is based on Elasticsearch, and n-gram tokenization (n = 2 to 4) is used for keyword tokenization. In the index structure, each document corresponds to a combined mapping of vector embeddings and full-text fields. The retrieval results are sorted according to the BM25 score. The top 50 full-text retrieval results are cross-compared with the top 50 vector retrieval results. The Jaccard similarity is used for cross-comparison, and the similarity threshold is set to 0.3. Finally, the intersection part is preferentially retained, and the non-intersection part is re-sorted by the dual-channel scoring mechanism. The scoring function is: Score = 0.6 × vector similarity + 0.4 × BM25 score. In this scoring mechanism, all scores are normalized to [0, 1], and the output of the hybrid retrieval data does not exceed 100 at most, and is arranged in descending order of scores.
[0025] Step S24: Based on the preset recursive reflection mechanism and the hybrid retrieval data, perform distributed reasoning to generate distributed reasoning data; In this embodiment, the recursive reflection mechanism realizes distributed reasoning by combining graph backtracking and link analysis. For each record in the hybrid retrieval data in step S23, extract its corresponding semantic graph node and obtain its upstream and downstream first-order connection paths. The path mining depth is set to a maximum of 3 hops, and the paths are sorted by path support. The support is defined as the product of the weights of each edge in the path, and the edge weights are derived from the edge confidence in the knowledge graph (calculated from the vector similarity in Node2Vec, and the weight range is [0, 1]). Select the top 10 high-support paths as candidate reasoning paths, and use heuristic search (such as the A* algorithm) to evaluate the relevance of the path nodes. The heuristic function is the cosine similarity between the target node and the end node of the path. The multi-node reasoning process is independently and parallelly completed on each index node, and the results are summarized by the master coordinator. All reasoning results are then subjected to a consistency check, and a triple confidence reconstruction strategy is adopted. If there are conflicts in the relationships within the path (such as the co-occurrence of "contains" and "excludes"), the conflicting paths are excluded. Finally, the top 20 conflict-free and high-confidence reasoning paths are retained and output as distributed reasoning data in the format of path sequences and their support scores.
[0026] Step S25: Perform data parallel detection based on the distributed reasoning data to obtain data parallel parameters.
[0027] In this embodiment, data parallel detection is based on the distributed inference data in step S24, and the purpose is to extract the model gradient fluctuation caused by each path inference. The knowledge unit embedding vectors corresponding to all inference paths are input into the target fine-tuning model, and the influence vector on the model output is obtained through forward propagation. The L2 norm operation is performed on each influence vector, and the result is used to judge the perturbation intensity of its direction of model update. All path embeddings are input into the tensor computing engine (such as TensorRT or ONNX Runtime) within the distributed node group, and each node uses 32 paths per batch as the processing unit. When the path perturbation intensity is higher than the set threshold (the threshold is set to 0.25), this path is marked as a gradient-sensitive path. The parallel detection output includes the number of each path, the perturbation intensity value, the node calculation time, the peak memory usage, and the influence probability of the final classification label. A data parallel parameter configuration table is established with paths as the unit, including: path number, recommended node number (sorted by the shortest calculation time), parallel batch number, memory occupancy per batch (in MB), inference trigger frequency (in times / minute). This parameter table is used for the scheduling configuration in the subsequent fine-tuning training process.
[0028] Preferably, step S24 is specifically as follows: Step S241: Perform result inconsistency detection on the mixed retrieval data according to the preset retrieval target information to obtain retrieval result abnormal data; In this embodiment, the mixed retrieval data is output by step S23, including the retrieval result entries, their vector similarity scores, and the final fusion score (range 0-1). The preset retrieval target information is defined by the task context, including the target entity category (such as "device type = infrared detector"), keyword constraints (such as "must contain 'active sensing'"), and context consistency templates (such as "if a device is involved, its physical property description must be involved"). Perform structured analysis on the mixed retrieval data, parse each retrieval result into a structure, and the fields include entity labels, main description fields, semantic vector labels, etc. Using Boolean logic rules, perform a matching operation on each result with the target information. For example, use regular expressions to detect whether the constraint keyword is included in the description field, and mark it as inconsistent if it is missing; then compare whether the entity category matches the target type, and mark it as inconsistent if it does not match. All data marked as "inconsistent" are summarized into a retrieval result abnormal data set, and each record is attached with an inconsistent type label (such as "keyword missing", "entity category conflict", "semantic mismatch"). This process does not rely on manual annotation and is completely based on structural rules. The maximum allowable abnormal ratio is set to 10%. If it exceeds this ratio, the overall retrieval task will be judged as abnormal.
[0029] Step S242: Dynamically correct the prompt words based on the retrieval result abnormal data to obtain prompt word correction data; In this embodiment, after obtaining the abnormal data of the search results, the original prompt words need to be dynamically corrected. The original prompt words are stored in the search initialization configuration and are natural language description sentences, such as: "Search for infrared monitoring devices with active perception capabilities". The prompt word correction operation is decomposed according to the "inconsistent type" label in the abnormal data. Taking "keyword missing" as an example, the keywords that frequently appear in the abnormal data but do not appear in the original prompt words are extracted, and the frequency statistics use the TF-IDF method, where TF calculation is based on the semantic text field of the abnormal data, and IDF is calculated with reference to the existing knowledge text in the training set. Only keywords with a threshold of TF-IDF ≥ 0.05 can enter the completion pool. Taking "entity category conflict" as an example, the correction strategy adopts the entity synonym classification standard library, maps the retrieved entities such as "thermal imager" to "infrared detector", and adds synonymous expansion words to the end of the prompt word. The corrected prompt words are recombined into a structured prompt word template, and the template structure includes: "task keyword + entity category phrase + function description phrase", and each part is connected by a connector "-". The final output prompt word correction data is recorded in JSON format, including the text before correction, the text after correction, the keyword weight table used, and the reason for correction.
[0030] Step S243: triggering a reflection condition according to abnormal data in the search result to obtain reflection condition data; In this embodiment, the abnormal data of the search results is analyzed to determine whether the reflection condition is met. The reflection condition can be set to trigger one of the three situations: 1) the proportion of the number of abnormal data entries in the total data of the mixed search is greater than the threshold (the threshold is set to 20%); 2) the proportion of "entity category conflict" in the abnormal type exceeds 50%; 3) the Jaccard similarity between the revised prompt word and the original prompt word is less than 0.5. The above three indicators are respectively implemented by the exact match counter, the classification proportion function, and the character-level Jaccard similarity calculation. Taking Jaccard similarity as an example, the prompt word is divided into keyword sets A and B according to the space, Jaccard (A, B) = | A ∩ B | / | A ∪ B |, and the similarity is lower than 0.5, which means that the semantic transfer is large and needs to be reflected. If any condition is met, the corresponding trigger type (such as "high abnormal proportion", "serious entity error", "dramatic prompt change") is recorded in the reflection condition data, and the corresponding indicator value and detection timestamp are attached. The reflection condition data is passed back as a standard structure.
[0031] Step S244: activating a preset recursive reflection mechanism based on the reflection condition data to obtain recursive reflection mechanism activation data; In this embodiment, after receiving the reflection condition data generated in step S243, the recursive reflection mechanism is activated according to the condition type. The startup process of the reflection mechanism is fixed and includes three parts: path backtracking, intention reconstruction, and semantic transformation enhancement. The path backtracking operation calls the reverse edge information of the knowledge graph relationship, starts from the entity node corresponding to the abnormal result, traces its nearest three upstream nodes, and constructs a set of reverse triples. The intention reconstruction operation is based on the difference between the original prompt and the corrected word, uses a syntactic analyzer (such as SpaCy) to extract the subject-predicate-object structure and modifying components, constructs an intention tree and performs logical correction, and retains the reconstructed prompt with consistent subject and coherent predicate logic. The semantic enhancement expands the prompt by setting a language template. For example, the original prompt "infrared monitoring" is expanded to "infrared monitoring device capable of object recognition in the dark", and this template is completed by the selection rules of the domain knowledge sentence library. Finally, all reflection operations form startup data, including a set of reverse triples, a prompt word intention tree, an enhanced text set, etc., and record the reflection number, activation time, and target inference round number.
[0032] Step S245: Re-inference execution is performed according to the recursive reflection mechanism startup data and the prompt word correction data to generate distributed inference data.
[0033] In this embodiment, based on the recursive reflection mechanism startup data generated in step S244 and the prompt word correction data in step S242, re-inference execution is performed. The re-inference does not reuse the previous path, but uses the newly corrected prompt word as the seed vector, generates a 128-dimensional prompt vector through the BERT embedding model, and inputs it into the initial vector channel of the distributed inference engine. In the path search stage, the Beam Search algorithm is used, with the width set to 5 and the maximum depth to 3 hops. Different from the previous time, the reverse triple data is added as the initial edge set for path search, and the priority is set to be 10% better than the general knowledge edge. In each hop, candidate paths with a similarity greater than 0.4 are retained, and the path similarity is defined as the cosine similarity between the current node embedding and the target entity embedding. During the inference process, the intention reconstruction tree provides a semantic alignment template, and all paths need to conform to the "target function + entity attribute" structure to be retained. Finally, after deduplication of the inference results, they are uniformly converted into a structured triple set and its score (weighted calculation of path support and semantic matching, with weights of 0.7 and 0.3 respectively), generating the final distributed inference data, and tagging it with "reflection execution" for use in subsequent steps.
[0034] Preferably, step S25 is specifically: Step S251: Extract the distributed inference path according to the distributed inference data; In this embodiment, the distributed inference data is generated by step S245, and the data format is a set of structured vectors containing triple paths. Each path consists of several entity nodes and edge types, and the format is [(entity A, relationship R1, entity B), (entity B, relationship R2, entity C),...]. The core operations of inference path extraction are path parsing and redundancy removal. First, read the entire triple set and use the path hashing rule to extract the unique path identifier. The hashing algorithm uses SHA-256, which is input after encoding the nodes and edges in the path in sequence to ensure the consistency of the path structure. Merge the paths with the same hash value, and retain the path with the highest average inference confidence (from the confidence field in the original path, with a numerical range of 0 to 1). Subsequently, perform label normalization on the nodes in the path, and use the preset entity normalization dictionary to unify synonyms (such as "infrared camera" and "infrared monitoring device") into standard terms. After processing, output the structured distributed inference path data set. Each path is represented as a JSON array, including a list of path nodes, a list of edge types, a cumulative confidence field (floating-point value), and a unique path ID.
[0035] Step S252: Identify the data flow dependency relationship based on the distributed inference path; In this embodiment, based on the distributed inference path, the data flow dependency relationship is identified. Each pair of consecutive triple entity nodes in the path constitutes a "dependency pair", that is, the output of the previous node is used as the input of the next node. Taking the triple path [(device identification, output signal), (signal processing, generate judgment), (judgment result, control feedback)] as an example, it is identified that the signal processing depends on the output of the device identification, and the judgment generation depends on the result of the signal processing, forming a dependency chain device identification → signal processing → judgment generation → control feedback. To achieve automatic identification, first construct a path dependency graph, which is represented by a directed graph structure. The nodes are entity actions or subtasks, and the edges are "depends on" relationships. After the graph is constructed, use the depth-first search (DFS) algorithm to find all paths from the starting node to the ending node. Each path is a data flow dependency chain. Each dependency chain is attached with the minimum confidence of the path used to construct it as a conservative estimate. To improve the accuracy of dependency identification, filter out the path segments with a confidence lower than 0.35 to ensure the practical usability of the data flow path. The final output format is: dependency chain number, dependency path structure (in JSON array format), path weight value, start and end identifiers.
[0036] Step S253: Calculate the task granularity factor according to the data flow dependency relationship; In this embodiment, according to the identified data stream dependency relationships, the task granularity factor corresponding to each task path is calculated. The task granularity factor is defined as a measure of the concurrency ability between tasks, and the calculation formula is: F = (1 - D / N) × (ΣT_i / M), where: D is the average dependency degree of the nodes in the path (the average value of the number of incoming edges of each node), N is the total number of nodes in the path, T_i is the processing complexity value of the i-th node (the value range is 1 to 10, preset according to the task type), and M is the maximum complexity value. The assignment of the processing complexity value is based on the task type mapping table: for example, "image parsing" is assigned 9, "semantic fusion" is assigned 7, "entity extraction" is assigned 5, and "simple logical judgment" is assigned 2. The D value is obtained by analyzing the dependency graph structure, and the number of incoming edges is divided by the number of nodes. After all path granularity factors are calculated, they are sorted from high to low according to the F value, which is used to guide the subsequent thread setting. The output content is: path ID, granularity factor value (floating point number), number of task nodes participating in the calculation, list of complexity mapping values of each node, and dependency degree table.
[0037] Step S254: Perform parallel thread setting according to the task granularity factor to obtain parallel thread data; In this embodiment, after receiving the calculation result of step S253, parallel thread setting is performed according to the task granularity factor. The parallel thread setting adopts a proportional mapping strategy, with the maximum upper limit of parallel threads set to 64 and the minimum to 8. The mapping rule is set as follows: if the granularity factor F ≥ 0.75, then 64 threads are allocated; if F ∈ [0.50, 0.75), then the number of threads is set according to a linear ratio: number of threads = round(8 + (F - 0.5) × (64 - 8) / 0.25); when F < 0.50, 8 threads are uniformly set. The thread configuration of each path records the thread value, the allocated target CPU core ID (matched according to the host thread load distribution table, with a maximum of 4 threads bound to each CPU), and the parallel section index (the non-dependent section of each path can be used as a parallel section), and records the configuration timestamp. The final output parallel thread data format is JSON, and each record includes the path ID, number of threads, list of node numbers allocated to each thread, corresponding CPU core ID, and expected concurrent execution time window (in ms).
[0038] Step S255: Perform task scheduling simulation based on the parallel thread data to obtain task scheduling data; In this embodiment, a task scheduling simulation system is constructed based on parallel thread data. The scheduling simulation adopts an event-driven timing simulation method. First, initialize the task scheduling table, and fill the event list according to the task segments assigned to each thread. Each event includes: task ID, start time, estimated execution time, and input data ready time. The estimated execution time of each task segment is provided by the task type mapping table. For example, "image parsing" is set to 180 ms, "semantic fusion" is set to 150 ms, and "entity matching" is set to 90 ms. The input data ready time is calculated based on the dependency path data flow graph, and the completion time of the predecessor task is the input time of the current task. The task simulation scheduling adopts the Time Wheel mechanism, advances in 10-ms wheel time slices, and executes the task scheduling status update. The simulation execution time is set to the total estimated time of the maximum task chain + 20%. During the simulation process, record the actual start time, completion time, waiting time, and the thread ID of each task. Finally, output the task scheduling data table, and each row record includes path ID, thread ID, task number, task type, estimated execution time, actual scheduling time window (start / finish), input waiting duration, and upstream and downstream dependency identifiers.
[0039] Step S256: Perform data parallelism detection based on the task scheduling data to obtain data parallelism parameters.
[0040] In this embodiment, based on the task scheduling data, perform data parallelism detection and extract data parallelism parameters. The detection logic is based on the analysis of the non-dependent overlapping segments between tasks. First, construct a Gantt chart of task scheduling, with the horizontal axis representing time (in ms) and the vertical axis representing thread IDs, and fill in all task scheduling time periods. Traverse all threads and check if there are more than two non-dependent task segments (identified by the dependency graph in step S252) within the same time window. If so, it is identified as a parallel potential segment. For each parallel potential segment, record the parallel level (number of concurrent tasks), average task granularity (from step S253), and thread utilization (= total task segment time / time window length). Set the output structure of the data parallelism parameters as: parallel segment ID, start and end times, list of involved thread IDs, number of parallel tasks, average granularity factor, thread utilization (range 0 to 1), and list of non-dependent task IDs in the parallel segment. If it is detected that the proportion of the total number of parallel segments in the total scheduling time period is less than 25%, record it as a low parallelism mark. All parallel parameters form a complete parallel execution evaluation set for subsequent inference scheduling optimization modules to call.
[0041] Preferably, in step S3, the node resource scheduling is specifically as follows: Perform load imbalance detection based on the data parallelism parameters to obtain load imbalance data; In this embodiment, based on the extraction of data parallel parameters, the execution time, data volume, and peak memory usage corresponding to each thread unit in the data parallel parameters are used as basic metrics, and the Standard Deviation Analysis method is used to calculate the load of each thread unit within the same time period. The load of each thread is defined as: Load Li = Data volume (MB) / Execution time (s). Based on this, the mean μ and standard deviation σ of the loads of all threads are statistically calculated. The standard deviation threshold δ = 1.2 is set, and if the condition |Li - μ| > δ × σ is met, the load of this thread is determined to be abnormal, and the corresponding thread number, its timestamp, and the node number where it is located are recorded to form load imbalance data. The data volume and execution time used are extracted through the training log collection module, with a recording frequency of 5 Hz per second, and the memory peak value is obtained through the system call getrusage(RUSAGE_SELF) in the resource monitoring system of each node.
[0042] Especially importantly, the load imbalance detection includes the following steps: Perform an analysis of the calculation resource allocation ratio based on the data parallel parameters to obtain the node calculation load distribution data; In this embodiment, in the multi-modal intelligent agent RAG-ReAct dual-engine collaborative training method, the training tasks are distributed to multiple computing nodes in a data parallel manner on the heterogeneous cluster computing architecture. To obtain the node calculation load distribution data, first, record the number of training samples actually allocated to each node in each training batch. The node types include but are not limited to GPU nodes, FPGA nodes, and TPU sub-arrays. Secondly, during the execution of each batch of training, use hardware-level event counting tools (such as NVIDIA Nsight Compute and Intel VTune Profiler) to monitor key parameters such as the call frequency of CUDA cores or equivalent processing units, the occupancy rate of streaming multiprocessors, and the utilization rate of tensor cores within each node, and set the sampling period to 200 milliseconds. Then, to quantify the processing ability of the node, pair the execution time window (in milliseconds) of the training batch with the number of samples borne by this node for analysis to form a load index describing the processing intensity of each node. Finally, by traversing all the computing nodes participating in the training, normalize and summarize the sample processing ability of each node to form a two-dimensional table structure composed of the node number and its corresponding load intensity, and output it as the node calculation load distribution data table, which provides a basic basis for subsequent processing rate analysis and abnormal discrimination of computing power scheduling behavior.
[0043] Calculate the task processing rate according to the node computing load distribution data; In this embodiment, the number of tasks processed per unit time of each node is extracted from the obtained node computing load distribution data, that is, the ratio of the number of batches completed by each node in a fixed period (such as every 10 seconds) to the processing period is calculated by division to obtain the task processing rate (unit: batch / s). By aggregating and extracting the execution logs recorded by each node and combining the Job completion timestamps collected in the system-level monitoring framework (such as Prometheus+Grafana), a task processing rate vector set R={r_1,r_2,...,r_n} is established for all computing units, where r_i represents the average task processing rate of the i-th node. The task processing rate also needs to be associated with the actual allocated operator type and operation density, and the Tensor calculation scheduling table is used to clarify the positions with load bottlenecks in the computing graph topology to ensure the unity of the rate measurement dimension.
[0044] Evaluate the computing power call frequency according to the task processing rate; In this embodiment, the task processing rate forms the basis for evaluating the computing power call frequency. Taking NVIDIA GPU as an example, the GPU Utilization rate and GPU Core Frequency (unit: MHz) data provided in the nvml library are read, and the frequency is normalized and mapped with the task processing rate to generate a computing power call frequency index matrix. For each node, a triple record is constructed: (processing rate r_i, call frequency f_i, power consumption P_i). Based on whether the frequency change increases linearly with the rate, the least squares fitting is used for all nodes to obtain the mean residual. If the residual of a certain node exceeds the set threshold (such as 5%), it is initially marked as an abnormal point of the computing power call frequency. All computing nodes are classified into the frequency call statistical table according to the node number, and a "frequency utilization deviation coefficient" is added to the consensus table for subsequent anomaly identification.
[0045] Identify abnormal chip calls based on the computing power call frequency to obtain abnormal chip call data; In this embodiment, the computing power call frequency statistics table is linked and verified with the task scheduling event log. For nodes with a frequency deviation greater than 0.05 (set value), continue to read the temperature rise gradient (unit: °C / s) and power load curve of the chip core. If there are frequent call events without obvious task loads, it is defined as an "abnormal chip call". Combine the event timestamps to screen for high-frequency and low-efficiency call behaviors, and form an "abnormal chip call data table" containing node numbers, time periods, and abnormal event numbers by tracking the GPU's Driver API logs, PCIe event conflict records, and memory-throttle records. The data table records call segments where the instruction call frequency within all chip abnormal cycles is more than 20% lower than the average value or the frequency soars more than 1.1 times the maximum frequency and lasts for more than 2 seconds.
[0046] Based on the abnormal chip call data, perform bus arbitration failure detection to obtain bus arbitration failure data; In this embodiment, select the chip numbers and cycles corresponding to the abnormal chip call data, and read the access records of the PCIe or inter-chip NoC bus. Through the Total Access Queue Length (unit: number of requests) and ArbitrationLatency (unit: ns) parameters, determine whether there are access conflicts within this time period. Set the arbitration delay threshold to 800 ns. If there are more than 3 consecutive timeout behaviors and the arbitration queue depth reaches more than 16 transactions, it is defined as an arbitration failure event. The event data is cross-validated using the internal state dump of the system control unit (such as the arbitration request response record of the AXI arbiter) to form a bus arbitration failure data table, and the fields include arbiter number, bus address, transaction type (read / write), delay value, and transaction blocking duration.
[0047] Based on the bus arbitration failure data, perform load imbalance analysis to obtain load imbalance data.
[0048] In this embodiment, based on the bus arbitration failure data, map it back to the corresponding task scheduling nodes and data parallel mapping table, and combine the node load index and call frequency deviation within the time window to construct an "imbalance index" calculation formula. This index is defined as: the ratio of the request waiting time to the average scheduling time of a certain node in the arbitration failure window. When the imbalance index of a certain node is greater than 2 and the corresponding number of arbitration failures exceeds 5 times, this node is marked as a load imbalance node. Aggregate the computing load, scheduling tasks, power consumption, arbitration delay, etc. data of these nodes, and output a load imbalance data table. The fields in the table include node number, imbalance index value, number of scheduling exceptions, computing power call offset, number of arbitration interruptions, etc. This table serves as the input data source for subsequent dynamic scheduling optimization strategies.
[0049] Identify load abnormal nodes based on the load imbalance data; In this embodiment, according to the data of uneven load, the abnormal frequency of the threads under each node is counted, and the abnormal frequency threshold τ = 3 is set. If the abnormal thread frequency of a certain node is greater than or equal to the threshold τ in the same training round, then this node is marked as a node with abnormal load. The abnormal frequency is statistically clustered by the thread timestamps and node IDs in the recording module, and each training cycle is uniformly 1000 steps (iterations). All nodes with abnormal load form a list of node numbers, and the corresponding details of abnormal threads are attached for the subsequent use of the scheduling module.
[0050] Especially importantly, identifying nodes with abnormal load includes the following steps: Calculate the unit processing time of the node according to the data of uneven load; In this embodiment, when performing the training task of the RAG-ReAct dual-engine model, each computing node deployed in the heterogeneous cluster system (including NVIDIA A100 GPU nodes, Intel Xeon CPU nodes, and Xilinx FPGA acceleration cards) respectively records the number of samples N_i allocated during the training process and the training processing time T_i of this node. All training batches use the same batch size, with 256 samples per batch, and the NCCL communication protocol is used for task synchronization during the training process. The processing time of each batch is recorded regularly (every 200ms cycle) through hardware performance counters such as NVIDIA Nsight Systems and Intel VTune Profiler. The unit processing time is calculated by T_i / N_i, with the unit of ms / sample. All calculation processes are completed on the control host through the NumPy library of Python and are uniformly written into the InfluxDB database for subsequent data processing.
[0051] Statistically calculate the CPU core temperature according to the unit processing time of the node; In this embodiment, a combination of the Linux lm-sensors service and the Intel PowerGadget tool is deployed on each computing node to collect the temperature data corresponding to each physical core of the CPU (Core 0 to Core N, where N varies according to the node) in real time. To ensure measurement consistency, 5 samples are taken per second, and the moving average value is taken. This temperature data is bound to the unit processing time according to the node correspondence and written into the Prometheus time series database with the "node ID + core number + timestamp" as the primary key. The temperature data uses the Celsius scale, and the typical normal temperature range is set to 45°C to 80°C, and a temperature higher than 85°C is determined as an abnormal warning temperature. All temperature collection and data binding operations are implemented through Python and the Prometheus Python Client library.
[0052] Based on the CPU core temperature, thermal imbalance detection is performed to obtain CPU thermal imbalance data; In this embodiment, the thermal imbalance detection adopts a dual judgment mechanism of "local temperature rise gradient + temperature difference between cores". First, within the same physical processor, the instantaneous temperature values of all cores are extracted, and the difference ΔT between the maximum temperature and the minimum temperature is calculated. If ΔT ≥ 15°C (this value is derived from Intel CPU Tjunction test data and the deviation tolerance of the chip's heat conduction layer), it is determined that the chip is in a thermal imbalance state. Second, if the temperature rise rate of a single core exceeds 10°C / s within three consecutive detection cycles (i.e., 600 ms), it is defined as a "thermal slope anomaly point". Both types of anomalies are recorded in the thermal imbalance data table, and the recorded content includes: node ID, core number, ΔT value, average temperature, timestamp, core power consumption (obtained through the RAPL interface), etc. The thermal imbalance detection logic runs on the main control node operating system through an embedded script written in C language.
[0053] Detect the abnormal heat dissipation path based on the CPU thermal imbalance data; In this embodiment, the coordinates of the area where the high-temperature core is located in the thermal imbalance data table are used to perform physical alignment in combination with the CAD drawing of the motherboard heat dissipation layout, and the radiator component numbers and thermal pad layout areas at the corresponding positions are mapped. At the hardware level, a thermal imaging device (such as a FLIR A615 infrared thermal imager) is equipped to collect the thermal field image of the CPU area for 30 seconds, with a resolution of 640×480 and a sampling frequency of 15 frames per second. The OpenCV library is used to enhance the temperature gradient image of the thermal map and identify the center coordinates of the locally overheated area. If the high-temperature area in the thermal field gradient map does not spread along the heat pipe direction, or there are high-temperature points but the temperatures at the fan and radiator outlets do not rise significantly, it is regarded as an abnormal heat dissipation path. All image processing steps are deployed on the NVIDIA Jetson TX2 edge platform for hardware acceleration to ensure real-time performance.
[0054] Perform power supply thermal pad contact aging detection based on the abnormal heat dissipation path to obtain thermal pad contact aging data; In this embodiment, an offline disassembly + online thermal field comparison combination method is required to detect the contact aging degree of the thermal pad in the power supply module area. During the online detection stage, based on the coincidence relationship between the high-temperature points in the aforementioned thermal image and the area covered by the thermal pad, it is judged that there is a potential risk of contact aging of the thermal pad. If no temperature drop trend is observed when the temperature exceeds 90 °C in the contact area between the VRM module (with IR35201 as the main control) and the heat dissipation module, and the fan speed is higher than 2500 RPM (directly read through the PWM signal), it is recorded as a suspected point of thermal pad aging. Further, an X-ray CT scanner is used to analyze the local three-dimensional structure of the main board to evaluate whether the air gap thickness of the thermal pad bonding layer exceeds 0.3 mm (defined as the minimum effective contact value according to the manufacturer's Datasheet). Finally, the thermal pad contact aging data table records the aging level, position number, air gap thickness value, temperature value, and the attenuation ratio of the heat conduction rate (obtained through on-site calibration and actual measurement).
[0055] Identify abnormal load nodes based on the thermal pad contact aging data.
[0056] In this embodiment, combined with the above thermal pad contact aging data, a hardware-level fault marking table is established. If any core on a node has a serious thermal pad aging level (air gap thickness > 0.5 mm, temperature > 95 °C for more than 600 seconds continuously), and the unit processing time of this node is 1.5 times higher than the average value of all nodes (the threshold is set according to the standard deviation range of the node processing time), then this node is marked as an "abnormal load node". A marking field "node_fault_flag = 1" is established in the Prometheus and InfluxDB databases so that the control system can exclude or reassign tasks to other nodes during training scheduling. The identification process is implemented using the Go language on the backend server, and the marking results are regularly written into the metadata interface of the Kubernetes scheduling system to achieve scheduling-level intervention for the training tasks of the RAG-ReAct agent.
[0057] Evaluate the node computing power based on the abnormal load nodes; calculate the node task communication overhead based on the abnormal load nodes; In this embodiment, after identifying the nodes with abnormal loads, the CPU main frequency (GHz), available memory (MB), and GPU floating-point computing power (FLOPS) of each node are extracted as the basic computing power metrics. The CPU main frequency and memory size are statically collected through the lscpu and free - m commands on the nodes, and the GPU computing power is obtained by calculating the performance parameters through the nvidia - smi command in combination with the GPU model. The nvidia - smi --query - gpu=name,clocks.sm,memory.total --format=csv command is used to collect in real - time and convert it to TFLOPS. The three ability metrics are respectively normalized to the interval [0, 1], and the weighted scoring method is used to evaluate the comprehensive ability of the nodes. Let the weight vector be w = [0.3, 0.2, 0.5] (corresponding to CPU, memory, and GPU respectively), then the computing power value of each node is: Cj = 0.3 * norm_CPUj+0.2 * norm_MEMj + 0.5 * norm_GPUj. After all nodes are evaluated, a node ability data table is generated by sorting in descending order of computing power. For each node with abnormal load, the cross - node communication frequency (fij) and communication data volume (dij, unit: MB) during data parallelism are extracted, and the communication overhead is calculated as: CommCostij = fij×dij×λ, where the communication delay coefficient λ depends on the network type: λ = 0.01 for Gigabit Ethernet and λ = 0.002 for InfiniBand network. The communication frequency is obtained by analyzing the task graph scheduling log, and the communication data volume is obtained by aggregating the number of bytes in the send / recv events sent by the monitoring thread. The communication overheads between nodes are summarized to form a communication overhead matrix CommMatrix[N][N], where N is the total number of nodes.
[0058] Node task allocation is performed according to the node computing power to obtain node task allocation data; In this embodiment, based on the node ability data table and the communication overhead matrix in step S264, a weighted task allocation strategy is adopted. First, the task load amount is quantified, the task execution duration Tj and its computing operation amount OPj (unit: GFLOP) are extracted, and the task unit load value is calculated: Wj = OPj / Tj. All tasks are sorted in descending order of Wj, and the tasks are allocated to the nodes with the highest computing power using the ability - first filling method, ensuring that the total task load of each node does not exceed its ability threshold, and the threshold is set as: θj = Cj×β, where β is a regulation factor, set to 0.95. The node task allocation results are recorded in a two - dimensional table, including information such as task ID, corresponding node ID, task load value, required communication path, etc.
[0059] The minimum - cost communication path selection is performed based on the node task communication overhead to obtain the minimum - cost communication path data; In this embodiment, a weighted directed graph is constructed among nodes by using the communication overhead matrix CommMatrix[N][N], and the edge weight is the communication overhead value. The Dijkstra algorithm is used to search for the minimum communication path. For each pair of communication-related node pairs (ni, nj), the minimum communication overhead path is determined, and detailed information such as the relay nodes, relay times, and total data volume included in the path is recorded. If there is a constraint on the communication link bandwidth, a bandwidth availability judgment logic is added to the path selection (for example, the bandwidth utilization rate shall not exceed 90%), and the required bandwidth utilization rate is collected in real time through / proc / net / dev. Finally, a data table of the minimum communication cost path is formed, and each row corresponds to the optimal path and path details of a pair of nodes.
[0060] Based on the data of the minimum path of the data communication cost for node task allocation, node resource scheduling is performed to obtain node load balancing data.
[0061] In this embodiment, after obtaining the node task allocation data and the minimum communication path data, a resource rescheduling strategy is adopted to improve the balance of task distribution. A scheduling optimization problem is constructed by minimizing the global load variance, and the objective function is defined as: Minimize Σ(Cj_actual - Cj_expected)^2, where Cj_actual is the total task load currently assigned to node j, and Cj_expected is the ideal average load = (Σ all tasks Wj) / total number of nodes. Under the constraint that the total communication path cost does not increase, the iterative migration method is used to gradually migrate some tasks of overloaded nodes to underloaded nodes, and the migration cost includes the communication path change and additional relay overhead caused by task reallocation. The scheduling result is recorded as node load balancing data, including node ID, actual load value, average load offset, task ID adjustment record, etc.
[0062] Preferably, the construction of the three-level reflection system in step S3 is specifically as follows: Based on the node load balancing data, user task reasoning is performed to obtain user task reasoning data; In this embodiment, based on the node load balancing data, the actual task assignment numbers of each node, the node resource usage (CPU usage rate, GPU computing thread activity, memory read / write frequency), and the task execution result data (intermediate representation tensor shape, computation path call log, response latency) are extracted. A task inference trigger mechanism is adopted. Whenever the node resource usage rate continuously exceeds the 95% threshold for more than a T = 200 ms period, backward inference is triggered for the completed tasks, and the input vector structure, inference stage execution path, node switching situation, etc. corresponding to the tasks are recorded as user task inference data. The inference information is implemented by calling the task backtracking component in the distributed training platform. The operator call log (exported by PyTorchProfiler) in the system is combined with the task assignment table for joint restoration to generate a complete user task inference data table. The table fields include: task ID, input modality type (text, image, text-image pair), task execution stage, node ID, path length (call depth), tensor shape change process, etc.
[0063] Based on the preset target inference data, inference difference analysis is performed on the user task inference data to obtain inference difference data, and the output generation policy parameters are corrected according to the inference difference data to obtain the result reflection layer data; In this embodiment, the standardized reference path, reference node call sequence, and target inference tensor features (tensor sparsity rate, activation mean, attention distribution entropy) are extracted from the target inference data imported in the system training initialization stage, and the corresponding fields in the user task inference data are compared item by item. The calculation of inference difference metrics includes: call path length deviation ΔL, node call sequence offset rate r_seq, tensor sparsity rate difference Δρ, and attention entropy offset ΔH. The judgment criterion is set as: ΔL > 2, r_seq > 15%, Δρ > 0.12, and ΔH > 0.25 are regarded as inference offsets. Based on the identified inference differences, the generation policy mapping table is called to correct the policy parameters. The correction items include: readjusting the attention focus coefficient α, with the correction range of α being [-0.05, +0.1]; adjusting the generation maximum step B, with the correction range of B being [+5, +20]; and reassigning the candidate answer sorting weight γ, with the update interval of γ being [-0.2, +0.4]. The corrected policy parameters are written into the policy parameter table and used as the result reflection layer data, including: inference difference type, inference metric offset, parameter correction content, and new parameter values.
[0064] Based on the user task inference data, the task execution decision tree path is identified to obtain the decision tree path data; based on the decision tree path data, inefficient nodes are identified and a caching mechanism is added to obtain the process reflection layer data; In this embodiment, an inference data analysis engine is called to map the module-level execution path in the user task inference data into a decision tree structure. The construction method is as follows: The input type is used as the root node, and branches are expanded sequentially according to different execution stages in the task flow. Each decision path node records the operator name, scheduling node number, and computing resource occupancy rate called at the current stage. Using the call order recorded in the task log file and combining with the scheduling path record (collected by the RAG execution engine log), a multi-level path tree structure is constructed in the graph database. If there are multiple execution paths under the same task modality, the longest execution path is used as the main path, and its branch count, call stack depth, and interruption count are recorded. The finally output decision tree path data includes: task type, main path length, branch path list, key node position (load peak node), and path stability score (calculated from the number of node scheduling switches). Perform per-path analysis on the decision tree path data, extract the average task execution delay Tj of each node, assume the node delay mean μT and standard deviation σT. If there is a node with Tj > μT + 2×σT, it is determined as an inefficient node. Add such nodes to the inefficient node pool and add a caching mechanism to them. The caching mechanism is implemented as follows: Before the inference task reaches this node, the weight parameters and intermediate activation values required for the next task are pre-read through the previous-level node and written into the local memory or GPU L2 cache. The cache module adopts a circular cache structure, and the cache capacity is set to 256MB. The replacement policy adopts the least recently used (LRU). The caching policy is written into the node configuration table, and the inefficient node and caching rule records are written into the process reflection layer data, including the inefficient node ID, delay metric, cache activation time, initial cache hit rate, etc.
[0065] Evaluate the long-term operation data based on the user task inference data, and identify the underlying cognitive biases based on the long-term operation data to obtain the underlying cognitive bias data; trigger the knowledge graph reconstruction based on the underlying cognitive bias data to obtain the policy reflection layer data; In this embodiment, the task tags recorded in the user task inference data (such as task objectives, task modalities), the actual generated content structure (such as the number of entity references in the response text, the localization box offset in the image inference task) are compared and analyzed with the historical operation logs. The operation logs include the response structure of the past 50 rounds of task execution, user feedback data (such as annotation consistency scores), task duration, generation failure rate, etc. The task performance indicators are deviation-fitted with the historical benchmarks. When the generation accuracy drops by more than 12%, the generation path length increases by more than 20%, or the entity reference deviation rate exceeds 18%, it is considered that there is a cognitive deviation. After identification, it is classified according to the deviation type, such as "insufficient fact recall", "path backtracking error", "cross-modal mapping misalignment", etc., and the underlying cognitive deviation data is output. The fields include: cognitive deviation type, deviation threshold trigger item, corresponding task tag and modality, and index offset degree. According to the cognitive deviation type list, each type of deviation is mapped to the relationship edge or entity node in the knowledge graph for marking. If the cognitive deviation is concentrated on a specific entity (such as a person or an organization), the attribute information of the node is reconstructed; if it is concentrated on the relationship edge (such as "belongs to", "is responsible for", etc.), the entity reference relationship in the text library and the image library is retrieved again for entity rebinding. The reconstruction process uses a rule-based entity extraction and graph structure update algorithm. When the frequency of the same deviation entity exceeds 30 times, local topological reconstruction of the graph is performed, and the edge weights, edge directions, entity classifications are updated again, and the path reachability is recalculated. After the reconstruction is completed, the data of the strategy reflection layer is generated, including: the ID of the reconstructed entity, the relationship type with deviation, the reconstructed edge change list, the change timestamp, etc.
[0066] Integrate the data of the result reflection layer, the process reflection layer, and the strategy reflection layer to construct a three-level reflection system.
[0067] In this embodiment, the data of the result reflection layer, the process reflection layer, and the strategy reflection layer are classified and archived according to the reflection types (output difference correction, path performance tuning, knowledge structure update), and a unified reflection index system is constructed. Each reflection record is set with a unique ID, and an association pointer with the original task inference data is established. A hash map (HashMap) is used to quickly map each reflection data entry to the source task. By configuring the metadata template (fields include: reflection trigger condition, reflection type, reflection time, execution node, correction parameter), a three-level reflection record structure is generated, which is respectively stored in the three-level reflection database, and a cross-layer reference relationship graph is established for subsequent traceability and adaptive strategy push.
[0068] Preferably, the cognitive distillation in step S3 is specifically: Structurally organize the three-level reflection data to obtain structured three-level reflection data; In this embodiment, the fields of the step result reflection layer data, the process reflection layer data, and the strategy reflection layer data are extracted and standardized. This process first loads three types of reflection data tables: Result_Reflection_DB, Process_Reflection_DB, and Strategy_Reflection_DB, and processes the structures of the fields through a field mapping table. For example, the "correction parameter" field in Result_Reflection_DB, the "cache strategy parameter" field in Process_Reflection_DB, and the "knowledge graph change entry" field in Strategy_Reflection_DB are all mapped to the "CorrectionField" field; the "reflection trigger condition" fields of the three tables are standardized as "TriggerKey"; the timestamp field is uniformly named "ReflectionTime". Use a structured conversion tool such as Spark SQL DataFrame to perform field mapping and merging operations, and the result is to construct a structured table "Structured_Reflection_Data", which contains the fields: ReflectionID (CHAR(32)), TriggerKey (VARCHAR(128)), CorrectionField (JSON), ReflectionType (ENUM), ReflectionTime (DATETIME), TaskPointer (CHAR(32)). Extract sub-fields from the JSON format fields using JSONPath and store them in a nested data table in a nested structure.
[0069] Extract common cognitive experiences based on the structured three-level reflection data; In this embodiment, in the stage of extracting common cognitive experience, by traversing the "CorrectionField" field in "Structured_Reflection_Data", a rule-based similarity matching algorithm is used to extract repetitive structures. For example, for the parameter correction items in multiple records, if multiple tasks repeatedly adjust the same parameter (such as the "BatchSize" field changing from 32 to 64) more than the set frequency threshold (set to 5 times), it is recorded as a common cognitive event. The extraction tool uses the FP-Growth frequent item set algorithm, with the minimum support set to 0.02 and the minimum confidence set to 0.8. The finally extracted common structures include: parameter name, correction mode, and trigger scenario. Taking "(parameter name = MaxNodeLoad, correction direction = decrease, trigger scenario = path time consumption > 100ms)" as the structural unit, it is uniformly stored in the "Common_Cognition_Experience" table. The fields of this table include: ExperienceID (UUID), ParamName (VARCHAR), AdjustmentPattern (ENUM), TriggerCondition (TEXT), SupportCount (INT).
[0070] Perform knowledge content compression based on the common cognitive experience to obtain compressed cognitive knowledge data; In this embodiment, during the knowledge content compression process, based on the records in the "Common_Cognition_Experience" table, common cognition is vector-reduced in the form of "trigger-adjustment-effect" triples. The compression method adopts a parameter merging compression algorithm, that is, merging triplets with high semantic repetition into abstract strategy groups. For example, for the following three records: "(parameter = LearningRate, adjustment = reduction, trigger = Loss large fluctuation)", "(parameter = LearningRate, adjustment = halved, trigger = Loss non-convergence)", "(parameter = LearningRate, adjustment = from 0.001 to 0.0005, trigger = Loss overfitting)", abstracted into a unified expression: "LearningRate↓ if LossPattern∈{Fluctuation, NonConvergence, Overfitting}". Use regular expressions to aggregate keywords, set the semantic matching Jaccard similarity threshold to 0.7, merge similar records and store the results in the "Compressed_Cognition_Knowledge" table. The fields include: CompressedID, CanonicalParam, GeneralTrigger, UnifiedAdjustment, and CompressionGroupID.
[0071] Abstract the compressed cognitive knowledge data into knowledge cognitive units; In this embodiment, during the abstraction stage of the knowledge cognition unit, according to the data in the "Compressed_Cognition_Knowledge" table, each compressed triple is logically split into separate cognition units. The structure of a cognition unit consists of three parts: CognitionFactor, ConditionLogic, and ActionTag. For example, the above triple is split into: CognitionFactor = "LearningRate", ConditionLogic = "The Loss fluctuation pattern belongs to the high-frequency category", and ActionTag = "Reduce LearningRate by half". The factor template matching table is used to normalize and label the parameters. The ConditionLogic field is expressed as a Boolean expression through the logical condition abstraction module (such as "LossStdDev>0.2 && TrainingEpoch>10"), and the ActionTag is uniformly recorded in the form of an operation code (such as "OP_DECAY_50" indicating that the parameter is halved). Finally, the "Knowledge_Cognition_Unit" table is constructed, and the fields include: UnitID (CHAR(32)), CognitionFactor (VARCHAR), ConditionLogic (TEXT), ActionTag (VARCHAR), and SourceGroupID (associated with the original compressed group ID).
[0072] Based on the knowledge cognition unit, semantic mapping is performed to obtain semantic mapping data; In this embodiment, the semantic mapping operation takes the knowledge cognition unit as the input, and automatically pairs the CognitionFactor with the task operation process fields through the semantic vector mapping rule. Using the field semantic vector construction scheme based on the WordPiece tokenizer, the CognitionFactor and the task process fields (such as scheduling parameters, algorithm component parameters, etc.) are vectorized respectively. The 300-dimensional GloVe embedding model is used as the underlying semantic encoding, and the cosine similarity threshold is set to 0.85 to screen and match the fields. If the similarity between the CognitionFactor "MaxNodeLoad" and the task process field "Task_NodeCapacity" reaches 0.92, a mapping pair is established. This mapping relationship is written into the "Semantic_Mapping_Data" table, and the fields include: MapID (CHAR(32)), Factor (VARCHAR), MappedField (VARCHAR), SimilarityScore (FLOAT), and AssociatedLogic (the corresponding ConditionLogic expression).
[0073] Based on the semantic mapping data, it is injected into the task process to generate cognitive distillation data.
[0074] In this embodiment, in the cognitive distillation data generation stage, the mapping fields in "Semantic_Mapping_Data" are fused and injected with the execution control parameters in the original task process data structure. The injection method is based on a conditional trigger mechanism, that is, when the task process is running, the system monitoring module reads the running values of the corresponding fields in real time and makes a boolean judgment based on the "AssociatedLogic" field expression. If the condition is true, the operation specified by "ActionTag" is executed, such as adjusting the "Task_NodeCapacity" field from 128 units to 64 units. The distilled data is recorded in the form of a log sequence to generate a "Cognition_Distillation_Trace" structure, and the fields include: TraceID (CHAR(32)), TargetField (VARCHAR), TriggerEval (TEXT), ActionExecuted (VARCHAR), Timestamp (DATETIME), CognitionUnitID (CHAR(32)), which serves as the input data source for the subsequent RAG-ReAct dual-engine adaptation training module.
[0075] Preferably, step S4 is specifically as follows: Step S41: Add type tags based on the cognitive distillation data to obtain cognitive label data; In this embodiment, when annotating the data in the "Cognition_Distillation_Trace" structure with cognition types, a cognition label type dictionary is first constructed, and six basic label types are divided: parameter adjustment (Code = 01), task assignment (Code = 02), conditional response (Code = 03), anomaly warning (Code = 04), data correction (Code = 05), and logic optimization (Code = 06). By performing regular matching on the operation codes in the "ActionExecuted" field to extract keywords, such as action instructions containing key substrings like "DECAY", "SPLIT", "CORRECT", etc., they are classified as the parameter adjustment class (01); operations containing "REASSIGN", "REBALANCE" are classified as the task assignment class (02). The lexical rule mapping table (a total of 62 rules are set) is used to match and mark all action instructions. The mapping result is added to the original data as the "CognitionLabelCode" field, obtaining the "Cognition_Label_Data" structure, and the fields include: TraceID (CHAR(32)), ActionExecuted (VARCHAR), CognitionLabelCode (CHAR(2)), CognitionUnitID (CHAR(32)), TriggerEval (TEXT). To improve consistency, a character normalization tool is used for all fields to perform operations such as unifying case, cleaning punctuation, and compressing spaces.
[0076] Step S42: Construct a knowledge display unit based on the cognition label data; In this embodiment, during the construction of the knowledge display unit, a data intermediate structure for graph drawing is generated based on the label classification and associated cognitive units in the "Cognition_Label_Data" structure. This structure is called "Knowledge_Display_Unit", and its fields include: DisplayID (CHAR(32)), MainLabel (VARCHAR), SubLabel (VARCHAR), LogicCondition (TEXT), ActionDesc (TEXT), UnitSourceID (CHAR(32)). The main label is mapped to the dictionary name by "CognitionLabelCode", for example, "01" is mapped to "parameter adjustment" and filled into the MainLabel field; the SubLabel field is taken from the specific parameter items in the cognitive factor (CognitionFactor) field, such as "MaxNodeLoad"; the LogicCondition field is directly copied from "TriggerEval". The ActionDesc field translates the standard action statements through a mapping table, for example, "OP_DECAY_50" is mapped to "halve the parameter value". The generated knowledge display unit data is persisted in JSON format as "Display_Unit_DB" for use in the thinking graph construction step. The lengths of all fields are restricted. The maximum length limits for the MainLabel and SubLabel fields are 64 characters, and the maximum limits for the LogicCondition and ActionDesc fields are 512 characters. The exceeded part uses the sliding window truncation strategy to preserve the expression integrity.
[0077] Step S43: Draw a cognitive thinking graph according to the knowledge display unit to obtain visualized thinking data; In this embodiment, during the stage of drawing the cognitive thinking map, the data in "Display_Unit_DB" is loaded and a graph structure is constructed. The "cognitive factor" is used as the node (Node), and the "logical trigger condition → action description" is used as the edge (Edge). During the construction process, the NetworkX graph structure library is used to construct a directed graph. The unique ID of each node is set as the value of the SubLabel field, the starting node of each edge is the condition factor involved in LogicCondition, and the ending node is the parameter target involved in ActionDesc. For example, if the condition is "LossStdDev>0.2" and the action is "LearningRate is reduced by 50%", then in the graph, the node "LossStdDev" points to the node "LearningRate", and the edge label is ">0.2→DECAY_50". The edge weight is set as the frequency of occurrence of the cognitive unit in Cognition_Label_Data, as the edge_weight attribute. The node attributes include: type (static / dynamic), label_type (MainLabel), unit_source (UnitSourceID). Finally, the "Cognitive_Graph_Structure" structure is generated, and the fields include: NodeID (CHAR(32)), NodeName (VARCHAR), NodeType (ENUM), EdgeStart (CHAR(32)), EdgeEnd (CHAR(32)), EdgeLogic (TEXT), EdgeWeight (INT). The graph structure is serialized into a GraphML format file, and an auxiliary index table is generated to provide the visualized thinking data "Visualized_Cognition_Data".
[0078] Step S44: Train the inference engine according to the visualized thinking data; In this embodiment, during the training of the inference engine, the graph structure in "Visualized_Cognition_Data" is loaded, and a sample sequence is constructed through a path heuristic mechanism for logical causal learning. Path sampling uses a directed path traversal algorithm (the DFS depth limit is 4), and each legal path is serialized into a set of logical relationship samples: starting factor → condition → intermediate operation → target factor. For example, the path "LossStdDev → LearningRate → BatchSize" constructs the logical expression: "If LossStdDev > 0.2, then LearningRate decreases by 50%, and if LearningRate < 0.0005, then BatchSize should be halved". This logical expression is constructed into a structured sample, including fields: InputFactor (TEXT), ConditionExpr (TEXT), IntermediateAdjustment (TEXT), TargetOutput (TEXT). All samples are combined into a training set, and a decision model based on a logical regression tree is used for training. The input dimension of the model depends on the types of InputFactor (a total of 187 types), the dimension of the conditional expression vector is 512, and IntermediateAdjustment and TargetOutput are encoded as label sequences. During the training process, the maximum tree depth is set to 6, the minimum sample splitting threshold is 10, the training batch is 100 rounds, and the cross-entropy loss function is used.
[0079] Step S45: Perform directional injection of model parameters into the inference engine according to the model fine-tuning gradient direction to obtain the agent inference model.
[0080] In this embodiment, during the model fine-tuning stage, the aforementioned trained inference engine parameter structure is loaded, and by analyzing the reverse gradient direction in the multi-modal agent, the node parameters showing high-bias paths in the model are corrected directionally. The specific operation is to extract the error gradients of the output results of the RAG and ReAct co-engines in the inference task, and map them to the intermediate nodes of the inference engine through the backpropagation path. During the process of extracting the error reverse gradients, the average gradient accumulation technique is used to average the gradient directions of 10 recent task sequences, and a threshold θ = 0.03 is set. If a factor node has the same gradient direction continuously in multiple tasks and the gradient magnitude exceeds this threshold, it is determined as a high-weight path node. Perform directional fine-tuning operations on the node parameters, adjusting its path coefficient by an increase or decrease in the range of 5% - 20% (the adjustment step size is 2.5%), and the adjustment rule is based on the positive or negative nature of the gradient direction. All adjustment operations are executed by mapping to the parameter table "Reasoning_Engine_Param_Set", and the corresponding numerical corrections are made to the fields: NodeWeight, EdgeScore, and BiasTerm respectively. Finally, a structurally consistent "Agent_Reasoning_Model" structure file is generated, containing all the inference node parameters and logical path weights after fine-tuning, providing a basis for the subsequent inference of the multi-modal RAG-ReAct dual engine.
[0081] Preferably, step S45 is specifically as follows: Step S451: Construct a gradient guidance signal according to the model fine-tuning gradient direction; In this embodiment, when constructing the gradient guidance signal, first, the parameter update processes of all neural network layers involved in the last five training batches of the multi-modal agent model are recorded. The recorded data includes the gradient values generated by the weights and biases in each fully connected layer and normalization layer during each backpropagation. These data are stored in tensor form, with dimensions corresponding to the layer number, parameter position, and training batch number. The average gradient of each parameter is obtained by averaging the gradient values generated by the parameter in five training batches. If the average gradient value of a parameter is greater than the set gradient threshold (for example, set to 0.015), then the position of this parameter is marked as a "guidance activation point". All activation point information is organized into a guidance mask map, which is a boolean matrix with the same structure as the model parameter structure, and is used to guide the fine-tuning direction subsequently. This guidance mask map is saved as a JSON structure file named "Gradient_Guide_Map".
[0082] Step S452: Construct a directional loss function based on the gradient guidance signal; In this embodiment, by reading the guiding mask graph generated in the previous step, all parameter positions marked as "activation points" are used as the target regions for constructing the directional loss function in this step. During the training process, in addition to retaining the loss function of the original classification or generation task, a penalty term is added to strengthen or suppress the offset of the parameters in the guiding region. The calculation of this penalty term is based on the degree of difference between the reference value and the current value of the parameter in the previous round of training, and is corrected by superimposing the weight marking of this position in the guiding mask. The penalty weight is set to 0.25 using the configuration script to maintain stability and limit the maximum impact of this term to no more than 10% of the overall loss. The finally constructed loss function is uniformly managed by the loss manager in the training controller and embedded into the weight update logic in the backpropagation stage to ensure that the guiding region can be strengthened and adjusted in the set direction.
[0083] Step S453: Determine the fine-tuning network layer based on the directional loss function; In this embodiment, according to the actual response of each network layer during the training process of the loss function in the previous step, the loss contribution of each layer is statistically analyzed layer by layer. The proportions of each layer in the guiding loss term are sorted, and the judgment threshold is set to 2%. When the guiding loss contribution of a certain layer exceeds this value, it is marked as the target layer that needs to be fine-tuned. This operation is implemented by the loss stratification analysis component in the loss processing module, and the analysis result is exported as a structured list, which includes the number, parameter activation ratio, and estimated adjustment strength of each target layer. This list file is named "Target_Layers_List" and is passed as input to the learning rate adjustment controller before the start of the training loop for the next step of dynamic strength setting.
[0084] Step S454: Dynamically adjust the fine-tuning strength based on the fine-tuning network layer to obtain a directional fine-tuning mechanism; In this embodiment, the target layer list is loaded into the fine-tuning controller, and different learning rates are assigned to each layer in the list according to its activation ratio. The base learning rate is set to 0.0001. The higher the activation ratio, the larger the learning rate assigned to this layer, but not exceeding the upper limit of 0.0004. The adjusted learning rates are set separately for each layer and injected through the parameter group mechanism of the optimizer, so that each layer has an independent update rate during the fine-tuning process. This dynamic learning rate mechanism is called the "directional fine-tuning mechanism" and is encapsulated in the fine-tuning scheduling module, recording the learning rate configuration of each layer in a dictionary structure. The final result is output as "Layer_LR_Profile" as the reference basis for the parameter injection module.
[0085] Step S455: Explicitly inject parameters into the inference engine based on the directional fine-tuning mechanism to obtain explicitly injected parameters; In this embodiment, during the explicit injection phase, the weight values of modules with clear paths and traceable structures in the inference engine (such as the path weights of the rule graph and the conditional transition matrix) are directly updated. The fine-tuning intensity information of each layer in Layer_LR_Profile is read, and combined with the guidance signal graph generated in the early stage, the parameters that need to be numerically adjusted are identified item by item, and the original parameters are directly updated forward or backward according to the gradient direction. This process does not involve random perturbations, but performs fixed increment corrections based on the average trend of the historical update directions. The update amplitude of each parameter is limited within 0.025 to prevent overfitting or excessive deviation. After the injection is completed, all changes are organized into a structure file "Explicit_Param_Set", recording the parameter names, the values before and after the update, and the network levels to which they belong, and written into the path graph weight cache of the inference engine.
[0086] Step S456: Perform implicit injection of parameters into the inference engine based on the directional fine-tuning mechanism to obtain implicitly injected parameters; In this embodiment, during the implicit injection process, the parameters of important components in the model that have non-direct output paths but affect information flow (such as the normalization layer and the embedding mapping layer) are adjusted. For the scale factor and offset parameters in the LayerNorm module, by sampling perturbation values corresponding to the fine-tuning intensity (using a zero-mean Gaussian distribution) and injecting them into the original parameters to form new adjusted values. The perturbation standard deviation is obtained by multiplying the corresponding learning rate in Layer_LR_Profile by the perturbation amplification factor, and the amplification factor is set to 0.5 to ensure that the update has a direction but does not destroy the original structure of the parameters. The results of all implicitly adjusted parameters are stored as "Implicit_Param_Set", including the parameter type, the original value, the perturbation value, and the updated value, and are synchronously entered into the normalization path cache area of the inference engine.
[0087] Step S457: Construct an agent inference model according to the explicitly injected parameters and the implicitly injected parameters.
[0088] In this embodiment, the Explicit_Param_Set and Implicit_Param_Set are parameter-fused by the unified scheduling module to generate a complete inference engine reconstruction configuration. First, all explicit parameters are injected into the logic graph path and the conditional judgment module to ensure the complete update of the symbolic inference chain. Subsequently, the implicit parameters are injected into the normalization module and the embedding layer of the feature flow path. After the parameter synchronization is completed, the model structure reconstruction process is triggered, including four steps: node reordering, cache flushing, weight rebinding, and logical consistency verification. All operations are executed by the "Reasoning_Engine_Reconstructor" module in the system under single-thread control to avoid race conditions. Finally, the constructed inference model can be seamlessly connected to other modules in the RAG-ReAct dual engine in terms of structure, forming a multi-modal intelligent agent structure with the capabilities of guiding memory, dynamic inference, and feature reconstruction.
[0089] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be embraced by the present invention.
[0090] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. The present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A collaborative training method for a multi-modal intelligent agent's RAG-ReAct dual engine, characterized in that, It includes the following steps: Step S1: Obtain multimodal data, and convert the multimodal data into high-dimensional vectors; construct a knowledge graph based on the high-dimensional vectors; use dynamic distillation technology to encode the semantic relationships of the knowledge graph into the model fine-tuning gradient direction; Step S2: Design a distributed architecture based on the high-dimensional vectors; Perform hybrid retrieval based on the distributed architecture to obtain hybrid retrieval data; perform distributed reasoning based on a preset recursive reflection mechanism and the hybrid retrieval data to generate distributed reasoning data; perform data parallelism detection based on the distributed reasoning data to obtain data parallelism parameters; Step S3: Perform node resource scheduling based on the data parallelism parameters to obtain node load balancing data; Construct a three-level reflection system based on the node load balancing data; perform cognitive three-level reflection based on the three-level reflection system to generate three-level reflection data; perform cognitive distillation on the three-level reflection data to obtain cognitive distillation data; Step S4: Visualize the cognitive distillation data to obtain visualized thinking data; train an inference engine based on the visualized thinking data; perform directional injection of model parameters into the inference engine according to the model fine-tuning gradient direction to obtain an intelligent agent inference model.
2. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein Specifically, Step S1 is as follows: Step S11: Obtain multimodal data, and extract text data and image data; Step S12: Identify key entities based on the text data; extract entity relationships according to the key entities; Step S13: Construct a semantic graph according to the key entities and entity relationships; Step S14: Identify image targets based on the image data, map the image targets to the semantic graph, and use a preset vector database to convert the semantic graph into high-dimensional vectors; Step S15: Construct a knowledge graph based on the high-dimensional vectors; Step S16: Use dynamic distillation technology to encode the semantic relationships of the knowledge graph into the model fine-tuning gradient direction.
3. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein Specifically, Step S2 is as follows: Step S21: Perform vector partitioning based on the high-dimensional vectors to obtain a partitioned vector space; Step S22: Design a distributed vector index structure according to the partitioned vector space; design a distributed architecture based on the distributed vector index structure; Step S23: Perform hybrid retrieval based on the distributed architecture, including vector retrieval, full-text retrieval, and re-ranking retrieval, to obtain hybrid retrieval data; Step S24: Perform distributed reasoning based on a preset recursive reflection mechanism and the hybrid retrieval data to generate distributed reasoning data; Step S25: Perform data parallelism detection based on the distributed reasoning data to obtain data parallelism parameters.
4. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 3, wherein, Specifically, Step S24 is as follows: Step S241: Perform result inconsistency detection on the hybrid retrieval data according to the preset retrieval target information to obtain retrieval result abnormal data; Step S242: Dynamically correct the prompt words based on the retrieval result abnormal data to obtain prompt word correction data; Step S243: Trigger a reflection condition according to the retrieval result abnormal data to obtain reflection condition data; Step S244: Activate a preset recursive reflection mechanism based on the reflection condition data to obtain recursive reflection mechanism startup data; Step S245: Perform re-reasoning execution according to the recursive reflection mechanism startup data and the prompt word correction data to generate distributed reasoning data.
5. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 3, wherein, Specifically, Step S25 is as follows: Step S251: Extract the distributed inference path according to the distributed inference data; Step S252: Identify the data flow dependency relationship based on the distributed inference path; Step S253: Calculate the task granularity factor according to the data flow dependency relationship; Step S254: Set parallel threads according to the task granularity factor to obtain parallel thread data; Step S255: Perform task scheduling simulation based on the parallel thread data to obtain task scheduling data; Step S256: Perform data parallel detection according to the task scheduling data to obtain data parallel parameters.
6. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein The node resource scheduling in Step S3 is specifically as follows: Detect load imbalance based on the data parallel parameters to obtain load imbalance data; Identify load abnormal nodes according to the load imbalance data; Evaluate the node computing power based on the load abnormal nodes; Calculate the node task communication overhead based on the load abnormal nodes; Allocate node tasks according to the node computing power to obtain node task allocation data; Select the minimum communication cost path based on the node task communication overhead to obtain the minimum communication cost path data; Perform node resource scheduling based on the node task allocation data and the minimum communication cost path data to obtain node load balancing data.
7. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein The specific construction of the three-level reflection system in Step S3 is as follows: Perform user task inference based on the node load balancing data to obtain user task inference data; Perform inference difference analysis on the user task inference data based on the preset target inference data to obtain inference difference data, and correct the output generation strategy parameters according to the inference difference data to obtain the result reflection layer data; Identify the decision tree path of the task execution according to the user task inference data to obtain the decision tree path data; Identify inefficient nodes based on the decision tree path data and add a caching mechanism to obtain the process reflection layer data; Evaluate the long-term operation data according to the user task inference data, and identify the underlying cognitive bias based on the long-term operation data; trigger the knowledge graph reconstruction based on the underlying cognitive bias data to obtain the policy reflection layer data; Integrate the result reflection layer data, the process reflection layer data, and the policy reflection layer data to construct a three-level reflection system.
8. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein The specific cognitive distillation in Step S3 is as follows: Structurally organize the three-level reflection data to obtain structured three-level reflection data; Extract common cognitive experiences based on the structured three-level reflection data; Perform knowledge content compression based on the common cognitive experiences to obtain compressed cognitive knowledge data; Abstract the compressed cognitive knowledge data into knowledge cognitive units; Perform semantic mapping based on the knowledge cognitive units to obtain semantic mapping data; Inject the semantic mapping data into the task process to generate cognitive distillation data.
9. The multimodal agent RAG-ReAct dual-engine collaborative training method according to claim 1, wherein Step S4 is specifically as follows: Step S41: Add type labels based on the cognitive distillation data to obtain cognitive label data; Step S42: Construct a knowledge display unit based on the cognitive label data; Step S43: Draw a cognitive thinking map according to the knowledge display unit to obtain visualized thinking data; Step S44: Train the inference engine according to the visualized thinking data; Step S45: Perform directional injection of model parameters into the inference engine according to the model fine-tuning gradient direction to obtain an agent inference model.
10. The multi-modal agent RAG-ReAct dual-engine collaborative training method according to claim 9, wherein, Step S45 is specifically as follows: Step S451: Construct a gradient guidance signal according to the gradient direction of model fine-tuning; Step S452: Construct an orientation loss function based on the gradient guidance signal; Step S453: Determine the fine-tuning network layer based on the orientation loss function; Step S454: Dynamically adjust the fine-tuning intensity based on the fine-tuning network layer to obtain an orientation fine-tuning mechanism; Step S455: Perform explicit parameter injection on the inference engine based on the orientation fine-tuning mechanism to obtain explicitly injected parameters; Step S456: Perform implicit parameter injection on the inference engine based on the orientation fine-tuning mechanism to obtain implicitly injected parameters; Step S457: Construct an agent inference model according to the explicitly injected parameters and the implicitly injected parameters.
Citation Information
Patent Citations
Multi-modal reasoning method and device based on large language model and knowledge graph
CN118193684A
Power field SQL intelligent agent construction method based on KMDI chain
CN119166662A
Intelligent maintenance reasoning method based on knowledge graph and large language model
CN119886334A
Generative pretraining of multimodal retrieval-augmented visual-language models
WO2024118578A1
Cited By
Multi-modal data dynamic reasoning system and method based on cognitive map
CN120471179A
Unmanned aerial vehicle autonomous obstacle avoidance and path planning method and system based on deep learning
CN120803005A
RAG-oriented embedded service flexible deployment method
CN120872615A
Multi-module collaborative agent for enhancing credibility of intelligence analysis, equipment and medium
CN120880722A
Electric power engineering multi-mode RAG system based on knowledge graph and multi-Agent cooperation
CN121051219A