Construction method of agent management platform supporting knowledge retrieval, generation and optimization
Through the deep semantic fusion of multimodal data and the optimization of the three-level reflection system, the shortcomings of the agent management platform in multimodal data fusion and resource scheduling are solved, the accuracy of knowledge retrieval and system execution efficiency are improved, and resource allocation and hardware collaborative scheduling are optimized.
Patent Information
- Application Number
- CN202510865368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Traditional agent management platforms have problems such as insufficient semantic understanding, uneven resource allocation, and low execution efficiency in multimodal data fusion, knowledge retrieval, resource scheduling and heterogeneous hardware collaborative scheduling, and it is difficult to cope with complex tasks and heterogeneous hardware environments.
Through the deep semantic fusion of multimodal data, a three-level reflection system is established, prompt word optimization and node load balancing are carried out, task chain reconstruction and heterogeneous hardware collaborative scheduling are realized, and knowledge retrieval strategies and resource allocation are optimized.
It improves the context relevance and accuracy of knowledge retrieval, avoids logic jumps and redundant generation, improves the accuracy of complex instruction responses, optimizes resource allocation and scheduling efficiency, and enhances system performance stability and expansion capabilities.
Smart Images

Figure CN120371999A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a construction method for an intelligent agent management platform that supports knowledge retrieval, generation, and optimization. Background Art
[0002] The semantic fusion ability between multi-modal data of traditional intelligent agent management platforms is limited, usually staying only at the stage of shallow feature splicing or rule mapping, unable to achieve deep semantic understanding across modalities, resulting in information deviation during knowledge retrieval or poor context relevance of retrieval results; lacking an effective reflection mechanism, unable to systematically evaluate and feedback optimize the knowledge generation results, especially in the face of complex tasks, lacking the ability to trace back the generation process and decision-making path, prone to problems such as logical jumps, semantic repetitions, or redundant generations; the process of optimizing prompt words relies on manual rules or static templates, unable to dynamically adjust the semantic control strategy in combination with the task context, resulting in insufficient generalization ability of prompt words and affecting the response accuracy of the system to complex instructions; the construction of task chains mostly adopts linear processes or static configuration methods, difficult to optimize in real-time according to the node resource load status, leading to waste of resources on some nodes while some nodes are overloaded, with problems of low scheduling efficiency and weak system throughput capacity; in the face of heterogeneous hardware resource environments, lacking a precise quantification mechanism for task resource types and a thread binding mechanism, unable to achieve efficient hardware co-scheduling, resulting in large fluctuations in the overall execution efficiency of the platform and high response latency, restricting the performance improvement of the intelligent agent management system in complex scenarios. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a construction method for an intelligent agent management platform that supports knowledge retrieval, generation, and optimization to solve at least one of the above technical problems.
[0004] To achieve the above object, a construction method for an intelligent agent management platform that supports knowledge retrieval, generation, and optimization includes the following steps: Step S1: Obtain multi-modal data, perform semantic fusion based on the multi-modal data to obtain semantic fusion data; perform knowledge retrieval based on the semantic fusion data to obtain knowledge retrieval data; Step S2: Establish a three-level reflection system according to the knowledge retrieval data, including a result reflection layer, a process reflection layer, and a strategy reflection layer, so as to obtain result reflection layer data, process reflection layer data, and strategy reflection layer data, and perform hierarchical system fusion to obtain three-level reflection system data; Step S3: Optimize the prompt words based on the data of the three-level reflection system to obtain optimized prompt word data; perform node load balancing according to the data of the three-level reflection system to obtain node load balancing data; reconstruct the task chain based on the node load balancing data to obtain task chain data; detect the resource consumption according to the task chain to obtain resource consumption data; Step S4: Perform heterogeneous hardware cooperative scheduling based on the resource consumption data to obtain heterogeneous hardware cooperative scheduling data; construct an agent management platform according to the heterogeneous hardware cooperative scheduling data and the optimized prompt word data, and perform adaptive adjustment on the knowledge retrieval strategy of the agent management platform to obtain adaptive strategy data.
[0005] Through the deep semantic fusion of multi-modal data, the present invention improves the understanding ability of cross-modal information, effectively solves the semantic deviation problem caused by the shallow feature splicing of traditional platforms, and ensures the context relevance and accuracy of knowledge retrieval. The established three-level reflection system realizes the comprehensive evaluation and adjustment of the knowledge generation results, construction processes, and strategies, avoids logical jumps and redundant generation, and improves the content logic. The optimization of the prompt words based on the reflection system realizes the dynamic adjustment of the prompt words, enhances their generalization ability and semantic control effect, and improves the accuracy of complex instruction responses. The introduction of node load balancing and task chain reconstruction solves the problem of uneven resource allocation, realizes node load balancing, optimizes the task execution process, improves the scheduling efficiency and system throughput. Through the resource consumption data for heterogeneous hardware cooperative scheduling, accurately quantify the task resource requirements and combine with the thread binding mechanism to achieve efficient hardware cooperation, reduce latency fluctuations, improve the execution efficiency and response speed in complex hardware environments, enhance the system performance stability and expansion ability, and comprehensively improve the comprehensive ability of the agent management platform. Brief Description of the Drawings
[0006] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent: Figure 1 It is a schematic flow chart of the steps of a method for constructing an agent management platform that supports knowledge retrieval, generation, and optimization according to the present invention; The realization of the objectives, functional characteristics, and advantages of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments
[0007] The technical method of the present invention for the patent will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0008] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0009] It should be understood that although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0010] To achieve the above object, please refer to Figure 1 , the present invention provides a method for constructing an intelligent agent management platform that supports knowledge retrieval, generation, and optimization. The method includes the following steps: Step S1: Obtain multimodal data, perform semantic fusion based on the multimodal data to obtain semantically fused data; perform knowledge retrieval based on the semantically fused data to obtain knowledge retrieval data; In this embodiment, the multimodal data includes text data, image data, audio data, and structured data. First, the corresponding preprocessing module is used to standardize various types of data. For text data, a word segmentation tool is used to segment the input text, and the selected dictionary library is a custom industry term dictionary to ensure that the word segmentation accuracy reaches over 95%. For image data, pixel normalization processing is adopted, the unified resolution is 224×224 pixels, and image features are extracted through a convolutional feature extraction algorithm. The extraction layer is set to the 4th convolutional layer, and the output dimension is fixed to a 512-dimensional vector. For audio data, the sampling rate is fixed at 16 kHz, and short-time Fourier transform (STFT) is used to extract time-frequency features with a window length of 25 ms and a step size of 10 ms, extracting 128-dimensional spectral features. The structured data is input in JSON format, and the fields are predefined to strictly follow the data standard. The hash mapping index method is adopted to quickly locate the key fields. Subsequently, multimodal alignment technology is adopted. Specifically, cross-modal feature alignment is performed through the Maximal Mutual Information criterion, calculating the mutual information between the features of each modality, and the threshold is set to 0.8 to ensure the alignment accuracy. After alignment, the features of each modality are fused through a weighted fusion method, and the weights are calculated by normalizing the mutual information values. The fusion result is converted into a 512-dimensional semantic vector through linear projection. Taking this fusion vector as the input, a multi-source knowledge base matching algorithm based on cosine similarity calculation is used to retrieve relevant knowledge fragments from the pre-constructed knowledge base. When retrieving, the similarity threshold is set to 0.75, the retrieval results are sorted in descending order of similarity, and the top 20 results are selected as knowledge retrieval data.
[0011] Step S2: Establish a three-level reflection system based on the knowledge retrieval data, including constructing a result reflection layer, a construction process reflection layer, and a construction strategy reflection layer, so as to obtain the data of the result reflection layer, the data of the process reflection layer, and the data of the strategy reflection layer, and perform hierarchical system fusion to obtain the data of the three-level reflection system; In this embodiment, in the result reflection layer, core keywords are extracted through text analysis technology. The TF-IDF algorithm is adopted, and the number of keywords is limited to the top 10. Based on the core keywords, a target semantic graph is constructed. The nodes of the semantic graph represent keywords, and the edge weights are calculated by the co-occurrence frequency of keywords. The threshold is set to 0.6, and the edges below the threshold are removed. According to the target semantic graph, an automatic question-answering generation process is executed, and a template-driven answer generation mechanism is used. All generated texts are recorded as actual output data. When calculating the keyword coverage rate, the number of keywords matched in the actual output is divided by the total number of core keywords. When the coverage rate is lower than 0.7, it is considered insufficient. The cosine similarity calculation method is used to calculate the semantic matching degree. When the matching degree is lower than 0.75, it is considered unmatched. The output error is calculated based on the keyword coverage rate and the semantic matching degree, and the error threshold is 0.3. The output error is used to extract language control parameters, including the answer length limit (maximum character count 500), the proportion of professional terms used (at least 15% of the total word count), and the semantic coherence control index. The language control parameters are converted into a structural output control template, which is defined in the form of a finite state machine. The nodes correspond to language generation rules, and the edges represent conversion conditions, finally forming the data of the result reflection layer. The process reflection layer is realized by analyzing the task execution path. The log analysis technology is adopted to parse the task node call sequence and extract the task execution path data. Based on the path data, a path decision tree is constructed. The nodes represent task nodes, and the tree height is limited to 10 layers to control the complexity. When identifying inefficient nodes, the average response time of each node is calculated, and the nodes exceeding 100 ms are determined to be inefficient. The node call frequency is counted, and the inefficient nodes exceeding the call count threshold of 500 times are marked as high-frequency inefficient nodes. Based on the high-frequency inefficient nodes, a result cache storage strategy is configured. The cache capacity is set to 1 GB, and the cache expiration time is set to 30 minutes. According to the cache strategy, the task process is reconstructed, and the execution priority of the inefficient nodes is adjusted to form the data of the process reflection layer. The strategy reflection layer is realized by extracting business rules. The rule format uniformly adopts JSON Schema, covering 20 business rules such as knowledge update and permission control. The historical decision output records of the last 3 months are obtained, and the storage format is a log file. By comparing the historical decisions with the business rules, the Bayesian network model is used to calculate the cognitive bias probability, and the bias threshold is 0.2. Based on the cognitive bias, the knowledge structure update requirements are generated and transformed into graph reconstruction requirements, including node addition and deletion, and relationship adjustment, with the adjustment range not exceeding 10%. After completing the adjustment of the knowledge graph structure, the data of the strategy reflection layer is formed. Finally, the three-layer reflection data are hierarchically fused, and the weighted fusion method is adopted, with the weight ratio being 40% for the result reflection layer, 35% for the process reflection layer, and 25% for the strategy reflection layer, to obtain the complete three-level reflection system data.
[0012] Step S3: Optimize the prompt words based on the data of the three-level reflection system to obtain optimized prompt word data; perform node load balancing according to the data of the three-level reflection system to obtain node load balancing data; reconstruct the task chain based on the node load balancing data to obtain task chain data; detect the resource consumption according to the task chain to obtain resource consumption data; In this embodiment, the task execution result data in the result reflection layer is analyzed, and the result reflection operation is performed to extract the key abnormal indicators in the task results. The anomaly detection algorithm is used for anomaly point identification. During the anomaly identification process, the statistical threshold setting method is adopted, and the anomaly threshold is set as the data points outside the range of the average value of the result data plus or minus 3 times the standard deviation. The standard deviation and the average value are dynamically calculated based on the numerical indicators in each task output result to ensure the accuracy of identification. According to the extracted anomaly point information, combined with the mapping path between the prompt words recorded in the result reflection path and the actual task feedback, the strategic deviation causing the anomaly is analyzed. During the deviation type identification process, the K-means clustering algorithm is used to cluster the abnormal task data. The clustering dimensions include the prompt word length, the nesting level, the semantic deviation value of the task response, and the information density index. The unit of the prompt word length is the number of words. The nesting level is parsed from the context structure. The semantic deviation value is obtained by calculating the word embedding cosine distance by comparing the historical task results. The information density index is obtained by dividing the number of high-frequency key entities appearing in the prompt words of unit length by the total number of words. The three clusters output by the clustering algorithm are classified as prompt word structure defect class, semantic deviation class, and information redundancy class. For the identified strategic deviation types, the context-free grammar is used to reorganize the structure of the prompt words, and a prompt word structure template containing non-terminal symbols and terminal symbols is constructed. The maximum parsing depth limit of the template is 5 layers. The terminal symbols include fields such as keywords, operation objectives, and task requirements. The non-terminal symbols include context-dependent semantics, context modifiers, and constraints. When generating the multi-level prompt word control structure, a control weight mechanism is introduced. The control weight is jointly calculated according to three types of characteristic parameters: part of speech (nouns, verbs, and adjectives are assigned 0.2, 0.5, and 0.3 respectively), semantic intensity (normalized to 0-1 based on the word frequency ranking), and context relevance (calculated by the word vector similarity between the prompt words and the context paragraphs, ranging from 0-1). The weighted average of the three is used as the final control weight, and the constraint range is from 0 to 1. After generating the optimized data of the prompt words, according to the historical execution data provided by the process reflection layer and the strategic reflection layer in the three-level reflection system, the execution status of all intelligent agent nodes in the current task system is counted. The node status is enumerated into four types: idle, running, waiting, and failed. Each status is automatically identified based on the task life cycle tags recorded in real time in the node logs. The calculation method of the node resource load rate is the weighted average of the CPU occupancy rate and the memory occupancy rate. The CPU occupancy rate is sampled once per second, and the unit is a percentage. The memory occupancy is measured in GB and converted to a percentage before participating in the calculation. The node scheduling priority is calculated according to the weight formula: priority = 0.7×(1 - CPU load rate) + 0.3×(1 - memory load rate). The final result is normalized to 0-1. The larger the value, the higher the node scheduling priority.According to the scheduling priority data, the priority-based round-robin scheduling algorithm is used to allocate the tasks to be executed to the nodes with appropriate resources, ensuring that the scheduling delay time does not exceed 50 milliseconds. Node load analysis is performed by calculating the mean and standard deviation of CPU and memory occupancy of all running nodes. If the variance value of the node load rate does not exceed 0.05, it is determined to be in a load-balanced state; otherwise, the scheduling order is readjusted to approach load balance. After the node allocation is completed, the dependency relationships between tasks are identified based on the historical dependency information of task execution. The dependency relationships are expressed in the form of a directed acyclic graph (DAG). The nodes in the graph are task identifiers, and the edges represent dependency relationships. The maximum depth of the graph does not exceed 8 layers. In the task chain reconstruction stage, the topological sorting method is used to sort all task nodes in the task graph. The sorting result is combined with the node load priority to dynamically adjust the execution order of specific tasks, so that high-priority nodes give priority to processing tasks with high dependency, and the task chain structure after reconstruction is completed. After reconstruction, the resource consumption during the overall execution process of the task chain is monitored. The resource items include three indicators: CPU usage time, memory peak value, and network bandwidth. The sampling period is 1 second. The maximum CPU occupancy of a single node shall not exceed 80%, the maximum memory usage is 4GB, and the network bandwidth limit is 100Mbps. The system continuously records the resource usage of each task chain node, generates resource consumption data, and synchronously writes it into the log and the database, providing data support for subsequent heterogeneous hardware scheduling and intelligent agent task dynamic configuration.
[0013] Step S4: Perform heterogeneous hardware collaborative scheduling based on the resource consumption data to obtain heterogeneous hardware collaborative scheduling data; construct an intelligent agent management platform according to the heterogeneous hardware collaborative scheduling data and the prompt word optimization data, and adaptively adjust the knowledge retrieval strategy of the intelligent agent management platform to obtain adaptive strategy data.
[0014] In this embodiment, according to the resource consumption data obtained in step S3, first identify the resource types required for the tasks, distinguish between compute-intensive tasks, storage-intensive tasks, and network-intensive tasks. The classification basis is that the CPU occupancy rate, memory usage rate, and network bandwidth occupancy rate exceed 60% respectively. Quantify the computing power requirements according to the task resource types. The calculation formula is computing power requirement = CPU occupancy rate × task priority, where the priority range is 0 to 1, and the unit of computing power requirement is GFLOPS. Set the calculation threshold to 10 GFLOPS. Match the heterogeneous hardware resources according to the computing power requirements. The hardware resources include CPU clusters, GPU clusters, and FPGA arrays, corresponding to computing power ranges of 100 GFLOPS, 1000 GFLOPS, and 500 GFLOPS respectively. The initial matching is based on the available capacity of the resources and the computing power requirements of the tasks, and select the hardware node that meets the requirements and has the lowest load. Perform task execution simulation based on the initial matching data, using the discrete event simulation method, with a simulation period of 10 seconds. Monitor the thread blocking points during the simulation. If the thread waiting time exceeds 100 ms, it is regarded as blocked, and count the blocking positions and times. According to the thread blocking data, adjust the thread binding of the hardware matching result, bind the high-blocking threads to the resource-free nodes, with the thread binding granularity at the thread level, and the binding rule is to avoid the number of threads on the same node exceeding 64. After completing the thread binding, execute the heterogeneous hardware collaborative scheduling. The scheduling algorithm uses a scheduler based on a priority queue, with a scheduling time slice of 50 ms, supporting task migration and dynamic load adjustment. According to the heterogeneous hardware scheduling data and the prompt word optimization data in step S3, construct the logical architecture of the intelligent agent management platform. The platform architecture adopts a modular design, and the modules communicate through RESTful interfaces, and the data transmission format is JSON. Finally, based on the platform construction result, execute the adaptive adjustment of the knowledge retrieval strategy.
[0015] Preferably, step S1 is specifically as follows: Step S11: Obtain multimodal data; In this embodiment, the multi-modal input types supported by the platform are clarified, including text modality, image modality, audio modality, and structured sensing data modality. The text modality data is extracted from the task log data and divided by lines. Each log contains fields such as timestamp, task ID, execution content, system response, etc.; the image modality data consists of monitoring screenshots uploaded by nodes, in JPEG format, with the resolution limited to 1920×1080; the audio modality data is intercepted from the system instruction broadcast content, in WAV format, with a sampling rate of 16 kHz and mono-channel; the structured data modality is the status monitoring metrics generated by each node, including CPU utilization (unit: %), memory occupancy (unit: GB), network throughput (unit: Mbps), etc., and is stored in CSV format at a sampling period of 1 second. All modality data needs to be labeled with a unified timestamp format "YYYY-MM-DD HH:MM:SS" for subsequent alignment operations. The data preprocessing tools are pandas and scipy in Python. The image data is normalized to the range of 0-1. The audio signal is denoised using a mean filter (window size set to 5). For text, special symbols are removed using regular expressions, and only letters, numbers, and common punctuation are retained.
[0016] Step S12: Extract intra-modal features based on the multi-modal data to obtain modal feature data; In this embodiment, separate feature extraction channels are constructed for each type of modality. The processing flow for the text modality is as follows: Use the n-gram statistical feature extraction method, where n ranges from 1 to 3, to generate a term frequency matrix (Term Frequency), then calculate the TF-IDF value (weight upper limit is 1.0), and generate a sparse matrix with a dimension of (number of samples × 1000) as the text feature; for the image modality, the Local Binary Pattern (LBP) algorithm is used to extract texture features, the LBP radius is set to 3, and the number of neighboring points is set to 8. Finally, the image is divided into 9×9 grid regions, and each region outputs a histogram, which is combined to generate an image feature vector with a dimension of (81×59); for the audio modality, Mel Frequency Cepstral Coefficients (MFCC) are extracted as features, the window size is 25 ms, the sliding step is 10 ms, 13-dimensional coefficients are generated and appended with first-order and second-order differences to obtain a feature vector sequence with a total dimension of 39; for the structured modality data, standardization processing (Z-score standardization) is used, and statistical features such as mean, standard deviation, maximum value, and minimum value are calculated for each second of data using a sliding window (window size 10 seconds, step 5 seconds) to obtain a feature vector with a dimension of 4. All modal features are managed with modal labels through a predefined dictionary structure and uniformly converted to the NumPy array format while maintaining the time alignment structure unchanged.
[0017] Step S13: Perform inter-modal alignment processing based on the modal feature data to obtain modal alignment data; In this embodiment, inter-modal temporal alignment and semantic structure alignment are performed on the in-modal features obtained in the previous step. In the temporal alignment process, interpolation and synchronization techniques are used to process all modalities. The minimum sampling interval is taken as the reference time axis (1 second for text and structured data, 5 seconds for images, and 10 seconds for audio). For sparsely sampled modalities, the linear interpolation method is used to fill in the missing time points. When the image modality is missing, the previous frame is copied. When the audio modality is missing, it is filled with silence features (all values are set to 0). For semantic structure alignment, cosine similarity is used to evaluate the semantic correlation between different modality vectors. For the text modality and the image modality, the text-image consistency score is calculated through co-occurring keyword matching. The matching threshold is set to 0.7. If it is lower than this threshold, it is considered that there is no consistency at the text-image level for the current sample, and it does not participate in the subsequent fusion operation. For audio modality alignment, the VAD (Voice Activity Detection) method is used to remove silent segments, and after retaining the speech segments, alignment is performed. Structured data is matched to the text log according to the task ID during alignment. If there is no match, this record is deleted. After alignment, all modalities are stored as multi-modal joint vectors of fixed length. Each sample is organized in dictionary form, and the fields include "timestamp", "modality identifier", and "feature vector content".
[0018] Step S14: Perform semantic association based on the modality-aligned data to obtain semantic association data; In this embodiment, semantic association processing is performed on the basis of the modality-aligned data. The processing process includes three steps: cross-modal association graph construction, edge weight calculation, and subgraph screening. First, a multi-modal semantic graph is constructed according to the aligned samples. The nodes of the graph are the feature representations of each modality, and the edges in the graph represent the semantic connection relationships between modalities. Each edge records the cosine similarity value between two modalities. The edge weight is set to 0 to 1, and edges with a weight lower than 0.3 are directly removed and regarded as unassociated. To ensure the closure of the semantic link, it is set that each node needs to have valid edge weights (edge weight value ≥ 0.5) with at least two other modalities. Otherwise, the node is removed from the semantic graph. In the subgraph screening step, the K-core algorithm is used with the core degree set to 3 to screen out the most central semantic core subgraph for subsequent semantic fusion operations. The graph data is constructed using the networkx tool and visually inspected to verify the rationality of the semantic graph structure. Finally, the semantic subgraph data structure corresponding to each sample is output in the format of a combination of an adjacency matrix and a node feature set.
[0019] Step S15: Perform semantic fusion according to the semantic association data to obtain semantic fusion data; In this embodiment, semantic fusion is performed based on the semantic sub-graph data structure and completed by using the maximum relevant path aggregation method based on graph traversal. The specific operation is to perform a depth-first traversal on each semantic sub-graph, find the maximum semantic weight path from each text modality node to other modality nodes, the path score is the weighted average of the weights of all edges on the path (the edge weight coefficient is set to 1.0), and the top 3 paths with the highest scores are retained. The feature vectors on the path nodes are combined using a weighted fusion method, where the weighting factor is the edge weight value between each node on the path and the starting node. The fused semantic vector is a vector structure with a fixed length (uniformly 512 dimensions), and all modalities are fused into a unified space. This operation is implemented using NumPy, and the vectors of all samples are normalized after fusion (mean is 0, standard deviation is 1). The fused semantic vector structure is organized into a database structure indexed by task ID, and each record stores the semantic vector content and timestamp index in the JSON format.
[0020] Step S16: Perform knowledge retrieval based on the semantic fusion data to obtain knowledge retrieval data.
[0021] In this embodiment, for the knowledge retrieval operation based on the semantic fusion data, a knowledge index library is first constructed. The data sources are the task operation descriptions extracted from the knowledge graph, the platform management rules, the prompt word template library, and the professional document texts collected from the external knowledge base. All knowledge contents are processed through text normalization to remove punctuation marks and stop words, an inverted index structure is constructed, and TF-IDF vectors (dimension is 2048) are generated for all text paragraphs, which is implemented using TfidfVectorizer in scikit-learn with the parameter settings of max_features = 2048 and ngram_range=(1,2). During the retrieval operation, the semantic fusion vector is used as the query vector, the cosine similarity with each TF-IDF vector in the index library is calculated, the retrieval threshold is set to 0.75, and the knowledge entries with similarity greater than this value are selected as the retrieval results. If there are no results meeting the threshold, the top 3 closest knowledge records are returned. The retrieval return structure contains fields "knowledge paragraph ID", "similarity score", "knowledge base type to which it belongs", and "original content text", which are used for subsequent generation or inference task calls. All retrieval results are stored in the Redis database cache with an expiration time of 10 minutes to prevent interference from outdated knowledge in the scheduling task.
[0022] Preferably, step S16 is specifically: Step S161: Construct a semantic vector based on the semantic fusion data to obtain semantic vector data; In this embodiment, text entities and their corresponding context semantic fragments are extracted from the semantic fusion data obtained in step S15. For each semantic fragment, a word segmentation tool (such as the Jieba word segmentation tool) is used to perform word element segmentation, and a domain dictionary (such as the manually constructed "Industrial Control Field Term Library", which contains 5,000 terms in total) is combined for part-of-speech tagging and term mapping to ensure that the extracted words are effective terms with semantic representativeness. Then, these terms are encoded in a standardized manner, and a pre-constructed "Semantic Label Mapping Table" is used to map each term to a unified semantic label, such as "Equipment_Failure_Type", "Process_Control_Parameter", etc., a total of 200 categories of labels. When constructing the semantic vector, a sparse vector space with a fixed dimension of 1024 is used, and the TF-IDF value of each term in the entire semantic fusion corpus is calculated to fill the semantic vector. The TF calculation method is the ratio of the number of occurrences of each term in this semantic fragment to the total number of words in this semantic fragment; the IDF value is obtained by taking the logarithm after dividing the total number of documents containing this term in the corpus by the number of documents in which the term appears. For example, for the term "fault current", if it appears 3 times in a certain semantic fragment and the total number of words is 100, its TF is 0.03; if it appears in 4,000 semantic fragments in the total corpus and the total number of semantic fragments is 200,000, then the IDF is log(200,000 / 4,000)=1.7, and the final TF-IDF is 0.03×1.7≈0.051. During the construction process, the sparse matrix structure (scipy.sparse.csr_matrix) in NumPy is used to implement the storage and subsequent calculation of the semantic vector.
[0023] Step S162: Perform multi-source knowledge base similarity matching based on the semantic vector data to obtain candidate knowledge fragment data; In this embodiment, the multi-source knowledge base includes manufacturing operation and maintenance manuals (converted from PDF to plain text, with a quantity of 1,600), industrial production knowledge graphs (stored in the form of triples, with a total number of nodes of 350,000), and device sensing data logs (structured text, totaling 2TB). For each knowledge fragment, the same method is used to construct a 1024-dimensional semantic vector to ensure the consistency of the vector space. The similarity matching uses the cosine similarity calculation method, and the calculation formula is: similarity=(A·B) / (||A||×||B||), where A is the current semantic vector and B is the semantic vector of the fragment in the knowledge base. To improve the matching efficiency, the vector index acceleration tool Faiss (Facebook AI Similarity Search) is used to construct a vector inverted index, and top-k = 100 is set to return the top 100 knowledge fragments with the highest similarity to the input semantic vector. The matching threshold is set to 0.75, that is, if the similarity of a certain fragment to the input semantic vector is lower than 0.75, it will be automatically discarded and not included in the candidates.
[0024] Step S163: Perform semantic relevance ranking based on the candidate knowledge fragment data to obtain ranked knowledge fragment data; In this embodiment, the ranking criteria mainly consist of three parts: (1) semantic similarity score S1, directly provided by cosine similarity; (2) context keyword overlap rate S2, calculated as the ratio of the number of intersecting entries to the number of union entries between the semantic fusion data and the candidate fragment at the keyword level; (3) knowledge timeliness score S3, defined as the interval (in days) between the current time and the creation or most recent update time of the knowledge fragment, and mapped to a standard score (for example, if the time difference is 10 days, then S3 = 1.0; for 300 days, then S3 = 0.1). The total ranking score Score is calculated by the following formula: Score = 0.5 × S1 + 0.3 × S2 + 0.2 × S3.
[0025] Specifically, for example, if the semantic similarity of a certain candidate fragment is 0.88, the keyword overlap rate is 0.5, and the knowledge timeliness score is 0.9, then the final ranking score of this fragment is 0.5 × 0.88 + 0.3 × 0.5 + 0.2 × 0.9 = 0.44 + 0.15 + 0.18 = 0.77. After calculating the scores for all 100 candidate fragments using pandas and sorting them in descending order by the Score field, the Top20 are retained as the final ranked knowledge fragment data and stored in a structured JSON format with fields including "fragment ID", "semantic content", "source knowledge base ID", and "semantic score".
[0026] Step S164: Refine knowledge based on the ranked knowledge fragment data to obtain knowledge retrieval data.
[0027] In this embodiment, first, syntactic analysis is performed on the sentences in each knowledge fragment. By using a dependency syntactic analysis tool (such as Stanford NLP Parser), the subject-predicate-object relationship and prepositional phrases are extracted. The syntactic tree is matched through a custom extraction rule library (which contains a total of 600 rules, covering common structures such as "device - function - operation", "parameter - upper and lower limits - impact", "fault - cause - solution", etc.). For example, in the sentence "When the motor temperature exceeds 85°C, the load should be reduced for operation.", through syntactic analysis, "motor temperature" is identified as the subject noun phrase, "exceeds" as the predicate verb, "85°C" as the numerical modifier, and "should reduce the load for operation" as the resulting behavior. According to the rule "entity - condition - behavior", the triple (motor temperature, exceeds 85°C, reduce the load for operation) is refined. All the extracted knowledge content is organized in the form of triples of "entity - attribute - value" or "entity - relationship - entity", and according to the knowledge graph construction specification, the extraction results are written into the Neo4j graph database. Each entity node in the graph database is marked with a type (such as "device", "parameter", "action"), the relationship edge is labeled with a relationship type, and a confidence field is attached (the confidence is set by the rule matching level, 1.0 for first-level rule matching, 0.8 for second-level rule matching, and the default minimum is 0.5). All the finally extracted structured triples are used as knowledge retrieval data for the platform to call.
[0028] Preferably, in step S2, the construction of the result reflection layer is specifically as follows: Extract core keywords based on the knowledge retrieval data to obtain core keyword data; In this embodiment, a specified natural language processing toolkit (such as Jieba or NLTK) is used to perform word segmentation on the obtained knowledge retrieval data. By using a custom dictionary combined with part-of-speech tagging, the entries with part-of-speech being noun (n), gerund (vn), professional term (nz), and adjective (a) are used as candidate keywords. Subsequently, based on the TF-IDF scoring criterion, the text weight of each entry is calculated. In the weight calculation, the word frequency threshold is set to 1, and the inverse document frequency threshold is set to 0.1. Entries below this threshold are removed to reduce the influence of irrelevant words. Then, through the co-occurrence frequency calculation method of the context window, the window size is set to 5 words, the word pairs with a co-occurrence frequency higher than 3 times are screened out, and a word network is constructed according to the co-occurrence frequency. The high-frequency keywords that appear simultaneously in multiple knowledge retrieval data are selected from it, and finally, a set of no more than 50 core keywords is output as the input for subsequent analysis.
[0029] Draw a target semantic graph based on the core keyword data; In this embodiment, to draw a target semantic graph based on the core keyword data, a node connection matrix needs to be constructed. The nodes are each entry in the keyword set, the edges are co-occurrence relationships, and the edge weights are taken as the co-occurrence frequencies. The NetworkX library is used to construct an undirected graph structure. The weight threshold of the connecting edges is set to 3, and the connecting edges with weights lower than 3 are not constructed, thereby controlling the complexity of the semantic graph. The visualization of the nodes uses the force-directed layout algorithm (Fruchterman-Reingold algorithm), with the layout iteration count set to 50, the attraction coefficient set to 0.1, and the repulsion coefficient set to 0.8. In the semantic graph, different semantic clusters are distinguished by colors. The clustering method uses the Louvain algorithm, and the minimum community resolution parameter is set to 0.5. Based on the edge weights between nodes, a modularity-maximizing partition is generated, thereby forming several semantic subgraphs, which respectively represent different topic areas. Finally, the graph structure data and its visualization image of the semantic graph are output as the subsequent generation reference structure.
[0030] Execute the answer generation process based on the target semantic graph and record it as the output text to obtain the actual output data; In this embodiment, the question text input by the user is tokenized, part-of-speech tagged, and TF-IDF scored in the same way, and the keywords that overlap with the nodes in the semantic graph are extracted, and the matching weights are calculated. For the nodes with matching weights greater than the set threshold of 0.2, expand to the high-weight nodes within 2 hops along the edge structure of the semantic graph, collect the node entries in all expansion paths, and use these entries and their edge weights as the semantic constraint structure to input into the answer generation process. The answer generation process does not use a general model, but calls a preset template set, and fills in the matching keywords and the entry fragments in their semantic subgraphs in the structured template. For example, when the keywords are "agent management", "knowledge retrieval", and "semantic graph", select the structure template "For the management requirements of XXX, organize the ZZZ structure in the YYY way", and fill in the above three types of keywords and their expansion fragments respectively to form a structurally complete response text. The output text needs to record the semantic paths and the entries used for the source of each paragraph for subsequent metric calculations.
[0031] Calculate the keyword coverage rate based on the target semantic graph and the actual output data; Calculate the semantic matching degree based on the target semantic graph and the actual output data; In this embodiment, the output text is segmented and stop words are removed to obtain an output word set. Then, by comparing with the node entry set in the semantic graph, the ratio of the number of covered entries to the total number of node entries is calculated, and the coverage rate is expressed as a percentage. Specifically, the total number of nodes N is set, the number of covered nodes M is counted, and the keyword coverage rate is (M / N)×100%. Here, N is counted from the aforementioned semantic graph node set, and M is the number of words in the output text that can match the node words. The matching requirement is exact word matching without fuzzy matching. The coverage rate data is recorded in the system log table for jointly evaluating the output error with the semantic matching degree. The path structure between nodes in the semantic graph is compared with the order of keyword occurrences in the output text. First, a word pair sequence is constructed according to the order of keyword occurrences in the output text, and it is checked whether there is an edge connection between adjacent entries in the semantic graph. If there is, it is recorded as one path match. If not, but there is an indirect connection with a co-occurrence relationship edge distance within 2, it is recorded as one weak match. The total number S of all path matches and weak matches is counted, and the matching degree value is calculated with the length T of the output word pair sequence, that is, (S / T)×100%. If T is 0, the matching degree is 0%. The matching degree threshold is set to 50%. If it is lower than this threshold, it is considered that the output structure deviates from the main path of the semantic graph. The matching degree result is recorded in floating-point form and passed to the subsequent module.
[0032] Determine the output error based on the keyword coverage rate and the semantic matching degree to obtain output error data; In this embodiment, the error is defined as the weighted deviation value of the coverage rate and the matching degree. The weights are set as the coverage rate weight 0.4 and the matching degree weight 0.6. The formula calculation is not reproduced here, only the actual operation steps are described: the coverage rate and the matching degree are respectively input into the operation module, and a single floating-point output error value is generated according to the weight configuration. The error classification standard is: an error value less than 20% is a low error, between 20% - 50% is a medium error, and higher than 50% is a high error. According to the error level, the system extracts the corresponding language control parameters from the preset rule table. For example, the high error level corresponds to enabling parameter sets such as "structural word adjustment", "term reinsertion", and "syntactic repair", and each parameter set consists of multiple operation instructions.
[0033] Extract language control parameters based on the output error data; In this embodiment, the parameter table in the configuration file is called according to the error level. The parameter table is in JSON format, and each item corresponds to a type of language control operation, including parameter name, applicable error range, enabling threshold, and scope of action. For example, the "syntactic structure adjustment" parameter is set to be enabled when the error is higher than 30%, and the scope of action is the replacement and order adjustment of keywords in all subject-predicate structures; the "term supplementation" parameter is set to be started when the coverage rate is lower than 40%, and the supplementary items are selected from the nodes not covered in the semantic graph according to the weight. Each parameter call operation needs to indicate the start order, and the control instructions are sorted and executed in the call order array. After the parameters are extracted, they are not directly used for generation, but are passed into the structural output template construction process.
[0034] Construct a structural output control template according to the language control parameters, and construct a result reflection layer to obtain the result reflection layer data.
[0035] In this embodiment, the control template is a multi-level rule structure, and each template includes a title sentence template, a main body paragraph template, a supplementary term template, a conjunction template, etc. In the construction process, the template type is selected according to the language control parameters. For example, when the parameter contains "term supplementation", the main body template structure with redundant term paragraphs is called; if it contains "syntactic structure adjustment", the subject-predicate structure transformation rules are added to the template, such as subject advance, verb-object order change, etc. The conjunction template inserts segmented words such as "in addition", "compared with" according to the semantic clustering boundary. The constructed template is filled with variables and then a new text is output. The result reflection layer is a structured data set that records the template type, used parameters, semantic graph path, number of covered nodes, number of syntactic adjustments, and final output text length generated each time, forming an execution trace log for subsequent training optimization, and the output is a structured JSON array, including all generation component call details.
[0036] Preferably, in step S2, the construction process reflection layer is specifically: Identify the task execution path according to the knowledge retrieval data to obtain the task execution path data; In this embodiment, the behavior record data in each knowledge request process is extracted from the stored knowledge retrieval data. This data includes, but is not limited to, fields such as user input content, matching keywords, matching paths, called component names, response generation timestamps, returned content lengths, and executed component numbers. The operation process uses the Logstash module in the ELK log processing engine to complete the extraction of structured behavior data, and sorts the call sequences in each request according to timestamps to form an ordered task execution linked list. For example, in a knowledge retrieval task, when the keyword "agent deployment" is input, the system calls the keyword extraction module (component number A1), the semantic clustering module (component number B3), and the template filling module (component number C2), records the execution order of these three steps as [A1, B3, C2], and associates it with the input content and the returned content. The above operations are based on Elasticsearch to build an index structure, using the user request ID as the primary key field, and each task execution linked list is stored as a document node, finally forming a complete set of task execution path data, including no less than 500 task samples as the initial data basis.
[0037] Build a path decision tree based on the task execution path data to obtain path decision tree data; In this embodiment, the representation format of the unified task path is standardized, and each task path is serialized into a fixed format string according to the component numbers. For example, the path "A1 - B3 - C2" represents a complete path. Subsequently, the sklearn.tree.DecisionTreeClassifier tool in Python is used to perform a structural modeling on the path data. The input data is the path sequence and its execution performance labels, such as processing time (milliseconds) and whether the return is successful (boolean value). During the construction process, "information gain" is used as the division criterion, the maximum depth of the tree is set to 10, and the minimum number of samples in the leaf nodes is set to 5 to ensure that the path node division has statistical significance. All path branches use the component number as the split node, and each layer of splitting represents whether a certain component in the task process is called, thus forming a complete decision structure. The path decision tree data is exported in JSON format, which includes the component number, condition judgment, sub-path branch number, and corresponding task execution label of each decision node. This data is used for subsequent inefficient node identification operations to ensure that each sub-path structure has clear path statistics support.
[0038] Identify inefficient nodes based on the path decision tree data to obtain inefficient node data; In this embodiment, the operation of identifying inefficient nodes based on path decision tree data includes two sub-steps: one is to calculate the average execution time and failure rate of the corresponding nodes in each decision path. The execution time is directly calculated through the start and end timestamp fields returned by the component, and the failure rate is realized by counting the proportion of the return results being empty or the status code being abnormal; the other is to set a threshold standard to identify inefficient nodes. Nodes with an average execution time greater than 500 ms and a failure rate exceeding 0.3 are set as inefficient nodes, and the node numbers are recorded in a list. For example, if the average execution time of component B3 in all occurrence paths is 634 ms and the failure rate is 0.37, it is identified as an inefficient node. The implementation logic uses Pandas for path data aggregation and grouped statistical operations, and combines a boolean filtering function to extract the list of node numbers that meet the above two criteria. Finally, the inefficient node data is output and saved in the form of a list, excluding duplicates and non-task process nodes, ensuring that only the component numbers that truly affect the task execution efficiency are extracted.
[0039] Calculate the node call frequency based on the inefficient node data to obtain high-frequency inefficient node data; In this embodiment, the operation of calculating the node call frequency based on the inefficient node data needs to re-traverse all task execution path data and count the total number of times each inefficient node appears in the path. This operation is realized by traversing all path sequences, matching the component numbers in the path string, and using collections.Counter in Python for frequency statistics. The frequency threshold is set to more than 30% of the total number of all task paths. For example, if the total number of task paths is 1000, the inefficient nodes with a call frequency exceeding 300 times are high-frequency inefficient nodes. During the statistical process, the context component number pairs where each high-frequency inefficient node appears in different comparison paths are also recorded to form node context data. Finally, the high-frequency inefficient node data is output in the form of key-value pairs, where the key is the node number, and the value is the number of occurrences and the corresponding path context list, which is used to judge the cache replacement granularity and coverage range when configuring the cache policy.
[0040] Configure the result cache storage policy based on the high-frequency inefficient node data to obtain result cache data; In this embodiment, a cache space is established for each high-frequency and low-efficiency node. The cache structure is implemented using a Redis database. The cache key is composed of the component number of the node and the hash value of its input parameters. The SHA-256 algorithm is used to hash the input parameters, and the cache value is the historical return result text of the component under specific input conditions. The cache time is set to 300 seconds, the cache hit policy adopts the LRU eviction mechanism, and the cache capacity limit is no more than 1000 results per node. During the configuration process, all input condition caches with a hit frequency exceeding 10 times are set to be forcibly pre-loaded. The cache pre-loading uses an initialization script to load the cache database during the startup phase of the task engine. All cache configuration parameters are written to the Redis configuration file, with the path / etc / redis / cache_config.json, which includes the cache key template corresponding to each component number, the maximum number of cached items, the expiration time, and the LRU weight value. Finally, structured result cache data is formed, including the cache key-value policy, hit frequency records, and component binding rules, ensuring that high-frequency and low-efficiency nodes do not perform repeated calculations in subsequent processes. All cached results record the calculation source path and result timestamp.
[0041] Reconstruct the task process according to the result cache data and build a process reflection layer to obtain process reflection layer data.
[0042] In this embodiment, first, modify the logic of the task process execution engine. Replace the entry function of each call to the high-frequency and low-efficiency node component with a cache query entry. Use the GET interface of Redis to retrieve the cache hit situation before execution. If a hit occurs, directly return the cache result and skip the actual component call. The execution order of the modified task process is recorded as a new version of the flowchart. Use the Mermaid flowchart syntax to automatically generate a visual flowchart and store it in the task record system. At the same time, to build the process reflection layer, all cache hit data, actual component call data, task execution path and its changes need to be recorded for each task execution process. The data structure of the reflection layer is stored using a MongoDB document database. Each reflection document contains fields: task ID, original path, new path, list of hit cache node numbers, hit time, hit parameter values, identification of the component call function skipped, etc. All reflection layer data is aggregated and statistically analyzed once a day, and analysis result fields such as cache coverage rate, component call savings, and overall process compression rate are output for the system scheduler to use for the next process scheduling optimization. Finally, the process reflection layer data is output in JSON format, with the data precision reserved to the millisecond-level timestamp, and the document structure level does not exceed three levels of nesting to ensure database indexing efficiency.
[0043] Preferably, in step S2, constructing the policy reflection layer specifically includes: Extract business rules from the knowledge retrieval data to obtain business rule data; In this embodiment, based on the content of the archived knowledge retrieval request, a structured business operation mode is extracted. First, regular expression rules are used to tokenize the retrieval request statement and extract grammatical phrases. For example, from the retrieval statement "Query the production data report for the first quarter of 2023 and sort it by workshop classification", operation intention phrases such as "Query", "the first quarter of 2023", "production data report", and "workshop classification sorting" are identified. Then, classification is performed by setting an operation intention template, and the template structure includes four fields: "operation type + object + qualification condition + display format". The operation type field is centered around verb phrases, such as "Query", "Generate", "Compare"; the object field is uniformly an entity keyword, such as "report", "fault record"; the qualification condition field is composed of words such as time, location, and status; the display format field extracts keywords such as "chart", "list", "ranking", etc. Each structured template generates a unique identification number after extraction. For example, the rule ID is BR-20230613-001, and all the extracted business rules are stored in JSON format. Each record contains the template structure fields, source retrieval content, call frequency, trigger component path, business response parameter type, and associated context entity. The business rule data is stored in a MongoDB collection named "BusinessRules", and a keyword inverted index is constructed to support fast retrieval.
[0044] Obtain historical decision output records; In this embodiment, the operation of obtaining historical decision output records uses the decision execution log of the agent platform as the data source, and extracts the output result records of all task decision modules from the platform database, including fields such as input parameters, selected paths, decision nodes, output responses, timestamps, trigger contexts, feedback score values, etc. for each decision task. In the operation, an SQL script is used to query the table named "DecisionLog" in the PostgreSQL database, and the query statement is limited to records where the decision output field is not empty, the feedback score field is not NULL, and the execution status field is "success" or "fail". After all fields are extracted, they are uniformly formatted into a JSON object structure. In the structure, "input_context" is the task context text, "decision_result" is the recommendation item given by the system, "response_time" is the millisecond timestamp, "feedback_score" is a floating point number between 0 and 1, and "path_id" is the system path number. This data is sorted in time series and de-duplicated to ensure no duplicate records, and finally a set of historical decision output record data with no less than 100,000 records is obtained. All data is checked for field mapping before being imported into the analysis environment to ensure consistent key names and unified unit standards (time in milliseconds, scores as decimals, paths as standard numbers), and is stored in a MongoDB collection named "DecisionHistory", with an index established using the path number as the primary key.
[0045] Cognitive bias analysis is performed based on the historical decision output records and business rule data to obtain cognitive bias data; In this embodiment, the operation process of cognitive bias analysis based on historical decision output records and business rule data is divided into two steps. The first step is to identify multiple historical decision outputs corresponding to the same business rule and count their output consistency in different contexts. The specific operation is to traverse the rule IDs of each rule in the business rule data, filter out all records corresponding to the rule ID in the historical decision records, and count the distribution frequency of the corresponding decision result field (decision_result). For example, if a certain rule ID has a decision output of "Solution A" 25 times, "Solution B" 15 times, and "Solution C" 10 times in 50 records, the output deviation degree calculation formula is: Deviation degree = 1 - (maximum frequency / total frequency) = 1 - 0.5 = 0.5. Set the deviation degree threshold to 0.3, and if it exceeds this threshold, it is considered that there is a cognitive bias. The second step is to perform semantic distance analysis on the deviation output records, call the Jaccard text similarity algorithm to compare the similarity between different input contexts. If the similarity is greater than 0.8 but the corresponding decision outputs are different, it is marked as a potential cognitive bias item. The deviation record data structure includes business rule ID, deviation degree value, context example pair, output difference type (suggestion conflict, execution path split, etc.) and occurrence times. All deviation items are recorded in the "CognitiveBias" set, and each record includes a unique number, the corresponding business rule ID, statistical timestamp, context difference description field, and output conflict example field.
[0046] Generate knowledge structure update requirements based on cognitive bias data to obtain knowledge structure update requirement data; In this embodiment, in the operation steps of generating knowledge structure update requirements based on cognitive bias data, each record in the cognitive bias data needs to be parsed into a unit that requires knowledge update. First, parse the output difference field in each cognitive bias record, extract the conflict items as explicit knowledge nodes. For example, "Solution A" and "Solution B" are used as decision output nodes, and construct the semantic conflict relationship between them to form a knowledge update requirement pair. Secondly, extract the keywords in the context field, and identify whether there are associated entity nodes in the current knowledge graph. If not, it is used as a new concept node requirement; if there are, but the connection edge attributes conflict, it is used as an edge attribute modification requirement. The knowledge structure update requirement data format is encoded using the RDF triple structure. The triple consists of "subject-predicate-object", where the subject is the context entity node, the predicate is the relationship such as "recommended output", "dependency rule", "generate conflict", etc., and the object is the target decision result or recommendation item. All update requirements are marked with update type (add, delete, replace), impact path, context source, and deviation number. The update requirement data is uniformly written into the pending processing area of the Neo4j graph database named "KnowledgeUpdateRequest", and the update flag is set to "not executed".
[0047] Convert the knowledge structure update requirement data into graph reconstruction requirement data; In this embodiment, perform a dependency path analysis on each knowledge update request record to identify the node range and edge connection structure affected in the graph structure. Use the depth-first search algorithm to traverse the graph nodes, calculate the path coverage range of the target nodes in the existing graph in the knowledge update request, set the maximum path depth to 3 levels, and automatically exclude nodes outside the range to ensure that the reconstruction requirements are controlled within the scope of the local subgraph. Secondly, label the edges with attribute conflicts and record the starting nodes, edge types, old attribute values, and new proposed values of the conflicting edges. Finally, encapsulate the IDs of all graph nodes to be modified, edge IDs, and operation types into a reconstruction task unit. The data structure includes: task ID, subgraph node set, edges to be deleted set, edges to be added set, edge attributes to be modified set, estimated update timestamp, and version control flag. The graph reconstruction requirement data is exported in YAML format, including clear reconstruction instructions, graph module numbers, namespaces, and relationship types to ensure parsing in the graph engine. All graph reconstruction tasks are written into a task scheduling table named "GraphRebuildTasks", and the status is initialized to "waiting to be executed". Each task needs to set a version rollback flag to support fault recovery.
[0048] Adjust the knowledge graph structure according to the graph reconstruction requirement data and construct a strategy reflection layer to obtain the strategy reflection layer data.
[0049] In this embodiment, call the reconstruction execution engine in the graph management module and execute the subgraph reconstruction operations one by one according to the reconstruction instructions recorded in the task scheduling table. Before each subgraph reconstruction execution, first perform a graph snapshot backup. The backup is completed using the export module provided by the Neo4j graph database, and the file is named "graph_snapshot_task ID.graph". Subsequently, submit transactions one by one in the graph database for the operations of adding edges, deleting edges, and modifying attributes listed in the reconstruction task. The transaction execution adopts the ACID mode to ensure structural consistency. Immediately after completing the graph structure reconstruction, construct the strategy reflection layer data. The reflection layer records the deviation numbers, reconstruction task IDs, the number of modified nodes, relationships, backup snapshot paths, trigger rule IDs, expected change ranges, and actual execution impact path lists involved in this reconstruction. All reflection layer data is uniformly written into the "StrategyReflectionLayer" collection in the MongoDB database, and each record is bound to a unique strategy number and graph update version number. Run the reflection layer data aggregation script every 24 hours to count evaluation dimensions such as the source frequency of various update requests, trigger rule coverage, and the number of path changes after the update for subsequent strategy fine-tuning and version control to ensure closed-loop management of the reconstruction and reflection processes in terms of structure.
[0050] Preferably, step S3 is specifically as follows: Step S31: Based on the result reflection analysis of the three-level reflection system data row results, obtain result reflection data; according to the result reflection data, restore the result reflection path to obtain result reflection path data; In this embodiment, the three-level reflection system data output from the strategy reflection layer is used as the input data source, including the first-level strategy reflection record (used to describe the original strategy output behavior), the second-level intermediate state monitoring record (used to depict the state changes of key nodes during strategy execution), and the third-level external feedback comparison data (used to record the differences between the external environment or standard output and the strategy output). The three-level data set is subjected to a structured mapping process, and the structured vector coding method is used to map it into a state vector S1, a behavior path vector S2, and a feedback error vector S3 respectively, and aligned in chronological order. Through bidirectional time window sliding analysis, the window size is set to 10 consecutive records, and the overlap rate is 60%. Each window unit performs a behavior-feedback matching ratio analysis, calculates the vector Euclidean distance between the strategy output and the feedback error, and records it as the error amplitude. The behavior paths with an error amplitude exceeding 0.7 (this threshold is determined by the 95% quantile statistically obtained from 100,000 historical data sets) are extracted as result reflection points. Then, through a path tracking mechanism (using the node number in the behavior path as the primary key and adopting a breadth-first search method), a path set from the current reflection point to the original decision starting point is constructed, named result reflection path data. This path set contains the behavior label, state feature vector, upstream and downstream connection node IDs, and time stamps of each node, and is stored in a directed graph structure to form a complete result reflection path data set.
[0051] Step S32: Based on the result reflection path data, identify strategy deviations to obtain strategy deviation data; according to the strategy deviation data, perform prompt word structure reorganization to obtain prompt word structure data; according to the prompt word structure data, fuse multi-level control semantics to obtain prompt word optimization data; In this embodiment, the result reflection path data output from step S31 is used as input, and deviations are first identified for all strategy branch nodes in the path. This identification is defined as the cosine similarity difference between the state vector of the actual behavior output and the expected target state vector by setting the strategy deviation index P_dev. The specific calculation method is: P_dev=1-cos(θ). When P_dev is greater than 0.25 (determined by the intersection of the historical deviation classification error curve), it is determined to be a strategy deviation node. The deviation nodes and their corresponding behavior semantic labels in all paths are counted to form a strategy deviation data set. For the strategy deviation data, the prompt word structure reorganization operation is performed. The deviation node semantic label is mapped with the upstream semantic control unit (such as "plan", "scheduling", "retrieval") to which it belongs, and the structure rewriting rules defined in the syntactic rule library are used. For example, if the deviation point is "resource allocation failure", the original prompt word structure "execute task allocation" is rewritten as "recalibrate the resource pool and then execute task allocation". All structure rewriting operations are performed based on the predefined dependency syntactic graph mapping rule library, without using a model, and the structure is adjusted after directly matching the semantic template. Finally, the reconstructed prompt word structure data is fused with multi-level control semantics (including three semantic dimensions such as operation class, constraint class, and target class) through the semantic fusion rule table to generate prompt word optimization data. The control semantic fusion process requires that each prompt word must have an "action-condition-target" triple expression at the same time, otherwise it will be marked as "semantically incomplete" and not output.
[0052] Step S33: Perform node load balancing according to the three-level reflection system data to obtain node load balancing data; In this embodiment, the system nodes involved in the execution of all strategies are classified into four categories according to their roles: decision processing nodes, information retrieval nodes, prompt generation nodes, and knowledge fusion nodes. The three load indicators of each node, namely, call frequency, processing time, and input and output data volume, are extracted and recorded as F (frequency), T (duration), and D (data volume), respectively, with units of times / hour, seconds, and MB. After normalization, the node load vector Vn=[F',T',D'] is constructed, where F'=F / Max(F_all), and the rest are analogous. The load vectors of all nodes are sent to the outlier detection module based on the standard deviation distribution, and the load fluctuation coefficient threshold is set to 0.15 (that is, if the standard deviation difference between a node and other nodes is greater than 0.15, it is marked as high load), node classification is performed, and a low-load node set and a high-load node set are output. The load transfer operation is performed on the high-load node. The transfer rule is: if the functional roles of the two nodes are consistent and logically replaceable (the judgment method is based on the node function label matching degree greater than 0.8), the original binding node identifier in the task flow is replaced with the target node identifier with a lower load, forming an adjusted task pointing graph structure and outputting the node load balancing data.
[0053] Step S34: Identify task nodes based on node load balancing data to obtain task node data; identify task dependency relationships based on the task node data to obtain task dependency relationship data; In this embodiment, according to the task pointing graph structure after load balancing adjustment in step S33, identify the logical start node and end node therein, and traverse the task graph by level. The depth-first search algorithm is used to start from the start node, record the visited nodes layer by layer, and mark the task operation label and task ID. Each time a complete path record is completed, a task node data unit is generated. Each task node data structure includes fields such as task ID, task operation type (such as "retrieval trigger", "rule analysis", "semantic fusion", etc.), the ID of the node to which it belongs, the set of predecessor task IDs, the set of successor task IDs, the processing duration threshold (calculated from the historical average plus the standard deviation), and the resource requirement level. On this basis, perform dependency analysis on all task nodes. Through the graph topological sorting method, judge the front and back constraint relationships between tasks. If task A has an output field that depends on the processing result of task B, record that A depends on B to form a task dependency relationship edge set. The dependency determination criteria include field name matching (edit distance less than 3) and field type consistency, and the value range of the field needs to have an intersection as a prerequisite for establishing the dependency relationship. Finally, output the task node data and the task dependency relationship data set.
[0054] Step S35: Reconstruct the task chain according to the task dependency relationship data to obtain task chain data; In this embodiment, based on the task dependency graph structure, identify all subgraphs with redundant branches and loop paths. The redundant branch identification rule is: if two task nodes A and B satisfy the output field consistency and the processing logic similarity is greater than 0.85 (judged according to the operation label matching table), and both finally point to the same target task C, they are merged into one branch. The loop path identification rule is: if there is a closed-loop structure in the graph where task A depends on B, B depends on C, and C depends on A, then set the node with the minimum processing duration as the loop break point and disconnect the one-way dependency. After cleaning up the redundancy and loops, reorder all task nodes, use the topological sorting + critical path algorithm to calculate the total processing time of each path, select the path with the longest time as the main task chain, and record the relative path structure of all branch task chains, mark the dependency level and trigger conditions, and finally output the task chain data set. The task chain structure is stored in JSON format, including field information such as chain ID, task order, time consumption per step, required nodes, and execution resource level.
[0055] Particularly importantly, step S35 includes the following steps: Step S351: Parse the task nodes according to the task dependency relationship data to obtain task chain node data; In this embodiment, it is necessary to obtain structured data containing all tasks and their dependencies. This data is usually stored in the form of a directed graph, where nodes represent specific tasks and edges represent the dependencies between tasks. During the parsing process, graph traversal techniques (such as depth-first traversal or breadth-first traversal) are used to traverse the task dependency graph to identify all independent task nodes and their predecessor and successor nodes. Each task node contains a unique task identifier, task type, estimated execution duration, input and output parameters, and a list of dependent tasks. The task node information is extracted through traversal operations and stored as a structured task chain node data table. The fields of this table include task ID, list of dependent task IDs, estimated execution time (in milliseconds), task type code, etc. During the parsing process, the data integrity is strictly verified to ensure that there are no isolated nodes and all dependencies are clear, and there are no loops. After the parsing is completed, the obtained task chain node data can be directly used for subsequent sorting and path analysis.
[0056] Step S352: Perform priority sorting based on the task chain node data to obtain sorted task node data; In this embodiment, first, a priority value is calculated for each task node. The priority calculation uses a weighted algorithm based on the task dependency depth and task urgency. The specific method is as follows: Calculate the dependency level depth of each node. The greater the level depth, the higher the priority. The urgency is calculated by the difference between the preset task deadline and the estimated task execution time. The smaller the difference, the higher the urgency. The priority value is the linear weighted sum of the level depth and urgency, with weights set to 0.7 and 0.3 respectively. The priority values of all task nodes are sorted. The sorting method uses the heap sort algorithm to ensure a time complexity of O(nlogn), where n is the number of task nodes. After the sorting is completed, a sorted task node data structure is generated, which contains the task ID and the corresponding priority value for subsequent path analysis. The priority sorting ensures that tasks with high dependency levels and approaching deadlines are processed first.
[0057] Step S353: Perform path analysis based on the sorted task node data to obtain task chain path data; In this embodiment, by using the task dependency graph combined with the priority sorting result, starting from the starting node, a topological sorting combined with dynamic programming method is used to determine the task execution path. Topological sorting is used to ensure that the execution order satisfies the dependency constraints, and dynamic programming is used to calculate the cumulative execution time and resource consumption of each path to ensure that the path selection takes into account both efficiency and resource utilization. During the analysis process, the sorted task node list is traversed, and the earliest start time and latest completion time of each node are updated in turn, and the predecessor task path of each task node is recorded. Using these time parameters, the critical tasks and bottleneck tasks in the path are identified. The path analysis result forms task chain path data, which includes fields such as task execution order, path duration, critical path identifier, etc. This data structure is stored in the form of a directed acyclic graph for subsequent execution order reconstruction.
[0058] Step S354: Reconstruct the execution order according to the task chain path data to obtain the reconstructed task chain data; In this embodiment, dependency verification and order optimization are performed on the path data. In the verification step, it is checked whether the execution order of all tasks satisfies the task dependency constraints, and the topological sorting algorithm is used to confirm the legality of the path. For order optimization, a task scheduling method based on a heuristic algorithm is adopted to adjust the arrangement order of tasks on the time axis, reducing idle waiting time and resource conflicts. The heuristic algorithm parameters include the task switching delay time (set to 5 milliseconds), the task execution priority weight (calculated based on step S352), and the task resource occupancy rate (obtained from the resource requirement data, and the threshold is set to 80% resource occupancy, which is regarded as high load). By repeatedly iterating and adjusting the task execution sequence, it is ensured that the execution time window of each task in the task chain maximally utilizes system resources, and finally, the optimized reconstructed task chain data is formed. This data details the optimized task execution order and the corresponding timestamps.
[0059] Step S355: Perform structure optimization according to the reconstructed task chain data to obtain the task chain data.
[0060] In this embodiment, when performing structure optimization according to the reconstructed task chain data, the focus is on adjusting the node connection relationship and the overall structure in the task chain, and the minimum spanning tree and task merging techniques in graph theory are adopted. Structure optimization first identifies redundant dependency relationships and task fragments in the task chain. By analyzing the task execution time and resource occupancy, the task nodes are merged, and the merging conditions include that the task execution time difference is less than 10 milliseconds and the dependency relationships are the same. Secondly, the minimum spanning tree algorithm is used to reconnect the task nodes, reducing the repeated dependency edges in the execution path and optimizing the task chain topological structure. The parameters used in the structure optimization process include the maximum allowable task delay (set to 15 milliseconds), the node merging threshold, and the dependency redundancy determination criterion (a dependency edge weight lower than 0.1 is regarded as redundant). After the optimization is completed, the task chain data is generated, which describes the optimized task chain nodes, dependency relationships, task execution order, and related performance indicators, supporting the efficient scheduling and execution of the intelligent agent management platform.
[0061] Step S36: Detect the resource consumption according to the task chain to obtain the resource consumption data.
[0062] In this embodiment, it is first preset that the resource types include four categories: CPU occupancy rate, memory usage, disk I / O read and write volume, and network bandwidth consumption. A basic unit consumption table is set for each type of resource. For example, the average CPU consumption per single "knowledge retrieval operation" is 0.5 cores, the memory consumption is 120 MB, the disk read and write is 10 MB / s, and the network bandwidth occupancy is 0.1 MB / s. This data is obtained by statistically averaging the hardware resource monitoring indicators of nearly ten thousand operation log records. Map the operation types of each task node in the task chain to obtain its standard resource consumption indicators, and correct them in combination with actual historical data (the correction factor depends on the server specifications to which the node belongs. If it is a high-performance node, the resource consumption correction coefficient is set to 0.85). Accumulate the usage amounts of the four types of resources in sequence according to the node order for the entire task chain, and store the results in the resource consumption data table. Each record includes fields such as task chain ID, total CPU core hours, total memory in MB, total I / O read and write volume in MB / s, and network bandwidth in MB / s. The resource detection results are used as input items for the resource pool allocation parameters in the subsequent scheduling optimization module.
[0063] Preferably, step S33 is specifically as follows: Step S331: According to the data of the three-level reflection system, statistically count the execution status of the nodes to obtain the node execution status data; In this embodiment, the system comprehensively statistically analyzes the data from the three-level reflection system through the operation log collection module deployed in each node of the intelligent agent management platform. The collected data includes fields such as task ID, task start time, completion time, execution status code, number of occupied CPU cores, memory usage, and I / O load. The log data is written into the centralized management database in JSON format. Through aggregating and analyzing the data of different node numbers according to time intervals, use the built-in data processing tools in Python to statistically operate on the number of tasks, completed tasks, failed tasks, average task duration, and average resource consumption of each node, and store the statistical results in an independent form named "node execution status data table". Each node contains at least five statistical indicators, including the total task volume, successful task volume, failed task volume, average execution duration, and average resource usage. This status data table will be used for load rate analysis in the next step.
[0064] Step S332: Calculate the node resource load rate based on the node execution status data; In this embodiment, the system calls the resource monitoring module deployed on each node to obtain the resource usage status of the node, including the current CPU usage rate, memory usage rate, and the number of currently active threads. This data is collected by polling every second for a continuous period of five minutes. The formed data structure uses the node number as the primary key, accompanied by time-series resource occupancy data. The platform uses the time-window sliding average method to process the collected data and extracts the five-minute average usage rates of the CPU and memory respectively. At the same time, in combination with the average task execution duration and the number of completed tasks obtained in step S331, the system compares the resource processing capacity and resource usage of each node within the time interval. The system marks this comparison value as the resource load rate of the node. After calculating the load rates of all nodes, they are uniformly stored in the node load rate dataset and prepared for subsequent priority calculations.
[0065] Step S333: Calculate the scheduling priority based on the node resource load rate to obtain the node scheduling priority data. In this embodiment, the system uses the node resource load rate data obtained in step S332 to divide the scheduling priorities of the nodes according to the preset load rate hierarchical intervals. The lower the load rate of a node, the higher the scheduling priority assigned to it. The system sets four hierarchical intervals and defines them in the platform configuration file. All nodes are automatically assigned scheduling priority levels according to this hierarchical standard. At the same time, the system will mark the current status labels of each node during the process of assigning scheduling priorities, including "idle", "lightly occupied", "heavily occupied", and "resource critical", and record the priority level and status label in the node scheduling priority dataset together. This dataset is used as a key reference table in the task allocation strategy and is set to be refreshed only when the resources change significantly.
[0066] Step S334: Perform agent task allocation according to the node scheduling priority data to obtain the agent task allocation data. In this embodiment, the system performs task allocation operations based on the node scheduling priority data and the list of tasks to be executed. The task list consists of three types: knowledge retrieval, content generation, and solution optimization. Each task is marked with the computing resource requirement level and the task completion time limit requirement during the task submission stage. The system enables the matching rule table in the task allocation module and preferentially matches the nodes that are currently idle or have a low load and have a higher priority level according to the resource requirement level and completion time requirement of the task. If the priority node can carry multiple tasks, a polling strategy is used for parallel distribution. If the matching fails, an attempt is made to recursively allocate to the nodes at the next lower scheduling level. After the task allocation is completed, the mapping relationship between all tasks and nodes is written into the task scheduling mapping table and synchronized to the task distribution module for the task scheduling engine to call.
[0067] Step S335: Perform node load analysis based on the agent task allocation data to obtain node load balancing data.
[0068] In this embodiment, after the task allocation is completed, the system monitors the actual task processing status of each node in real time, re-collects the node resource status data before and after the task is issued, and performs a differential analysis operation. The analyzed fields include CPU usage change, memory usage change, thread growth number, and average task execution duration change. The resource fluctuation data of all nodes is processed by an analysis script to form a node load deviation data set. The system calculates the current load distribution deviation level of each node by comparing the task quantity and resource usage deviation value of each node, and finally forms a node load balancing data table. This data table is used to detect whether there is a problem of uneven node load. If there is a deviation node exceeding the preset threshold, the node number and the current task ID set will be recorded for subsequent reference by the load migration mechanism or the task fallback mechanism.
[0069] Preferably, step S4 is specifically as follows: Step S41: Identify the task resource type based on the resource consumption data to obtain task resource type data; In this embodiment, the system first calls the resource consumption data stored in the resource monitoring module. This data includes detailed indicators such as task ID, CPU occupancy time (unit: millisecond), GPU occupancy time (unit: millisecond), memory usage (unit: MB), and network bandwidth occupancy (unit: Mbps). By parsing each indicator of each task in the resource consumption data, the resource type determination rules set by the rule engine are used to identify the task resource type. The specific rules are as follows: A task with a CPU occupancy time exceeding 500 milliseconds and a GPU occupancy time lower than 100 milliseconds is determined to be CPU-intensive; a task with a GPU occupancy time exceeding 500 milliseconds and a memory usage lower than 512 MB is determined to be GPU-intensive; a task with a memory usage exceeding 1024 MB is determined to be memory-intensive; a task with a network bandwidth occupancy exceeding 100 Mbps is determined to be network-intensive. The system automatically generates task resource type data according to these rules, records the mapping relationship between the task ID and the resource type, and the data format is JSON. Each record contains the task ID, the resource type identification field, and the corresponding threshold indicators for the next step of computing power demand quantification.
[0070] Step S42: Quantify the computing power demand according to the task resource type data to obtain task computing power demand data; In this embodiment, the system quantifies the computing power requirement by using the empirical parameter method based on the task resource type data obtained in step S41, combined with the task execution duration and the peak resource consumption data. Taking CPU-intensive tasks as an example, the computing power requirement is quantified by calculating the product of the CPU occupancy time (milliseconds) and the node CPU frequency (GHz) to obtain the computing power requirement value. The threshold is set as 1000 GHz·ms for low demand, 3000 GHz·ms for medium demand, and above 5000 GHz·ms for high demand; for GPU-intensive tasks, the computing power requirement is calculated by the product of the GPU occupancy time and the GPU core frequency. For memory-intensive tasks, they are classified according to the peak memory usage. Less than 2048 MB is low demand, 2048 MB to 8192 MB is medium demand, and more than 8192 MB is high demand; for network-intensive tasks, the computing power requirement is calculated by the bandwidth occupancy and the task transmission data volume. Less than 200 Mbps·MB is low demand, 400 Mbps·MB is medium demand, and above 800 Mbps·MB is high demand. The system stores the quantified results of the computing power requirements of all tasks in the task computing power requirement data table. The data fields include the task ID, the computing power requirement value, and the demand level, and are stored in CSV format for quick query and scheduling.
[0071] Step S43: Perform heterogeneous hardware resource matching according to the task computing power requirement data to obtain preliminary hardware matching data; In this embodiment, the system uses the task computing power requirement data generated in step S42 to match the heterogeneous hardware resource library supported by the agent management platform. This resource library details information such as the CPU model, frequency, number of cores, GPU model, frequency, number of CUDA cores, memory capacity and bandwidth, and network interface bandwidth of each node. The matching rule filters the resource library in sequence according to the task computing power requirement level. First, it screens the hardware configuration nodes that meet the demand level; for CPU-intensive tasks, it screens the nodes with a CPU frequency of not less than 3.0 GHz and the number of cores of not less than 8; for GPU-intensive tasks, it screens the nodes with a CUDA computing power support of not less than 7.0 and the video memory of not less than 8 GB; for memory-intensive tasks, it screens the nodes with a memory capacity greater than 16 GB; for network-intensive tasks, it screens the nodes with a network bandwidth greater than 1 Gbps. The system automatically extracts the set of nodes that meet the conditions from the resource library through a query statement to form preliminary hardware matching data. The data storage structure is a table, and the fields include the task ID, the matching node ID, and the snapshot of the matching hardware parameters (CPU frequency, number of cores, GPU model, memory size, etc.), which are used for subsequent simulation analysis.
[0072] Particularly importantly, step S43 includes the following steps: Step S431: Perform hardware resource capacity analysis according to the task computing power requirement data to obtain hardware resource capacity data; In this embodiment, key performance indicators are extracted from the task computing power demand data, including but not limited to the number of CPU cores, GPU computing power (measured by floating-point operation performance GFLOPS), memory bandwidth (in GB / s), storage I / O speed (in MB / s), and network transmission rate (in Gbps). The actual performance parameters of all current heterogeneous hardware resources are collected using system resource monitoring interfaces such as the / proc file system in Linux or hardware management interfaces (such as IPMI). Based on these parameters, compared with the threshold criteria required by the task, such as the number of CPU cores not less than 8 cores, GPU computing performance not less than 5 TFLOPS, memory bandwidth at least reaching 100 GB / s, storage I / O speed greater than 500 MB / s, and network transmission rate at least 10 Gbps, the hardware resources that meet the task computing power requirements are screened out to form a hardware resource capability data table. The data table details the model, performance indicator values, current usage rate, and idle resource capacity of each hardware resource for subsequent screening.
[0073] Step S432: Screen available heterogeneous hardware resources according to the hardware resource capability data to obtain a list of candidate hardware resources; In this embodiment, according to the matching degree between the performance indicators recorded in the hardware resource capability data and the task computing power requirements, screening rules are set, including performance threshold satisfaction, resource idle rate greater than 30%, load rate lower than 70%, etc. as screening criteria. The traversal algorithm is used to evaluate the hardware resource capability data item by item, and the hardware resources that meet all screening rules are added to the list of candidate hardware resources. To ensure load balancing, the physical location and network topology of the resources are also considered during screening, and nodes with low latency and high bandwidth are preferred. The screening results are stored in a list form, including hardware resource identifiers, specific performance indicators, current resource status, percentage of matching with task requirements, and physical topology location for subsequent matching.
[0074] Step S433: Perform resource performance matching according to the list of candidate hardware resources to obtain resource matching data; In this embodiment, the performance indicators of each resource in the list of candidate hardware resources are compared item by item with the task computing power demand data, and a weighted scoring algorithm is used to calculate the matching degree. The weight assignment is based on the importance of different performance indicators for the task. For example, the CPU performance weight is 0.3, the GPU performance weight is 0.4, the memory bandwidth weight is 0.2, and the storage I / O weight is 0.1. During the matching degree calculation process, the actual hardware performance value is normalized to the 0-1 interval, compared with the normalized task requirement value, the matching score for each indicator is obtained, and then weighted summation is performed by multiplying by the weight. Finally, the resource matching score is obtained, ranging from 0 to 1. After completing this calculation for all candidate hardware resources, the results are summarized to form a resource matching data table, recording the hardware resource identifier, matching scores for each performance indicator, and comprehensive matching score.
[0075] Step S434: Perform priority sorting on the hardware resources according to the resource matching data to obtain sorted hardware resource data; In this embodiment, a descending sorting algorithm based on the resource matching score is adopted to sort each hardware resource in the resource matching data table in descending order of the comprehensive matching score. The sorting algorithm uses quicksort with a time complexity of O(nlogn) to ensure high efficiency when processing a large number of hardware resources. After sorting, the result forms a sorted hardware resource data structure, which includes information such as the unique identifier of the hardware resource, the corresponding matching score, the details of the performance indicators, and the current load status of the resource, preparing for subsequent resource allocation.
[0076] Step S435: Perform resource allocation planning according to the sorted hardware resource data to obtain preliminary hardware matching data.
[0077] In this embodiment, the hardware resources are allocated in sequence according to the sorting result, and the resources with higher comprehensive matching scores are preferentially allocated to the task sub-modules with higher computing power requirements. The allocation process follows the principles of resource capacity limitation and load balancing to ensure that the load rate of a single hardware node does not exceed 85%. The greedy allocation algorithm is adopted. According to the numerical value of the task computing power requirement, the capacity of the allocated resources is gradually subtracted until the computing power requirements of all task modules are met. After the allocation is completed, preliminary hardware matching data is formed, which details the hardware resource identifier, the allocated computing power share, and the remaining resource capacity corresponding to each task module, ensuring that this matching scheme can be directly called during the next task execution simulation.
[0078] Step S44: Perform task execution simulation based on the preliminary hardware matching data to obtain task execution data; perform thread blocking point analysis based on the task execution data to obtain thread blocking data; In this embodiment, the system performs task execution simulation based on the preliminary hardware matching data in Step S43. During the simulation process, the node simulator is called to simulate the task execution process relying on the real hardware parameters, including details such as the task start time, execution time slice allocation, and thread scheduling. The discrete event simulation technology is used for the simulation, and the input parameters include the task computing load (computing power requirement value), the number of threads, the node hardware configuration, etc., and the output indicators include the task execution time, the CPU / GPU time occupied by each thread, and the waiting time, etc. Through the simulation results, the system analyzes the thread blocking points, collects the blocking time, blocking reasons (such as resource contention, waiting for I / O, etc.) and occurrence frequencies of each thread during the simulated execution. The blocking point analysis forms thread blocking data by scanning the thread execution log, extracting the blocking event timestamp and context, and combining the CPU core occupancy and memory access conflict information. The data structure includes the thread ID, blocking type, blocking duration, frequency, and the node where the blocking occurs, and the storage format is XML for easy subsequent parsing.
[0079] Step S45: Adjust the thread binding of the preliminary hardware matching data according to the thread blocking data to obtain thread binding data; perform heterogeneous hardware collaborative scheduling according to the thread binding data to obtain heterogeneous hardware collaborative scheduling data; In this embodiment, the system adjusts the thread binding of the preliminary hardware matching data according to the thread blocking data obtained in step S44. During the adjustment process, the system analyzes the mapping relationship between the blocking points and the hardware resources, and for the threads with a higher blocking rate, rebinds them to the computing cores or independent hardware units with a lower blocking rate to reduce resource contention. The binding rule depends on the blocking duration threshold, and the threads blocked for more than 50 milliseconds must be rebound. The binding operations include thread migration and priority adjustment. After completing the thread binding adjustment, the system performs heterogeneous hardware collaborative scheduling on the adjusted thread allocation status. The scheduling module, based on the thread binding data, comprehensively considers the hardware performance metrics and thread priorities to achieve the coordination of task execution across nodes and multiple hardware resources. The collaborative scheduling process records the scheduling path, resource allocation ratio, and load balancing metrics, forming a heterogeneous hardware collaborative scheduling data table, which includes the task ID, thread ID, bound hardware unit, and scheduling time slice distribution, and is stored in JSON format for management.
[0080] Step S46: Construct an agent management platform according to the heterogeneous hardware collaborative scheduling data and the prompt word optimization data, and adaptively adjust the knowledge retrieval strategy of the agent management platform to obtain adaptive strategy data.
[0081] In this embodiment, the platform maps the hardware resource utilization rate and thread scheduling efficiency metrics in the scheduling data to the semantic hierarchical information in the prompt word structure, defines a mapping rule table, and uses the node response time and load balancing degree in the scheduling data as trigger conditions to dynamically adjust the retrieval range and depth of the knowledge retrieval module. The adaptive adjustment process includes parsing the scheduling data, extracting key performance metrics, combining the multi-level semantic fusion results in the prompt word optimization data, and reconfiguring the retrieval index weights and filtering thresholds. The adjustment result generates an adaptive strategy data file, which includes the strategy version number, adjustment parameter values, corresponding task categories, and timestamps, and the file format is YAML, which is used for the real-time scheduling of subsequent tasks and the update of the retrieval strategy in the agent management platform.
[0082] Preferably, step S46 is specifically: Step S461: Analyze the node distribution according to the heterogeneous hardware collaborative scheduling data to obtain node distribution data; divide the scheduling units based on the node distribution data to obtain scheduling unit data; perform module resource mapping according to the scheduling unit data to obtain module mapping data; In this embodiment, the system first performs node distribution analysis for heterogeneous hardware collaborative scheduling data execution. The data used includes node identifiers, task assignment situations, hardware resource utilization rates, and time series load change information. The node distribution analysis calculates the node load balancing index by counting the number of tasks, resource occupancy rates, and task execution times of each node. The load balancing index is calculated by the weighted average of the node CPU utilization rate, GPU utilization rate, and memory usage rate, with the weights set to 0.5, 0.3, and 0.2 respectively. When the node load balancing index exceeds 0.8, it is marked as a high-load node, and when it is lower than 0.3, it is marked as a low-load node. Based on the node load situation and the geographical network topology relationship, a clustering algorithm (such as K-means, with the number of clusters set to the integer part of the square root of the total number of nodes) is used to group the nodes to obtain node distribution data. The data structure is indexed by the node group ID and records the node list, the load balancing index of the nodes within the group, and the average network delay. Subsequently, the system performs scheduling unit division based on the node distribution data. The scheduling unit division rule is based on the load balance within the node group and the network delay threshold. Nodes with a network delay not exceeding 5 milliseconds and a load index difference not exceeding 0.15 are divided into the same scheduling unit. The scheduling unit data records information such as the unit ID, the set of nodes it belongs to, and the total resources within the unit (number of CPU cores, number of GPU cores, total memory). This data format is JSON for easy subsequent processing. Finally, the system performs module resource mapping based on the scheduling unit data. The mapping process involves docking the resource specifications required by the functional module with the resource capabilities of the scheduling unit. The resource specifications include the number of CPU cores, the number of GPU cores, the memory capacity, and the network bandwidth. The specific values are determined from the functional module design document. For example, the search module requires 4 CPU cores, 1 GPU unit, and 8GB of memory, and the generation module requires 6 CPU cores, 2 GPU units, and 16GB of memory. The system traverses all scheduling units and matches the scheduling units that meet the resource specifications as the module mapping targets, and outputs module mapping data with fields including the module ID, the mapped scheduling unit ID, and detailed resource matching metrics. The data structure is in tabular form to support fast querying.
[0083] Step S462: Perform semantic detection based on the prompt word-optimized data to obtain semantic data; In this embodiment, the system performs semantic detection based on the prompt word optimized data. The prompt word optimized data includes a hierarchical vocabulary list, part-of-speech tagging, a semantic relationship graph, and an optimized historical version. The semantic detection process uses a method that combines lexical semantic matching and dependency syntactic analysis. First, word segmentation is performed on the prompt word text. The word segmentation tool uses a dictionary- and rule-based word segmenter, and the stop word list has a length of 5000 to ensure the removal of invalid words. Then, part-of-speech tagging is performed on the word segmentation result. A statistical language model is used to calculate the part-of-speech probability distribution, and the part of speech with the highest probability is selected as the final part of speech. Subsequently, the system constructs a dependency syntactic tree, analyzes the grammatical relationships between words, and generates a dependency relationship graph. Based on the calculation of the dependency graph and lexical semantic similarity, the system detects semantic clustering, and uses a cosine similarity threshold of 0.75 as the judgment criterion to aggregate phrases with similar semantics. Finally, the detection results are formed into semantic data, which includes lexical entities, semantic categories, dependency relationships, and semantic clustering identifiers. The storage format uses XML to facilitate subsequent parsing and intention extraction.
[0084] Step S463: Parse the behavioral intention according to the semantic data to obtain the behavioral intention data; select the functional module according to the behavioral intention data, so as to obtain the functional module data; In this embodiment, the behavioral intention parsing uses a method that combines semantic rule matching and comparison with a behavioral pattern library. The behavioral pattern library contains predefined behavioral templates, which clearly describe the set of intention keywords, triggering conditions, and corresponding behavioral codes within the template. For example, a semantic clustering containing words such as "query", "retrieve", "optimize", etc. is mapped to the "information retrieval" behavioral intention. The system scans the lexical entities and dependency relationships in the semantic data, and judges the intention matching degree according to the keyword coverage threshold (set to 70%) and the semantic clustering integrity. If the matching degree exceeds the threshold, the corresponding behavioral intention is confirmed. After parsing, behavioral intention data is generated, and the fields include behavioral intention code, intention description, and a list of matching keywords. Subsequently, the functional module selection is performed according to the behavioral intention data. The module selection rule is based on the behavioral intention code mapping table, which maps different intention codes to the corresponding functional module IDs. For example, "information retrieval" corresponds to the retrieval engine module, and "knowledge generation" corresponds to the text generation module. The module selection data includes the functional module ID, module name, and module description, and the format is JSON to facilitate subsequent calls by the platform.
[0085] Step S464: Build the platform logical architecture according to the module mapping data and the functional module data, so as to obtain the platform architecture data; In this embodiment, the system constructs the platform logic architecture by combining the module mapping data obtained in step S461 with the functional module data in step S463. During the construction process, first, determine the deployment locations of each functional module according to the module mapping data, that is, map them to the corresponding scheduling units. Based on the call dependency relationships and data flow paths between functional modules, design the logical connection structure to ensure that inter-module communication meets the requirements of low latency and high throughput. The communication protocol adopts the message queue mechanism, which supports asynchronous communication and load balancing. The message queue configuration parameters include the queue length limit (default 1000 messages), the message timeout threshold (set to 5 seconds), and the number of retry attempts (3 times). The logical architecture design also considers the redundancy backup mechanism. Key modules are deployed with dual instances, and a heartbeat detection mechanism is used between instances. The heartbeat interval is 1 second, and the heartbeat timeout threshold is set to 3 seconds to trigger a failover. Finally, generate the platform architecture data, including module deployment nodes, inter-module communication topologies, backup strategies, and fault detection parameters. The data is saved in YAML format for easy version control and dynamic adjustment.
[0086] Step S465: Construct the agent management platform according to the platform architecture data, and adaptively adjust the knowledge retrieval strategy of the agent management platform to obtain the adaptive strategy data.
[0087] In this embodiment, the agent management platform constructed based on step S464 is deployed and initialized. During the construction process, first, configure the running environments of each module on physical or virtual servers according to the platform architecture data, allocate computing resources, and set network access permissions. The deployment uses automated scripts, and the script parameters include the module deployment path, the version numbers of dependent libraries, and startup parameters (such as setting the number of threads to 8 and the memory limit to 16GB). Subsequently, the platform starts the modules and loads the initial knowledge base and prompt word optimization data. The adaptive adjustment of the knowledge retrieval strategy is based on the aforementioned heterogeneous hardware collaborative scheduling data and prompt word optimization data, combining the scheduling load, response latency, and prompt word level weights to adjust the retrieval index structure and weight distribution. The adjustment process includes reallocating index nodes, setting the index shard size (default 10MB), adjusting the retrieval filtering threshold (set to 0.85 cosine similarity), and updating the retrieval cache policy (cache validity period 300 seconds). The adjustment result outputs the adaptive strategy data, including the strategy version number, the details of parameter adjustment, and the effective timestamp, which is stored in JSON format. This strategy data is continuously used to dynamically optimize the retrieval performance and resource utilization during the operation of the platform.
Claims
1. A method for constructing an agent management platform that supports knowledge retrieval, generation, and optimization, characterized in that It includes the following steps: Step S1: Obtain multimodal data, perform semantic fusion based on the multimodal data to obtain semantically fused data; perform knowledge retrieval based on the semantically fused data to obtain knowledge retrieval data; Step S2: Establish a three-level reflection system according to the knowledge retrieval data, including constructing a result reflection layer, a process reflection layer, and a strategy reflection layer, so as to obtain result reflection layer data, process reflection layer data, and strategy reflection layer data, and perform hierarchical system fusion to obtain three-level reflection system data; Step S3: Optimize the prompt words based on the three-level reflection system data to obtain optimized prompt word data; Perform node load balancing according to the three-level reflection system data to obtain node load balancing data; Reconstruct the task chain based on the node load balancing data to obtain task chain data; Detect the resource consumption according to the task chain to obtain resource consumption data; Step S4: Perform heterogeneous hardware collaborative scheduling based on the resource consumption data to obtain heterogeneous hardware collaborative scheduling data; Construct an agent management platform according to the heterogeneous hardware collaborative scheduling data and the optimized prompt word data, and perform adaptive adjustment of the knowledge retrieval strategy for the agent management platform to obtain adaptive strategy data.
2. The construction method of the intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 1, characterized in that, Specifically, Step S1 is as follows: Step S11: Obtain multimodal data; Step S12: Extract intra-modal features based on the multimodal data to obtain modal feature data; Step S13: Perform inter-modal alignment processing according to the modal feature data to obtain modal alignment data; Step S14: Perform semantic association based on the modal alignment data to obtain semantic association data; Step S15: Perform semantic fusion according to the semantic association data to obtain semantically fused data; Step S16: Perform knowledge retrieval according to the semantically fused data to obtain knowledge retrieval data.
3. The construction method of the intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 2, characterized in that, Specifically, Step S16 is as follows: Step S161: Construct a semantic vector according to the semantically fused data to obtain semantic vector data; Step S162: Perform similarity matching of multi-source knowledge bases based on the semantic vector data to obtain candidate knowledge fragment data; Step S163: Perform semantic relevance ranking based on the candidate knowledge fragment data to obtain ranked knowledge fragment data; Step S164: Perform knowledge refinement according to the ranked knowledge fragment data to obtain knowledge retrieval data.
4. The construction method of the intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 1, characterized in that, Specifically, in Step S2, constructing the result reflection layer is as follows: Extract core keywords according to the knowledge retrieval data to obtain core keyword data; Draw a target semantic graph according to the core keyword data; Execute the answer generation process based on the target semantic graph and record it as the output text to obtain actual output data; Calculate the keyword coverage rate according to the target semantic graph and the actual output data; Calculate the semantic matching degree according to the target semantic graph and the actual output data; Determine the output error according to the keyword coverage rate and the semantic matching degree to obtain output error data; Extract language control parameters based on the output error data; Construct a structured output control template according to the language control parameters and construct the result reflection layer to obtain result reflection layer data.
5. The construction method of the intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 1, characterized in that, Specifically, in Step S2, constructing the process reflection layer is as follows: Identify the task execution path according to the knowledge retrieval data to obtain task execution path data; Construct a path decision tree based on the task execution path data to obtain path decision tree data; Identify inefficient nodes according to the path decision tree data to obtain inefficient node data; Calculate the node call frequency based on the inefficient node data to obtain high-frequency inefficient node data; Configure the result cache storage policy based on the high-frequency inefficient node data to obtain result cache data; Reconstruct the task process according to the result cache data and build a process reflection layer to obtain process reflection layer data.
6. The construction method of the intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 1, characterized in that In step S2, the construction of the strategy reflection layer is specifically as follows: Extract business rules according to the knowledge retrieval data to obtain business rule data; Obtain the historical decision output record; Conduct cognitive bias analysis based on the historical decision output record and the business rule data to obtain cognitive bias data; Generate knowledge structure update requirements based on the cognitive bias data to obtain knowledge structure update requirement data; Convert the knowledge structure update requirement data into graph reconstruction requirement data; Adjust the knowledge graph structure according to the graph reconstruction requirement data and build a strategy reflection layer to obtain strategy reflection layer data.
7. The construction method of the intelligent agent management platform for supporting knowledge retrieval, generation and optimization according to claim 1, characterized in that Step S3 is specifically as follows: Step S31: Conduct result reflection analysis based on the three-level reflection system data to obtain result reflection data; Restore the result reflection path according to the result reflection data to obtain result reflection path data; Step S32: Identify strategy deviations based on the result reflection path data to obtain strategy deviation data; Execute the prompt structure reorganization according to the strategy deviation data to obtain prompt structure data; Integrate multi-level control semantics according to the prompt structure data to obtain optimized prompt data; Step S33: Conduct node load balancing according to the three-level reflection system data to obtain node load balancing data; Step S34: Identify task nodes based on the node load balancing data to obtain task node data; Identify task dependencies based on the task node data to obtain task dependency data; Step S35: Reconstruct the task chain according to the task dependency data to obtain task chain data; Step S36: Detect resource consumption according to the task chain to obtain resource consumption data.
8. The method for constructing an intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 7, characterized in that Step S33 is specifically as follows: Step S331: Count the node execution status according to the three-level reflection system data to obtain node execution status data; Step S332: Calculate the node resource load rate based on the node execution status data; Step S333: Calculate the scheduling priority according to the node resource load rate to obtain node scheduling priority data; Step S334: Conduct agent task allocation according to the node scheduling priority data to obtain agent task allocation data; Step S335: Conduct node load analysis based on the agent task allocation data to obtain node load balancing data.
9. The construction method of the intelligent agent management platform for supporting knowledge retrieval, generation and optimization according to claim 1, wherein Step S4 is specifically as follows: Step S41: Identify task resource types based on the resource consumption data to obtain task resource type data; Step S42: Quantify the computing power requirements according to the task resource type data to obtain task computing power requirement data; Step S43: Match heterogeneous hardware resources according to the task computing power requirement data to obtain preliminary hardware matching data; Step S44: Conduct task execution simulation based on the preliminary hardware matching data to obtain task execution data; Conduct thread blocking point analysis based on the task execution data to obtain thread blocking data; Step S45: Adjust the thread binding of the preliminary hardware matching data according to the thread blocking data to obtain thread binding data; Perform heterogeneous hardware collaborative scheduling based on the thread binding data to obtain heterogeneous hardware collaborative scheduling data; Step S46: Construct an agent management platform according to the heterogeneous hardware collaborative scheduling data and the prompt word optimization data, and adaptively adjust the knowledge retrieval strategy of the agent management platform to obtain adaptive strategy data.
10. The method for constructing an intelligent agent management platform supporting knowledge retrieval, generation and optimization according to claim 9, wherein, Step S46 is specifically as follows: Step S461: Analyze the node distribution according to the heterogeneous hardware collaborative scheduling data to obtain node distribution data; divide the scheduling unit based on the node distribution data to obtain scheduling unit data; Perform module resource mapping according to the scheduling unit data to obtain module mapping data; Step S462: Perform semantic detection based on the prompt word optimization data to obtain semantic data; Step S463: Parse the behavior intention according to the semantic data to obtain behavior intention data; Select a functional module according to the behavior intention data to obtain functional module data; Step S464: Build a platform logic architecture according to the module mapping data and the functional module data to obtain platform architecture data; Step S465: Construct an agent management platform according to the platform architecture data, and adaptively adjust the knowledge retrieval strategy of the agent management platform to obtain adaptive strategy data.
Citation Information
Patent Citations
Large model intelligent agent system based on thinking clustering planning and domain knowledge retrieval
CN118861305A
Intelligent question answering method and system based on multi-module collaborative optimization
CN119557409A
Intelligent processing method and system for business process decision nodes based on AI Agent
CN119761796A
Knowledge Graph Extraction
US20250131289A1
Cited By
Intelligent agent high-order relation modeling method based on side attention weight
CN120611644A
Intelligent agent interpretable retrieval path generation system and verification method
CN120654839A
Vehicle control method and device, vehicle control application system and test method thereof
CN121019203A
Project whole-process monitoring method and system based on artificial intelligence
CN121094509A
A project end-to-end monitoring method and system based on artificial intelligence
CN121094509B