Agent Automatic Configuration Method and System Based on Large Language Model, Knowledge Base and Tools
By collaboratively optimizing the inference model, Embedding model and reordering model in the agent, the problem of waste of computing resources and insufficient accuracy of search results in high-dimensional data processing is solved, and more efficient and accurate data processing and search results are achieved.
Patent Information
- Application Number
- CN202510134433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art is prone to waste of computing resources and insufficient retrieval result accuracy when processing high-dimensional data, especially in complex data scenarios such as knowledge base Q&A.
The automatic configuration method of agents based on large models, knowledge bases and tools is adopted to solve the problem of inconsistency model, Embedding model and reordering model, and optimize data processing efficiency and the accuracy of retrieval results through collaborative optimization of inference model, Embedding model and reordering model.
It significantly improves the retrieval accuracy and efficiency, is suitable for knowledge base Q&A and complex data processing scenarios, and reduces the waste of computing resources caused by feature redundancy and unbalanced distribution.
Smart Images

Figure CN119578559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model data processing, and particularly to an intelligent agent automatic configuration method and system based on a large model, a knowledge base, and tools. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, knowledge-based intelligent question-answering systems and complex task processing systems have been widely applied in multiple fields, such as technical document retrieval, knowledge Q&A, intelligent decision-making support, etc. However, in the prior art, there are still many challenges in the automatic configuration of intelligent agents and multi-model collaborative processing:
[0003] In the multi-stage processing of task parsing, semantic retrieval, and result optimization, the input and output formats between different models often do not match. For example, the task results generated by an inference model may not be suitable for the vectorization requirements of an Embedding model, and the retrieval results of the Embedding model may also affect the performance of a re-ranking model.
[0004] When existing solutions process high-dimensional data, it is easy to waste computing resources due to feature redundancy and inefficient optimization, especially in complex data scenarios (such as knowledge base Q&A).
[0005] For example, Chinese Patent Application No. CN118819617A discloses a configuration method and device for an intelligent agent of a large model application. The method includes receiving an intelligent agent configuration instruction; and the intelligent agent configuration instruction is obtained based on web front-end configuration; configuring corresponding intelligent agent metadata information according to the intelligent agent configuration instruction and storing it in a metadata service; and the intelligent agent metadata information includes knowledge base metadata information, plugin metadata information, workflow metadata information, and database metadata information; calling a large model intelligent agent service to create an intelligent agent framework, and binding the intelligent agent metadata information to the intelligent agent framework to create an intelligent agent. This prior art can greatly reduce the difficulty of developing intelligent agents during the configuration process through a web visualization configuration mechanism, shielding underlying technologies such as large language models and knowledge bases, enabling non-technical business personnel to easily develop intelligent agents.
[0006] The above prior arts all have the problems raised in this background art: it is easy to waste computing resources and the accuracy of retrieval results is insufficient when processing high-dimensional data. To solve the above problems, this application designs an intelligent agent automatic configuration method and system based on a large model, a knowledge base, and tools. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide an intelligent agent automatic configuration method and system based on a large model, a knowledge base, and tools in view of the deficiencies of the prior art. Training data is obtained, natural language is input into an inference model to generate a set of query tasks, and after optimization, it is input into an Embedding model to generate a set of preliminary retrieval results; the optimized results are input into a re-ranking model to generate final retrieval results; when the training conditions are met, the inference model, the Embedding model, and the re-ranking model are encapsulated into an intelligent agent. The system consists of an inference model configuration module, a retrieval model configuration module, and a re-ranking model configuration module, which cooperate to achieve task parsing, semantic retrieval, and result optimization. The present invention improves the retrieval accuracy and efficiency and is applicable to knowledge base question answering and complex data processing scenarios.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] An intelligent agent automatic configuration method based on a large model, a knowledge base, and tools, the method comprising:
[0010] Obtain training data, the training data including natural language content input by multiple users;
[0011] Input each natural language content into a preset inference model, and fix the output of the inference model as a set of query tasks manually annotated;
[0012] Input the set of query tasks manually annotated into a preset Embedding model, the Embedding model being configured to be able to call a preset knowledge base, and fix the output of the Embedding model as a set of preliminary retrieval results manually annotated;
[0013] Input the set of preliminary retrieval results manually annotated into a preset re-ranking model, and fix the output of the re-ranking model as optimized retrieval results manually annotated;
[0014] After the training reaches a set number of times, or after the similarity between the test results and the optimized retrieval results manually annotated is higher than a set similarity threshold, encapsulate the final inference model, the Embedding model, and the re-ranking model into an intelligent agent.
[0015] The step of fixing the output of the inference model as a set of query tasks manually annotated further includes:
[0016] Generate an inference output, the inference output being a set of preliminary query tasks;
[0017] Calculate the semantic similarity between the inference output and the manually annotated query task set, and optimize the parameters of the inference model through iterative training according to the semantic similarity, so that the output of the inference model gradually approaches the manually annotated query task set until the semantic similarity reaches the set semantic threshold;
[0018] Use the optimized inference model for subsequent encapsulation.
[0019] The step of fixing the output of the Embedding model as the manually annotated preliminary retrieval result set further includes:
[0020] Generate a semantic vector representation corresponding to each query task, and perform vector retrieval in the preset knowledge base to recall the preliminary retrieval results semantically related to the query task;
[0021] Calculate the retrieval coverage rate between the preliminary retrieval results and the manually annotated preliminary retrieval result set, and update the semantic vector generation method of the Embedding model through rule optimization constraints according to the retrieval coverage rate;
[0022] When the retrieval coverage rate between the retrieval result set of the Embedding model and the manually annotated preliminary retrieval result set reaches the set retrieval threshold, terminate the optimization training and use the optimized Embedding model for subsequent encapsulation.
[0023] Before inputting the manually annotated query task set into the preset Embedding model, the method further includes:
[0024] Perform semantic structure optimization on the output of the inference model;
[0025] According to the query task after semantic structure optimization, divide the semantic content of the query task into high-frequency features and low-frequency features, and assign dynamic weights to different frequency features according to the vectorization requirements of the Embedding model to generate an input feature set;
[0026] According to the input feature set, generate a specific format for the input requirements of the Embedding model through an input converter.
[0027] The semantic structure optimization includes:
[0028] Decompose the query task into fine-grained subtasks according to semantic parsing rules, and the subtasks include key semantic units and attribute descriptions;
[0029] Through the context information completion algorithm, add context information semantically related to the preset knowledge base to the fine-grained subtasks to update the query task.
[0030] Before inputting the set of preliminary retrieval results with manual annotations into a preset re-ranking model, the method further includes:
[0031] Extract multi-dimensional features for each retrieval item in the set of preliminary retrieval results, calculate the feature distribution balance of the retrieval item according to the multi-dimensional features, map the retrieval item with a feature distribution balance lower than the distribution threshold to an independent retrieval space, and delete the corresponding retrieval item from the set of preliminary retrieval results;
[0032] Adjust the feature density of the set of preliminary retrieval results according to the feature distribution balance;
[0033] In the independent retrieval space, reconstruct the retrieval item through a feature compression algorithm to generate a low-dimensional feature representation, which can be adapted to the sorting format of the re-ranking model.
[0034] The feature density adjustment includes:
[0035] Calculate a feature distribution matrix according to the feature distribution balance of the retrieval items in the set of preliminary retrieval results;
[0036] Calculate the feature co-information amount of different retrieval items in the set of preliminary retrieval results according to the feature distribution matrix, and adjust the density of the set of preliminary retrieval results according to the feature co-information amount.
[0037] The step of fixing the output of the re-ranking model as the optimized retrieval result with manual annotations further includes:
[0038] Construct a feature mapping matrix for the low-dimensional feature representation;
[0039] Assist in the sorting process of the re-ranking model according to the feature mapping matrix;
[0040] Fuse the optimization process of the re-ranking model according to the feature mapping matrix.
[0041] The method further includes:
[0042] Generate a corresponding query requirement through an inference model in the intelligent agent according to the real-time query task input by the user;
[0043] Input the query requirement into the Embedding model of the intelligent agent to retrieve relevant knowledge base content and generate preliminary retrieval results;
[0044] Optimize the preliminary retrieval results through the re-ranking model in the intelligent agent to generate a sorted set of retrieval results;
[0045] When the semantic relevance of the retrieval result set is lower than the result threshold, the inference model is called again, and the query requirements are regenerated by analyzing the semantic distribution of the retrieval result set and the user input.
[0046] An intelligent agent automatic configuration system based on a large model, a knowledge base, and tools, the system includes an inference model configuration module, a retrieval model configuration module, and a re-ranking model configuration module;
[0047] The inference model configuration module is used to receive the natural language content input by the user, parse it into a structured query task, and optimize the output of the inference model according to the manually annotated query task set to gradually match it with the manually annotated query task set;
[0048] The retrieval model configuration module is used to vectorize the structured query task by using an Embedding model, retrieve a preset knowledge base and generate a preliminary retrieval result set, and perform feature optimization and density adjustment on the preliminary retrieval result set;
[0049] The re-ranking model configuration module is used to perform sorting logic processing on the optimized preliminary retrieval result set, and adjust the sorting weight in combination with the feature mapping matrix to generate a final retrieval result set consistent with the manually annotated optimization result.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1. Through the collaborative optimization of the inference model, Embedding model, and re-ranking model, the present invention solves the problem of input-output mismatch between models, and significantly improves the data processing efficiency and the accuracy of retrieval results;
[0052] 2. By performing structured optimization and feature adaptation on the output during the model training process, the present invention ensures the efficient data flow between models, and at the same time reduces the waste of computing resources caused by feature redundancy and uneven distribution. Description of the Drawings
[0053] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:
[0054] Figure 1 It is a flowchart of the intelligent agent automatic configuration method based on a large model, a knowledge base, and tools in Embodiment 1 of the present invention;
[0055] Figure 2 It is a data flow diagram of intelligent agent training in Embodiment 1 of the present invention;
[0056] Figure 3 It is a module diagram of the intelligent agent automatic configuration system based on a large model, a knowledge base, and tools in Embodiment 2 of the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0058] Embodiment 1
[0059] Please refer to Figure 1 , an embodiment provided by the present invention: an intelligent agent automatic configuration method based on a large model, a knowledge base, and tools, and the specific steps of the method are as follows:
[0060] S1: Obtain training data;
[0061] In this step, natural language content of users is collected from various sources as training data. The data sources can include interaction logs, common question-and-answer sets, and domain-specific corpora, etc., and the collected data is preprocessed, including data cleaning, duplicate removal, and filling in missing information. Finally, the data is classified and labeled according to the task requirements to support the optimization of the subsequent model for specific scenarios.
[0062] S2: Input the training data into a preset inference model, output a query task set, and optimize the query task set;
[0063] In this step, the preset inference model is the core of the intelligent agent, responsible for parsing the natural language input of users and generating structured query tasks. This model can identify the intentions, keywords, and context logic in the user input, and complete preliminary task decomposition and semantic optimization;
[0064] Furthermore, large language models such as OpenAI and GPT-4 can be used as the kernel of the inference model;
[0065] Specifically, through natural language processing technology, the natural language input of users is converted into structured query tasks. Based on rule-based semantic verification and parsing algorithms, semantic consistency checks are performed on the generated tasks to ensure the clarity and operability of the tasks. Combining the user's historical interaction data and the current input, more context-adaptive query tasks are generated.
[0066] S3: Input the query task set into a preset Embedding model, output a preliminary retrieval result set, and optimize the preliminary retrieval result set;
[0067] In this step, the preset Embedding model serves as the core semantic mapping module of the intelligent agent, responsible for converting query tasks and knowledge base content into high-dimensional vector representations to support efficient semantic retrieval. By capturing the deep semantic relationships between texts, the Embedding model achieves precise matching of user queries and knowledge base content;
[0068] Furthermore, the content of the Embedding model can adopt Embedding API, Text EmbeddingInference, etc., to capture semantic features through the trained vectorization mapping model. The preset model supports multi-language modes and adapts to knowledge base content in Chinese, English, or multiple languages;
[0069] Specifically, the Embedding model is generally used in knowledge base question answering to match user queries with relevant document fragments. For example, it helps retrieve paragraphs related to the user's question in technical documents. The Embedding model preprocesses the query task, which includes content segmentation, dividing it into sentences or text blocks to adapt to the limitations of the model's context window, converting the query task into a semantic vector representation. At the same time, it preprocesses the knowledge base content into high-dimensional vectors, calculates the semantic similarity between the user query vector and the knowledge base text vector through a vector similarity algorithm (such as cosine similarity), and sets a threshold to filter the results to ensure that the recalled document fragments are highly relevant to the query task.
[0070] S4: Input the initial retrieval result set into the preset re-ranking model and output the optimized retrieval result;
[0071] In this step, the re-ranking model serves as the optimization module of the intelligent agent, performing sorting logic processing on the initial retrieval result set generated by the Embedding model. By evaluating the relevance, importance, and context coordination relationship of the results, the re-ranking model ensures that the final retrieval result is closer to the user's needs.
[0072] Furthermore, the re-ranking model can adopt the Rerank model (such as the re-ranking module in Text Embedding Inference), optimizing the sorting logic based on deep learning technology. The model reprocesses the features of the initial retrieval results and generates an optimized sorting result by combining semantic relevance and context importance.
[0073] Specifically, the preset re-ranking model is used for search engine result optimization, intelligent question answering systems, and knowledge base relevance ranking. Exemplarily, when the user's input query task is "How to read a CSV file in Python?", the intelligent agent needs to retrieve relevant content from the technical document library and sort it by priority to ensure that the user can quickly obtain the most useful information.
[0074] S5: After the training reaches the set number of times, or after the similarity between the test result and the optimized retrieval result manually annotated is higher than the set similarity threshold, package the final inference model, the Embedding model, and the re-ranking model into an agent;
[0075] The selection of the threshold is based on experiments and actual business requirements, such as being set to 90%.
[0076] Please refer to Figure 2 , the agent training data flow diagram of the embodiment of the present invention. Further, in this embodiment, mainly by training a large language model, through the coupling between models, the inference model, the Embedding model, and the re-ranking model are deeply coupled to construct an efficient and accurate agent. During the training process, by fixing the output of manual annotation, the goal of the model at each stage is clarified, providing a reliable comparison benchmark for the parameter optimization and performance improvement of the model.
[0077] Specifically, in the training of the inference model, the manually annotated query task set is used as the target output to guide the query tasks generated by the model to gradually approach the real semantic requirements. This process can effectively reduce the ambiguity and redundant information in the generated query tasks and improve the semantic parsing ability of the model. In the training of the Embedding model, the preliminary retrieval result set manually annotated is fixed to ensure that the retrieval coverage rate and matching accuracy meet the business requirements. At the same time, by optimizing the vector representation, the ability of the model to process complex semantic relationships in the high-dimensional vector space is improved. In the training of the re-ranking model, the optimized retrieval result set manually annotated is used to guide the optimization of the ranking logic to ensure that the ranking of the output results is consistent with the user's expectations and enhance the adaptability of the model to the actual application scenario.
[0078] Further, through the modular design and coupling optimization method, the efficient cooperation between the inference, retrieval, and ranking models is realized, significantly reducing the redundant operations in the calculation process of a single model and optimizing the utilization rate of computing resources. In addition, during the training process, problems of input-output mismatch between different modules in model training may occur. By performing additional processing on the output of each stage, the overall system performance can be optimized, especially reducing the calculation loss in complex data processing scenarios.
[0079] Specifically, optimizing the output of the inference model is to enhance its structural degree and semantic clarity, ensuring that the generated tasks have clear vectorization requirements when input into the Embedding model; the initial retrieval result set generated by the Embedding model also needs to optimize its feature distribution through feature extraction, denoising, and density adjustment, so as to reduce the redundancy of feature dimensions and the consumption of computing resources when input into the re-ranking model. In addition, the re-ranking model further executes the ranking logic based on the optimized retrieval results to ensure that the final output result is highly consistent with the target set of manual annotations. This processing method not only applies to scenarios with limited computer resources, effectively reducing the model's computational loss and memory occupancy, but also improves the training efficiency and shortens the training time; at the same time, it ensures the efficient transfer of the output of each model and its adaptation to the input of the next model, forming an efficient collaborative workflow from task generation to retrieval ranking. This multi-stage optimization and training method is particularly suitable for high-complexity intelligent application scenarios such as large-scale knowledge base question answering and technical document retrieval.
[0080] The specific steps of S2 are as follows:
[0081] S2.1: Input the cleaned and annotated training data (including the natural language content input by users and its corresponding manually annotated query task set) into the inference model;
[0082] In this embodiment, the training data is the basis for model optimization and is cleaned and annotated to ensure its quality. Untreated data may contain noise, redundancy, or inconsistent information, which will reduce the model's performance. Using the manually annotated query task set as a high-quality supervision signal to guide the inference model to generate highly adaptable query tasks.
[0083] Specifically, remove duplicate or invalid data (such as garbled characters, stop words, etc.) through data cleaning to improve the effectiveness and consistency of the data. Through manual annotation, align the natural language input with the standardized query tasks to form a training set for supervised learning. Input the cleaned and annotated data into the inference model in a batch loading and chunking manner.
[0084] S2.2: Convert the user input into a structured query task through natural language processing technology to generate an inference output, and the inference output is an initial query task set;
[0085] Natural language input is usually unstructured, and directly using it will make it difficult for subsequent models (such as the Embedding model) to process. Converting natural language into a structured query task can provide a standardized input for downstream models.
[0086] Specifically, the inference model performs word segmentation, entity recognition, and intent analysis operations on the user input to extract key elements, where the key elements include the goal, constraints, and context information. Then, according to the predefined template, the parsed key elements are organized into a structured set of query tasks, including clear task goals, conditional constraints, and supplementary information.
[0087] S2.3: Calculate the semantic similarity between the inference output and the manually annotated set of query tasks, and based on the semantic similarity, optimize the parameters of the inference model through iterative training, so that the output of the inference model gradually approaches the manually annotated set of query tasks until the semantic similarity reaches the set semantic threshold;
[0088] The initial output of the inference model may deviate from the manually annotated set of query tasks, and it is necessary to continuously optimize the model parameters through supervised learning to make the generated output gradually approach the manually annotated target. The semantic similarity, as an evaluation metric, can quantify the gap between the model output and the manual annotation, providing a feedback basis for iterative training.
[0089] Specifically, calculate the semantic similarity between the two through cosine similarity, define a loss function, use the similarity difference as the optimization goal, and then adjust the parameters of the inference model according to the similarity through the gradient descent algorithm to narrow the semantic gap between the model output and the manual annotation. Repeat the training and evaluation process to gradually improve the semantic consistency of the model output. When the similarity reaches the set threshold, stop the training.
[0090] S2.4: Record the optimized inference model and process the output of the inference model;
[0091] The optimized inference model needs to be saved in the final version for subsequent packaging as an agent for use. Processing the output of the inference model is to address possible missing details or incompatibility with the input of the downstream Embedding model.
[0092] The specific steps of S2.4 are as follows:
[0093] S2.4.1: Decompose the query task into fine-grained subtasks according to semantic parsing rules, and the subtasks include key semantic units and attribute descriptions;
[0094] The query task may contain multiple intents or complex semantic structures. Directly inputting it into the subsequent model may lead to information conflicts or incompatibilities. Decomposing the query task into fine-grained subtasks helps to extract the core semantics, clarify the attributes and goals of the task, and provide clear input for the downstream module;
[0095] Specifically, through a preset semantic parsing rule library, the logical structure of the query task is parsed. The rule library includes common grammar patterns (such as "goal - constraint - context") and domain - specific semantic templates. Using a semantic parsing algorithm, the query task is decomposed into subtasks, each subtask consisting of key semantic units (such as "read a CSV file") and property descriptions (such as "using Python"). Decomposing complex tasks into easily - handled subtasks improves the parsing efficiency of subsequent models.
[0096] S2.4.2: Through a context information completion algorithm, add context information semantically related to the preset knowledge base to the fine - grained subtask, and update the query task;
[0097] Fine - grained subtasks usually only contain core semantic units and basic property descriptions, while ignoring context - related information (such as domain knowledge or background constraints), resulting in incomplete semantic information. Completing context information can enhance the semantic integrity of the task and its adaptability to subsequent retrieval tasks.
[0098] Specifically, screen the most relevant background knowledge from the preset knowledge base according to semantic similarity, extract context information related to the subtask, merge the extracted context information with the fine - grained subtask, and use a weighted fusion method to ensure that the priority of background information does not override the core semantics of the subtask, enhancing the semantic integrity and expression ability of the query task. At the same time, verify whether the completed task conforms to the intention of the original query task, and remove context information that may cause ambiguity or conflict. Improve the depth of understanding of the query task by the Embedding model during vectorization, and reduce the decline in retrieval accuracy caused by semantic loss.
[0099] S2.4.3: According to the query task optimized by the semantic structure, divide the semantic content of the query task into high - frequency features and low - frequency features, and assign dynamic weights to different frequency features according to the vectorization requirements of the Embedding model to generate an input feature set;
[0100] Different semantic features of the query task have different effects on the vectorization of the Embedding model. High - frequency features usually represent the main goal of the task, while low - frequency features may contain important but not prominent auxiliary information. By assigning dynamic weights, the expression ability of vectorization can be effectively enhanced.
[0101] Specifically, through the feature frequency analysis algorithm, high-frequency features and low-frequency features in the query task are extracted. High-frequency features (such as "read CSV") reflect the core semantics of the task, and low-frequency features (such as "memory optimization") supplement background information. The feature weights are adjusted according to the frequency of the features and the requirements of the Embedding model. Higher weights are assigned to high-frequency features to ensure that the model focuses on the core task. The weights of low-frequency features are dynamically adjusted according to their semantic importance to avoid losing key information and reduce the computational burden of the Embedding model when processing low-correlation features. The feature set with dynamic weights is combined into a unified input format to adapt to the processing requirements of the Embedding model, improve the accuracy of the vectorization result, and enhance the expression ability of the core objective of the query task.
[0102] The calculation formula for the dynamic weight is:
[0103] ,
[0104] where, represents the dynamic weight of feature , represents the semantic similarity function, represents the semantic similarity between feature and the current context, represents the semantic frequency normalization function, represents the semantic frequency normalization value of feature , represents the correlation function of the vectorization requirement, represents the correlation value between feature and the vectorization requirement, represents the non-linear attenuation parameter, which is used to control the sensitivity to the adaptation score to smooth the weight distribution. j represents other individual features, and n represents the total number of other features;
[0105] where, the calculation formula for the correlation function between feature and the vectorization requirement is:
[0106] ,
[0107] where, represents the rank of the solution matrix, represents the rank of the matrix corresponding to feature after spatial projection, represents matrix multiplication, represents the vector matrix of the feature semantic space projection, represents the projection matrix of the Embedding model vectorization requirement space. The number of rows of the projection matrix is equal to the number of columns of the vector matrix, represents the cosine similarity calculation function, The reference center vector representing the vectorization requirements of the Embedding model, represents a feature and the reference center vector of the cosine similarity;
[0108] wherein, the feature the calculation formula of the semantic frequency normalization function is:
[0109] ,
[0110] wherein, represents the logarithm to the base 2, represents the feature frequency, represents the transposed matrix, represents the vector matrix of the overall semantic space projection of the query task, represents a constant greater than zero.
[0111] S2.4.4: According to the input feature set, generate a specific format for the input requirements of the Embedding model through an input converter.
[0112] The specific steps of S3 are as follows:
[0113] S3.1: Generate a semantic vector representation corresponding to each query task, and perform vector retrieval in the preset knowledge base to recall the preliminary retrieval results semantically related to the query task, aiming to efficiently retrieve the content semantically related to the query task from a large number of documents;
[0114] The core function of the Embedding model is to convert unstructured natural language data into high-dimensional vector representations, enabling it to efficiently match with the content in the knowledge base in the vector space. Vector retrieval is an important means to quickly locate the content semantically related to the query task from a large-scale knowledge base, which can significantly improve the retrieval efficiency.
[0115] Specifically, use the pre-trained Embedding model to convert the query task into a semantic vector representation, and at the same time convert the documents in the knowledge base into vector forms in segments. Through vector retrieval, it is possible to quickly lock the documents related to the query task in the massive knowledge base, significantly shortening the retrieval time. Calculate the similarity between the query task vector and the knowledge base vector through the vector similarity algorithm, and recall the most relevant document segments in descending order of similarity. Configure retrieval parameters (such as the TopK value and similarity threshold) to screen high-quality retrieval results to construct a preliminary retrieval result set. Precise semantic matching ensures that the recalled content is highly relevant to the query task semantics, laying a foundation for subsequent ranking optimization.
[0116] S3.2: Calculate the retrieval coverage rate between the preliminary retrieval results and the manually annotated preliminary retrieval result set, and update the semantic vector generation method of the Embedding model through rule optimization constraints according to the retrieval coverage rate;
[0117] The coverage rate of the preliminary retrieval results directly affects the retrieval performance of the system. If the coverage rate is insufficient, important content may be missed. Through coverage rate evaluation, it is possible to guide the Embedding model to optimize the semantic vector generation method, thereby improving the retrieval accuracy.
[0118] Specifically, compare the preliminary retrieval results with the manually annotated set, evaluate the proportion of relevant segments recalled, and adjust the rules for the Embedding model to generate semantic vectors according to the coverage rate results, including:
[0119] Enhance the weight of low-frequency terms;
[0120] Adjust the feature distribution weight in vector generation;
[0121] Update the domain-specific semantic rules in vector generation;
[0122] Optimizing the coverage rate ensures that the retrieval results can comprehensively reflect the semantic requirements of the query task and reduce the risk of missing key content. Through rule adjustment, the Embedding model performs more stably in subsequent retrieval tasks.
[0123] S3.3: When the retrieval coverage rate between the retrieval result set of the Embedding model and the manually annotated preliminary retrieval result set reaches the set retrieval threshold, terminate the optimization training and use the optimized Embedding model for subsequent packaging;
[0124] S3.4: Extract multi-dimensional features for each retrieval item in the preliminary retrieval result set, calculate the feature distribution balance of the retrieval item according to the multi-dimensional features, map the retrieval item with a feature distribution balance lower than the distribution threshold to an independent retrieval space, and delete the corresponding retrieval item from the preliminary retrieval result set;
[0125] In the initial retrieval result set, there may be significant imbalances in the feature distributions of different retrieval terms. For example, the features of some terms are densely distributed in a few dimensions, while the feature distributions of others are too sparse. Such imbalanced feature distributions can lead to computational biases in subsequent model processing. For example, some features may be over-amplified or ignored, thus affecting the ranking effect of the re-ranking model. Feature distribution refers to how the feature values of each retrieval term are distributed across multiple dimensions in the initial retrieval result set, usually generated by an Embedding model, representing the numerical representation of each retrieval term in dimensions such as semantic relevance, context consistency, and global weight. Feature distribution directly reflects the expression of the attributes of retrieval terms in the multi-dimensional feature space.
[0126] Specifically, for each retrieval term, extract feature vectors from the initial retrieval results, including semantic relevance scores, context consistency metrics, and global citation weights. Conduct distribution analysis on the feature vectors, calculate the mean, variance, and frequency distribution of each dimension to form a feature distribution matrix. Map the retrieval terms with balance lower than the set threshold to an independent retrieval space and delete these retrieval terms from the initial retrieval result set.
[0127] S3.5: Adjust the feature density of the initial retrieval result set according to the feature distribution balance;
[0128] Feature density adjustment can redistribute the proportion of features among different retrieval terms, making the overall distribution of the result set more balanced and avoiding wasting computational resources.
[0129] Specifically, the feature density adjustment includes:
[0130] Calculate a feature distribution matrix according to the feature distribution balance of the retrieval terms in the initial retrieval result set;
[0131] According to the feature distribution matrix, calculate the feature co-information amount of different retrieval terms in the initial retrieval result set, and perform density adjustment on the initial retrieval result set according to the feature co-information amount.
[0132] S3.6: In the independent retrieval space, reconstruct the retrieval terms through a feature compression algorithm to generate a low-dimensional feature representation that can be adapted to the ranking format of the re-ranking model.
[0133] Since the feature distribution balance of some retrieval terms is low, directly participating in the optimization of the initial retrieval results may lead to a decline in the quality of the overall retrieval set. Therefore, they are separately mapped to an independent retrieval space for processing. In the independent retrieval space, the features of these retrieval terms need to be further compressed and reconstructed to remove redundant information, ensuring that even low-balance features can generate a low-dimensional feature representation adapted to the re-ranking model.
[0134] Furthermore, the retrieval terms in the independent retrieval space usually have the characteristics of high dimension and low correlation. Directly inputting them into the re-ranking model will increase the computational cost and affect the accuracy of the ranking results. By generating low-dimensional feature representations through feature compression, the input format requirements of the re-ranking model can be met, and the processing efficiency can be improved at the same time.
[0135] Specifically, extract the principal components of the feature vectors, select the first few principal components whose cumulative variance contribution rate reaches a set threshold (such as 90%), and compress the low-correlation dimensions. The generated low-dimensional feature representations retain the core information related to the query task of the retrieval terms, while significantly reducing the dimension. Construct an encoder-decoder structure to perform non-linear mapping on the feature vectors and compress the high-dimensional features into the target low-dimensional space. During the training process, use the ranking format of the re-ranking model as the decoding target to ensure that the generated low-dimensional feature representations are adapted to the subsequent model. By calculating the importance weights of each feature dimension (such as based on information gain or gradient contribution), select the high-weight features as the components of the low-dimensional feature representations. Reconstruct the low-equilibrium retrieval terms through the feature compression algorithm to generate low-dimensional feature representations to adapt to the re-ranking model. Even the retrieval terms with uneven feature distributions can play their effective roles in the subsequent model, while reducing the overall computational loss of the system.
[0136] The specific steps of S4 are as follows:
[0137] S4.1: Receive the preliminary retrieval result set and its low-dimensional feature representation;
[0138] When the re-ranking model receives the preliminary retrieval result set, it simultaneously accepts the low-dimensional feature representations generated in the independent retrieval space. The preliminary retrieval result set comes from the vector retrieval results of the Embedding model and contains high-dimensional features such as semantic relevance and context consistency after feature optimization. The low-dimensional feature representations are the results generated by the feature compression algorithm in the independent retrieval space, which condense the key information of the low-equilibrium retrieval terms and have compactness and high adaptability. Process these two parts of data in parallel to ensure that the combination of high-dimensional features and low-dimensional features provides sufficient support for subsequent ranking.
[0139] S4.2: Construct the feature mapping matrix of the low-dimensional feature representation;
[0140] To solve the compatibility problem between low-dimensional features and high-dimensional features in the re-ranking model, construct a feature mapping matrix. The mapping matrix not only realizes the conversion between features but also enhances the role of low-dimensional features in the ranking logic;
[0141] S4.3: Assist in processing the ranking process of the re-ranking model according to the feature mapping matrix;
[0142] During the sorting optimization process of the re-ranking model, potential irrelevant items in the initial retrieval result set are identified and down-weighted through the collaborative information provided by the low-dimensional feature representation, thereby reducing sorting interference;
[0143] Provide the global feature distribution trend after feature compression for the re-ranking model, which is used to optimize the global feature bias calibration of the ranking model and improve the overall consistency of the ranking results;
[0144] Utilize the compactness of the low-dimensional feature representation to reduce the complexity of the re-ranking model in feature interaction calculations, thereby improving the computational efficiency of the model.
[0145] S4.4: Perform fusion processing on the optimization process of the re-ranking model according to the feature mapping matrix;
[0146] During the feature optimization process of the initial retrieval result set, use the low-dimensional feature representation as supplementary features and combine them with the high-dimensional feature representation of the initial retrieval result set;
[0147] According to the feature density information provided by the low-dimensional feature representation, adjust the attention weight of the re-ranking model for specific feature dimensions, thereby enhancing the rationality and relevance of the feature distribution;
[0148] Enhance the relevance of the retrieval items in the initial retrieval result set through the low-dimensional feature representation, optimize the collaboration of the input features, and make the sorting logic of the re-ranking model closer to the optimization standard of manual annotation.
[0149] The method further includes:
[0150] According to the real-time query task input by the user, generate corresponding query requirements through the inference model in the intelligent agent;
[0151] Input the query requirements into the Embedding model of the intelligent agent, retrieve relevant knowledge base content and generate initial retrieval results;
[0152] Optimize the initial retrieval results through the re-ranking model in the intelligent agent to generate a sorted retrieval result set;
[0153] When the semantic relevance of the retrieval result set is lower than the result threshold, re-call the inference model, and re-generate query requirements by analyzing the semantic distribution of the retrieval result set and the user input.
[0154] Exemplarily, the embodiments of the present invention provide a water conservancy flood control intelligent agent, which aims to provide users with accurate query and analysis functions. By accessing various data resources such as historical meteorology and hydrology, it supports the query and dynamic analysis of key flood control information. Inside the intelligent agent, the inference model, Embedding model, and re-ranking model work together to achieve the full process from natural language queries of users to the return of structured data.
[0155] S101: The user inputs a query task through the system, for example: "Query the typhoon movement path and rainfall distribution in South China in the past 10 years.
[0156] S102: The inference model converts the natural language query input by the user into a structured query task, for example:
[0157] Time range: in the past 10 years;
[0158] Geographical range: South China;
[0159] Query target: typhoon movement path and rainfall distribution.
[0160] S103: The inference model splits the complex query into multiple subtasks, including:
[0161] Query historical typhoon path data;
[0162] Query rainfall distribution data related to typhoons.
[0163] S103: Combine domain rules and historical query records to optimize the query task, for example, supplement that the "typhoon path" data should include fields such as "starting point, ending point, and affected area".
[0164] S104: Convert the query task into a high-dimensional semantic vector. For example, "typhoon path" is mapped to a set of spatial features, including geographical coordinates of path points, time series, and moving speed, etc.
[0165] S105: The Embedding model uses a vector similarity algorithm (such as cosine similarity) to retrieve content matching the input vector in the knowledge base. The knowledge base includes various historical meteorological data sources, for example:
[0166] Historical typhoon path data;
[0167] Historical rainfall distribution maps;
[0168] Historical flood area distribution data.
[0169] S106: Return multiple relevant documents or data segments, including records of historical typhoon paths and a collection of pictures of rainfall distributions, and conduct an analysis on the balance of feature distributions for the preliminary retrieval results. For example, filter out some typhoon records whose time spans do not meet the range of "the past 10 years", and retain highly relevant retrieval items such as "hotspot map of rainfall in South China in the past 10 years".
[0170] S107: The re-ranking model receives the preliminary retrieval results and their relevant features and conducts priority ranking. Specifically, according to the important fields of the user's query task (such as "time range" and "geographical range"), higher weights are assigned to the results with high relevance, and the typhoon records at adjacent time points are sorted logically through the context consistency rule. For example:
[0171] Give priority to returning the complete typhoon path (complete record from generation to dissipation).
[0172] S108: Supplement missing information in combination with the specific requirements input by the user (such as "rainfall distribution"). For example, predict the associated rainfall data from the path data without marked rainfall. Specifically, the results output by the re-ranking model are returned in the following format:
[0173] Typhoon path data: Geographical coordinates and time series of specific path points;
[0174] Rainfall distribution map: Hotspot distribution map matching the scope of South China.
[0175] Embodiment 2
[0176] Please refer to Figure 3 , the present invention provides an embodiment: an intelligent agent automatic configuration system based on a large model, a knowledge base, and tools. The system includes an inference model configuration module, a retrieval model configuration module, and a re-ranking model configuration module;
[0177] The inference model configuration module is used to receive the natural language content input by the user, parse it into a structured query task, and optimize the output of the inference model according to the manually annotated query task set to gradually match it with the manually annotated query task set;
[0178] The retrieval model configuration module is used to vectorize the structured query task by using an Embedding model, retrieve a preset knowledge base, generate a set of preliminary retrieval results, and conduct feature optimization and density adjustment on the set of preliminary retrieval results;
[0179] The re-ranking model configuration module is used to perform sorting logic processing on the optimized set of preliminary retrieval results, and adjust the sorting weights in combination with a feature mapping matrix to generate a set of final retrieval results consistent with the manually annotated optimization results.
[0180] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An agent automatic configuration method based on a large model, knowledge base and tools, characterized in that: The method comprises: Acquire training data, wherein the training data includes natural language content input by multiple users; Input each natural language content into a preset reasoning model, and fix the output of the reasoning model to a manually annotated query task set; Optimize the semantic structure of the output of the reasoning model; According to the query task after the semantic structure optimization, the semantic content of the query task is divided into high-frequency features and low-frequency features, and dynamic weights are assigned to different frequency features according to the vectorization requirements of the Embedding model to generate an input feature set; According to the input feature set, a format for the input requirements of the Embedding model is generated through an input converter; Inputting the manually annotated query task set after format conversion into a preset Embedding model, wherein the Embedding model is configured to call a preset knowledge base, and fixing the output of the Embedding model to a manually annotated preliminary search result set; Inputting the manually annotated preliminary search result set into a preset re-ranking model, and fixing the output of the re-ranking model to the manually annotated optimized search result; After the training reaches a set number of times, or the similarity between the test result and the manually annotated optimized retrieval result is higher than the set similarity threshold, the final reasoning model, the Embedding model and the re-ranking model are encapsulated as an intelligent agent.
2. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 1, characterized in that: The step of fixing the output of the inference model to a manually labeled query task set further includes: Generate an inference output, wherein the inference output is a preliminary query task set; Calculating the semantic similarity between the inference output and the manually annotated query task set, and optimizing the parameters of the inference model through iterative training according to the semantic similarity, so that the output of the inference model gradually approaches the manually annotated query task set until the semantic similarity reaches a set semantic threshold; The optimized inference model is used for subsequent packaging.
3. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 2, characterized in that: The method of fixing the output of the Embedding model to a manually annotated preliminary search result set further includes: Generate a semantic vector representation corresponding to each query task, and perform vector retrieval in the preset knowledge base to recall preliminary retrieval results related to the semantics of the query task; Calculate the search coverage between the preliminary search results and the manually annotated preliminary search result set, and update the semantic vector generation method of the Embedding model through rule optimization constraints according to the search coverage; When the retrieval result set of the Embedding model and the manually annotated preliminary retrieval result set reach a set retrieval threshold in terms of retrieval coverage, the optimization training is terminated and the optimized Embedding model is used for subsequent packaging.
4. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 1, characterized in that: The semantic structure optimization includes: Decomposing the query task into fine-grained subtasks according to semantic parsing rules, wherein the subtasks include key semantic units and attribute descriptions; Through the context information completion algorithm, context information related to the semantics of the preset knowledge base is added to the fine-grained subtask to update the query task.
5. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 1, characterized in that: Before inputting the manually annotated preliminary search result set into the preset re-ranking model, the method further includes: Extract multidimensional features from each search item in the preliminary search result set, calculate the feature distribution balance of the search items based on the multidimensional features, map the search items whose feature distribution balance is lower than the distribution threshold to an independent search space, and delete the corresponding search items from the preliminary search result set; Adjusting the feature density of the preliminary search result set according to the feature distribution balance; In the independent retrieval space, the retrieval items are reconstructed by a feature compression algorithm to generate a low-dimensional feature representation, which can be adapted to the ranking format of the re-ranking model.
6. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 5, characterized in that: The feature density adjustment includes: Calculate the feature distribution matrix according to the feature distribution balance of the search items in the preliminary search result set; According to the feature distribution matrix, the feature collaborative information amount of different search items in the preliminary search result set is calculated, and the density of the preliminary search result set is adjusted according to the feature collaborative information amount.
7. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 6, characterized in that: The step of fixing the output of the re-ranking model to the manually annotated optimized search result further includes: Constructing a feature mapping matrix of the low-dimensional feature representation; Performing auxiliary processing on the sorting process of the re-sorting model according to the feature mapping matrix; The optimization process of the reordering model is fused according to the feature mapping matrix.
8. The method for automatic configuration of intelligent agents based on large models, knowledge bases and tools according to claim 1, characterized in that: The method further comprises: According to the real-time query tasks input by users, the corresponding query requirements are generated through the reasoning model in the intelligent agent; Input the query requirements into the embedding model of the intelligent agent, retrieve relevant knowledge base content and generate preliminary search results; Optimizing the preliminary search results through a re-ranking model in the agent to generate a ranked search result set; When the semantic relevance of the retrieval result set is lower than the result threshold, the reasoning model is recalled to regenerate the query requirement by analyzing the semantic distribution of the retrieval result set and the user input.
9. An agent automatic configuration system based on a large model, a knowledge base and a tool, implemented based on an agent automatic configuration method based on a large model, a knowledge base and a tool as claimed in any one of claims 1 to 8, characterized in that: The system includes a reasoning model configuration module, a retrieval model configuration module and a re-ranking model configuration module; The inference model configuration module is used to receive the natural language content input by the user, parse it into a structured query task, and optimize the output of the inference model according to the manually annotated query task set so that it gradually matches the manually annotated query task set; The retrieval model configuration module is used to vectorize the structured query task using the Embedding model, search the preset knowledge base and generate a preliminary retrieval result set, and perform feature optimization and density adjustment on the preliminary retrieval result set; The re-ranking model configuration module is used to perform a ranking logic process on the optimized preliminary search result set, and adjust the ranking weight in combination with the feature mapping matrix to generate a final search result set that is consistent with the manually labeled optimization result.
Citation Information
Patent Citations
Configuration method and device of large model application agent
CN118819617A
Method and system for enhancing large language model generation by using network search
CN119271893A