Large language model optimization method and system based on external tool
Through the large language model optimization method based on external tools, the large language model is optimized using historical training data and external tools to call to optimize the large language model, the problem of excessive convergence during the training process is solved, and the accuracy and reliability of the model are improved.
Patent Information
- Application Number
- CN202510262137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
Large language models are prone to excessive convergence during training, resulting in poor training results and inaccurate responses.
A large language model optimization method based on external tools is proposed. By obtaining historical training data, keyword extraction is performed, training is performed as input to the large language model, and external tool calls are made based on keywords to obtain an alternative solution set, a reply set is generated and similarity verification is performed. If the similarity of the target reply is greater than the similarity threshold, the large language model is updated.
By using historical training data for keyword extraction and external tool calls, the model's ability to understand problems can be enhanced, the accuracy and reliability of the model can be improved, and the model can be effectively updated and optimized.
Smart Images

Figure CN120196764A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large model training, and specifically relates to a method and system for optimizing a large language model based on external tools. Background Art
[0002] Large language models are an important breakthrough in the field of artificial intelligence in recent years. Through deep learning techniques, especially models based on the Transformer architecture, they are trained on a vast amount of text data, thus possessing powerful natural language understanding and generation capabilities.
[0003] Patent CN117149989A discloses a method, a text processing method and a device for training a large language model; specifically, first, a training sample set is obtained; the training sample set includes a plurality of training samples; the plurality of training samples include a plurality of first training samples and a plurality of second training samples; the first training sample is a training sample with a prediction accuracy greater than a preset threshold; the second training sample is a training sample with a prediction accuracy less than the preset threshold; then, the initial reward model is trained based on the training sample set to obtain a trained reward model; then, the pre-trained large language model is trained based on the reward model to obtain a trained large language model. The embodiments of the present application provide a richer and better-trained data basis for training a large language model with better performance, and better meet the actual application requirements. Although some problems are solved in the above technology, there are still some problems, for example: during the training process of the large language model, over-convergence is likely to occur, resulting in poor training effects and inaccurate responses. Summary of the Invention
[0004] The object of the present invention is to solve the problem that during the training process of the large language model, over-convergence is likely to occur, resulting in poor training effects and inaccurate responses, and to propose a method and system for optimizing a large language model based on external tools.
[0005] In the first aspect of the implementation of the present invention, a method for optimizing a large language model based on external tools is first proposed, and the method includes:
[0006] Obtain historical training data, and perform keyword extraction on the historical training data to obtain a historical training data set; the historical training data includes: historical question data and historical response data;
[0007] Use the historical training data set as the input of the large language model for training, and call external tools according to keywords to obtain a set of alternative solutions;
[0008] Generate a set of responses according to the set of alternative solutions, and perform similarity verification on the target response in the set of responses;
[0009] If the similarity of the target response is greater than the similarity threshold, update the large language model.
[0010] Optionally, before using the historical training dataset as the input for training the large language model, it further includes:
[0011] Convert the text in the historical training data into a text graph structure, and identify the entities in the text to obtain the first target entities;
[0012] Match the first target entities in the knowledge graph to obtain the second target entities, and retrieve subgraphs from the knowledge graph according to the second target entities to obtain knowledge graph subgraphs;
[0013] Merge the text graph structure and the knowledge graph subgraphs to obtain an enhanced graph, and encode the enhanced graph into vectors;
[0014] Convert the vectors into a sequence form, and convert the sequence form into a text sequence to be input into the large language model for training.
[0015] Optionally, obtaining a set of alternative solutions by calling external tools according to keywords includes:
[0016] Obtain lexical data and an external tool set, and perform part-of-speech classification on the lexical data to obtain a part-of-speech classification result;
[0017] Determine a query index set according to the part-of-speech classification result, and construct a query database according to the query index set and the external tool set;
[0018] Determine the target query index of the keyword, perform a retrieval operation in the query database according to the target query index, and obtain a target external tool set;
[0019] Optimize the path of the target external tool set through the ant colony algorithm until the preset conditions are met or the maximum number of iterations is reached, then output the external tool combination, and use this external tool combination as the set of alternative solutions.
[0020] Optionally, optimizing the path of the target external tool set through the ant colony algorithm includes:
[0021] Update the ant colony algorithm through the state transition probability formula, record the tree diagram during the path optimization of the ant colony algorithm, and evaluate the effectiveness of each external tool combination in the tree diagram;
[0022] If the effectiveness score of the external tool combination is greater than the effectiveness threshold, determine that this external tool combination is effective;
[0023] Use the effective external tool combinations as common paths to update the model memory area in the large language model;
[0024] State transition probability formula:
[0025]
[0026] where τ ij is the pheromone concentration on the path i→j, η ij is the heuristic information, usually the reciprocal of the distance between two points, α and β are the weight coefficients of the pheromone and the heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when at the current node i.
[0027] Optionally, the similarity verification for the target reply in the reply set includes:
[0028] Performing word segmentation on the target reply to obtain a number of words, converting the words into word vectors and aggregating the word vectors to obtain a target sentence vector;
[0029] Determining a standard sentence vector comparison table according to the historical reply data, and calculating the cosine similarity between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain a similarity table;
[0030] Determining the reply weight according to the numerical size of the similarity in the similarity table, and outputting the reply with the largest reply weight as the target reply.
[0031] In the second aspect of the implementation of the present invention, a large language model optimization system based on external tools is proposed, including: a training data module, a training input module, a similarity verification module, and a model update module:
[0032] The training data module is used to obtain historical training data, and perform keyword extraction on the historical training data to obtain a historical training data set; the historical training data includes: historical question data and historical reply data;
[0033] The training input module is used to use the historical training data set as the input of the large language model for training, and call external tools according to keywords to obtain a set of alternative solutions;
[0034] The similarity verification module is used to generate a reply set according to the set of alternative solutions, and perform similarity verification on the target reply in the reply set;
[0035] The model update module is used to update the large language model if the similarity of the target reply is greater than the similarity threshold.
[0036] Optionally, the system further includes: a text conversion module, an entity matching module, an image enhancement module, and a model training module:
[0037] The text conversion module is used to convert the text in the historical training data into a text graph structure and identify the entities in the text to obtain the first target entity;
[0038] The entity matching module is used to match the first target entity in the knowledge graph to obtain the second target entity, and retrieve the subgraph in the knowledge graph according to the second target entity to obtain the knowledge graph subgraph;
[0039] The image enhancement module is used to merge the text graph structure and the knowledge graph subgraph to obtain the enhanced graph, and encode the enhanced graph into a vector;
[0040] The model training module is used to convert the vector into a sequence form, and convert the sequence form into a text sequence and input it into the large language model for training.
[0041] Optionally, the training input module further includes: a part-of-speech classification module, a database establishment module, an index query module, and a path optimization module:
[0042] The part-of-speech classification module is used to obtain the lexical data and the external tool set, and perform part-of-speech classification on the lexical data to obtain the part-of-speech classification result;
[0043] The database establishment module is used to determine the query index set according to the part-of-speech classification result, and construct a query database according to the query index set and the external tool set;
[0044] The index query module is used to determine the target query index of the keyword, and perform a retrieval operation in the query database according to the target query index to obtain the target external tool set;
[0045] The path optimization module is used to optimize the path of the target external tool set through the ant colony algorithm until the preset condition is met or the maximum number of iterations is reached, then output the external tool combination, and use this external tool combination as the backup solution set.
[0046] Optionally, the path optimization module further includes: a tree diagram recording module, a validity evaluation module, and a model memory area update module:
[0047] The tree diagram recording module is used to update the ant colony algorithm, record the tree diagram during the path optimization of the ant colony algorithm, and evaluate the validity of each external tool combination in the tree diagram. The ant colony algorithm is updated using the state transition probability formula;
[0048] The validity evaluation module is used to determine that the external tool combination is valid if the validity score of the external tool combination is greater than the validity threshold;
[0049] The model memory area update module is used to update the model memory area in the large language model by using the valid external tool combination as the common path;
[0050] State transition probability formula:
[0051]
[0052] where τ ij is the pheromone concentration on the path i→j, η ij is the heuristic information, usually the reciprocal of the distance between two points, α and β are the weight coefficients of the pheromone and the heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when at the current node i.
[0053] Optionally, the similarity verification module includes: a sentence vector generation module, a similarity calculation module, and a reply output module:
[0054] The sentence vector generation module is used to tokenize the target reply to obtain several words, convert the words into word vectors, and aggregate the word vectors to obtain the target sentence vector;
[0055] The similarity calculation module is used to determine the standard sentence vector comparison table according to the historical reply data, and calculate the cosine similarity between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain the similarity table;
[0056] The reply output module is used to determine the reply weight according to the similarity value in the similarity table, and output the reply with the largest reply weight as the target reply.
[0057] The present invention proposes an optimization method for a large language model based on external tools. By obtaining historical training data, keyword extraction is performed on the historical training data to obtain a historical training data set; the historical training data set is used as the input of the large language model for training, and external tool calls are made according to the keywords to obtain a set of alternative solutions; a set of replies is generated according to the set of alternative solutions, and similarity verification is performed on the target reply in the set of replies; if the similarity of the target reply is greater than the similarity threshold, the large language model is updated; by using historical training data for keyword extraction and using the extracted data set as the input of the large language model for training, the model's ability to understand problems can be enhanced. By obtaining a set of alternative solutions through external tool calls and performing similarity verification, the accuracy and reliability of the model can be improved, and effective update and optimization of the model can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described below with reference to the accompanying drawings.
[0059] Figure 1The embodiment of the present invention provides a flowchart of a method for optimizing a large language model based on external tools;
[0060] Figure 2 The embodiment of the present invention provides a framework diagram of a system for optimizing a large language model based on external tools. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The term "and / or" in this article is only a description of the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and B can represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations. In addition, the descriptions such as "first" and "second" in the present invention are only for the purpose of description, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of the technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0062] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0063] The embodiment of the present invention provides a method for optimizing a large language model based on external tools. Refer to Figure 1 , Figure 1 which is a flowchart of a method for optimizing a large language model based on external tools provided by the embodiment of the present invention. The method includes the following steps:
[0064] S101, obtain historical training data, and extract keywords from the historical training data to obtain a historical training data set;
[0065] S102, use the historical training data set as the input of the large language model for training, and call external tools according to the keywords to obtain a set of alternative solutions;
[0066] S103, generate a set of responses according to the set of alternative solutions, and verify the similarity of the target response in the set of responses;
[0067] S104, if the similarity of the target response is greater than the similarity threshold, update the large language model;
[0068] The historical training data includes: historical question data and historical reply data;
[0069] Based on an optimization method for large language models using external tools provided by an embodiment of the present invention, by extracting keywords using historical training data and using the extracted dataset as the input for training the large language model, the model's ability to understand questions can be enhanced. By calling external tools to obtain a set of alternative solutions and performing similarity verification, the accuracy and reliability of the model can be improved, and effective update and optimization of the model can be achieved.
[0070] In one implementation, by extracting keywords from the historical training data, it is possible to more precisely focus on the core information in the data, that is, to determine the semantic relationships of the historical question data in the historical training data, making the training dataset more concise and targeted, which helps to improve the accuracy and efficiency of the large language model in processing similar questions.
[0071] In one implementation, during the training process, calling external tools based on keywords can provide additional information and knowledge sources for the large language model, thereby generating a more abundant and accurate set of alternative solutions; the set of alternative solutions contains several alternative solutions, and each alternative solution corresponds to a reply; it can enhance the functionality and generalization ability of the large language model, and also improve its ability to handle complex questions; the external tool set is a comprehensive resource library (containing several external tools), which includes basic tools obtained from various channels, as well as tool path combinations discovered during the complex multi-tool joint reasoning process of the large language model and having wide application value. The external tools and their combinations are mainly expanded by network retrieval, manual addition, and synchronous update with the memory area of the large language model. An external tool is an independent software or program that assists the large language model in completing a specific task or reasoning process, and the reasonable application of the external tool combination can improve the processing ability and efficiency of the large language model for questions.
[0072] In one implementation, a reply set is generated based on the set of alternative solutions, and similarity verification is performed on the target reply in the reply set. For the questions in the historical question data, the large language model can generate several replies for each question, and each reply is compared with the standard reply in the historical reply data (similarity verification) to determine whether each reply of the large language model is reasonable. If it is reasonable, the model can be updated; if it is not reasonable, further training is required; by performing similarity verification on the target reply in the reply set, duplicate or low-quality replies can be discovered and filtered out in a timely manner, ensuring the diversity and quality of the reply set, which helps to improve the innovation and user satisfaction of the large language model when generating replies.
[0073] In one implementation, when the similarity of the target response is greater than the set similarity threshold, the large language model is updated, and the performance of the model can be continuously optimized based on new data and feedback. This iterative update mechanism enables the large language model to continuously learn and adapt, thereby improving the accuracy of the response. The self-attention mechanism can be expressed by the following formula: where Q, K, and V are the Query, Key, and Value matrices respectively, and d k is the dimension of the key vector, and softmax is a normalization function used to calculate the attention weights.
[0074] In one implementation, when training a language model, the cross-entropy loss function is usually used to optimize the model parameters. For a given input sequence X and target sequence Y, the loss function can be expressed as L = -∑logP(y i |X), where y i is the i-th word in the target sequence, and P(y i |X) is the probability that the model predicts this word.
[0075] In one implementation, the actual application scenarios of the above method include: 1. Optimization of intelligent customer service systems: In the field of intelligent customer service, the performance of customer service robots can be significantly improved through the above method. After keyword extraction from the historical question data and historical reply data in the historical training data, they are used as the input of the large language model, enabling it to more accurately understand the core semantics of user questions. Combining with the set of alternative solutions generated by external tool calls, the customer service robot can provide richer and more accurate replies. By screening high-quality replies through similarity verification and further optimizing the model, the accuracy and user satisfaction of the customer service system can be improved, reducing the intervention of human customer service and enhancing the customer service efficiency. 2. Intelligent tutoring systems in the education field: In the education field, this method can be used to develop intelligent tutoring systems. By processing and training historical teaching data (such as student questions, teacher answers, etc.), the model can more accurately understand students' questions and provide targeted answers. External tool calls can introduce more educational resources and generate more comprehensive alternative teaching plans. Similarity verification can ensure that the generated answers are consistent with high-quality teaching standards, while avoiding repetition and low-quality content. This helps to improve the teaching effect of the intelligent tutoring system and provide a more personalized learning experience for students. 3. Medical health consultation systems: In the field of medical health consultation, this method can be used to optimize intelligent health consultation systems. By processing and training historical medical data (such as patient consultation questions, doctor answers, etc.), the model can more accurately understand patients' health problems. External tool calls can introduce medical databases and professional tools and generate more accurate alternative diagnostic suggestions. Similarity verification can ensure that the generated suggestions are consistent with high-quality medical standards, while avoiding repetition and low-quality content. This helps to improve the reliability of the medical consultation system and provide more professional health advice for patients. 4. Enterprise knowledge management systems In the field of enterprise knowledge management, this method can be used to optimize the internal knowledge management system of enterprises. By processing and training historical enterprise data (such as employee questions, expert answers, etc.), the model can more accurately understand employees' business problems. External tool calls can introduce the internal knowledge base and external resources of the enterprise and generate more comprehensive alternative solutions. Similarity verification can ensure that the generated content is consistent with high-quality knowledge standards, while avoiding repetition and low-quality content. This helps to improve the efficiency and quality of the knowledge management system and provide more accurate knowledge support for enterprise employees.
[0076] In one embodiment, before step S102, it further includes:
[0077] Convert the text in the historical training data into a text graph structure and identify the entities in the text to obtain the first target entity;
[0078] Match the first target entity in the knowledge graph to obtain the second target entity, and retrieve the knowledge graph subgraph according to the second target entity in the knowledge graph;
[0079] Merge the text graph structure with the knowledge graph sub - graph to obtain an enhanced graph, and encode the enhanced graph into a vector;
[0080] Convert the vector into a sequence form, and convert the sequence form into a text sequence to be input into the large - language model for training.
[0081] In one implementation, for the conversion from text to graph structure, perform syntactic analysis on the input text to identify the grammatical structures in the sentence, such as: subject, predicate, and object, etc.; use named - entity recognition (NER) technology to identify the entities in the text, where entities are, for example: person names, locations, organizations, or other important noun phrases. Based on the results of syntactic analysis and entity recognition, predict the relationships between entities to form the edges in the graph structure, use the recognized entities as nodes, and the predicted relationships as edges to construct a graph structure, which represents the semantic information of the text.
[0082] In one implementation, link (match) the entities in the text with the corresponding entities in the knowledge graph to ensure that the entities in the text correspond to the entities in the knowledge graph, that is, to ensure that the entities in the text can correspond to the structured information in the knowledge graph. For example: if the text mentions "apple", entity linking will associate this mention with the specific node representing apple in the knowledge graph. Incorporate the additional information linked to the knowledge - graph entities into the graph structure, and the above - mentioned additional information may include the attributes of the entities, relationships with other entities, etc. Extract the relevant knowledge - graph sub - graph from the knowledge graph, where the knowledge - graph sub - graph contains the structured knowledge related to the text entities, and merge the extracted knowledge - graph sub - graph with the text's graph structure to form a new enhanced graph, which contains both the semantic information of the text and the structured knowledge of the knowledge graph.
[0083] In one implementation, integrate the factual knowledge in the knowledge graph into the large - language model in different ways to improve the model's performance on tasks that require factual knowledge and structured reasoning. In this way, the large - language model can use explicit factual knowledge for training, improve its factual - reasoning ability when generating text, and provide more accurate responses.
[0084] In one embodiment, step S102 includes:
[0085] Obtain lexical data and an external tool set, and perform part - of - speech classification on the lexical data to obtain the part - of - speech classification result;
[0086] Determine the query index set according to the part - of - speech classification result, and construct a query database according to the query index set and the external tool set;
[0087] Determine the target query index of the keyword, and perform a retrieval operation in the query database according to the target query index to obtain the target external tool set;
[0088] Optimize the path of the target external tool set through the ant colony algorithm until the preset condition is met or the maximum number of iterations is reached, then output the external tool combination, and use this external tool combination as the backup solution set.
[0089] In one implementation, by classifying the part of speech of the lexical data (for example: machine learning algorithms learn the mapping relationship between words and parts of speech by training a large amount of labeled data, so as to be able to accurately classify the part of speech of new lexical data), that is, judging the subject, predicate, object, etc., the function and expressive meaning of the word in the sentence or text can be understood more precisely, which helps the large language model to better grasp the semantic structure of the language, thereby improving the model's understanding and processing ability of the text content.
[0090] In one implementation, the query index set constructed according to the part-of-speech classification results, combined with the external tool set, can build an efficient query database. This database can quickly respond to keyword queries and provide relevant external tool sets, thus accelerating the speed at which the large language model obtains relevant resources and information when processing problems.
[0091] In one implementation, by optimizing the path of the target external tool set through the ant colony algorithm, external tools can be intelligently selected and combined; among them, the external tool combination is obtained by matching external tools, which is equivalent to decomposing the problem, matching external tools for each part, and solving the problem through different permutations and combinations. The above optimization process can ensure that the model can call the most suitable tools and resources when dealing with complex problems, improving the model's problem-solving ability and efficiency.
[0092] In one implementation, the external tool set is regarded as a node network, where each tool corresponds to a node, and the association or dependency relationship between tools is regarded as the connection between nodes. The goal is to find the optimal path from the starting tool to the target tool through the ant colony algorithm. According to the characteristics of the external tool set and the requirements of path optimization, parameters of the ant colony algorithm are set, such as the number of ants, the number of iterations, the pheromone evaporation coefficient, etc. The initial positions of the ants are randomly generated, and the pheromone matrix and the heuristic information matrix are initialized. The heuristic information can be set according to the degree of association or importance between tools. Each ant selects the next tool node to visit based on the pheromone concentration and the heuristic information of the current node. In this process, ants tend to choose nodes with high pheromone concentration and high degree of association. The pheromone concentration on the path is updated according to the path length of the ant and the pheromone evaporation coefficient. The increase in pheromone on the shorter path is relatively more to attract more ants to choose this path. Through multiple iterations, the ant colony will gradually converge to the optimal path, thus finding the optimal path from the starting tool to the target tool, and then outputting the optimal path and the corresponding tool sequence. Through the positive feedback mechanism of pheromone, the ant colony algorithm can gradually converge to the global optimal solution or an approximate optimal solution. Several ants can perform path search simultaneously, with strong parallelism and distributed characteristics, suitable for solving large-scale problems. The algorithm can automatically adjust the path selection strategy according to environmental changes, with strong self-adaptability. Due to the diversity of the ant colony, the algorithm has a certain robustness to noise and interference in the environment.
[0093] In one embodiment, path optimization is performed on the target external tool set through the ant colony algorithm, including:
[0094] The ant colony algorithm is updated through the state transition probability formula, and the tree diagram during the path optimization of the ant colony algorithm is recorded, and the effectiveness of each external tool combination in the tree diagram is evaluated;
[0095] If the effectiveness score of the external tool combination is greater than the effectiveness threshold, it is determined that the external tool combination is effective;
[0096] The effective external tool combination is used as the common path to update the model memory area in the large language model;
[0097] State transition probability formula:
[0098]
[0099] where τ ij is the pheromone concentration on path i→j, η ij is the heuristic information, usually the reciprocal of the distance between two points, and α and β are the weight coefficients of the pheromone and the heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when at the current node i.
[0100] In one implementation, the effectiveness of each external tool combination in the tree diagram is evaluated, where the effectiveness evaluation is carried out through success rate and efficiency. Success rate: It refers to the proportion of successfully solving problems using a specific tool path. Record the number of times of solving problems using this tool path and calculate its ratio to the total number of attempts. Efficiency: It refers to the time or resource consumption required to solve problems using a specific tool path. Quantify the efficiency by timing the total time required from the start to the solution of the problem, or by evaluating the computing resources (such as CPU, memory, etc.) used during the execution of the path. Effectiveness evaluation = α * success rate + β * efficiency, where α and β are constant proportionality coefficients. When the external tool combination is effective, it means that the external tool combination has a good effect. By using the ant colony algorithm to optimize the path of the external tool combination, it can intelligently explore and discover the optimal or near-optimal tool combination path. The above optimization process reduces the time for the model to search and select tools when processing tasks, and updating the model memory area in the language model can improve the overall efficiency of the application of the tool combination.
[0101] In one implementation, record the tree diagram when the ant colony algorithm performs path optimization, which can display each step of decision-making and path selection in the search process of the algorithm, help backtrack and analyze the performance and effect of the algorithm, and also provide valuable external tool combinations for subsequent tool combination optimization. The state transition probability formula (the probability P ij ) for an ant to move from the current node i to the next node j: τ ij is the pheromone concentration on the path i→j, η ij is the heuristic information, usually the reciprocal of the distance between two points. α and β are the weight coefficients of the pheromone and the heuristic information. In the state transition probability formula of the ant colony algorithm, k∈allowed represents the set of next nodes that the ant can choose when at the current node i. allowed is a constraint condition used to ensure that the ant does not repeatedly visit the nodes that have already been traversed when choosing the next node, thus ensuring the feasibility and integrity of the path. The pheromone update is divided into two parts: evaporation and enhancement. Pheromone evaporation (after each iteration, the pheromone on all paths will evaporate at a certain rate): τ ij ←(1 - ρ)·τ ij , where ρ is the pheromone evaporation rate. Pheromone enhancement (after a round of search, the pheromone on the excellent paths will increase): where m is the number of ants, and Δτ ij (k) is the pheromone released by ant k on the path i→j, usually determined by the path quality (such as path length). The ant colony algorithm runs through the following steps: Initialization: Set the initial pheromone concentration τ ij and the heuristic information ηij Path selection: Ants select a path according to the state transition probability formula. Pheromone update: Update the pheromone according to the evaporation and enhancement rules. Iterative optimization: Repeat path selection and pheromone update until the termination condition is met.
[0102] In one implementation, for the state transition probability formula, assume there is a graph with 4 nodes, and the node numbers are 0, 1, 2, and 3;
[0103] Pheromone concentration τ ij :
[0104] Heuristic information η ij :
[0105] When α is 1, β is 2, the current node is 0, the visited node is 0, and the optional nodes are {1, 2, 3}, calculate the transition probabilities from node 0 to other nodes:
[0106] Node 1:
[0107] Node 2:
[0108] Node 3:
[0109] In one implementation, evaluate the effectiveness of each external tool combination in the tree diagram and set an effectiveness threshold, which can ensure that only high-quality tool combinations are retained and used as common paths. This screening mechanism improves the accuracy and reliability of the model in calling tool combinations when processing tasks.
[0110] In one embodiment, step S103 includes:
[0111] Segment the target reply to obtain several words, convert the words into word vectors, and aggregate the word vectors to obtain the target sentence vector;
[0112] Determine the standard sentence vector comparison table according to the historical reply data, and calculate the cosine similarity between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain the similarity table;
[0113] Determine the reply weight according to the numerical size of the similarity in the similarity table, and output the reply with the largest reply weight as the target reply.
[0114] In one implementation, by converting the words in the target response into word vectors, the system can represent the semantic information of the words in the form of numerical vectors. This representation method not only preserves the semantic relationships between words, but also enables the system to process and understand natural language texts more flexibly; by aggregating several word vectors, the system can generate a sentence vector representing the entire target response. This aggregation process effectively integrates the semantic information of each word in the sentence, enabling the target sentence vector to more comprehensively reflect the overall characteristics of the response. For example: the sentence vector representations of two responses are A and B respectively, and they are composed of the following elements: A = (x1, x2, x3,..., xn) and B = (y1, y2, y3,..., yn), and the dot product calculation formula is: A * B = x1y1 + x2y2 + … + x n y n , calculate the modulus length: The modulus length of a vector is the square root of the sum of the squares of each element of the vector. The modulus length calculation formulas for vector A and vector B are respectively: Calculate the cosine similarity: The cosine similarity is the ratio of the dot product to the product of the moduli of the two vectors. The calculation formula is: The closer the value of the similarity is to 1, the more similar the two vectors are; the closer it is to -1, the less similar they are; when the value is 0, it means that the two vectors are orthogonal, that is, there is no correlation.
[0115] In one implementation, by calculating the cosine similarity between the target sentence vector and each standard sentence vector in the standard sentence vector comparison table, the system can quantify the similarity between the target response and the historical response. This calculation method is not only highly accurate, but also has high calculation efficiency, and can quickly screen out the historical response closest to the target response.
[0116] In one implementation, through steps such as the conversion from words to word vectors, the aggregation of target sentence vectors, the calculation of cosine similarity, and the determination of response weights, the precise matching and efficient screening of natural language responses are achieved, thereby improving the accuracy and efficiency of the large language model in processing natural language tasks.
[0117] Based on the same inventive concept, the embodiments of the present invention also provide a large language model optimization system based on external tools. Refer to Figure 2 , Figure 2 which is the structural schematic diagram of a large language model optimization system based on external tools provided by the embodiments of the present invention, including: a training data module, a training input module, a similarity verification module, and a model update module:
[0118] The training data module is used to obtain historical training data and extract keywords from the historical training data to obtain a historical training data set; the historical training data includes: historical question data and historical response data;
[0119] The training input module is used to train the large language model with the historical training dataset as the input, and call external tools according to keywords to obtain a set of alternative solutions;
[0120] The similarity verification module is used to generate a set of responses according to the set of alternative solutions, and perform similarity verification on the target response in the set of responses;
[0121] The model update module is used to update the large language model if the similarity of the target response is greater than the similarity threshold.
[0122] Based on an optimization system for a large language model based on external tools provided by an embodiment of the present invention, by using historical training data for keyword extraction and using the extracted dataset as the input for training the large language model, the model's ability to understand problems can be enhanced. By calling external tools to obtain a set of alternative solutions and performing similarity verification, the accuracy and reliability of the model can be improved, and effective update and optimization of the model can be achieved.
[0123] In one embodiment, the system further includes: a text conversion module, an entity matching module, an image enhancement module, and a model training module:
[0124] The text conversion module is used to convert the text in the historical training data into a text graph structure, and identify the entities in the text to obtain the first target entity;
[0125] The entity matching module is used to match the first target entity in the knowledge graph to obtain the second target entity, and perform subgraph retrieval in the knowledge graph according to the second target entity to obtain a knowledge graph subgraph;
[0126] The image enhancement module is used to merge the text graph structure and the knowledge graph subgraph to obtain an enhanced graph, and encode the enhanced graph into a vector;
[0127] The model training module is used to convert the vector into a sequence form, and convert the sequence form into a text sequence and input it into the large language model for training.
[0128] In one embodiment, the training input module further includes: a part-of-speech classification module, a database establishment module, an index query module, and a path optimization module:
[0129] The part-of-speech classification module is used to obtain lexical data and an external tool set, and perform part-of-speech classification on the lexical data to obtain a part-of-speech classification result;
[0130] The database establishment module is used to determine a query index set according to the part-of-speech classification result, and construct a query database according to the query index set and the external tool set;
[0131] The index query module is used to determine the target query index of the keyword, perform a retrieval operation in the query database according to the target query index, and obtain the target external tool set;
[0132] The path optimization module is used to optimize the path of the target external tool set through the ant colony algorithm until the preset condition is met or the maximum number of iterations is reached, then output the external tool combination, and use this external tool combination as the alternative solution set.
[0133] In one embodiment, the path optimization module further includes: a tree diagram recording module, a validity evaluation module, and a model memory area update module:
[0134] The tree diagram recording module is used to update the ant colony algorithm through the state transition probability formula, record the tree diagram during the path optimization of the ant colony algorithm, and evaluate the validity of each external tool combination in the tree diagram;
[0135] The validity evaluation module is used to determine that the external tool combination is valid if the validity score of the external tool combination is greater than the validity threshold;
[0136] The model memory area update module is used to use the valid external tool combination as the common path to update the model memory area in the large language model;
[0137] State transition probability formula:
[0138]
[0139] where τ ij is the pheromone concentration on path i→j, η ij is the heuristic information, usually the reciprocal of the distance between two points, α and β are the weight coefficients of the pheromone and heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when at the current node i.
[0140] In one embodiment, the similarity verification module includes: a sentence vector generation module, a similarity calculation module, and a reply output module:
[0141] The sentence vector generation module is used to tokenize the target reply to obtain several words, convert the words into word vectors, and aggregate the word vectors to obtain the target sentence vector;
[0142] The similarity calculation module is used to determine the standard sentence vector comparison table according to the historical reply data, and calculate the cosine similarity between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain the similarity table;
[0143] The reply output module is used to determine the reply weight according to the numerical value of the similarity in the similarity table, and output the reply with the largest reply weight as the target reply.
[0144] The above has described in detail an embodiment of the present invention. However, the content is only a preferred embodiment of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made in accordance with the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A large language model optimization method based on external tools, characterized in that: The method comprises: Acquire historical training data, and extract keywords from the historical training data to obtain a historical training data set; the historical training data includes: historical question data and historical reply data; The historical training data set is used as an input of a large language model for training, and an external tool is called according to keywords to obtain a set of backup solutions; Generate a response set according to the backup solution set, and perform similarity verification on the target response in the response set; If the similarity of the target response is greater than a similarity threshold, the large language model is updated.
2. The large language model optimization method based on external tools according to claim 1, characterized in that: Before using the historical training data set as input for training the large language model, the following steps need to be performed: Converting the text in the historical training data into a text graph structure, and identifying entities in the text to obtain a first target entity; Match the first target entity in the knowledge graph to obtain the second target entity, and perform subgraph retrieval in the knowledge graph based on the second target entity to obtain a knowledge graph subgraph; Merging the text graph structure with the knowledge graph subgraph to obtain an enhanced graph, and encoding the enhanced graph into a vector; The vector is converted into a sequence form, and the sequence form is converted into a text sequence, which is then input into a large language model for training.
3. The large language model optimization method based on external tools according to claim 1, characterized in that: The backup solution set obtained by calling external tools based on keywords includes: Acquire vocabulary data and an external tool set, and perform part-of-speech classification on the vocabulary data to obtain a part-of-speech classification result; Determine a query index set according to the part-of-speech classification result, and construct a query database according to the query index set and the external tool set; Determine a target query index of the keyword, perform a search operation in the query database according to the target query index, and obtain a target external tool set; The target external tool set is optimized by using an ant colony algorithm until a preset condition is met or a maximum number of iterations is reached, then an external tool combination is output and used as a backup solution set.
4. The large language model optimization method based on external tools according to claim 3, characterized in that: Optimizing the path of the target external tool set by using the ant colony algorithm includes: The ant colony algorithm is updated through the state transition probability formula, the tree diagram when the ant colony algorithm performs path optimization is recorded, and the effectiveness of each external tool combination in the tree diagram is evaluated; If the effectiveness score of the external tool combination is greater than the effectiveness threshold, the external tool combination is determined to be effective; Use a combination of effective external tools as a common path to update the model memory area in the large language model; The state transition probability formula is: Among them, τ ij is the pheromone concentration on path i→j, η ij is the heuristic information, usually the inverse of the distance between two points, α and β are the weight coefficients of pheromone and heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when it is at the current node i.
5. The large language model optimization method based on external tools according to claim 1, characterized in that: The similarity verification for the target response in the response set includes: Segment the target response to obtain several words, convert the words into word vectors and aggregate the word vectors to obtain the target sentence vector; Determine a standard sentence vector comparison table according to the historical reply data, and calculate the cosine similarity between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain a similarity table; The reply weight is determined according to the numerical value of the similarity in the similarity table, and the reply with the largest reply weight is output as the target reply.
6. A large language model optimization system based on external tools, characterized in that: The system includes: a training data module, a training input module, a similarity verification module and a model update module: The training data module is used to obtain historical training data, and extract keywords from the historical training data to obtain a historical training data set; the historical training data includes: historical question data and historical reply data; The training input module is used to use the historical training data set as the input of the large language model for training, and to call external tools according to keywords to obtain a set of backup solutions; The similarity verification module is used to generate a response set according to the backup solution set, and perform similarity verification on the target response in the response set; The model updating module is used to update the large language model if the similarity of the target response is greater than a similarity threshold.
7. The large language model optimization system based on external tools according to claim 6, characterized in that: The system also includes: a text conversion module, an entity matching module, an image enhancement module and a model training module: The text conversion module is used to convert the text in the historical training data into a text graph structure, and identify entities in the text to obtain a first target entity; The entity matching module is used to match the first target entity in the knowledge graph to obtain the second target entity, and perform subgraph retrieval in the knowledge graph according to the second target entity to obtain the knowledge graph subgraph; The image enhancement module is used to merge the text graph structure with the knowledge graph subgraph to obtain an enhanced graph, and encode the enhanced graph into a vector; The model training module is used to convert the vector into a sequence form, and convert the sequence form into a text sequence and input it into a large language model for training.
8. The large language model optimization system based on external tools according to claim 6, characterized in that: The training input module also includes: a part-of-speech classification module, a database establishment module, an index query module and a path optimization module: The part-of-speech classification module is used to obtain vocabulary data and an external tool set, and to perform part-of-speech classification on the vocabulary data to obtain a part-of-speech classification result; The database building module is used to determine a query index set according to the part-of-speech classification result, and to build a query database according to the query index set and the external tool set; The index query module is used to determine a target query index of a keyword, and perform a search operation in the query database according to the target query index to obtain a target external tool set; The path optimization module is used to optimize the path of the target external tool set by using an ant colony algorithm until a preset condition is met or a maximum number of iterations is reached, then an external tool combination is output and used as a backup solution set.
9. The large language model optimization system based on external tools according to claim 8, characterized in that: The path optimization module also includes: a tree diagram recording module, a validity evaluation module and a model memory area updating module: The tree diagram recording module is used to update the ant colony algorithm, record the tree diagram when the ant colony algorithm performs path optimization, evaluate the effectiveness of each external tool combination in the tree diagram, and the ant colony algorithm is updated using the state transition probability formula; The effectiveness evaluation module is used to determine that the external tool combination is effective if the effectiveness score of the external tool combination is greater than the effectiveness threshold; The model memory area update module is used to update the model memory area in the large language model by combining valid external tools as a common path; State transition probability formula: Among them, τ ij is the pheromone concentration on path i→j, η ij is the heuristic information, usually the inverse of the distance between two points, α and β are the weight coefficients of pheromone and heuristic information, and k∈allowed represents the set of next nodes that the ant can choose when it is at the current node i.
10. The large language model optimization system based on external tools according to claim 6, characterized in that: The similarity verification module includes: a sentence vector generation module, a similarity calculation module and a reply output module: The sentence vector generation module is used to segment the target response to obtain a number of words, convert the words into word vectors and aggregate the word vectors to obtain the target sentence vector; The similarity calculation module is used to determine a standard sentence vector comparison table according to the historical reply data, and perform cosine similarity calculation between each standard sentence vector in the standard sentence vector comparison table and the target sentence vector to obtain a similarity table; The reply output module is used to determine the reply weight according to the numerical value of the similarity in the similarity table, and output the reply with the largest reply weight as the target reply.
Citation Information
Patent Citations
Large language model training method and device and text processing method and device
CN117149989A