Large language model networking query method, system and equipment based on active triggering
By introducing a trigger token mechanism into the large language model, dynamically adjusting the timing and content of network searches, the problem that large language model is difficult to deeply explore user intentions in the network search task is solved, and more efficient and intelligent information acquisition and answer generation are achieved.
Patent Information
- Application Number
- CN202510577985.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When performing network search tasks, it is difficult for large language models to deeply explore the real intentions and potential needs of users, resulting in a deviation from user expectations. The network search process of traditional methods is fixed and cannot be dynamically adjusted according to the real-time needs of the model.
Through pre-training or prompt word engineering, large language models can learn to identify the context locations that need to trigger network searches during the process of generating answer text, and automatically generate and insert trigger tokens. After the token is detected, the answer generation is paused, the network search is performed, the summary is extracted, and the summary information is fused with the original question and historical generated content through a dynamic gate mechanism to form new inputs and continue to generate the answer text.
It realizes the autonomy of large language models in information acquisition decision-making, can respond to complex and diverse user problems more flexibly and intelligently, significantly enhances the universality and adaptability of model applications, and improves the accuracy and practicality of answer content.
Smart Images

Figure CN120087483A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and specifically relates to a method, system, and device for querying the Internet by a large language model based on active triggering. Background Art
[0002] When current large language models (LLMs) execute Internet search tasks, they generally follow a fixed process pattern: First, the user input question is initially rewritten to generate a standardized query statement; subsequently, this query statement is used to retrieve relevant network information from external information sources; finally, the retrieved results are re-input into the model for summarization and generation of a response.
[0003] However, this traditional method exposes many deficiencies in practical applications: (1) The generation process of the query statement mainly relies on a superficial understanding of the user's initial question, making it difficult to deeply explore the user's true intentions and potential needs, resulting in a deviation between the query results and the user's expectations.
[0004] (2) Due to the limitations of the query rewriting rules, the generated query statement may not accurately express the user's intentions, thus leading to a mismatch between the retrieved information and the user's actual needs, affecting the accuracy and effectiveness of the answer.
[0005] (3) The Internet search process of the traditional method is usually preset and fixed, and cannot dynamically adjust the timing and content of the Internet query according to the real-time needs of the model during the text generation process, thereby limiting the improvement space of the answer quality. Summary of the Invention
[0006] In view of the above problems existing when the large language model (LLM) executes the Internet search task, the present invention provides a method, system, and device for querying the Internet by a large language model based on active triggering.
[0007] In a first aspect, the technical solution of the present invention provides a method for querying the Internet by a large language model based on active triggering, including the following steps: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger Internet search during the generation of the answer text, and autonomously generate and insert trigger tokens; a token is the basic unit constituting the text, and a phrase or expression used to trigger Internet query is defined as a trigger token; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, execute an Internet search to obtain retrieval results, and extract a summary of the retrieval results; S3. Fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; S4. Re-enter the new input into the large language model and continue to generate the response text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0008] As a further limitation of the technical solution of the present invention, in S1, training the large language model includes: S11. Construct a training data set to enable the large language model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training data set includes user questions, response texts annotated with trigger tokens, search decision annotations, and search results; S12. Design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the response text generation process; S13. Optimize the search decision-making ability of the large language model based on reinforcement learning.
[0009] As a further limitation of the technical solution of the present invention, the loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function. Among them, the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0010] As a further limitation of the technical solution of the present invention, the comprehensive loss function ; Language modeling loss function ; Trigger classification loss function ; Single-position query generation loss function ; Overall query generation loss formula ; In the formula, is the probability predicted by the model, indicating that given the context x and the previously generated text , the probability that the next token is ; is the probability predicted by the model, indicating the probability of inserting a trigger token at position t; The m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, that is ; is the probability predicted by the model, indicating that given the context , the previously generated text and the previous part of the query , the probability that the next query token is The probability; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, is the true label.
[0011] As a further limitation of the technical solution of the present invention, in S13, the search decision-making ability of the large language model is dynamically adjusted through a hierarchical reward function and human feedback, specifically including: Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality; Adjust the model parameters based on user feedback or automatic scoring; The reward function is: ; In the formula, is the search trigger reward, is the query quality reward, is the answer quality reward.
[0012] As a further limitation of the technical solution of the present invention, in S2, the information integration steps for multiple retrievals include: Use a pre-trained embedding model to convert the retrieval results into vectors; Calculate the cosine similarity between the new retrieval results and the existing information; If the similarity is greater than the set similarity threshold, delete duplicate content; Use K-means to cluster the deduplicated retrieval results to generate multiple text clusters; Generate summaries for each cluster based on the large language model.
[0013] As a further limitation of the technical solution of the present invention, the formula for information fusion in S3 is as follows:
[0014]
[0015] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, W is the trainable parameter. The summary information is fused with the original question and the historical generated content through a dynamic gating mechanism, enabling the model to intelligently adjust the weights of each part of the information according to the specific situation, fully utilize the coherence of the existing generated content and the new information obtained from the search, effectively avoid information redundancy or conflict, form a new input with logical coherence and rich content, and thus assist the large language model to generate more comprehensive, accurate, and fluent answer texts.
[0016] As a further limitation of the technical solution of the present invention, when triggering the network search in step S2, the number of network queries N is synchronously recorded; where the initial N = 0, and each time the network search is triggered N = N + 1; Step S5 specifically includes: S51. Determine whether the large language model generates a trigger token again; If not, directly output the generated text; If so, execute S52; S52. Determine whether N reaches the preset maximum query times Nmax; If N ≥ Nmax, terminate the network search query and output the current text; If N < Nmax, execute S53; S53. Determine whether the comprehensive matching degree between the newly added abstract and the existing answer is lower than the preset threshold; If so, terminate the network search query and output the current text; If not, return to execute step S2.
[0017] In a second aspect, the technical solution of the present invention provides a large language model network query system based on active triggering, including an input parsing module, a large language model processing module, a network search control module, a retrieval processing module, an information fusion module, and a loop control module; The input parsing module is used to receive the question and context information input by the user and transmit it to the large language model processing module; The large language model processing module is used to generate text output, and identify the context position that needs to trigger the network search during the process of generating the answer text, and autonomously generate and insert a trigger token, where the trigger token contains the query content; The network search control module is used to pause the text generation when detecting the trigger token, trigger the network search, and record the number of network queries; The retrieval processing module is used to execute the network search to obtain the retrieval result and extract the abstract of the retrieval result; The information fusion module is used to fuse the abstract information with the original question and historical generated content through a dynamic gating mechanism to form a new input; input the new input into the large language model again to continue generating the answer text; The loop control module is used to monitor the number of network searches, and when the number of network searches reaches the preset maximum value, control to stop triggering the network search.
[0018] As a further limitation of the technical solution of the present invention, the system further includes a model training module, which is specifically used to construct a training data set to enable the large language model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the answer text generation process; and optimize the search decision-making ability of the large language model based on reinforcement learning.
[0019] Construct a training data set containing multiple elements, design a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, and optimize the search decision-making ability based on reinforcement learning to train and optimize the large language model multi-dimensionally and systematically. Conduct targeted optimization at multiple levels such as language generation, search trigger judgment, and query content generation, effectively improving the comprehensive performance of the model in the online query task and enabling it to better adapt to complex task requirements.
[0020] As a further limitation of the technical solution of the present invention, the loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, where the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0021] Comprehensive loss function ; Language modeling loss function ; Trigger classification loss function ; Single-position query generation loss function ; Overall query generation loss formula ; In the formula, is the probability predicted by the model, indicating that given the context x and the previously generated text , the probability that the next token is , is the probability predicted by the model, indicating the probability of inserting a trigger token at position t, the m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, that is ; is the probability predicted by the model, indicating that given the context , the previously generated text and the previous part of the query , the probability that the next query token is The probability; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, is the true label.
[0022] As a further limitation of the technical solution of the present invention, the model training module dynamically adjusts the search decision-making ability of the large language model through a hierarchical reward function and human feedback, specifically including: Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality; Adjust the model parameters based on user feedback or automatic scoring; The reward function is: ; is the search trigger reward, is the query quality reward, is the answer quality reward.
[0023] By evaluating the search trigger timing, query content relevance, and final answer quality through a hierarchical reward function, and dynamically adjusting the model parameters based on user feedback or automatic scoring, a closed-loop mechanism for model self-optimization is established. This mechanism enables the model to continuously learn and improve in practical applications, continuously enhancing the search decision-making ability and answer quality, and ensuring the dynamic adaptation of the model performance to user needs and actual application scenarios.
[0024] As a further limitation of the technical solution of the present invention, the retrieval processing module integrates the information retrieved multiple times, specifically including: using a pre-trained embedding model to convert the retrieval results into vectors; calculating the cosine similarity between the new retrieval results and the existing information; deleting duplicate content if the similarity is greater than the set similarity threshold; using K-means to cluster the de-duplicated retrieval results to generate multiple text clusters; generating summaries for each cluster based on the large language model.
[0025] The formula for the information fusion module to perform information fusion is as follows:
[0026]
[0027] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is the trainable parameter.
[0028] The networked search control module includes: A counting sub-module for maintaining the networked query count N; A termination judgment sub-module, used to N ≥ N When it reaches max or the comprehensive matching degree between the newly added abstract and the existing answers is lower than the preset threshold, forcefully terminate the online query.
[0029] Synchronously record the number of online queries and set the maximum number of queries, while ensuring sufficient information is obtained, effectively avoiding problems such as resource waste and reduced answer efficiency caused by excessive searching, reasonably controlling the query process, enabling the model to achieve a good balance between information acquisition and answer generation efficiency, improving the user experience and system operation efficiency.
[0030] In a third aspect, the technical solution of the present invention further provides an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for querying the large language model for networking based on active triggering as described in the first aspect.
[0031] From the above technical solutions, it can be seen that the present application has the following advantages: Through pre-training or prompt engineering, the large language model can autonomously identify the context position where online search needs to be triggered and insert a trigger token, realizing the transformation from passively receiving instructions to actively judging search needs, greatly enhancing the autonomy of the model in information acquisition decision-making, being able to respond to complex and diverse user questions more flexibly and intelligently, avoiding blind search or excessive reliance on manual intervention, and significantly enhancing the versatility and adaptability of model applications. When the model recognizes the trigger token, it pauses answer generation, performs an online search and extracts an abstract, ensuring that the obtained information is the latest and highly relevant to the question, making up for the deficiency of limited timeliness of the training data of the large language model. Especially when answering questions in fields such as real-time news and the latest research results, it can significantly improve the accuracy and practicality of the answer content and provide more valuable information services for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the present application, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention.
[0034] Figure 2 It is a block diagram of the system provided by the embodiment of the present invention. Detailed Implementation Manner
[0035] To make the application purpose, features, and advantages of the present application more obvious and understandable, the following will use specific embodiments and accompanying drawings to clearly and completely describe the technical solutions protected by the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0036] As Figure 1 shown, an embodiment of the present invention provides a method for querying the Internet of a large language model based on active triggering, including the following steps: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger Internet search during the process of generating answer texts, and autonomously generate and insert trigger tokens; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, execute an Internet search to obtain retrieval results, and perform abstract extraction on the retrieval results; S3. Fuse the abstract information with the original question and historical generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0037] In some embodiments, in S1, the training of the large language model includes: S11. Construct a training data set to enable the large language model to learn to dynamically insert trigger tokens during generation and generate corresponding search queries. The training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; To enable the large language model (LLM) to have the ability to actively trigger Internet search, the present application designs a specific training data construction method. The training data set includes the following key information: Question: The original question input by the user.
[0038] Answer: The answer text generated by the model, with the trigger token [search_web("query")] inserted at the position where Internet search is required.
[0039] Search decision annotation: Annotate which parts of the answer need to trigger a search and the specific search query content.
[0040] Search Results: For each query, annotate the retrieved information and its summary.
[0041] Example: User question: "Please introduce the major breakthroughs in AI technology in 2024 and provide links to relevant papers." The answer in the training data is annotated as: The breakthroughs in AI technology in 2024 include [search_web("2024 breakthroughs in AI technology")], and relevant papers can be referred to the following [search_web("AI technology papers in 2024")].
[0042] This method of dataset construction enables the model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. Traditional model training data only focuses on text generation and does not contain explicit annotations for search decisions. The present invention trains the model to have the ability to autonomously judge and generate search instructions by annotating search trigger points and query content.
[0043] S12. Design a loss function for multi-task learning to dynamically trigger searches during the generation of answer texts by the large language model; To train the model to generate both fluent text and accurate search decisions simultaneously, this application adopts a loss function for multi-task learning, including the following parts: Language Modeling Loss (LM Loss): Optimize the grammar and semantic consistency of the answer. Calculated using cross-entropy loss, formula: ; In the formula : The t-th token in the sequence. : All tokens before the t-th token in the sequence, i.e., . : Input context. : The probability predicted by the model, indicating the probability that the next token is given the context and the previous tokens.
[0044] Search Decision Loss: Divided into two parts: Trigger Classification Loss: Judge whether a trigger token needs to be inserted at each token position, and use cross-entropy loss to implement a binary classification task. Formula:
[0045] In the formula, True label. : A trigger token needs to be inserted at position t. : No insertion is required at position t. : The probability predicted by the model, representing the likelihood of inserting a trigger token at position (t).
[0046] Query generation loss: To predict the correct search query content, sequence generation loss is calculated. Assume that at position (t), a query needs to be generated , where (M) is the length of the query. This is a sequence generation task, similar to language modeling.
[0047] Single-position query generation loss formula:
[0048] In the formula, : The m-th token in the query sequence. : All tokens before the m-th token in the query sequence, i.e., . : The probability predicted by the model, representing that given the context , the previously generated text and the previous part of the query , the probability that the next query token is .
[0049] Overall query generation loss formula:
[0050] In the formula, is the set of all positions where a trigger token needs to be inserted. For all positions where a trigger token is inserted, the query generation loss at each position is accumulated to obtain the total query generation loss.
[0051] The combined loss function is:
[0052] In the formula, is the language modeling loss, ensuring that the generated text is fluent. is the trigger classification loss, ensuring that the model accurately determines whether to insert a trigger token. is the query generation loss, ensuring that the generated query content is accurate. and are hyperparameters used to adjust the relative importance of each part of the loss. For example, increasing will make the model pay more attention to the trigger decision, and increasing will make the model pay more attention to the query generation quality.
[0053] Traditional models only use language modeling loss and focus on text generation. This application introduces search decision loss, enabling the model to trigger searches dynamically during the generation process, enhancing intelligence and flexibility.
[0054] S13. Optimize the search decision-making ability of large language models based on reinforcement learning.
[0055] Reinforcement Learning with Human Feedback (RLHF) is used in the present invention to optimize the search decision-making ability of the model.
[0056] The reinforcement learning framework includes the following core elements: State: The state is composed of the context of the current conversation, including the question raised by the user, the answer fragments generated by the model, the retrieved external information, and the conversation history. This multi-dimensional state design enables the model to comprehensively perceive the current task requirements.
[0057] Action: The actions of the model include two categories: Whether to insert the trigger token [search_web("query")] to trigger an external search.
[0058] If the search is triggered, generate specific search query content (e.g., "Global temperature change data in 2023").
[0059] Reward: The reward signal is used to evaluate the quality of the model's decision-making.
[0060] The reward function is the key to reinforcement learning and directly affects the optimization direction of the model. The present invention designs a multi-level reward mechanism: Search trigger reward ( ) If the question requires external information and the model correctly triggers the search, the reward is positive (e.g., +1).
[0061] If the question does not require a search but the model wrongly triggers it, or if a search is needed but not triggered, the reward is negative (e.g., -0.5).
[0062] Query quality reward ( ) The relevance between the query content and the user's question is evaluated by manual annotation or a semantic similarity model (e.g., cosine similarity based on BERT). If the query is precise and efficient, the reward is high (e.g., +0.8); if the query is vague or irrelevant, the reward is low (e.g., -0.3).
[0063] Answer quality reward ( ) The fluency, accuracy, and completeness of the final answer are calculated through user feedback or automatic evaluation (e.g., BLEU score, ROUGE score) and used as part of the long-term reward.
[0064] The comprehensive reward formula is as follows:
[0065] This hierarchical design ensures that the model not only learns when to search but also optimizes the search content and the quality of answers. Traditional large language models are usually based on pre-training and fine-tuning, lacking dynamic optimization of search decisions. Through RLHF in this application, the model can adjust its behavior according to real-time feedback, especially showing higher intelligence in complex and information-deficient questions.
[0066] The formula for the information fusion module to perform information fusion is as follows:
[0067]
[0068] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is the trainable parameter.
[0069] The specific training steps are as follows: Construct a training dataset that includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results. This dataset is used to let the large language model learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries.
[0070] Initialize the parameters of the large language model and set hyperparameters, such as α and β in the comprehensive loss function, and the trainable parameter W in the information fusion formula.
[0071] Based on the set loss function, perform the following operations in each training cycle: Input the user question x in the training dataset, and the large language model starts to generate an answer text. During the generation process, the model attempts to dynamically insert trigger tokens and generate corresponding search queries.
[0072] Calculate the prediction probability at each time step: For the language modeling loss function, calculate where is the true token at the current time step, and are all the tokens generated previously.
[0073] For the trigger classification loss function, calculate , indicating at time stept The probability of triggering a search is the search decision annotation.
[0074] For the query generation loss function, calculate , where is the m th token in the query sequence, is all the tokens before the m th token in the query sequence.
[0075] According to the above predicted probability, calculate the comprehensive loss function .
[0076] Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters.
[0077] Use an optimization algorithm (such as Stochastic Gradient Descent, Adam, etc.) to update the model parameters according to the calculated gradient.
[0078] Dynamically adjust the search decision-making ability of the large language model through a hierarchical reward function and human feedback: Design a reward function:
[0079] In the formula, is the search trigger reward, which evaluates whether the search trigger timing is appropriate. is the query quality reward, which evaluates the relevance of the query content to the question. is the answer quality reward, which evaluates the quality of the final answer.
[0080] During the training process, it is also necessary to train the information fusion step: after the model generates a trigger token, perform an online search and obtain retrieval results, and extract a summary of the retrieval results.
[0081] Calculate the new input , and input it into the large language model again to continue generating answer text. In this process, it is also necessary to calculate the relevant loss and perform backpropagation and parameter update to optimize the trainable parameters W .
[0082] In some embodiments, in S2, deduplicate the results of multiple rounds of retrieval to avoid the same information from appearing repeatedly in the final answer. The information integration steps for multiple retrievals include: Use a pre-trained embedding model (such as BERT or Sentence-BERT) to convert the retrieval results into vectors; Calculate the cosine similarity between the new retrieval results and the existing information;
[0083] If the similarity is greater than the set similarity threshold, duplicate content is deleted; the set similarity threshold (for example, 0.9), when the similarity exceeds this value, duplicate information is eliminated. This method effectively reduces redundancy while retaining diversity.
[0084] The retrieved information needs to be sorted according to importance. The scoring formula is:
[0085] Relevance: Use classic information retrieval algorithms (such as TF-IDF or BM25) to calculate the matching degree between the information and the user's question.
[0086] Credibility: Score according to the authority of the source. For example: academic journals, government websites: high score (0.9). Ordinary blogs, forums: medium to low scores (0.4 - 0.6).
[0087] Timeliness: Calculate according to the release time. Recent information (such as in the past 1 year) scores higher, and the weight can be dynamically adjusted. The weight is optimized through experiments to ensure that the scoring results meet the actual requirements.
[0088] Use a large language model to hierarchically summarize the multi-round search results, making the final answer well-organized. To process the multi-round retrieval results, this application adopts a hierarchical method: Topic clustering: Use clustering algorithms (such as K-means or LDA) to group the retrieval results by topic.
[0089] Summary generation: Use a large language model (LLM) to generate a concise summary for each topic cluster. For example, compress 500-word information into 50 words.
[0090] Structured output: Organize the summaries by topic to form a clear hierarchical structure. For example: Topic 1: Climate change data; Summary: The global temperature rose by 0.8°C in 2023...; Topic 2: Policy responses; Summary: Many countries have launched carbon neutrality plans...; Traditional methods usually directly splice the retrieval results, which easily leads to information redundancy and chaos. This application ensures the efficient integration and clear presentation of information through deduplication, scoring, and hierarchical summarization.
[0091] In some embodiments, each time the LLM is called, the previous output, the content retrieved from the network, and the summary information are used as new context inputs. To maintain the coherence of the conversation, this application uses the following content as input each time the LLM is called: Original user question: Ensure that the answer always focuses on the core question.
[0092] Answer generated by the model: Avoid repetition or contradiction.
[0093] Retrieved information summary: Provide external support.
[0094] Search query history: Record the triggering situation of the trigger token [search_web("query")].
[0095] This information is input into the model through concatenation or embedding, and the input length is limited by the maximum context window of the model (e.g., 4096 tokens).
[0096] To balance the weights of new information and historical information, this application introduces a gating mechanism: Gating calculated value
[0097] Where is the sigmoid function, W is the trainable parameter, and the input is the vector concatenation of the query, new information, and history.
[0098] Intelligently fuse information through the dynamic gating mechanism to ensure the logic and consistency of the answer.
[0099]
[0100]
[0101] In the formula, is the new input, is the summary information, history is the historically generated content, and query is the original question. Traditional methods usually simply concatenate the context, which easily leads to information overload or incoherence.
[0102] In some embodiments, when triggering an online search in step S2, the number of online queries N is synchronously recorded; where the initial N = 0, and each time an online search is triggered N = N + 1; Step S5 specifically includes: S51. Determine whether the large language model generates the trigger token again; If not, directly output the generated text; If so, execute S52; S52. Determine whether N reaches the preset maximum query times Nmax; If N ≥ Nmax, terminate the online search query and output the current text; If N < Nmax, execute S53; S53. Determine whether the comprehensive matching degree between the newly added abstract and the existing answers is lower than a preset threshold; If so, terminate the online search query and output the current text; If not, return to execute step S2.
[0103] In the embodiments of the present invention, before executing step S53, it is necessary to calculate the comprehensive matching degree between the newly added abstract and the existing answers. The specific process is as follows: Encode the existing answers and the newly added abstract into semantic vectors respectively; Calculate the cosine similarity and keyword overlap degree of the two groups of semantic vectors; Obtain the comprehensive matching degree by weighted summation of the cosine similarity and the keyword overlap degree.
[0104] In the embodiments of the present invention, the existing answer A: all the texts generated by the large language model; the newly retrieved abstract D: the abstract information extracted after the current online retrieval.
[0105] Use a pre-trained language model (such as BERT, Sentence-BERT) to convert A and D into semantic vectors: ,
[0106] If the text is too long, split it by sentences and take the vector mean; Cosine similarity
[0107] Extract the named entities in A and D respectively to generate entity sets and , calculate the Jaccard similarity of the two entity sets to obtain the keyword overlap degree; Jaccard similarity ; Comprehensive matching degree
[0108] Among them, is a weight parameter, defaulting to 0.7.
[0109] Preset threshold (set to 0.1 in the embodiments of the present invention), if ≥ , the matching degree between the latest search result and the existing answers is too low, then reject the new online search to avoid invalid queries.
[0110] Such as Figure 2As shown in the figure, an embodiment of the present invention provides a large language model network query system based on active triggering, including an input parsing module, a large language model processing module, a network search control module, a retrieval processing module, an information fusion module, and a loop control module; The input parsing module is used to receive the question and context information input by the user and transfer it to the large language model processing module; The large language model processing module is used to generate text output, and identify the context position that needs to trigger network search during the process of generating the answer text, and autonomously generate and insert a trigger token, where the trigger token contains the query content; The network search control module is used to pause text generation when the trigger token is detected, trigger network search, and record the number of network queries; The retrieval processing module is used to execute network search to obtain retrieval results and extract summaries from the retrieval results; The information fusion module is used to fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; input the new input into the large language model again to continue generating the answer text; The loop control module is used to monitor the number of network searches. When the number of network searches reaches the preset maximum value, it controls to stop triggering network search.
[0111] In some embodiments, the system further includes a model training module, which is specifically used to construct a training data set so that the large language model can learn to dynamically insert trigger tokens during generation and generate corresponding search queries. The training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the process of generating answer texts; optimize the search decision-making ability of the large language model based on reinforcement learning.
[0112] In some embodiments, the loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, where the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0113] Comprehensive loss function ; Language modeling loss function ; Trigger classification loss function ; Single-position query generation loss function ; Overall query generation loss formula ; Where, The probability predicted by the model, representing the probability that the next token is given the context x and the previously generated text . The probability predicted by the model, representing the probability of inserting a trigger token at position t, the m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, i.e., ; The probability predicted by the model, representing that given the context , the previously generated text and the previous part of the query , the probability that the next query token is ; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, is the true label.
[0114] In some embodiments, the model training module dynamically adjusts the large language model's search decision-making ability through a hierarchical reward function and human feedback, specifically including: Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality; Adjust the model parameters based on user feedback or automatic scoring; The reward function is: ; In the formula, is the search trigger reward, is the query quality reward, is the answer quality reward.
[0115] In some embodiments, the retrieval processing module integrates the information retrieved multiple times, specifically including: using a pre-trained embedding model to convert the retrieval results into vectors; calculating the cosine similarity between the new retrieval results and the existing information; deleting duplicate content if the similarity is greater than a set similarity threshold; using K-means to cluster the deduplicated retrieval results to generate multiple text clusters; generating summaries for each cluster based on the large language model.
[0116] The formula for the information fusion module to perform information fusion is as follows:
[0117]
[0118] In the formula, is the new input, The abstract information is denoted as abstract, the historically generated content as history, and the original question as query. is the sigmoid function, and W are the trainable parameters.
[0119] The networked search control module includes: A counting sub-module for maintaining the number of networked queries N; A termination judgment sub-module for N ≥ N When max or when the comprehensive matching degree between the newly added abstract and the existing answer is lower than the preset threshold, the networked query is forcibly terminated.
[0120] In the embodiments of the present invention, by setting a preset maximum number of queries, the query process can be effectively controlled, unnecessary repeated queries and resource waste can be avoided, and the query efficiency and system stability can be improved.
[0121] By introducing a trigger token (such as [search_web(“query”)]), the model can autonomously decide when to perform a networked retrieval during the generation process. When the trigger token is detected, the system pauses the model output, performs a networked query, summarizes the retrieval results, and updates the new information to the model input, making the answering process more intelligent, coherent, and meeting the user's needs.
[0122] The embodiments of the present invention also provide an electronic device, which includes: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following methods: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger a networked search during the process of generating an answer text, and autonomously generate and insert a trigger token; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform a networked search to obtain the retrieval results, and extract a summary of the retrieval results; S3. Fuse the summary information with the original question and the historically generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates the trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0123] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0124] Embodiments of the present invention provide a non-transitory computer-readable storage medium that stores computer instructions, and these computer instructions cause a computer to execute the methods provided in the above method embodiments. For example, it includes: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger network search during the generation of the answer text, and autonomously generate and insert trigger tokens; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform network search to obtain retrieval results, and extract summaries from the retrieval results; S3. Fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0125] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A large language model online query method based on active triggering, characterized in that: It includes the following steps: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger networked search during the process of generating answer texts, and autonomously generate and insert trigger tokens; Phrases or expressions used to trigger networked queries are defined as trigger tokens; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform networked search to obtain retrieval results, and extract summaries from the retrieval results; S3. Fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until any of the following termination conditions is reached: a) The number of networked queries reaches the preset maximum value; b) The comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold.
2. The method for online querying of a large language model based on active triggering according to claim 1 is characterized in that: In S1, the training of the large language model includes: S11. Construct a training dataset to enable the large language model to learn to dynamically insert trigger tokens and generate corresponding search queries during the generation process. The training dataset includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; S12. Design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the answer text generation process; S13. Optimize the search decision-making ability of the large language model based on reinforcement learning.
3. The method for online querying of a large language model based on active triggering according to claim 2 is characterized in that: The loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function. Among them, the search decision loss function includes a trigger classification loss function and a query generation loss function.
4. The method for online querying of a large language model based on active triggering according to claim 3 is characterized in that: Comprehensive loss function ; Language Modeling Loss Function ; Triggering classification loss function ; Single-position query generation loss function ; Overall query generation loss formula ; In the formula, is the probability predicted by the model, indicating the text generated before given the context x In the case of The probability of is the probability predicted by the model, indicating the probability of inserting a trigger token at position t. Query the mth token in the sequence; is all tokens before the mth token in the query sequence, that is ; is the probability predicted by the model, indicating that , previously generated text and the query front part In the case of , the next query token is The probability of; M is the total number of tokens in the query sequence; A collection of all locations where trigger tokens need to be inserted; and is a hyperparameter, is the true label.
5. The method for online querying of a large language model based on active triggering according to claim 4 is characterized in that: In S13, the search decision-making ability of the large language model is dynamically adjusted through a hierarchical reward function and human feedback, specifically including: Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality; Adjust the model parameters based on user feedback or automatic scoring; The reward function is: ; In the formula, Trigger rewards for searches, To reward the quality of the query, Reward for quality of answer.
6. The method for online querying of a large language model based on active triggering according to claim 5 is characterized in that: The formula for information fusion in S3 is as follows: In the formula, For new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is a trainable parameter.
7. The method for online querying of a large language model based on active triggering according to claim 6 is characterized in that: When the online search is triggered in step S2, the number of online queries N is recorded synchronously; N =0, trigger network search every time N = N +1; Step S5 specifically includes: S51. Determine whether the large language model generates a trigger token again; If not, directly output the generated text; If so, execute S52; S52. Determine whether N reaches the preset maximum query times Nmax; If N≥Nmax, terminate the networked search query and output the current text; If N<Nmax, execute S53; S53. Determine whether the comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold; If so, terminate the networked search query and output the current text; If not, return to execute step S2.
8. A large language model online query system based on active triggering, characterized in that: It includes an input parsing module, a large language model processing module, a networked search control module, a retrieval processing module, an information fusion module, and a loop control module; The input parsing module is used to receive the questions and context information input by the user and transfer them to the large language model processing module; The large language model processing module is used to generate text output, and identify the context positions that need to trigger networked search during the process of generating answer texts, and autonomously generate and insert trigger tokens, where the trigger tokens contain query content; An online search control module, used to suspend text generation, trigger online search, and record the number of online queries when the trigger token is detected; A retrieval processing module is used to perform online search to obtain retrieval results and extract summary of the retrieval results; The information fusion module is used to fuse the summary information with the original question and the historical generated content through a dynamic gating mechanism to form a new input; the new input is input into the large language model again to continue to generate the answer text; The loop control module is used to monitor the number of network searches. When the number of network searches reaches a preset maximum value, the control stops triggering the network search.
9. The large language model network query system based on active triggering according to claim 8 is characterized in that: The network search control module includes: The counting submodule is used to maintain the number of online queries N; The termination judgment submodule is used to N ≥ N When the combined matching degree between the max or newly added summary and the existing answer is lower than the preset threshold, the online query is terminated forcibly.
10. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the large language model network query method based on active triggering as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Networking question and answer method based on large language model
CN118227890A
Scientific research assisting method based on large language model
CN119202331A
Network search enhanced question answering method, system, device, medium and program product
CN119739920A
An on-device AI NAS system equipped with a multimodal LLM, and a method for providing search and dialogue functions in the system
KR102763677B1
Cited By
Full-link AI depth reasoning container configuration method
CN120930791A
Voice interaction method and system based on multi-stage large model
CN121011180A
Large-model thinking-enhanced networking search question and answer method, system and device and medium
CN121658702A