Large Language Model Internet Query Method, System and Device Based on Active Trigger
By introducing trigger tokens and dynamic search decisions in the large language model, the problems of inaccurate user intention expression and fixed query process are solved, and smarter and more accurate networked query and answer generation are achieved.
Patent Information
- Application Number
- CN202510577985.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing large language model is difficult to deeply explore the user's true intentions in the network search task. The generated query statements cannot accurately express user needs, and the query process cannot be adjusted dynamically, which affects the accuracy and effectiveness of the answers.
Through pre-training or prompt word engineering, large language models are recognized and inserted into trigger tokens, the generation and execution of network searches are paused, search results and historical content are dynamically integrated, search decisions are optimized by multi-task learning and reinforcement learning, and termination conditions are set to control the number of queries.
It improves the autonomy and flexibility of large language models in information acquisition decision-making, generates more accurate and smooth answers, adapts to complex and diverse user questions, and enhances the universality and adaptability of the model.
Smart Images

Figure CN120087483B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and specifically relates to a method, system, and device for querying the Internet based on an actively triggered large language model. Background Art
[0002] When current large language models (LLMs) execute Internet search tasks, they generally follow a fixed process pattern: First, the user's input question is initially rewritten to generate a standardized query statement; subsequently, this query statement is used to retrieve relevant network information from external information sources; finally, the retrieved results are re-input into the model for summarization and generation of a response.
[0003] However, this traditional method exposes many deficiencies in practical applications:
[0004] (1) The generation process of the query statement mainly relies on a superficial understanding of the user's initial question, making it difficult to deeply explore the user's true intentions and potential needs, resulting in a deviation between the query results and the user's expectations.
[0005] (2) Due to the limitations of the query rewriting rules, the generated query statement may not accurately express the user's intentions, thereby causing the retrieved information to not match the user's actual needs and affecting the accuracy and effectiveness of the answer.
[0006] (3) The Internet search process of the traditional method is usually preset and fixed, and cannot dynamically adjust the timing and content of the Internet query according to the real-time needs of the model during the text generation process, thus limiting the room for improvement of the answer quality. Summary of the Invention
[0007] In view of the above problems existing when large language models (LLMs) execute Internet search tasks, the present invention provides a method, system, and device for querying the Internet based on an actively triggered large language model.
[0008] In a first aspect, the technical solution of the present invention provides a method for querying the Internet based on an actively triggered large language model, including the following steps:
[0009] S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger Internet search during the generation of the answer text, and autonomously generate and insert trigger tokens; a token is the basic unit that constitutes the text, and a phrase or expression used to trigger Internet query is defined as a trigger token;
[0010] S2. After detecting the trigger token, pause the generation of the answer text by the large language model, execute an Internet search to obtain retrieval results, and extract a summary of the retrieval results;
[0011] S3. Integrate the summary information with the original question and the historical generated content through a dynamic gating mechanism to form a new input;
[0012] S4. Input the new input into the large language model again to continue generating the answer text;
[0013] S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0014] As a further limitation of the technical solution of the present invention, in S1, training the large language model includes:
[0015] S11. Construct a training data set to enable the large language model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results;
[0016] S12. Design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the answer text generation process;
[0017] S13. Optimize the search decision-making ability of the large language model based on reinforcement learning.
[0018] As a further limitation of the technical solution of the present invention, the loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, where the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0019] As a further limitation of the technical solution of the present invention, the comprehensive loss function ;
[0020] The language modeling loss function ;
[0021] The trigger classification loss function ;
[0022] The single-position query generation loss function ;
[0023] The overall query generation loss formula ;
[0024] In the formula, is the probability predicted by the model, indicating that given the context x and the previously generated text , the probability that the next token is , is the probability predicted by the model, indicating the probability of inserting a trigger token at position t, Query the m-th token in the sequence; is all the tokens before the m-th token in the query sequence, that is ; is the probability predicted by the model, indicating that in the given context , the previously generated text and the previous part of the query the probability that the next query token is ; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, is the true label.
[0025] As a further limitation of the technical solution of the present invention, in S13, the search decision-making ability of the large language model is dynamically adjusted through a hierarchical reward function and human feedback, specifically including:
[0026] Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality;
[0027] Adjust the model parameters based on user feedback or automatic scoring;
[0028] The reward function is: ;
[0029] In the formula, is the search trigger reward, is the query quality reward, is the answer quality reward.
[0030] As a further limitation of the technical solution of the present invention, in S2, the information integration steps for multiple retrievals include:
[0031] Use a pre-trained embedding model to convert the retrieval results into vectors;
[0032] Calculate the cosine similarity between the new retrieval results and the existing information;
[0033] If the similarity is greater than the set similarity threshold, delete the duplicate content;
[0034] Use K-means to cluster the deduplicated retrieval results to generate multiple text clusters;
[0035] Generate summaries for each cluster based on the large language model.
[0036] As a further limitation of the technical solution of the present invention, the formula for information fusion in S3 is as follows:
[0037]
[0038]
[0039] Wherein, is the new input, is the summary information, history is the historically generated content, and query is the original question. is the sigmoid function, and W is the trainable parameter. The summary information, the original question, and the historically generated content are fused through a dynamic gating mechanism, enabling the model to intelligently adjust the weights of each part of the information according to the specific situation, making full use of the coherence of the existing generated content and the new information obtained from the search, effectively avoiding information redundancy or conflicts, and forming a new input with logical coherence and rich content, thereby assisting the large language model in generating more comprehensive, accurate, and fluent response texts.
[0040] As a further limitation of the technical solution of the present invention, when triggering the network search in step S2, the network query count N is synchronously recorded; wherein, the initial N = 0, and each time the network search is triggered N = N + 1;
[0041] Step S5 specifically includes:
[0042] S51. Determine whether the large language model generates a trigger token again;
[0043] If not, directly output the generated text;
[0044] If so, execute S52;
[0045] S52. Determine whether N reaches the preset maximum query count Nmax;
[0046] If N ≥ Nmax, terminate the network search query and output the current text;
[0047] If N < Nmax, execute S53;
[0048] S53. Determine whether the comprehensive matching degree of the newly added summary and the existing answer is lower than the preset threshold;
[0049] If so, terminate the network search query and output the current text;
[0050] If not, return to execute step S2.
[0051] In a second aspect, the technical solution of the present invention provides a large language model network query system based on active triggering, including an input parsing module, a large language model processing module, a network search control module, a retrieval processing module, an information fusion module, and a loop control module;
[0052] An input parsing module, configured to receive questions and context information input by a user, and transmit same to a large language model processing module;
[0053] A large language model processing module, configured to generate text output, and identify context positions that need to trigger an online search during the process of generating an answer text, autonomously generate and insert a trigger token, where the trigger token contains query content;
[0054] An online search control module, configured to pause text generation when detecting the trigger token, trigger an online search, and record the number of online query times;
[0055] A retrieval processing module, configured to perform an online search to obtain retrieval results, and extract summaries from the retrieval results;
[0056] An information fusion module, configured to fuse summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; input the new input into the large language model again to continue generating an answer text;
[0057] A loop control module, configured to monitor the number of online search times, and when the number of online search times reaches a preset maximum value, control to stop triggering an online search.
[0058] As a further limitation of the technical solution of the present invention, the system further includes a model training module, specifically configured to construct a training data set, enable the large language model to learn to dynamically insert a trigger token during generation and generate corresponding search queries, where the training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; design a loss function for multi-task learning to enable the large language model to dynamically trigger a search during the process of generating an answer text; optimize the search decision-making ability of the large language model based on reinforcement learning.
[0059] Construct a training data set containing multiple elements, design a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, and optimize the search decision-making ability based on reinforcement learning, training and optimizing the large language model in a multi-dimensional and systematic manner. Conduct targeted optimization from multiple levels such as language generation, search trigger judgment, and query content generation, effectively improving the comprehensive performance of the model in the online query task and enabling it to better adapt to complex task requirements.
[0060] As a further limitation of the technical solution of the present invention, the loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function, where the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0061] Comprehensive loss function ;
[0062] Language modeling loss function ;
[0063] Trigger classification loss function ;
[0064] Single-position query generation loss function ;
[0065] Overall query generation loss formula ;
[0066] In the formula, is the probability predicted by the model, representing the probability that the next token is given the context x and the previously generated text ; is the probability predicted by the model, representing the probability of inserting a trigger token at position t, the m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, i.e., ; is the probability predicted by the model, representing the probability that the next query token is given the context the previously generated text and the previous part of the query ; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, is the true label.
[0067] As a further limitation of the technical solution of the present invention, the model training module dynamically adjusts the large language model search decision-making ability through a hierarchical reward function and human feedback, specifically including:
[0068] Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality;
[0069] Adjust the model parameters based on user feedback or automatic scoring;
[0070] The reward function is: ;
[0071] is the search trigger reward, is the query quality reward, is the answer quality reward.
[0072] The search trigger timing, query content relevance, and final answer quality are evaluated through a hierarchical reward function, and the model parameters are dynamically adjusted based on user feedback or automatic scoring, establishing a closed-loop mechanism for model self-optimization. This mechanism enables the model to continuously learn and improve in practical applications, continuously enhancing the search decision-making ability and answer quality, and ensuring the dynamic adaptation of the model performance to user needs and actual application scenarios.
[0073] As a further limitation of the technical solution of the present invention, the retrieval processing module integrates the information retrieved multiple times, specifically including: converting the retrieval results into vectors using a pre-trained embedding model; calculating the cosine similarity between the new retrieval results and the existing information; deleting duplicate content if the similarity is greater than the set similarity threshold; using K-means to cluster the de-duplicated retrieval results to generate multiple text clusters; and generating summaries for each cluster based on a large language model.
[0074] The formula for information fusion by the information fusion module is as follows:
[0075]
[0076]
[0077] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is the trainable parameter.
[0078] The networked search control module includes:
[0079] A counting sub-module for maintaining the networked query count N;
[0080] A termination judgment sub-module for forcibly terminating the networked query when N ≥ N max or when the comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold.
[0081] Synchronously record the networked query count and set the maximum query count, effectively avoiding problems such as resource waste and reduced answer efficiency caused by excessive searching while ensuring sufficient information is obtained, reasonably controlling the query process, achieving a good balance between information acquisition and answer generation efficiency of the model, and improving the user experience and system operation efficiency.
[0082] In a third aspect, the technical solution of the present invention further provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor to enable the at least one processor to execute the method for querying the large language model network connection based on active triggering as described in the first aspect.
[0083] As can be seen from the above technical solutions, the present application has the following advantages: Through pre-training or prompt engineering, the large language model can autonomously identify the context position where network search needs to be triggered and insert a trigger token, realizing the transformation from passive receiving of instructions to active judgment of search requirements, greatly enhancing the autonomy of the model in information acquisition decision-making, being able to respond to complex and diverse user questions more flexibly and intelligently, avoiding blind search or excessive reliance on manual intervention, and significantly enhancing the generality and adaptability of model applications. When the model recognizes the trigger token, it pauses the answer generation, performs network search and extracts the abstract, ensuring that the obtained information is the latest and highly relevant to the question, making up for the deficiency of the limited timeliness of the training data of the large language model. Especially when answering questions in fields such as real-time news and the latest research results, it can significantly improve the accuracy and practicality of the answer content and provide more valuable information services for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the technical solutions of the present application, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0085] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention.
[0086] Figure 2 It is a block diagram of the system provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] In order to make the application purpose, features, and advantages of the present application more obvious and understandable, the technical solutions protected by the present application will be clearly and completely described below by using specific embodiments and the accompanying drawings. Obviously, the embodiments described below are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0088] Such as Figure 1As shown in the figure, an embodiment of the present invention provides a method for querying the Internet by a large language model based on active triggering, including the following steps:
[0089] S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger Internet search during the process of generating answer text, and autonomously generate and insert trigger tokens;
[0090] S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform an Internet search to obtain retrieval results, and extract a summary of the retrieval results;
[0091] S3. Integrate the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input;
[0092] S4. Input the new input into the large language model again to continue generating the answer text;
[0093] S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0094] In some embodiments, in S1, the training of the large language model includes:
[0095] S11. Construct a training data set to enable the large language model to learn to dynamically insert trigger tokens during generation and generate corresponding search queries. The training data set includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results;
[0096] To enable the large language model (LLM) to have the ability to autonomously trigger Internet search, this application designs a specific training data construction method. The training data set includes the following key information:
[0097] Question: The original question input by the user.
[0098] Answer: The answer text generated by the model, in which the trigger token [search_web("query")] is inserted at the position where Internet search is required.
[0099] Search decision annotation: Annotate which parts of the answer need to trigger a search and the specific search query content.
[0100] Search result: For each query, annotate the retrieved information and its summary.
[0101] Example: User question: "Please introduce the main breakthroughs in AI technology in 2024 and provide relevant paper links."
[0102] The answers in the training data are labeled as: The breakthroughs of AI technology in 2024 include [search_web("Breakthroughs of AI Technology in 2024")], and relevant papers can refer to the following content [search_web("Papers on AI Technology in 2024")].
[0103] This way of constructing the dataset enables the model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training data of traditional models only focuses on text generation and does not contain explicit annotations for search decisions. The present invention trains the model to have the ability to autonomously judge and generate search instructions by annotating search trigger points and query content.
[0104] S12. Design the loss function for multi-task learning to dynamically trigger search during the generation of the response text by the large language model;
[0105] To train the model to generate both fluent text and accurate search decisions simultaneously, this application adopts the loss function of multi-task learning, including the following parts:
[0106] Language Modeling Loss (LM Loss): Optimize the grammar and semantic consistency of the response. Calculate using cross-entropy loss, formula: ;
[0107] In the formula : The (t)-th token in the sequence. : All tokens before the (t)-th token in the sequence, that is . : Input context. : The probability predicted by the model, indicating that given the context and the previous tokens, the probability that the next token is .
[0108] Search Decision Loss: Divided into two parts:
[0109] Trigger Classification Loss: Judge whether a trigger token needs to be inserted at each token position, and use cross-entropy loss to implement a binary classification task. Formula:
[0110]
[0111] In the formula, True label. : A trigger token needs to be inserted at position (t). : No insertion is required at position (t). : The probability predicted by the model, indicating the possibility of inserting a trigger token at position (t).
[0112] Query generation loss: Predicting the correct search query content and calculating it using sequence generation loss. Assume that at position (t), the query needs to be generated , where (M) is the length of the query. This is a sequence generation task, similar to language modeling.
[0113] Single-position query generation loss formula:
[0114]
[0115] In the formula, : The m-th token in the query sequence. : All tokens before the m-th token in the query sequence, i.e., . : The probability predicted by the model, indicating that given the context , the previously generated text and the previous part of the query , the probability that the next query token is .
[0116] Overall query generation loss formula:
[0117]
[0118] In the formula, is the set of all positions where trigger tokens need to be inserted. For all positions where trigger tokens are inserted, the query generation loss at each position is accumulated to obtain the total query generation loss.
[0119] The comprehensive loss function is:
[0120] In the formula, is the language modeling loss, ensuring that the generated text is fluent. is the trigger classification loss, ensuring that the model accurately judges whether to insert a trigger token. is the query generation loss, ensuring that the generated query content is accurate. and are hyperparameters used to adjust the relative importance of each part of the loss. For example, increasing will make the model pay more attention to the trigger decision, and increasing will make the model pay more attention to the query generation quality.
[0121] Traditional models only use language modeling loss and focus on text generation. This application introduces search decision loss, enabling the model to dynamically trigger searches during the generation process, improving intelligence and flexibility.
[0122] S13. Optimize the search decision-making ability of the large language model based on reinforcement learning.
[0123] Reinforcement Learning with Human Feedback (RLHF) is used in the present invention to optimize the search and decision-making ability of the model.
[0124] The reinforcement learning framework includes the following core elements:
[0125] State: The state is composed of the context of the current conversation, including the question raised by the user, the answer fragments generated by the model, the retrieved external information, and the conversation history. This multi-dimensional state design enables the model to comprehensively perceive the current task requirements.
[0126] Action: The actions of the model include two categories:
[0127] Whether to insert the trigger token [search_web("query")] to trigger an external search.
[0128] If the search is triggered, generate the specific search query content (e.g., "2023 global temperature change data").
[0129] Reward: The reward signal is used to evaluate the quality of the model's decision-making.
[0130] The reward function is the key to reinforcement learning and directly affects the optimization direction of the model. The present invention designs a multi-level reward mechanism:
[0131] Search trigger reward ( )
[0132] If the question requires external information and the model correctly triggers the search, the reward is positive (e.g., +1).
[0133] If the question does not require a search but the model wrongly triggers it, or if a search is needed but not triggered, the reward is negative (e.g., -0.5).
[0134] Query quality reward ( )
[0135] The relevance between the query content and the user's question is evaluated by manual annotation or a semantic similarity model (e.g., cosine similarity based on BERT). If the query is accurate and efficient, the reward is high (e.g., +0.8); if the query is vague or irrelevant, the reward is low (e.g., -0.3).
[0136] Answer quality reward ( )
[0137] The fluency, accuracy, and completeness of the final answer are calculated through user feedback or automatic evaluation (such as BLEU score, ROUGE score) as part of the long-term reward.
[0138] The comprehensive reward formula is:
[0139]
[0140] This hierarchical design ensures that the model not only learns when to search but also optimizes the search content and the quality of the answer. Traditional large language models are usually based on pre-training and fine-tuning, lacking dynamic optimization of search decisions. Through RLHF, this application enables the model to adjust its behavior according to real-time feedback, showing higher intelligence especially in complex and information-insufficient questions.
[0141] The formula for the information fusion module to perform information fusion is as follows:
[0142]
[0143]
[0144] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is the trainable parameter.
[0145] The specific training steps are as follows:
[0146] Construct a training dataset that includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results. This dataset is used to let the large language model learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries.
[0147] Initialize the parameters of the large language model and set hyperparameters, such as α and β in the comprehensive loss function, as well as the trainable parameter W in the information fusion formula.
[0148] Based on the set loss function, perform the following operations in each training cycle:
[0149] Input the user question x in the training dataset, and the large language model starts to generate the answer text. During the generation process, the model tries to dynamically insert trigger tokens and generate corresponding search queries.
[0150] Calculate the prediction probability at each time step:
[0151] For the language modeling loss function, calculate , where is the true token at the current time step, is all the tokens generated previously.
[0152] For the trigger classification loss function, calculate , representing the probability of triggering a search at time step t , is the search decision annotation.
[0153] For the query generation loss function, calculate , where is the m -th token in the query sequence, is all the tokens before the m -th token in the query sequence.
[0154] According to the above predicted probabilities, calculate the comprehensive loss function .
[0155] Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters.
[0156] Use an optimization algorithm (such as stochastic gradient descent, Adam, etc.) to update the model parameters according to the calculated gradient.
[0157] Dynamically adjust the large language model's search decision-making ability through a hierarchical reward function and human feedback:
[0158] Design a reward function:
[0159] In the formula, is the search trigger reward, evaluating whether the search trigger timing is appropriate. is the query quality reward, evaluating the relevance of the query content to the question. is the answer quality reward, evaluating the quality of the final answer.
[0160] During the training process, it is also necessary to train the information fusion step: after the model generates a trigger token, perform an online search and obtain retrieval results, and extract a summary of the retrieval results.
[0161] Calculate the new input , and input it into the large language model again to continue generating the answer text. In this process, it is also necessary to calculate the relevant losses and perform backpropagation and parameter updates to optimize the trainable parameters W .
[0162] In some embodiments, in S2, the results of multiple rounds of retrieval are de-duplicated to avoid the same information from appearing repeatedly in the final answer. The steps for integrating the information retrieved multiple times include:
[0163] Use a pre-trained embedding model (such as BERT or Sentence-BERT) to convert the retrieval results into vectors;
[0164] Calculate the cosine similarity between the new retrieval results and the existing information;
[0165]
[0166] If the similarity is greater than the set similarity threshold, delete the duplicate content; the set similarity threshold (for example, 0.9), when the similarity exceeds then, eliminate the duplicate information. This method effectively reduces redundancy while retaining diversity.
[0167] The retrieved information needs to be sorted according to importance, and the scoring formula is:
[0168]
[0169] Relevance: Use classical information retrieval algorithms (such as TF-IDF or BM25) to calculate the matching degree between the information and the user's question.
[0170] Credibility: Score according to the authority of the source. For example: academic journals, government websites: high score (0.9). Ordinary blogs, forums: medium to low scores (0.4 - 0.6).
[0171] Timeliness: Calculate according to the release time. Information from the recent period (such as in the past 1 year) scores higher, and the weight can be dynamically adjusted. The weight is optimized through experiments to ensure that the scoring results meet the actual requirements.
[0172] Use a large language model to perform hierarchical summarization on the results of multiple rounds of search, making the final answer well-organized. To process the results of multiple rounds of retrieval, this application adopts a hierarchical method:
[0173] Topic clustering: Use clustering algorithms (such as K-means or LDA) to group the retrieval results by topic.
[0174] Summary generation: Use a large language model (LLM) to generate a concise summary for each topic cluster. For example, compress 500 words of information into 50 words.
[0175] Structured output: Organize the summaries by topic to form a clear hierarchical structure. For example:
[0176] Subject 1: Climate Change Data; Abstract: The global temperature rose by 0.8 °C in 2023...;
[0177] Subject 2: Policy Responses; Abstract: Many countries have launched carbon neutrality plans...;
[0178] Traditional methods usually directly splice search results, which easily leads to information redundancy and chaos. This application ensures the efficient integration and clear presentation of information through deduplication, scoring, and hierarchical summarization.
[0179] In some embodiments, each time the LLM is called, the previous output, the content retrieved from the network, and the summary information are used as new context inputs. To maintain the coherence of the conversation, this application uses the following as inputs each time the LLM is called:
[0180] User's original question: Ensure that the answer always focuses on the core question.
[0181] Answer generated by the model: Avoid repetition or contradiction.
[0182] Summary of retrieved information: Provide external support.
[0183] Search query history: Record the triggering situation of the trigger token [search_web("query")].
[0184] These information are input into the model in a splicing or embedding manner, and the input length is limited by the maximum context window of the model (e.g., 4096 tokens).
[0185] To balance the weights of new information and historical information, this application introduces a gating mechanism:
[0186] Gating calculation value
[0187] Among them, is the sigmoid function, W is the trainable parameter, and the input is the vector splicing of the query, new information, and history.
[0188] Intelligently fuse information through the dynamic gating mechanism to ensure the logic and consistency of the answer.
[0189]
[0190]
[0191] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question,
[0192] Traditional methods usually simply concatenate the context, which easily leads to information overload or incoherence.
[0193] In some embodiments, when triggering an online search in step S2, the number of online queries N is synchronously recorded; where the initial N = 0, and each time an online search is triggered N = N + 1;
[0194] Step S5 specifically includes:
[0195] S51. Determine whether the large language model generates a trigger token again;
[0196] If not, directly output the generated text;
[0197] If so, execute S52;
[0198] S52. Determine whether N reaches the preset maximum query number Nmax;
[0199] If N ≥ Nmax, terminate the online search query and output the current text;
[0200] If N < Nmax, execute S53;
[0201] S53. Determine whether the comprehensive matching degree between the newly added abstract and the existing answer is lower than the preset threshold;
[0202] If so, terminate the online search query and output the current text;
[0203] If not, return to execute step S2.
[0204] In the embodiments of the present invention, before executing step S53, it is necessary to calculate the comprehensive matching degree between the newly added abstract and the existing answer. The specific process is as follows:
[0205] Encode the existing answer and the newly added abstract into semantic vectors respectively;
[0206] Calculate the cosine similarity and keyword overlap degree of the two groups of semantic vectors;
[0207] Obtain the comprehensive matching degree by weighted summing the cosine similarity and the keyword overlap degree.
[0208] In the embodiments of the present invention, the existing answer A: all the text generated by the large language model; the newly retrieved abstract D: the abstract information extracted after the current online retrieval.
[0209] Use a pre-trained language model (such as BERT, Sentence-BERT) to convert A and D into semantic vectors:
[0210] ,
[0211] If the text is too long, split it by sentences and take the average of the vectors;
[0212] Cosine similarity
[0213] Extract the named entities in A and D respectively to generate entity sets and , calculate the Jaccard similarity of the two entity sets to obtain the keyword overlap;
[0214] Jaccard similarity ;
[0215] Comprehensive matching degree
[0216] where is a weight parameter, defaulting to 0.7.
[0217] Preset threshold (set to 0.1 in the embodiments of the present invention), if ≥ , the match degree between the latest search result and the existing answer is too low, then reject the new online search to avoid invalid queries.
[0218] As Figure 2 shown, the embodiments of the present invention provide a large language model online query system based on active trigger, including an input parsing module, a large language model processing module, an online search control module, a retrieval processing module, an information fusion module and a loop control module;
[0219] The input parsing module is used to receive the question and context information input by the user and transfer it to the large language model processing module;
[0220] The large language model processing module is used to generate text output, and identify the context position that needs to trigger online search during the process of generating the answer text, and autonomously generate and insert a trigger token, where the trigger token contains the query content;
[0221] The online search control module is used to pause text generation when the trigger token is detected, trigger online search, and record the number of online queries;
[0222] The retrieval processing module is used to perform online search to obtain retrieval results and extract summaries from the retrieval results;
[0223] The information fusion module is used to fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; input the new input into the large language model again to continue generating the answer text;
[0224] A loop control module for monitoring the number of online searches. When the number of online searches reaches a preset maximum value, it controls to stop triggering online searches.
[0225] In some embodiments, the system further includes a model training module, specifically for constructing a training dataset to enable the large language model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training dataset includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; designing a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the answer text generation process; and optimizing the search decision-making ability of the large language model based on reinforcement learning.
[0226] In some embodiments, the loss function for multi-task learning is a combined loss function weighted by a language modeling loss function and a search decision loss function. Among them, the search decision loss function includes a trigger classification loss function and a query generation loss function.
[0227] Combined loss function ;
[0228] Language modeling loss function ;
[0229] Trigger classification loss function ;
[0230] Single-position query generation loss function ;
[0231] Overall query generation loss formula ;
[0232] Wherein, is the probability predicted by the model, indicating the probability that the next token is given the context x and the previously generated text ; is the probability predicted by the model, indicating the probability of inserting a trigger token at position t, the m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, that is, ; is the probability predicted by the model, indicating the probability that the next query token is given the context , the previously generated text and the previous part of the query ; M is the total number of tokens in the query sequence; is the set of all positions where trigger tokens need to be inserted; and are hyperparameters, and
[0233] In some embodiments, the model training module dynamically adjusts the search decision-making ability of the large language model through a hierarchical reward function and human feedback, specifically including:
[0234] Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality;
[0235] Adjust the model parameters based on user feedback or automatic scoring;
[0236] The reward function is: ;
[0237] In the formula, is the search trigger reward, is the query quality reward, is the answer quality reward.
[0238] In some embodiments, the retrieval processing module integrates the information retrieved multiple times, specifically including: using a pre-trained embedding model to convert the retrieval results into vectors; calculating the cosine similarity between the new retrieval results and the existing information; if the similarity is greater than the set similarity threshold, deleting duplicate content; using K-means to cluster the de-duplicated retrieval results to generate multiple text clusters; generating summaries for each cluster based on the large language model.
[0239] The formula for the information fusion module to perform information fusion is as follows:
[0240]
[0241]
[0242] In the formula, is the new input, is the summary information, history is the historical generated content, query is the original question, is the sigmoid function, and W is the trainable parameter.
[0243] The networked search control module includes:
[0244] A counting sub-module for maintaining the number of networked queries N;
[0245] A termination judgment sub-module for N ≥ N max or when the comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold, forcibly terminate the networked query.
[0246] In the embodiments of the present invention, setting a preset maximum number of queries can effectively control the query process, avoid unnecessary repeated queries and resource waste, and improve query efficiency and system stability.
[0247] By introducing a trigger token (such as [search_web(“query”)]), the model can autonomously decide when to perform an online search during the generation process. When the trigger token is detected, the system pauses the model output, performs an online query, summarizes the retrieval results, and updates the new information to the model input, making the answering process more intelligent, coherent, and meeting user needs.
[0248] The embodiments of the present invention also provide an electronic device, which includes: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following methods: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger an online search during the generation of the answer text, and autonomously generate and insert the trigger token; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform an online search to obtain the retrieval results, and extract the summary of the retrieval results; S3. Fuse the summary information with the original question and the historical generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates the trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0249] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0250] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the method provided in the above method embodiment. For example, it includes: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger network search during the process of generating answer text, and autonomously generate and insert trigger tokens; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, execute network search to obtain retrieval results, and perform summary extraction on the retrieval results; S3. Fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until the termination condition is reached.
[0251] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for querying the Internet by a large language model based on active triggering, characterized in that, It includes the following steps: S1. Through pre-training or prompt engineering, enable the large language model to learn to identify the context positions that need to trigger an online search during the process of generating answer texts, and autonomously generate and insert trigger tokens; Phrases or expressions used to trigger online queries are defined as trigger tokens; S2. After detecting the trigger token, pause the generation of the answer text by the large language model, perform an online search to obtain retrieval results, and extract summaries from the retrieval results; S3. Fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; S4. Input the new input into the large language model again to continue generating the answer text; S5. When the large language model generates a trigger token again, repeat steps S2 to S4 until any of the following termination conditions is reached: a) The number of online queries reaches the preset maximum value; b) The comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold.
2. The method for querying the large language model network connection based on active triggering according to claim 1, wherein, In S1, the training of the large language model includes: S11. Construct a training dataset to enable the large language model to learn to dynamically insert trigger tokens during the generation process and generate corresponding search queries. The training dataset includes user questions, answer texts annotated with trigger tokens, search decision annotations, and search results; S12. Design a loss function for multi-task learning to enable the large language model to dynamically trigger searches during the answer text generation process; S13. Optimize the search decision-making ability of the large language model based on reinforcement learning.
3. The method for querying the large language model network connection based on active triggering according to claim 2, wherein, The loss function for multi-task learning is a comprehensive loss function weighted by a language modeling loss function and a search decision loss function. Among them, the search decision loss function includes a trigger classification loss function and a query generation loss function.
4. The method for querying the large language model network connection based on active triggering according to claim 3, wherein, Comprehensive loss function ; Language modeling loss function ; Trigger classification loss function ; Single-location query generation loss function ; Overall query generation loss formula ; wherein, is the probability predicted by the model, indicating that given the context x and the previously generated text , the probability that the next token is ; is the probability predicted by the model, indicating the probability of inserting a trigger token at position t; the m-th token in the query sequence; is all the tokens before the m-th token in the query sequence, i.e., ; is the probability predicted by the model, indicating that given the context , the previously generated text and the previous part of the query , the probability that the next query token is ; M is the total number of tokens in the query sequence; is the set of all positions where a trigger token needs to be inserted; and are hyperparameters, is the true label.
5. The method for querying the large language model network connection based on active triggering according to claim 4, wherein, In S13, the search decision-making ability of the large language model is dynamically adjusted through a hierarchical reward function and human feedback, specifically including: Design a reward function to evaluate the search trigger timing, query content relevance, and final answer quality; Adjust model parameters based on user feedback or automatic scoring; The reward function is as follows: ; Wherein, is the search trigger reward, is the query quality reward, is the answer quality reward.
6. The method for querying the large language model network connection based on active triggering according to claim 5, wherein The formula for information fusion in S3 is as follows: In the formula, is the new input, is the abstract information, history is the historically generated content, and query is the original question, is the sigmoid function, and W is the trainable parameter.
7. The method for querying the large language model network connection based on active triggering according to claim 6, wherein When triggering an online search in step S2, synchronously record the number of online queries N; where the initial N = 0, and each time an online search is triggered N = N + 1; Step S5 specifically includes: S51. Determine whether the large language model generates a trigger token again; If not, directly output the generated text; If so, execute S52; S52. Determine whether N reaches the preset maximum query times Nmax; If N≥Nmax, terminate the online search query and output the current text; If N<Nmax, execute S53; S53. Determine whether the comprehensive matching degree between the newly added summary and the existing answer is lower than the preset threshold; If so, terminate the online search query and output the current text; If not, return to execute step S2.
8. A large language model network query system based on active triggering, characterized in that, It includes an input parsing module, a large language model processing module, an online search control module, a retrieval processing module, an information fusion module, and a loop control module; The input parsing module is used to receive the user input question and context information and transfer them to the large language model processing module; The large language model processing module is used to generate text output, and identify the context positions that need to trigger an online search during the process of generating the answer text, and autonomously generate and insert trigger tokens, where the trigger tokens contain query content; The networked search control module is used to pause text generation when the trigger token is detected, trigger a networked search, and record the number of networked queries; The retrieval processing module is used to perform a networked search to obtain retrieval results and extract a summary of the retrieval results; The information fusion module is used to fuse the summary information with the original question and historical generated content through a dynamic gating mechanism to form a new input; input the new input into the large language model again to continue generating the answer text; The loop control module is used to monitor the number of networked searches. When the number of networked searches reaches the preset maximum value, it controls to stop triggering the networked search.
9. The large language model network query system based on active triggering according to claim 8, wherein The networked search control module includes: A counting sub-module for maintaining the number of networked queries N; The termination judgment sub-module is used to N ≥ N When max or the comprehensive matching degree of the newly added abstract and the existing answer is lower than the preset threshold, the online query is forcibly terminated.
10. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the method for querying the large language model through active triggering according to any one of claims 1 to 7.
Citation Information
Patent Citations
Networking question and answer method based on large language model
CN118227890A
Scientific research assisting method based on large language model
CN119202331A