Interactive inquiry method and device based on front end
By acquiring user questions on the front end and using a large language model for intent recognition and merging, and parsing streaming data in real time, the problem of low display rate of user questions on the front end is solved, and a more efficient display of thought chains is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, the front-end display speed of the user's thought process corresponding to the question in the large model-based question-answering or public opinion analysis system is relatively low, resulting in a long waiting time for users.
By acquiring user questions at the front end, using a large language model for intent recognition and merging, receiving streaming data and performing incremental parsing, the system displays the updating process of the target question and thought chain.
It improves the speed at which the front-end displays the thought process corresponding to the user's question, allowing users to select more targeted questions in real time and improving the efficiency of displaying the thought chain.
Smart Images

Figure CN121858699A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent question answering technology, and in particular to a front-end-based interactive questioning method and apparatus. Background Technology
[0002] With the continuous development of artificial intelligence technology, question-answering or public opinion analysis systems based on large models have been widely used in various fields. The main purpose is to improve real-time interactivity through interactive queries between the front end and the back end.
[0003] In related technologies, most question-answering or public opinion analysis systems based on large models concentrate the parsing and display of "model reasoning" and "thinking information" on the back end, with the front end only acting as a passive display. From the time the user inputs a question to the time the front end presents the model's thought process, there is a noticeable delay, and the user wait time is relatively long. There is also the problem that the front end displays the thought process corresponding to the user's question at a low rate. Summary of the Invention
[0004] The purpose of this application is to provide a front-end-based interactive inquiry method, device, electronic device, and storage medium to solve the problem of low speed of front-end display of user's thought process corresponding to the question.
[0005] To solve the above-mentioned technical problems, the embodiments of this application are implemented as follows: In a first aspect, embodiments of this application provide a front-end-based interactive query method applied to a client. The method includes: acquiring a user question from a target user and inputting the user question into a large language model on a server; receiving a set of candidate questions output by the large language model after performing intent recognition on the user question; responding to one or more candidate questions selected by the target user from the set of candidate questions, inputting one or more of the candidate questions into the large language model to obtain a target question, wherein the target question is obtained by merging one or more of the candidate questions by the large language model; receiving and displaying streaming data output by the large language model on the server, wherein the streaming data includes: a thought chain corresponding to the target question; performing incremental parsing processing on the thought chain through a parser to obtain a target thought chain corresponding to the target question, and displaying the target question and the target thought chain to the target user.
[0006] Secondly, embodiments of this application provide a front-end-based interactive query method applied to a client. The method includes: receiving a user question from a target user sent by the client; performing intent recognition on the user question using a large language model to obtain a candidate question set, and sending this set to the client; the candidate question set includes multiple candidate questions in an ordered manner; receiving one or more candidate questions selected by the target user from the candidate question set; merging the one or more candidate questions to obtain a target question; performing reasoning processing on the target question to obtain a thought chain corresponding to the target question; and sending the target question and the thought chain corresponding to the target question as streaming data to the client.
[0007] Thirdly, embodiments of this application provide a front-end-based interactive inquiry device, comprising: an acquisition module for acquiring user questions from a target user and inputting the user questions into a large language model on a server; a receiving module for receiving a set of candidate questions output by the large language model after performing intent recognition on the user questions; a response module for responding to one or more candidate questions selected by the target user from the set of candidate questions, inputting one or more of the candidate questions into the large language model to obtain a target question, wherein the target question is obtained by merging one or more of the candidate questions by the large language model; an output module for receiving and displaying streaming data output by the large language model on the server, wherein the streaming data includes a thought chain corresponding to the target question; and a display module for performing incremental parsing processing on the thought chain through a parser to obtain a target thought chain corresponding to the target question, and displaying the target question and the target thought chain to the target user.
[0008] Fourthly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement a front-end-based interactive query method as described above.
[0009] Fifthly, embodiments of this application provide a readable storage medium storing a program or instructions, wherein the program or instructions, when executed by a processor, constitute the aforementioned front-end-based interactive query method.
[0010] Sixthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement a front-end-based interactive query method as described above.
[0011] The technical solution adopted in this application embodiment is applied to a client. It obtains the user's question from the target user and inputs the user's question into a large language model on the server. It receives a set of candidate questions output by the large language model after intent recognition of the user's question. In response to the target user selecting one or more candidate questions from the candidate question set, it inputs one or more candidate questions into the large language model to obtain the target question, which is obtained by merging one or more candidate questions by the large language model. It receives and displays streaming data output by the server's large language model, including: the thought chain corresponding to the target question. It performs incremental parsing processing on the thought chain through a parser to obtain the target thought chain corresponding to the target question, and displays the target question and the target thought chain to the target user. As can be seen, by having the target user select the desired candidate question from the candidate question set, the large language model can output a more explicit target question. Furthermore, the client does not need to wait for the server to update the entire thought chain of the target question before displaying it. Instead, it can parse and process the increments in the thought chain in real time and display the update process of the thought chain, ultimately obtaining the target thought chain corresponding to the target question. This allows users to select more targeted target questions and improves the efficiency of displaying the target thought chain, solving the problem of the low speed of front-end display of the user's thought process corresponding to the question. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating a front-end-based interactive query method according to an embodiment of this application; Figure 2 This is a flowchart illustrating another front-end-based interactive query method provided according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a front-end-based interactive query system provided according to an embodiment of this application; Figure 4 This is a flowchart illustrating another front-end-based interactive query method provided according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a front-end-based interactive query device according to an embodiment of this application; Figure 6 This is a schematic diagram of another front-end-based interactive query device provided according to an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] The front-end-based interactive query method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0016] Figure 1 This illustration shows an embodiment of the present invention providing a front-end-based interactive query method. This method can be executed by a client, such as an in-vehicle terminal or a mobile terminal. The method includes the following steps: S102, obtain the user's question from the target user and input the user's question into the large language model on the server.
[0017] The target users include users who consult the client, which is the front end, such as software programs running on a mobile phone.
[0018] User questions include those posed by the target user. These can be fuzzy questions that a large language model cannot accurately determine, or questions that a large language model can accurately identify. For example, "What will Apple's price trend be in November?" The target user might be asking about the price trend of Apple phones, or it could be asking about the price trend of fruit. Such questions are classified as fuzzy questions.
[0019] Large language models include: deep learning models that handle multiple tasks; and models used for intent recognition of user questions and for obtaining corresponding thought processes through reasoning analysis.
[0020] The server-side is the backend. The large language model is integrated into the backend to conduct in-depth analysis of user problems and obtain the thought process and answers to solve user problems.
[0021] S104 receives the set of candidate questions output by the large language model after performing intent recognition on the user's question.
[0022] The candidate question set includes a set of multiple candidate questions related to the user's question, obtained by performing intent recognition and reasoning analysis on the user's question using a large language model.
[0023] Specifically, after receiving user questions sent by the client, the large language model uses deep learning and natural language processing techniques to map the user questions to predefined intent categories, resulting in multiple candidate questions, which are then used as a candidate question set.
[0024] The candidate question set can be used as streaming data and output through a large language model. There is no need to wait for all candidate questions in the set to be updated before outputting.
[0025] S106, responding to one or more candidate questions selected by the target user from the candidate question set, inputting one or more candidate questions into the large language model to obtain the target question, which is obtained by merging one or more candidate questions by the large language model.
[0026] The target user selects one or more candidate questions from the candidate question set that meet the expectations. The client responds with one or more candidate questions selected by the target user and inputs them into the large language model.
[0027] The large language model merges one or more candidate questions selected by the target user to obtain the target question. Specifically, the large language model can iteratively process the candidate questions through interaction with the client and the user to arrive at the target question.
[0028] As an example, the large language model merges one or more candidate questions selected by the target user to obtain the target candidate question. The target candidate question can be one or more. When the large language model outputs only one target candidate question, it is sent to the client as the target question. When the large language model outputs multiple target candidate questions, they are merged to obtain the target question. Alternatively, the large language model can iteratively process the candidate questions to obtain the target question. Specifically, multiple target candidate questions are sent as a set to the client. The client displays the target candidate questions to the target user, who selects one or more target candidate questions and inputs them as new candidate questions into the large language model. This process continues until the large language model merges and analyzes the new candidate questions to obtain the target question, which is then used as the question required by the target user. The corresponding thought chain for the target question is also sent to the client. The number of candidate questions selected by the target user and the number of candidate questions in the set generated by the large language model can be preset.
[0029] S108 receives and displays streaming data output from the server's large language model. The streaming data includes the thought process chain corresponding to the target problem.
[0030] The large language model performs reasoning and parsing on the target problem to obtain the corresponding thought chain, which is then sent to the client as streaming data. After receiving the streaming data, the client displays the thought chain to the target user. It should be understood that the thought chain corresponding to the target problem includes: data after the large language model has not fully reasoned and analyzed the target problem, but it can display the data of reasoning and parsing the target problem at the current moment.
[0031] Streaming data includes data output using a streaming method. Specifically, when a large language model generates content, the output data is broken down into multiple data chunks (called chunks or tokens), which are gradually transmitted and displayed to the target user through the client, like a stream of water, rather than returning the complete output data all at once. Its core lies in the gradual nature of data transmission: the server immediately pushes each small piece of text (e.g., a word or a sentence) generated to the client, allowing the client to receive and update the interface for the target user in real time, without waiting for the server to finish generating all the content.
[0032] The thought chain corresponding to the target problem refers to the chain of reasoning steps generated by the large language model to solve the target problem. This chain can be continuously updated as the large language model analyzes the target problem more deeply to obtain the final answer.
[0033] As can be seen, the client can receive streaming data output from the large language model. For example, it can receive data blocks at preset intervals using a streaming method. These data blocks include: a continuously updated thought chain corresponding to the target question, and potentially a continuously expanding set of candidate questions. It should be understood that the candidate questions in the received set can be continuously updated, and the thought chain corresponding to the target question can also be continuously updated. There is no need to wait for the server to perform intent recognition or reasoning analysis on the user's question before sending the complete set of candidate questions and the target thought chain corresponding to the target question to the client; instead, it can be output as streaming data.
[0034] S110 uses a parser to perform incremental parsing of the thought chain to obtain the target thought chain corresponding to the target problem, and then displays the target problem and the target thought chain to the target user.
[0035] The parser is used to generate the minimum set of changes through an algorithm (such as Myers' difference algorithm). In other words, the client uses the parser to parse and process the incremental data in the thought chain, determine the changed data in the thought chain, add the changed data to the thought chain, and obtain the target thought chain.
[0036] Specifically, the parser analyzes the incremental data of the thought chain corresponding to the target problem in the streaming data to obtain the change data of the thought chain. The incremental data of the thought chain is continuously acquired according to a preset time period. The incremental data is then subjected to incremental parsing processing to finally generate a new thought chain that is different from the thought chain that has not undergone incremental parsing, which serves as the target thought chain.
[0037] The target thought chain corresponding to the target problem refers to the thought chain obtained by the large language model based on the generated target problem and through reasoning analysis. The server performs reasoning analysis on the target problem, and the resulting thought chain is sent to the client as streaming data. The client performs incremental parsing processing on the obtained thought chain to obtain the target thought chain corresponding to the target problem. The target thought chain also includes the final answer obtained through multiple reasoning steps, which can break down complex problems into logically related sub-steps and ultimately integrate them to form a complete chain.
[0038] The technical solution adopted in this application embodiment is applied to a client. It obtains the user's question from the target user and inputs the user's question into a large language model on the server. It receives a set of candidate questions output by the large language model after intent recognition of the user's question. In response to the target user selecting one or more candidate questions from the candidate question set, it inputs one or more candidate questions into the large language model to obtain the target question, which is obtained by merging one or more candidate questions by the large language model. It receives and displays streaming data output by the server's large language model, including: the thought chain corresponding to the target question. It performs incremental parsing processing on the thought chain through a parser to obtain the target thought chain corresponding to the target question, and displays the target question and the target thought chain to the target user. As can be seen, by having the target user select the desired candidate question from the candidate question set, the large language model can output a more explicit target question. Furthermore, the client does not need to wait for the server to update the entire thought chain of the target question before displaying it. Instead, it can parse and process the increments in the thought chain in real time and display the update process of the thought chain, ultimately obtaining the target thought chain corresponding to the target question. This allows users to select more targeted target questions and improves the efficiency of displaying the target thought chain, solving the problem of the low speed of front-end display of the user's thought process corresponding to the question.
[0039] In one embodiment, receiving and displaying the streaming data output by the large language model from the server (i.e., S108) can be accomplished by performing the following step A: Step A: Receive and display the streaming data output by the server's large language model based on a predefined streaming structured response protocol and the target question; wherein, the large language model is used to perform intent recognition and reasoning analysis on the target question to obtain the thought chain corresponding to the target question.
[0040] The streaming structured response protocol is a data transmission mechanism that allows the server to send data to the client incrementally as it is being generated, rather than waiting for all the data to be ready before transmitting it all at once. By transmitting in chunks, latency is reduced and real-time performance is improved. Specifically, the streaming structured response protocol can use a predefined JSON / serialized structure to output data inferred from the user's question.
[0041] The large language model receives one or more candidate questions selected by the client, obtains the target question through intent recognition, and derives the corresponding thought chain through reasoning analysis of the target question. Based on a predefined streaming structured response protocol, the large language model can stream the reasoned thought chain, providing the client with streaming data. The client receives and displays the streaming data output by the large language model, allowing the client to display corresponding data based on changes in the output streaming data. For example, when the streaming data is a thought chain, the client receives the thought chain; upon receiving incremental data from the thought chain, the incremental data is added to the thought chain and displayed, until the target thought chain is obtained.
[0042] As an example, the streaming structured response protocol can be set as a JSON fragment: {"response_id":"r123","streaming":true,"segments":[{"id":"s1","type":"chain_step","content":"inference step text or token fragment","start_token":10,"end_token":27,"timestamp":1700000000000","source":"lm_v1","confidence":0.82,"evidence_refs":["doc:987:span:12-45"]}],"candidate_questions":[{"qid":"q1","text"]}] :“Is Y caused by X?”,“similarity”:0.92,“intent”:“cause_inquiry”},{“qid”:“q2”,“text”:“What is the scope of the impact of event A?”,“similarity”:0.89,“intent”:“impact_inquiry”}],“intents”:[{“label”:“investigate_cause”,“score”:0.87},{“label”:“monitor_trend”,“score”:0.45}],“provenance”:{“signature”:“sha256:abcd…”,“generated_at”:“2025-12-11T10:00:00Z”}}.
[0043] The large language model receives user questions sent by the client, performs intent recognition and reasoning analysis on the user questions, and outputs streaming data based on a predefined streaming structured response protocol. The streaming data includes: the thought chain corresponding to the target question, and also includes: a set of candidate questions after intent recognition of the user questions.
[0044] It should be noted that the streaming data includes: the thought chain corresponding to the target question, as well as: a set of candidate questions, intent labels for user questions identified by the large language model, intent labels for each candidate question, evidence citations for the target thought chain, and confidence scores for each candidate question. The evidence citations include: timestamps and source identification information.
[0045] In this embodiment, the large language model, based on a predefined streaming structured response protocol, outputs the results of reasoning about user questions in the form of streaming data, which can minimize the burden on the main thread and ensure timely response of the interface.
[0046] In one embodiment, incremental parsing of the thought chain is performed by a parser to obtain the target thought chain (i.e., S110) corresponding to the target problem. The following steps B1-B3 can be executed: Step B1: The parser performs incremental parsing on the thought chain in the streaming data to obtain the parsing result. The parsing result is then stored as a chain node to obtain the changed chain node.
[0047] The client uses a parser to process incremental data in the thought chain, generating parsing results showing the changes in the thought chain, and storing these results as chain nodes. Specifically, by comparing changes in the data within the thought chain, the changed portions of the data are parsed and updated to obtain the parsing results. This significantly reduces computational resource consumption, making it particularly suitable for scenarios with frequent updates.
[0048] The parsing results include: data after parsing the updated portion of the streaming data.
[0049] The changed chain node refers to the chain node after the client stores the parsing results. The parsing results include data after incremental parsing of streaming data.
[0050] Step B2: Obtain the chain nodes that have been changed in batches by using the identification information of each changed chain node, and perform differential rendering processing on the chain nodes that have been changed in batches to obtain the rendered chain nodes.
[0051] Incremental parsing also includes differential rendering, which involves updating the canvas based on differential algorithms for changed chain nodes. For example, by recording incremental changes in chain node attributes using a differential array, the operation of batch updating chain nodes can be optimized to perform canvas / DOM / Canvas updates, thereby reducing repaint costs and main thread usage.
[0052] The identification information includes: a unique identifier and version number for each chain node, used to minimize DOM lookups and replacements. The streaming data includes the identification information for each chain node in the thought chain.
[0053] Obtaining chain nodes that have been changed in batches involves using a data block-level chunk-level merging strategy to obtain the changed chain nodes in batches.
[0054] Differential rendering refers to updating the canvas by using differential algorithms to detect changed chain nodes. These algorithms include detecting changed chain nodes by calculating the pixel difference between two consecutive frames, or identifying changed chain nodes through color channel differences. The changed chain nodes are then rendered and updated to obtain the rendered chain nodes.
[0055] Step B3: Based on the thought chain and the rendered chain nodes, determine the target thought chain corresponding to the target problem.
[0056] The thought chain refers to the thought chain corresponding to the target problem. The rendered chain node refers to the chain node obtained after splitting and rendering the chain node after incremental parsing in step B2.
[0057] Based on the thought chain and the rendered chain nodes, the target thought chain corresponding to the target problem is obtained. It can be seen that the target thought chain includes: the thought chain corresponding to the initial target problem, as well as the chain nodes obtained after incremental parsing and differential rendering.
[0058] Specifically, the parser analyzes the incremental data of the thought chain in the streaming data to obtain the parsing results. These results are stored as chain nodes. The changed chain nodes undergo differential rendering to obtain differentially rendered chain nodes, generating the target thought chain corresponding to the target problem. The changed chain nodes include both altered nodes in the original thought chain and newly added chain nodes. The target thought chain includes the original thought chain and the chain nodes after differential rendering of the changed chain nodes, generating a new thought chain different from the thought chain without incremental parsing processing, which serves as the target thought chain.
[0059] As an example, the parsing and rendering pseudocode is as follows: `onStreamChunk(chunk) { parsed = Parser.parseChunk(chunk) / / Deserialize the stream fragment into segment nodes; Cache.append(parsed) deltas = DiffEngine.computeDeltas(parsed, UIState) UIRenderer.applyDeltas(deltas) / / Only update changed nodes if (parsed.candidate_questions) UI.showCandidateList(parsed.candidate_questions)} onUserSelect(selectedQids) { selected = Cache.lookupQuestions(selectedQids) mergedQuery = Merger.mergeByStrategy(selected) if (mergedQuery.confidence)` <CLIENT_THRESHOLD){NetworkAdapter.sendForRefinement(mergedQuery)}else{NetworkAdapter.sendToModel(mergedQuery)}}。
[0060] In this embodiment, by performing differential rendering on the changed chain nodes in batches on the client side, the incrementally parsed chain nodes can be displayed quickly, improving the speed at which users obtain information, reducing redrawing costs and main thread usage, and saving resources.
[0061] In one embodiment, the streaming data further includes: evidence citations of the thought chain, and the method may also perform the following step C: Step C: Obtain the evidence references carried by each chain node in the thought chain of the streaming data, and display the evidence references of the thought chain to the target user; wherein, the evidence references include one or more of the following: timestamp, signature information and corresponding original text.
[0062] Specifically, each node in the MindChain carries evidence refs, which can be expanded on demand by the client's page to display the corresponding evidence refs for the MindChain, such as raw data (original text) requested from the server or cached locally. Each evidence ref also carries a timestamp and signature information to ensure the auditability of the MindChain.
[0063] As an example, each chain_step carries evidence_refs, which are used to point to the document's ID and location, as well as a field for signature information (based on the returned content hash). When a user expands the evidence reference of a chain node in the client's page, they can view the original text and signature information to verify that it has not been tampered with.
[0064] Preferably, in offline or disconnected network scenarios, the most recent preset number of evidence references are cached locally, and operation logs are recorded to support subsequent replay and auditing.
[0065] In this embodiment, the streaming data output by the large language model also includes evidence citations for the thought chain. Based on the timestamps and signature information in the evidence citations, it is ensured that the thought chain has not been tampered with. The original text in the evidence citations provides the original data for the target user's reference and ensures the auditability of the data chain. Through the evidence citations carried by each thought chain segment, the corresponding original evidence can be loaded and displayed on demand when requested by the target user. The target user can quickly trace the evidence chain on the client side without initiating additional requests.
[0066] Figure 2 This invention illustrates another front-end-based interactive query method, applied to a server, comprising the following steps: S202: Receive the user question from the target user sent by the client, perform intent recognition on the user question through a large language model, obtain a candidate question set, and send it to the client. The candidate question set includes multiple candidate questions in order.
[0067] The target users include users who ask questions to the client.
[0068] User questions include: questions raised by the target users, which can be ambiguous questions that large language models cannot accurately judge, or questions that large language models can accurately identify.
[0069] Large language models include deep learning models that handle multiple tasks.
[0070] On the server side, the large language model performs intent recognition on the received user question, resulting in multiple candidate questions. These candidate questions are then sorted to obtain a candidate question set. The sorting process includes calculating the similarity, intent confidence, and diversity score of each candidate question, ranking them from highest to lowest total score to obtain ordered candidate questions. This candidate question set is then sent to the client as streaming data. The streaming data generated by the large language model also includes: the intent label of the user question identified by the large language model, the intent label of each candidate question, and the confidence score of each candidate question.
[0071] The diversity score includes a score calculated based on the maximum marginal diversity (MMD) or rearrangement penalty.
[0072] S204, Receive one or more candidate questions selected by the target user from the candidate question set by the client.
[0073] The client receives one or more candidate questions selected by the target user from the candidate question set and sends them to the server.
[0074] The server receives one or more candidate questions selected by the target user from the candidate question set, sent by the client.
[0075] S206: Merge one or more candidate questions to obtain the target question, and perform reasoning on the target question to obtain the thought chain corresponding to the target question.
[0076] The merging process includes merging a preset number of candidate questions into one question.
[0077] The target problem includes: the large language model iteratively processes candidate problems through interactions with clients and users, and finally obtains the problem.
[0078] The server uses a large language model to merge one or more candidate questions and iterates through the candidate questions multiple times to obtain the target question. By performing reasoning analysis on the target question, the corresponding thought process chain is obtained.
[0079] S208 sends the target problem and its corresponding thought chain as streaming data to the client.
[0080] After the server sends the target question and its corresponding thought chain as streaming data to the client, the client can generate a target thought chain based on the thought chain corresponding to the target question and display it to the target user.
[0081] The technical solution adopted in this application embodiment is applied to the server. It receives user questions from the client, performs intent recognition on the user questions using a large language model, obtains a candidate question set, and sends it to the client. The candidate question set includes multiple candidate questions in an order. The server receives one or more candidate questions selected by the target user from the candidate question set. It merges the one or more candidate questions to obtain the target question, performs reasoning processing on the target question to obtain the corresponding thought chain. The target question and its corresponding thought chain are then sent to the client as streaming data. As can be seen, the server, by performing intent recognition on the user questions sent by the client, obtains a candidate question set with an order and sends it to the client for the user to choose from. Based on the target user's selection of candidate questions, it performs merging processing to generate a more targeted target question and a corresponding thought chain, which are then sent to the client as streaming data. This eliminates the need to generate complete data before sending it to the client, improving the accuracy of user question recognition and enabling the front-end to continuously receive streaming data, thus improving the efficiency of displaying the final target thought chain of the user question.
[0082] In one embodiment, merging one or more candidate problems to obtain the target problem (i.e., S206) can be performed by the following steps D1-D3: Step D1: Based on a predefined template, concatenate one or more candidate questions into a compound query statement.
[0083] The predefined template allows users to combine a preset number of candidate questions into a single compound query. For example, a predefined template could be "Based on question A and question B, please explain...".
[0084] The server-side classifier uses a predefined template to concatenate one or more candidate questions to obtain a composite query. Furthermore, the candidate questions selected by the target user can be pre-sorted, for example, by combining similarity, confidence of intent, and diversity scores to obtain sorted candidate questions, which are then concatenated into the composite query.
[0085] Preferably, a first number of candidate questions selected by the target user is preset, and a compound query statement can be generated based on the first number of candidate questions; alternatively, a second number of candidate questions selected by the target user can be preset, and multiple compound query statements can be generated based on the second number of candidate questions; wherein the first number is less than the second number.
[0086] Step D2: Map each candidate question in the compound query statement to a vector to generate the query statement vector.
[0087] Step D3: Generate the target question based on the query statement vector and the target user's selection weight.
[0088] Based on the query statement vector generated in step D2, and the target user's selection weight for each candidate question in the query statement vector, a weighted average is calculated according to the target user's selection weight, and then analyzed using the nearest... The neighbor search can be used to retrieve or generate the final target question.
[0089] If the text similarity of candidate questions is higher than a threshold, then candidate questions with text similarity higher than the threshold will be merged, and the corresponding evidence citations will be merged.
[0090] In this embodiment, one or more candidate questions selected by the target user are concatenated according to a predefined template to obtain a composite query statement. Each candidate question in the composite query statement is mapped to a vector, and the target question is generated based on the target user's selection weight. This allows for the generation of a more accurate target question based on the preferences of one or more candidate questions selected by the target user.
[0091] In one embodiment, based on the query vector and the target user's selection weights, the target question is generated (i.e., step D3), and the following steps D31-D35 can be executed: Step D31: Use a lightweight classifier to perform intent recognition on the candidate questions corresponding to the query statement vector to obtain intent recognition results. Based on the intent recognition results and the target user's selection weight, determine the intent label and confidence level corresponding to the target question.
[0092] Lightweight classifiers include miniaturized Transformers or Support Vector Machines.
[0093] The candidate questions corresponding to the query statement vector include: the candidate questions selected by the target user.
[0094] A lightweight classifier is used to identify the intent of candidate questions corresponding to the query vector, yielding intent identification results. This involves classifying the intent of each candidate question to obtain an intent label, which is then used to determine the category of the candidate question, facilitating reasoning and analysis of the thought process. Finally, the confidence level of the target question is determined based on the target user's selection weight for each candidate question.
[0095] Step D32: When the confidence level is greater than a preset threshold, generate the target question.
[0096] Step D33: When the confidence level is less than or equal to a preset threshold, output a target candidate question set based on one or more candidate questions and send it to the client, and receive one or more candidate questions selected by the target user from the target candidate set from the client.
[0097] The preset threshold can be 0.8. When the confidence level is less than or equal to the preset threshold, the large language model performs merging processing based on one or more candidate questions selected by the user, resulting in multiple merged candidate questions. These multiple merged candidate questions are then used as the target candidate question set and sent to the client.
[0098] The client displays the target candidate question set to the target user, responds to one or more candidate questions selected by the target user from the target candidate question set, and sends them to the server.
[0099] Step D34 involves concatenating one or more candidate questions selected by the target user from the target candidate set to obtain the corresponding compound query statement.
[0100] Specifically, for one or more candidate questions selected by the target user from the target candidate set, the query is concatenated according to a preset template to obtain the corresponding compound query statement.
[0101] Step D35: Generate the target question when the confidence level of the candidate questions in the compound query statement is greater than a preset threshold.
[0102] The candidate questions are iteratively processed until the confidence level of the candidate questions in the resulting compound query statement is greater than a preset threshold. Based on the compound query statement, the target question is then determined.
[0103] It should be noted that in high-latency or low-bandwidth environments, the number of candidate questions can be limited to N (e.g., N=3), and a coarser rendering (e.g., only text summaries instead of complete evidence fragments) can be used. Furthermore, the preset threshold can be adjusted based on the target scenario.
[0104] In this embodiment, by identifying the intent of each candidate question in the query statement, the intent label and confidence level corresponding to the target question are obtained. By judging the size of the confidence level, it is determined whether to merge the candidate questions selected by the target user to obtain a target candidate question set. The target user selects candidate questions again from the target candidate question set so that the confidence level of the obtained target question meets the preset threshold. Compared with related technologies, which often only return a single answer or require the target user to re-enter, this can efficiently guide the user to obtain a more accurate requirement through the iteration of candidate questions.
[0105] Figure 3 This is a schematic diagram of the structure of a front-end-based interactive query system according to an embodiment of this application, as shown below. Figure 3As shown, the system includes: a Stream Parser module for client-side parsing of incremental thought chains; an Incremental Renderer module for client-side differential rendering of incremental thought chains; an Intent Engine module for server-side intent recognition processing of user input questions and candidate questions using a large language model; a Question Selector / Merger module for server-side merging of candidate questions using a large language model; a Network Adapter module for obtaining network information for configuration data; and a Cache module for caching chain nodes and evidence references within the thought chains.
[0106] It should be noted that the interactive query system dynamically adjusts the rendering quality and the number of sampled candidate questions based on network latency / bandwidth and device performance metrics. This significantly reduces end-to-end interaction latency, minimizes unnecessary backend calls and bandwidth consumption, improves the efficiency of target users in identifying intent and clarifying questions, and enhances the interpretability and auditability of the output.
[0107] Figure 4 This is a schematic diagram of a scenario flow for another front-end-based interactive query method provided in the embodiments of this application, applied to a system, such as... Figure 4 As shown, the method includes the following steps: S401, the client obtains the user's question from the target user and inputs the user's question into the server's large language model.
[0108] S402, the server uses a large language model to identify the user's intent in the question, obtains a set of candidate questions, and sends them to the client.
[0109] S403: The client receives the set of candidate questions output by the large language model after it performs intent recognition on the user's question and displays it to the target user.
[0110] S404, the client responds to the target user by selecting one or more candidate questions from the candidate question set and inputs one or more candidate questions into the large language model.
[0111] S405, the large language model, is based on a predefined template. It concatenates one or more candidate questions into a compound query statement and maps each candidate question in the compound query statement into a vector to generate a query statement vector.
[0112] S406 uses a lightweight classifier to perform intent recognition on the candidate questions corresponding to the query statement vector, obtains intent recognition results, and determines the intent label and confidence level corresponding to the target question based on the intent recognition results and the target user's selection weight.
[0113] S407: When the confidence level is greater than the preset threshold, generate the target question.
[0114] S408, when the confidence level is less than or equal to a preset threshold, output a target candidate question set based on one or more candidate questions and send it to the client, and receive one or more candidate questions selected by the target user from the target candidate set sent by the client.
[0115] S409: Concatenate one or more candidate questions selected by the target user from the target candidate set to obtain the corresponding compound query statement.
[0116] S410, until the confidence level of the candidate questions in the compound query statement is greater than the preset threshold, generate the target question, and perform reasoning processing on the target question to obtain the thought chain corresponding to the target question.
[0117] S411 sends the target problem and its corresponding thought chain as streaming data to the client.
[0118] S412: The client receives and displays the server's large language model based on a predefined streaming structured response protocol and target question, outputting streaming data.
[0119] S413, the parser performs incremental parsing on the thought chain in the streaming data, obtains the parsing result, stores the parsing result as a chain node, and obtains the changed chain node.
[0120] S414: Obtain the batch of changed chain nodes by using the identification information of each changed chain node, and perform differential rendering processing on the batch of changed chain nodes to obtain the rendered chain nodes.
[0121] S415, Based on the thought chain and the rendered chain nodes, determine the target thought chain corresponding to the target problem.
[0122] S416 presents the target problem, the target thought chain corresponding to the target problem, and the evidence citations for the target thought chain corresponding to the target problem to the target user.
[0123] The specific processes from S401 to S416 described above have been explained in detail in the above embodiments and will not be repeated here.
[0124] The technical solution adopted in this application uses an incremental parsing method on the client side to process the streaming data output by the large language model. Real-time parsing of the incremental data in the thought process chain improves the efficiency of obtaining the target thought process chain. Furthermore, users can select more targeted target questions, solving the problem of low display speed of the thought process corresponding to user questions on the front end. The server obtains the user questions sent by the client to generate streaming data, eliminating the need to generate complete data before sending it to the client. Moreover, the candidate questions in the candidate question set can be sorted for user selection. The selections of the target user are merged to generate more targeted target questions, and the corresponding thought processes for the target questions are generated and sent to the client, improving the accuracy of user question identification.
[0125] It should be noted that the interactive query method based on the front end provided in this application can be executed by an interactive query device or a control module within that interactive query device for executing the interactive query method based on the front end. This application uses the execution of the interactive query method based on the front end by an interactive query device as an example to illustrate the interactive query device provided in this application.
[0126] Figure 5 This is a schematic diagram of a front-end-based interactive query device according to an embodiment of the present invention. It is applied to a client, such as... Figure 5 As shown, the interactive query device includes: an acquisition module 51, a receiving module 52, a response module 53, an output module 54, and a display module 55. The acquisition module 51 is used to acquire user questions from the target user and input the user questions into the large language model on the server. The receiving module 52 is used to receive the set of candidate questions output by the large language model after performing intent recognition on the user's question; The response module 53 is used to respond to one or more candidate questions selected by the target user from the candidate question set, and input one or more candidate questions into the large language model to obtain the target question. The target question is obtained by merging one or more candidate questions by the large language model. Output module 54 is used to receive and display streaming data output by the large language model on the server. The streaming data includes: the thought chain corresponding to the target question. The display module 55 is used to perform incremental parsing of the thought chain through the parser, obtain the target thought chain corresponding to the target problem, and display the target problem and the target thought chain to the target user.
[0127] In one embodiment, the output module 54 is specifically used to receive and display the streaming data output by the server's large language model based on a predefined streaming structured response protocol and the target question; wherein, the large language model is used to perform reasoning analysis on the target question to obtain the thought chain corresponding to the target question.
[0128] In one embodiment, the display module 55 is specifically used to perform incremental parsing processing on the thought chain in the streaming data through a parser to obtain the parsing result, store the parsing result as a chain node to obtain the changed chain node; obtain the batch changed chain node through the identification information of each changed chain node, and perform differential rendering processing on the batch changed chain node to obtain the rendered chain node; and determine the target thought chain corresponding to the target problem based on the thought chain and the rendered chain node.
[0129] In one embodiment, the streaming data further includes: evidence references for the thought chain. The device is also used to acquire evidence references carried by each chain node of the thought chain in the streaming data and display the evidence references of the thought chain to the target user. The evidence references include one or more of the following: timestamp, signature information, and corresponding original text.
[0130] The technical solution adopted in this application embodiment is applied to a client. It obtains the user's question from the target user and inputs the user's question into a large language model on the server. It receives a set of candidate questions output by the large language model after intent recognition of the user's question. In response to the target user selecting one or more candidate questions from the candidate question set, it inputs one or more candidate questions into the large language model to obtain the target question, which is obtained by merging one or more candidate questions by the large language model. It receives and displays streaming data output by the server's large language model, including: the thought chain corresponding to the target question. It performs incremental parsing processing on the thought chain through a parser to obtain the target thought chain corresponding to the target question, and displays the target question and the target thought chain to the target user. As can be seen, by having the target user select the desired candidate question from the candidate question set, the large language model can output a more explicit target question. Furthermore, the client does not need to wait for the server to update the entire thought chain of the target question before displaying it. Instead, it can parse and process the increments in the thought chain in real time and display the update process of the thought chain, ultimately obtaining the target thought chain corresponding to the target question. This allows users to select more targeted target questions and improves the efficiency of displaying the target thought chain, solving the problem of the low speed of front-end display of the user's thought process corresponding to the question.
[0131] Figure 6 This is a schematic diagram of another front-end-based interactive query device according to an embodiment of the present invention. It is applied to the client-server side, such as... Figure 6As shown, the interactive query device includes: a generation module 61, a selection module 62, a reasoning module 63, and a sending module 64. The generation module 61 is used to receive user questions from the target user sent by the client, perform intent recognition on the user questions through a large language model, obtain a set of candidate questions, and send them to the client. The set of candidate questions includes multiple candidate questions in order. Selection module 62 is used to receive one or more candidate questions selected by the target user from a set of candidate questions sent by the client; The reasoning module 63 is used to merge one or more candidate questions to obtain the target question, and to perform reasoning on the target question to obtain the thought chain corresponding to the target question. The sending module 64 is used to send the target problem and the corresponding thought chain as streaming data to the client.
[0132] In one embodiment, the inference module 63 is used to concatenate one or more candidate questions into a composite query statement based on a predefined template; map each candidate question in the composite query statement into a vector to generate a query statement vector; and generate a target question based on the query statement vector and the target user's selection weight.
[0133] In one embodiment, the inference module 63 is specifically configured to perform intent recognition on the candidate questions corresponding to the query statement vector using a lightweight classifier, obtain intent recognition results, determine the intent label and confidence level corresponding to the target question based on the intent recognition results and the target user's selection weight; generate the target question when the confidence level is greater than a preset threshold; output a target candidate question set and send it to the client based on one or more candidate questions, and receive one or more candidate questions selected by the target user from the target candidate set from the client; concatenate the one or more candidate questions selected by the target user from the target candidate set to obtain the corresponding composite query statement; generate the target question until the confidence level of the candidate questions in the composite query statement is greater than the preset threshold.
[0134] The technical solution adopted in this application embodiment is applied to the server. It receives user questions from the client, performs intent recognition on the user questions using a large language model, obtains a candidate question set, and sends it to the client. The candidate question set includes multiple candidate questions in an order. The server receives one or more candidate questions selected by the target user from the candidate question set. It merges the one or more candidate questions to obtain the target question, performs reasoning processing on the target question to obtain the corresponding thought chain. The target question and its corresponding thought chain are then sent to the client as streaming data. As can be seen, the server, by performing intent recognition on the user questions sent by the client, obtains a candidate question set with an order and sends it to the client for the user to choose from. Based on the target user's selection of candidate questions, it performs merging processing to generate a more targeted target question and a corresponding thought chain, which are then sent to the client as streaming data. This eliminates the need to generate complete data before sending it to the client, improving the accuracy of user question recognition and enabling the front-end to continuously receive streaming data, thus improving the efficiency of displaying the final target thought chain of the user question.
[0135] The interactive query device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0136] The interactive query device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0137] The interactive query device provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the method embodiments are not described in detail here to avoid repetition.
[0138] Based on the same technical concept, embodiments of this application also provide an electronic device for executing the aforementioned front-end-based interactive query method. Figure 7 This is a schematic diagram of the structure of an electronic device to implement various embodiments of this application. The electronic device can vary significantly due to differences in configuration or performance, and may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call a computer program stored in the memory 730 and executable on the processor 710 to perform the following steps: Obtain user questions from the target users and input them into the large language model on the server. Receive the set of candidate questions output by the large language model after it performs intent recognition on the user's question; In response to one or more candidate questions selected by the target user from the candidate question set, the one or more candidate questions are input into the large language model to obtain the target question, which is obtained by merging the one or more candidate questions by the large language model; Receive and display streaming data output from the server's large language model. The streaming data includes: the thought process chain corresponding to the target question. The parser performs incremental parsing of the thought chain to obtain the target thought chain corresponding to the target problem, and then displays the target problem and the target thought chain to the target user.
[0139] The technical solution adopted in this application embodiment is applied to a client. It obtains the user's question from the target user and inputs the user's question into a large language model on the server. It receives a set of candidate questions output by the large language model after intent recognition of the user's question. In response to the target user selecting one or more candidate questions from the candidate question set, it inputs one or more candidate questions into the large language model to obtain the target question, which is obtained by merging one or more candidate questions by the large language model. It receives and displays streaming data output by the server's large language model, including: the thought chain corresponding to the target question. It performs incremental parsing processing on the thought chain through a parser to obtain the target thought chain corresponding to the target question, and displays the target question and the target thought chain to the target user. As can be seen, by having the target user select the desired candidate question from the candidate question set, the large language model can output a more explicit target question. Furthermore, the client does not need to wait for the server to update the entire thought chain of the target question before displaying it. Instead, it can parse and process the increments in the thought chain in real time and display the update process of the thought chain, ultimately obtaining the target thought chain corresponding to the target question. This allows users to select more targeted target questions and improves the efficiency of displaying the target thought chain, solving the problem of the low speed of front-end display of the user's thought process corresponding to the question.
[0140] When applied to the server side, the following steps can be performed: The system receives user questions from the target user sent by the client, performs intent recognition on the user questions using a large language model, obtains a set of candidate questions, and sends it to the client. The set of candidate questions includes multiple candidate questions in an ordered manner. Receive one or more candidate questions selected by the target user from a set of candidate questions from the client. One or more candidate questions are merged to obtain the target question, and the target question is reasoned to obtain the thought chain corresponding to the target question; The target problem and its corresponding thought process chain are sent to the client as streaming data.
[0141] The technical solution adopted in this application embodiment is applied to the server. It receives user questions from the client, performs intent recognition on the user questions using a large language model, obtains a candidate question set, and sends it to the client. The candidate question set includes multiple candidate questions in an order. The server receives one or more candidate questions selected by the target user from the candidate question set. It merges the one or more candidate questions to obtain the target question, performs reasoning processing on the target question to obtain the corresponding thought chain. The target question and its corresponding thought chain are then sent to the client as streaming data. As can be seen, the server, by performing intent recognition on the user questions sent by the client, obtains a candidate question set with an order and sends it to the client for the user to choose from. Based on the target user's selection of candidate questions, it performs merging processing to generate a more targeted target question and a corresponding thought chain, which are then sent to the client as streaming data. This eliminates the need to generate complete data before sending it to the client, improving the accuracy of user question recognition and enabling the front-end to continuously receive streaming data, thus improving the efficiency of displaying the final target thought chain of the user question.
[0142] The specific execution steps can be found in the various steps of the above-described front-end-based interactive query method embodiment, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0143] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.
[0144] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.
[0145] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0146] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0147] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described front-end-based interactive query method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0148] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0149] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described front-end-based interactive query method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0150] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0151] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0153] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A front-end-based interactive query method, characterized in that, Applied to a client, the method includes: Obtain the user questions from the target user and input the user questions into the large language model on the server. Receive the set of candidate questions output by the large language model after performing intent recognition on the user's question; In response to the target user selecting one or more candidate questions from the candidate question set, the one or more candidate questions are input into the large language model to obtain the target question, which is obtained by the large language model through merging the one or more candidate questions; Receive and display streaming data output by the large language model on the server, the streaming data including: the thought chain corresponding to the target question; The parser performs incremental parsing on the thought chain to obtain the target thought chain corresponding to the target problem, and then displays the target problem and the target thought chain to the target user.
2. The method according to claim 1, characterized in that, The process of receiving and displaying the streaming data output by the large language model from the server includes: The system receives and displays the streaming data output by the large language model on the server side based on a predefined streaming structured response protocol and the target question; wherein the large language model is used to perform reasoning analysis on the target question to obtain the thought chain corresponding to the target question.
3. The method according to claim 1, characterized in that, The step of incrementally parsing the thought chain through a parser to obtain the target thought chain corresponding to the target problem includes: The parser performs incremental parsing on the thought chain in the streaming data to obtain the parsing result, and stores the parsing result as a chain node to obtain the changed chain node. By using the identification information of each changed chain node, the chain nodes that have been changed in batches are obtained, and differential rendering processing is performed on the chain nodes that have been changed in batches to obtain the rendered chain nodes. Based on the thought chain and the rendered chain nodes, the target thought chain corresponding to the target problem is determined.
4. The method according to claim 1, characterized in that, The streaming data also includes: evidence citations of the thought chain; Obtain the evidence reference carried by each chain node of the thought chain in the streaming data, and display the evidence reference of the thought chain to the target user; wherein, the evidence reference includes one or more of the following: timestamp, signature information and corresponding original text.
5. A front-end-based interactive query method, characterized in that, Applied to the server side, the method includes: The system receives user questions from the target user sent by the client, performs intent recognition on the user questions using a large language model, obtains a set of candidate questions, and sends it to the client. The set of candidate questions includes multiple candidate questions in an ordered manner. Receive one or more candidate questions selected by the target user from the candidate question set, sent by the client; One or more candidate questions are merged to obtain a target question, and the target question is reasoned to obtain the thought chain corresponding to the target question. The target problem and the corresponding thought chain are sent as streaming data to the client.
6. The method according to claim 5, characterized in that, The process of merging one or more candidate problems to obtain the target problem includes: Based on a predefined template, one or more of the candidate questions are concatenated into a compound query statement; Map each candidate question in the composite query statement to a vector to generate a query statement vector; The target question is generated based on the query statement vector and the selection weight of the target user.
7. The method according to claim 6, characterized in that, The step of generating the target question based on the query statement vector and the target user's selection weight includes: A lightweight classifier is used to perform intent recognition on the candidate questions corresponding to the query statement vector to obtain intent recognition results. Based on the intent recognition results and the selection weight of the target user, the intent label and confidence level corresponding to the target question are determined. When the confidence level is greater than a preset threshold, the target question is generated; When the confidence level is less than or equal to the preset threshold, a target candidate question set is output based on one or more of the candidate questions and sent to the client, and the client receives one or more of the candidate questions selected by the target user from the target candidate set. The candidate questions selected by the target user from the target candidate set are concatenated to obtain the corresponding composite query statement; The target question is generated when the confidence level of the candidate question in the compound query statement is greater than the preset threshold.
8. A front-end-based interactive query device, characterized in that, Applied to the client side, including: The acquisition module is used to acquire user questions from the target user and input the user questions into the large language model on the server. The receiving module is used to receive the set of candidate questions output by the large language model after performing intent recognition on the user question; The response module is used to respond to one or more candidate questions selected by the target user from the candidate question set, and input one or more of the candidate questions into the large language model to obtain the target question, wherein the target question is obtained by the large language model through merging processing of one or more of the candidate questions; The output module is used to receive and display the streaming data output by the large language model on the server, the streaming data including: the thought chain corresponding to the target question; The display module is used to perform incremental parsing processing on the thought chain through a parser to obtain the target thought chain corresponding to the target problem, and to display the target problem and the target thought chain to the target user.
9. An electronic device, characterized in that, The system includes a processor and a memory electrically connected to the processor, the memory storing a computer program, and the processor being configured to call and execute the computer program from the memory to implement a front-end-based interactive query method as described in claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium is used to store a computer program that can be executed by a processor to implement a front-end-based interactive query method as described in claims 1-7.