Data processing method and device and model training method and device

Through the small model, multiple rounds of problem updates and clarification interactions are carried out, and combined with the multi-agent model to identify and handle exceptions, the ambiguity and incomplete problems of the large model when processing user input is solved, achieving more efficient and accurate response content generation, and improving user satisfaction and system performance.

CN120297331APending Publication Date: 2025-07-11GUANGZHOU DULING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329729.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The large model has ambiguity, incompleteness or lengthy problems when processing user input, resulting in reduced correlation and accuracy of output results, decreased user satisfaction, and high computing resource consumption, lack of user feedback optimization and multiple rounds of interaction, resulting in incomplete or uncorrelated answers.

Method used

Multiple rounds of problem updates are performed through small models, clarification of inquiry information and obtain feedback information, combined with multiple proxy models to identify and handle exceptions, use the collaborative work of small models and large models to optimize input quality and reduce computing resource consumption, and guide user interaction to improve input accuracy and completeness.

Benefits of technology

It improves the clarity and integrity of user input, improves the correlation and integrity of model output, reduces computing resource consumption, and enhances user satisfaction and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297331A_ABST
    Figure CN120297331A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device and a model training method and device, and relates to the field of artificial intelligence, in particular to the fields of large models, natural language processing, deep learning and the like. According to the specific implementation scheme, according to an input problem, a first model is adopted to conduct at least one round of problem updating of the problem; in any round of question updating process, displaying clarification inquiry information for at least one time, and obtaining corresponding clarification feedback information; performing question updating according to the at least one time of clarification inquiry information and clarification feedback information corresponding to the at least one time of clarification inquiry information; and reply contents are generated according to the updated questions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technologies such as large models, natural language processing, and deep learning. In particular, it relates to a data processing method, apparatus, and model training method and apparatus. Background Art

[0002] In related technologies, the input of large models often has problems such as ambiguity, incompleteness, or verbosity. This makes it difficult for large models to accurately understand the user's intention, thereby affecting the relevance and accuracy of the model output results and reducing user satisfaction. Summary of the Invention

[0003] The present disclosure provides a data processing method, apparatus, and model training method and apparatus.

[0004] According to a first aspect of the present disclosure, a data processing method is provided, including:

[0005] Performing at least one round of question update on the question using a first model according to the input question;

[0006] During any round of question update, presenting at least one clarification query message and obtaining the corresponding clarification feedback message;

[0007] Updating the question according to the at least one clarification query message and the clarification feedback message corresponding to the at least one clarification query message;

[0008] Generating a response content based on the updated question.

[0009] According to a second aspect of the present disclosure, a model training method is provided, including:

[0010] Obtaining multiple groups of training samples, where the training samples include the input question, at least one clarification query message for any round of question update and the corresponding clarification feedback message, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content;

[0011] Training a first model based on the multiple groups of training samples.

[0012] According to a third aspect of the present disclosure, a data processing apparatus is provided, including:

[0013] A first processing module for performing at least one round of question update on the question using a first model according to the input question;

[0014] A second processing module for presenting at least one clarification query message and obtaining the corresponding clarification feedback message during any round of question update;

[0015] An update module, configured to update the question according to the at least one clarification query information and the clarification feedback information corresponding to the at least one clarification query information;

[0016] A generation module, configured to generate a reply content based on the updated question.

[0017] According to a fourth aspect of the present disclosure, there is provided a model training device, including:

[0018] A second acquisition module, configured to acquire multiple groups of training samples, where the training samples include an input question, at least one clarification query information for any round of question update and the corresponding clarification feedback information, a reply content generated based on the updated question, and a satisfaction feedback corresponding to the reply content;

[0019] A first training module, configured to train a first model based on the multiple groups of training samples.

[0020] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data processing method as described in the first aspect, or execute the model training method as described in the second aspect.

[0024] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the data processing method as described in the first aspect, or execute the model training method as described in the second aspect.

[0025] According to a seventh aspect of the present disclosure, there is provided a computer program product, including computer instructions, where the computer instructions, when executed by a processor, implement the steps of the data processing method as described in the first aspect, or implement the steps of the model training method as described in the second aspect.

[0026] A data processing method, device, and model training method and device provided by the present disclosure have the following beneficial effects:

[0027] According to the input question, at least one round of question update is performed on the question using the first model; during any round of question update, at least one clarification inquiry message is displayed, and the corresponding clarification feedback message is obtained; the question is updated according to at least one clarification inquiry message and the clarification feedback message corresponding to at least one clarification inquiry message; a reply content is generated based on the updated question. In the present disclosure, by performing multiple clarification interactions on the input question using the first model, the updated question becomes more accurate and complete, and thus the reply content generated based on the updated question is significantly improved in terms of relevance and integrity, solving the problem in the related art that the answer is off-topic or incomplete due to unclear questioning, and enhancing user satisfaction.

[0028] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0030] Figure 1 is a schematic flowchart of a data processing method according to a first embodiment of the present disclosure;

[0031] Figure 2 is a schematic flowchart of a data processing method according to a second embodiment of the present disclosure;

[0032] Figure 3 is a schematic diagram of model collaboration optimization prompts according to a third embodiment of the present disclosure;

[0033] Figure 4 is a schematic flowchart of a data processing method according to a fourth embodiment of the present disclosure;

[0034] Figure 5 is a schematic flowchart of a model training method according to a fifth embodiment of the present disclosure;

[0035] Figure 6 is a schematic structural diagram of a data processing device according to a sixth embodiment of the present disclosure;

[0036] Figure 7 is a schematic structural diagram of a data processing device according to a seventh embodiment of the present disclosure;

[0037] Figure 8 shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0039] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information is carried out on the premise of obtaining the user's consent, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0040] In related technologies, large Transformer architecture language models (LLMs) (also known as large language models or large models), such as Generative Pre-trained Transformer (GPT), Bidirectional Encoder Representations from Transformers (BERT), etc., face many challenges in practical applications. For example, the input quality is uneven, the single-round input method is inefficient, the inference resource consumption is high, the computing resources are wasted, the response integrity and accuracy are insufficient, the user feedback is not fully utilized, and the input optimization strategy is single and inefficient. Therefore, it is urgent to optimize the input link to improve performance and reduce costs.

[0041] Uneven input quality: The input provided by users often has problems such as ambiguity, incompleteness, or verbosity, which makes it difficult for large models to accurately understand the intention, thereby affecting the relevance and accuracy of the answers and reducing user satisfaction. For example, when there are multiple possible interpretations of the input provided by the user, the large model often has difficulty identifying and clarifying the ambiguity in a timely manner, resulting in an unsatisfactory answer and affecting the user experience.

[0042] Inefficient single-round input method: The single-round input method lacks a mechanism for interactive confirmation with users, which may lead to misunderstandings of user needs or omission of important information, thereby affecting the relevance and accuracy of answers and reducing user satisfaction. For example, most interactions with large models are in the "question and answer" single-round mode, where user inputs are directly handed over to the large model for generating answers without special processing. If the user's question is unclear, the large model can only try to answer based on incomplete information, often resulting in unsatisfactory results. Although users can iterate their questions multiple times to improve the results themselves, this requires users to have certain skills in prompt design and patience. In other words, the existing process completely shifts the burden of prompt optimization to users and lacks systematic assistance and guidance. Moreover, without guidance, users may need to repeatedly trial and error, wasting both time and increasing the cost of multiple calls to the large model.

[0043] High inference resource consumption: Large models have a huge number of parameters and require a large amount of computing resources and memory space during inference, resulting in significant computational overhead for each call. When the length or complexity of the input content is long, the processing overhead of the large model increases exponentially. For example, some Retrieval-Augmented Generation (RAG) processes need to put a large amount of context into the prompt, which will significantly increase the processing volume of the large model. Therefore, how to reduce the number of calls and processing load of the large model while ensuring performance is a key requirement.

[0044] Waste of computing resources: The large model call strategy is often "one-size-fits-all". Whether the user input is simple or complex, it directly uses the large model for processing. This approach causes waste of computing power when dealing with simple queries (many simple tasks that could be completed by small models or rule engines consume the precious inference computing power of the large model). For example, in resource-constrained scenarios or real-time services, directly calling the large model to process all inputs will lead to increased computing costs and excessive latency. In addition, without pre-screening and streamlining of inputs, the large model may process long and irrelevant information, increasing the unnecessary token burden. Therefore, how to elastically slice computing resources according to needs so that the large model can achieve an economical and intelligent response is a key requirement.

[0045] Insufficient response integrity and accuracy: Due to input problems, the answers of large models are sometimes not perfect, which usually manifests as omitting key points, lack of coherence in context, and even possibly containing "hallucination" content that does not conform to facts. For example, when the user's question exceeds the knowledge of the model training or the provided information is insufficient, the large model tends to generate untrue content (i.e., the so-called hallucination). In addition, if the question involves complex structures or multi-level information, the large model may ignore some details in a single generation and the output does not meet expectations. This indicates that relying solely on the large model itself is difficult to ensure that the answers are both complete and relevant, and optimization is needed from the input stage.

[0046] Insufficient utilization of user feedback: Large language models rarely use user feedback in real time to optimize subsequent answers. Usually, if users are dissatisfied with an answer, they can only manually adjust the question and ask again. The large language model does not remember whether the previous answer was accurate, resulting in the inability to transfer experience. Although offline Reinforced Learning with Human Feedback (RLHF) can align with human preferences during the model training phase, there is no closed-loop for dynamically using user feedback in the online system. Each conversation is like starting anew, and the model cannot improve itself based on user responses. This fragmentation limits the model's ability to continuously improve the quality of question answering.

[0047] Single and inefficient input optimization strategies: Related technologies attempt to rewrite user input through methods such as prompt templates and keyword hints, but usually stay at static rules or simple rewriting, lacking intelligence and interaction. For example, in Prompt Engineering, fixed templates are pre-designed and then users are asked to apply them. This method relies heavily on experience and is difficult to adapt to the ever-changing question requirements. In addition, related technologies have also proposed allowing the model to reflect on and improve the prompt by itself, but without introducing external information or user participation, the model may fall into a self-loop with limited effects. In short, input optimization in related technologies lacks multi-round interaction and automatic data driving, making it difficult to fully improve efficiency and quality.

[0048] In view of the above problems, the present disclosure provides a data processing method, apparatus, as well as a model training method and apparatus, to improve the quality of user questions, reduce unnecessary computational overhead, and enhance the integrity and relevance of answers.

[0049] The data processing method, apparatus, as well as the model training method and apparatus according to embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0050] It should be noted that the execution subject of the data processing method in this embodiment is a data processing apparatus, which can be implemented in software and / or hardware and can be configured in an electronic device.

[0051] Figure 1 is a schematic flowchart of the data processing method according to the first embodiment of the present disclosure.

[0052] As Figure 1 shown, the data processing method includes:

[0053] Step 101, according to the input question, use a first model to perform at least one round of question update on the question.

[0054] In the embodiments of the present disclosure, the input problem is a problem provided by the user that needs to be processed or answered, which is the original problem input by the user in the system interface. For example, it can be "I want to buy a new mobile phone. Do you have any suggestions?" or "I want to plan a trip.", etc.

[0055] In the embodiments of the present disclosure, the first model is a pre-trained model that can analyze and process the input problem. The number of the first models can be one or more, and the embodiments of the present disclosure do not limit this.

[0056] In an alternative embodiment, the specific type and function of each first model depend on the specific problem it is trained to solve. For example, the first model can include a model for improving the clarity of the problem, a model for adding context background information, a model for correcting ambiguous wording, a model for compressing problems containing redundant information, and so on.

[0057] In an alternative embodiment, in order to improve the processing efficiency and reduce the computational cost, a small model (a model with far fewer model parameters than a large model) can be used as the first model.

[0058] In the embodiments of the present disclosure, since the input problem (the user's original problem) may be ambiguous, incomplete, or verbose, therefore, for the input problem, a corresponding reply will not be directly given. Instead, the first model is used to perform at least one round of problem update on the input problem to further clarify or refine the user's problem.

[0059] Step 102, in any round of problem update process, display at least one clarification query message and obtain the corresponding clarification feedback message.

[0060] In the embodiments of the present disclosure, the clarification query message is determined based on the problem to be updated in this round and is used to better understand the user's intention. For example, if the problem to be updated in this round is "I want to plan a trip.", the clarification query message can be "Can you provide the destination or budget of the trip?"

[0061] In the embodiments of the present disclosure, the clarification feedback message is the user's answer to the displayed clarification query message. For example, if the clarification query message is "Can you provide the destination or budget of the trip?" and the user's answer to this question is "The approximate budget is 10,000 yuan and I want to go to an island.", then "The approximate budget is 10,000 yuan and I want to go to an island." is the clarification feedback message corresponding to the clarification query message "Can you provide the destination or budget of the trip?".

[0062] In the embodiments of the present disclosure, since the problem to be updated in any round of problem update process may be ambiguous, incomplete or lengthy, it is necessary to generate and display at least one clarification inquiry message based on the problem to be updated in any round of problem update process, and obtain the corresponding clarification feedback message to further clarify or refine the user's problem.

[0063] In an alternative embodiment, the problem to be updated in any round of problem update process has at least one of the following problems: being ambiguous, incomplete, or lengthy, so that at least one clarification inquiry message can be displayed according to the problems existing in the problem to be updated.

[0064] In an alternative embodiment, displaying at least one clarification inquiry message according to the problems existing in the problem to be updated may include at least one of the following:

[0065] Generate and display the corresponding clarification inquiry message according to the ambiguous information existing in the problem to be updated;

[0066] Determine the information to be supplemented according to the problem to be updated, and generate and display the corresponding clarification inquiry message according to the information to be supplemented;

[0067] Determine the redundant information to be compressed according to the problem to be updated, and generate and display the corresponding clarification inquiry message according to the redundant information to be compressed.

[0068] As an example but not a limitation, if the problem to be updated is "What equipment should I bring for the outdoor activities this weekend?", at this time, the ambiguous information in the problem to be updated is ("the outdoor activities this weekend"), so that the following clarification inquiry message can be generated and displayed according to the ambiguous information in this problem ("the outdoor activities this weekend"): "What specific outdoor activities do you mean by the outdoor activities this weekend, hiking, camping or other activities?"

[0069] As an example but not a limitation, if the problem to be updated is "I want to plan a trip.", at this time, the problem information to be updated is incomplete, so that the information to be supplemented can be determined according to this problem, including the travel destination, budget, time, etc., and then the following clarification inquiry messages can be generated and displayed according to the information to be supplemented: "Can you provide the travel destination or budget?"、"When do you plan to travel?" etc.

[0070] As an example and not by way of limitation, if the question to be updated is "I plan to travel to country A next month. What materials do I need to prepare for the visa application? Also, how early do I need to apply for the visa? Additionally, I heard that there are some places in country A that are worth visiting. Can you recommend some?", there is redundant information in the question to be updated at this time, and the following clarification inquiry information is generated and displayed: "Can the question be simplified to an inquiry about the visa application materials and time for country A, and recommendations for tourist attractions in country A?"

[0071] In the embodiments of the present disclosure, each time the clarification inquiry information is displayed, the corresponding clarification feedback information can be obtained.

[0072] As an example and not by way of limitation, for the example where the question to be updated is "I want to plan a trip.", the following clarification inquiry information can be displayed first: "Can you provide the destination or budget of the trip?" Assuming the obtained clarification feedback information is "The approximate budget is 10,000 yuan and I want to go to an island", the following clarification inquiry information can be continued to be displayed: "Okay. When do you plan to travel?" and continue to obtain the clarification feedback information.

[0073] Step 103, update the question according to at least one piece of clarification inquiry information and the clarification feedback information corresponding to at least one piece of clarification inquiry information.

[0074] In the embodiments of the present disclosure, in any round of question update process, based on at least one piece of clarification inquiry information and the clarification feedback information corresponding to at least one piece of clarification inquiry information in this round of question update process, the question to be updated in this round can be updated to obtain the updated question in this round.

[0075] As an example and not by way of limitation, for the example where the question to be updated is "I want to plan a trip.", assuming the following clarification inquiry information is displayed: "Can you provide the destination or budget of the trip?", the obtained clarification feedback information is "The approximate budget is 10,000 yuan and I want to go to an island", and the following clarification inquiry information is displayed: "Okay. When do you plan to travel?", and the obtained clarification feedback information is "Next Spring Festival", then the following updated question can be obtained: "The user plans a trip, with a budget of about 10,000 yuan, prefers an island, and the travel time is Spring Festival of xx year. Please provide the best travel plan."

[0076] Step 104, generate a response content based on the updated question.

[0077] In the embodiments of the present disclosure, if the updated question information in any round is sufficient, the expression is concise and clear, and there is no need to continue with the clarification inquiry, the question update can be stopped, and a response content can be generated and displayed based on the updated question in this round.

[0078] In an alternative embodiment, if the question after any round of update has no ambiguity, is complete in information, and has no redundant information, then the response content can be generated and displayed based on the updated question.

[0079] In the embodiments of the present disclosure, according to the input question, a first model is used to perform at least one round of question update on the question; during any round of question update, at least one clarification query message is displayed, and the corresponding clarification feedback message is obtained; the question is updated according to at least one clarification query message and the clarification feedback message corresponding to at least one clarification query message; the response content is generated based on the updated question. In the present disclosure, by using the first model to perform multiple clarification interactions on the input question, the updated question is made more accurate and complete, so that the response content generated based on the updated question is significantly improved in terms of relevance and integrity, solving the problem in the related art that the answer is off-topic or incomplete due to unclear questioning, and improving user satisfaction.

[0080] Figure 2 It is a schematic flowchart of a data processing method according to the second embodiment of the present disclosure.

[0081] Such as Figure 2 shown, the data processing method includes:

[0082] Step 201, according to the input question, use the first model to identify whether the question has an abnormality.

[0083] In an alternative embodiment, for the input question, the first model can be used to identify at least one of whether the information of the question is sufficient, whether the question lacks context background information, whether the question has ambiguous wording, and whether the question has redundant information.

[0084] As an example rather than a limitation, multiple dedicated small model Agents can be used, each improving the input quality from different aspects. For example, there can be a "clarification Agent" to improve the clarity of the question, a "context association Agent" to add relevant background information, a "conciseness Agent" to compress redundancy, an "accuracy Agent" to correct ambiguous wording, etc.

[0085] In an alternative embodiment, the first model includes multiple types of agents, so that according to the input problem, the first agent can determine the corresponding second agent for the problem and distribute the problem to the corresponding second agent; the second agent is used to identify whether there is an abnormality in the problem, and in the case where there is an abnormality in the problem, determine the type of abnormality to which the problem belongs. In this process, the first agent can quickly and accurately determine the second agent suitable for processing the problem based on the characteristics of the input problem, and the second agent focusing on a specific field can also quickly and accurately identify whether there is an abnormality in the problem and determine the type of abnormality to which the problem belongs. Therefore, through the cooperation mechanism of the first agent and the second agent, the accuracy and efficiency of problem processing are effectively improved.

[0086] In an alternative embodiment, the type of abnormality includes at least one of the following:

[0087] The information of the problem does not meet the corresponding information sufficiency standard;

[0088] The problem lacks context background information;

[0089] The problem has ambiguous wording;

[0090] The problem has redundant information.

[0091] Therefore, through different types of abnormalities, the accuracy, pertinence and efficiency of problem processing can be improved.

[0092] In an alternative embodiment, the second agent includes at least one of a clarity agent, a relevance agent, a precision agent, and a conciseness agent; the second agent is used to identify whether there is an abnormality in the problem, and in the case where there is an abnormality in the problem, determine that the type of abnormality to which the problem belongs includes at least one of the following:

[0093] Use the clarity agent to determine whether the information of the problem meets the corresponding information sufficiency standard, so as to determine that the type of abnormality to which the problem belongs is that the information of the problem does not meet the corresponding information sufficiency standard in the case where the information of the problem does not meet the corresponding information sufficiency standard;

[0094] Use the relevance agent to determine whether the problem lacks context background information, so as to determine that the type of abnormality to which the problem belongs is that the problem lacks context background information in the case where the problem lacks context background information;

[0095] Use the precision agent to determine whether the problem has ambiguous wording, so as to determine that the type of abnormality to which the problem belongs is that the problem has ambiguous wording in the case where the problem has ambiguous wording;

[0096] Use the conciseness agent to determine whether the problem has redundant information, so as to determine that the type of abnormality to which the problem belongs is that the problem has redundant information in the case where the problem has redundant information.

[0097] Thus, each second agent focuses on identifying and handling specific types of exceptions. This division of labor enables each agent to more deeply understand and respond to the type of exceptions it is responsible for, thereby improving the accuracy and speed of identification.

[0098] By way of example and not limitation, Figure 3 is a schematic diagram of optimizing prompt for model collaboration provided according to the third embodiment of the present disclosure. As Figure 3 shown, the prompt: Write a short mythological story based on xx is the input question. The prompt generator agent is the first agent, which is used to determine the corresponding second agent according to the input question and distribute the question to the corresponding second agent. The clarity agent, creativity agent, relevance agent, precision agent, and conciseness agent are the second agents, which are used to identify whether there are exceptions in the question and, in the case where the question has exceptions, determine the type of exception to which the question belongs.

[0099] It should be noted that for different questions, the corresponding second agents are different. Figure 3 In, since the question is to write a mythological story, the prompt generator agent distributes the question to the clarity agent, creativity agent, relevance agent, precision agent, and conciseness agent. If the question is to plan a trip, the corresponding second agents may include at least one of the clarity agent, relevance agent, precision agent, and conciseness agent, and will not include the creativity agent.

[0100] It should be noted that in the embodiments of the present disclosure, in addition to determining the type of exception to which the question belongs, the second agent is also used to display the corresponding clarification inquiry information and obtain the corresponding clarification feedback information.

[0101] Figure 3 In, the integration agent is used to integrate the clarification inquiry information displayed by the second agent and the clarification feedback information obtained, and the integrity agent is used to integrate the information integrated by the integration agent with the question distributed by the prompt generator to obtain an updated question.

[0102] Step 202, in the case where the question has exceptions, perform at least one round of question update according to the type of exception identified by the first model.

[0103] It should be noted that in the embodiments of the present disclosure, only when the input question has exceptions, at least one round of question update is performed. If the input question has no exceptions, that is, the input question information is sufficient and the expression is concise and clear, and no clarification inquiry is required, the response content is directly generated based on the input question.

[0104] Step 203, in any round of question update process, according to the exception type to which the question to be updated in this round belongs, display the corresponding clarification inquiry information and obtain the corresponding clarification feedback information.

[0105] In an optional embodiment, in the first round of question update process, based on the exception type identified by the first model (i.e., the exception type to which the question to be updated in the first round belongs, that is, the exception type to which the input question belongs), display the corresponding clarification inquiry information and obtain the corresponding clarification feedback information. Then, based on the clarification inquiry information and the clarification feedback information, perform question update to obtain the question updated in the first round.

[0106] Then, use the first model to identify whether there is an exception in the question updated in the first round. In the case where there is an exception in the question updated in the first round, based on the exception type identified by the first model (i.e., the exception type to which the question to be updated in the second round belongs, that is, the exception type to which the question updated in the first round belongs), perform the second round of question update.

[0107] In the second round of question update process, based on the exception type identified by the first model (i.e., the exception type to which the question to be updated in the first round belongs, that is, the exception type to which the input question belongs), display the corresponding clarification inquiry information and obtain the corresponding clarification feedback information. Then, based on the clarification inquiry information and the clarification feedback information, perform question update to obtain the question updated in the second round.

[0108] Then, use the first model to identify whether there is an exception in the question updated in the second round, and so on, until in a certain round of question update process, use the first model to identify that there is no exception in the question updated in that round, and stop the question update.

[0109] It can be understood that if there is no exception in the question updated in a certain round, it means that the question information is sufficient and the expression is concise and clear, and there is no need to conduct a clarification inquiry, and the reply content can be directly generated.

[0110] Thus, by displaying the clarification inquiry information matching the exception type, it is possible to guide the user to provide the required information more accurately, reduce the repeated inquiries and clarifications caused by insufficient information or unclear expression, and thus improve the efficiency of question processing.

[0111] In an optional embodiment, in any round of question update process, according to the exception type to which the question to be updated in this round belongs, display the corresponding clarification inquiry information, including at least one of the following:

[0112] In any round of question update process, in the case where the exception type to which the question to be updated in this round belongs is that the information of the question does not meet the corresponding information sufficiency standard, display the clarification inquiry information about the information to be supplemented;

[0113] When the exception type to which the problem to be updated in this round belongs is that the problem lacks context background information, display the clarification inquiry information about the context background information to be added;

[0114] When the exception type to which the problem to be updated in this round belongs is that the problem has ambiguous wording, display the clarification inquiry information about the ambiguous wording to be corrected;

[0115] When the exception type to which the problem to be updated in this round belongs is that the problem has redundant information, display the clarification inquiry information about the redundant information to be compressed.

[0116] Thus, each clarification inquiry information is generated and displayed based on a specific exception type, enabling the questioner to clearly know which information needs to be provided or confirmed, and reducing misunderstandings caused by insufficient information or unclear expressions.

[0117] It should be noted that different problems have different information sufficiency criteria. For example, when the user inputs "I want to buy a mobile phone. Do you have any suggestions?", the corresponding information sufficiency criteria include the purchase budget, purchase requirements, and brand preference. Another example is when the user inputs "I want to plan a trip.", the corresponding information sufficiency criteria include the travel destination, budget, and travel time.

[0118] As an example but not a limitation, if the problem is "Please help me write a project report.", at this time the information of the problem does not meet the corresponding information sufficiency criteria (the corresponding information sufficiency criteria for this problem include the project name, project content, etc.), so the clarification inquiry information about the information to be supplemented can be "Can you provide the name and main content of this project?";

[0119] If the problem is "I encountered a technical problem. How can I solve it?", at this time the problem has ambiguous wording (specifically what the technical problem refers to), and lacks specific context background information, such as the environment or scenario where the technical problem appears, the solutions that have been tried, etc. Therefore, the clarification inquiry information about the ambiguous wording to be corrected can be "What is the technical problem you encountered?", and the clarification inquiry information about the context background information to be added can be "In what environment or scenario did this technical problem appear?" and "Have you tried any solutions? What specific solutions have you tried?";

[0120] If the question is "I plan to travel to Country A next month. May I ask what materials are needed to apply for a visa? Also, how early do I need to apply for the visa? Additionally, I heard that there are some places in Country A that are worth visiting. Can you recommend some?", there is redundant information in the question. Therefore, the clarification query information regarding the redundant information to be compressed can be "Can the question be simplified to inquiries about the visa application materials and time for Country A, as well as recommendations for tourist attractions in Country A?".

[0121] Step 204, update the question according to at least one piece of clarification query information and the corresponding clarification feedback information for at least one piece of clarification query information.

[0122] Step 205, generate a response content based on the updated question.

[0123] It should be noted that the explanations of Steps 204 to 205 can be found in the relevant descriptions in any embodiment of the present disclosure and will not be elaborated here.

[0124] In the embodiments of the present disclosure, according to the input question, a first model is used to identify whether there is an abnormality in the question; in the case where there is an abnormality in the question, according to the type of abnormality identified by the first model, at least one round of question update is performed. In the present disclosure, targeted question update according to the type of abnormality helps to more accurately understand the user's intention, thereby generating a response content that better meets the user's needs.

[0125] Figure 4 It is a schematic flowchart of a data processing method according to the fourth embodiment of the present disclosure.

[0126] As Figure 4 shown, the data processing method includes:

[0127] Step 401, according to the input question, use a first model to perform at least one round of question update on the question.

[0128] Step 402, during any round of question update, display at least one piece of clarification query information and obtain the corresponding clarification feedback information.

[0129] Step 403, semantically fuse at least one piece of clarification query information, the corresponding clarification feedback information for at least one piece of clarification query information, and the question to be updated in this round to obtain an updated question.

[0130] By way of example and not limitation, assume that the question to be updated this round is an example of "I want to plan a trip.", and the following clarification inquiry information is shown: "Can you provide the destination or budget of the trip?", and the obtained clarification feedback information is "The approximate budget is 10,000 yuan and I want to go to an island.", and the following clarification inquiry information is shown: "Okay. When do you plan to travel?", and the obtained clarification feedback information is "Next Spring Festival", then the semantic integration of "Can you provide the destination or budget of the trip?", "The approximate budget is 10,000 yuan and I want to go to an island.", "Okay. When do you plan to travel?", "Next Spring Festival" and "I want to plan a trip." can be carried out to obtain the following updated question: "The user plans a trip, with a budget of about 10,000 yuan, prefers an island, and the travel time is the Spring Festival of xx year. Please provide the best travel plan."

[0131] Step 404, when there is no abnormality in the updated question, select one of the first model and the second model, and generate a reply content based on the updated question.

[0132] Among them, the second model includes a large model, and the first model includes a small model with a model layer depth less than that of the large model.

[0133] In an optional embodiment, the first model can be used to generate a candidate reply content based on the updated question; an evaluation model is used to evaluate the candidate reply content to obtain the confidence level of the candidate reply content; when the confidence level is greater than the set threshold, the candidate reply content is determined as the reply content; when the confidence level is not greater than the set threshold, the second model is used to generate a reply content based on the updated question. Thus, through the cooperation of the first model and the second model, part of the questions are generated with reply contents by the first model, thereby reducing the call frequency and processing load of the second model and avoiding waste of computing resources.

[0134] Step 405, obtain the satisfaction feedback corresponding to the reply content.

[0135] In the embodiments of the present disclosure, the user can give feedback on whether they are satisfied with the reply content. If the user is satisfied, this process ends; if the user is not satisfied or has new supplementary information, the first model can be used to trigger a new round of optimization based on the user's feedback. For example, when the user points out that there are omissions in the reply content, the first model can be used to ask the user or the context for the omissions, and then update the question for re-answering. This mechanism allows for re-optimization after answering, and can form a closed loop to improve the response integrity and accuracy.

[0136] It should be noted that in most cases, the updated question obtained after at least one round of question update is already high-quality input, and the reply content generated based on the updated question is also relatively perfect, and it is rarely necessary to trigger a new round of optimization.

[0137] It should be noted that the explanations of steps 401 to 402 can be found in the relevant descriptions of any embodiment of the present disclosure, and will not be elaborated here.

[0138] In the embodiments of the present disclosure, at least one clarification inquiry message, the clarification feedback message corresponding to at least one clarification inquiry message, and the question to be updated in this round are semantically fused to obtain an updated question; when there is no abnormality in the updated question, one of the first model and the second model is used to generate a response content based on the updated question; and the satisfaction feedback corresponding to the response content is obtained. In the present disclosure, through semantic fusion, the updated question can be more comprehensive and clear than the question to be updated, which helps the model to more accurately understand the question and generate a response content. By using the first model or the second model to generate a response content based on the updated question, the collaborative processing of the first model and the second model can be realized, and the waste of computing resources can be avoided. By obtaining the satisfaction feedback of the response content, when the user is not satisfied with the response content or has new supplementary information, the first model can be used to perform a new round of optimization based on the user's feedback, improving the response integrity and accuracy.

[0139] The data processing method of the embodiments of the present disclosure will be exemplified below in a large model input optimization system based on small model preprocessing and multi-round feedback. In this system, the small model and the large model are cooperated to jointly implement the methods mentioned in the foregoing embodiments.

[0140] In an alternative embodiment, the structure of the system includes:

[0141] A small model, which is used to receive the question input by the user, perform at least one round of question update on the input question, and input the updated question into the large model;

[0142] A large model, connected in series with the small model, and generates a response content based on the updated question output by the small model.

[0143] It should be noted that the model parameters of the small model are much smaller than those of the large model.

[0144] In an alternative embodiment, the small model identifies whether there is an abnormality in the input question, and when there is an abnormality in the input question, at least one round of question update is performed according to the abnormal type identified by the small model.

[0145] In an alternative embodiment, during any round of question update, according to the abnormal type to which the question to be updated in this round belongs, at least one clarification inquiry message is displayed to the user, and the clarification feedback message corresponding to at least one clarification inquiry message input by the user is obtained, so as to update the question to be updated in this round according to at least one clarification inquiry message and the clarification feedback message corresponding to at least one clarification inquiry message, and obtain the question updated in this round.

[0146] For each question to be updated in each round, the small model shows at least one clarification query message to the user according to the exception type to which the question to be updated belongs, and obtains the clarification feedback message corresponding to at least one clarification query message input by the user. In this process, the small model can have multiple rounds of conversations with the user, and these multiple rounds of conversations can provide more context information, guiding the small model to gradually optimize the question input by the user, so that the response content generated by the large model is more in line with the user's expectations. Thus, in this system, the answer of the large model is no longer final, but continuously improves the question finally input to the large model under the drive of user feedback, ensuring that the response accurately meets the requirements.

[0147] In an optional embodiment, the small model continues to identify whether each updated question is abnormal, and in the case that any updated question is abnormal, continues to update the question.

[0148] As an example but not a limitation, the user inputs the question: "Please help me write a project report." The small model identifies that this question is abnormal (this question lacks necessary information: project name, project content), and shows the clarification query message: "Can you provide the name and main content of this project?" Obtain the clarification feedback message corresponding to this clarification query message input by the user: "The development of the system optimized by the small model is adopted, and the main content includes the background and technical solutions." After receiving this clarification feedback message, the small model updates the question input by the user to obtain the updated question, such as: "Write a project report on 'The development of the system optimized by the small model', which needs to include the background and technical solutions." Then the small model identifies whether the updated question is abnormal. If there is no abnormality, the updated question is input to the large model, and the large model generates the response content. Thus, by updating the question through the small model, the question finally input to the large model is more specific and complete than the question input by the user, which not only reduces the processing burden of the large model, but also reduces the invalid answers caused by poor input quality, improves the accuracy and integrity of the generation result of the large model, and improves the overall efficiency of the system.

[0149] In an optional embodiment, the system can save the interaction data (including the input question, at least one clarification query message, the clarification feedback message corresponding to at least one clarification query message, the updated question, the response content, the satisfaction feedback corresponding to the response content, etc.). Compared with the artificially constructed data set, the real interaction data can better reflect the real-time needs of users and the defects of the model, and is more conducive to the improvement of the model.

[0150] In an alternative embodiment, the system may periodically use the saved interaction data to retrain or fine-tune the small model and / or the large model, enhance the model capabilities, improve the user experience, thereby attracting more users to use the system and generating new interaction data, further promoting model improvement, and so on in a cycle, forming self-reinforcement of the system.

[0151] By way of example and not limitation, after the small model performs at least one round of question update on a question input by a user, the large model generates a response content, and obtains a satisfaction feedback corresponding to the response content input by the user. The system may save the interaction data in this process (including the input question, at least one clarification query information, the clarification feedback information corresponding to at least one clarification query information, the updated question, the response content, the satisfaction feedback corresponding to the response content, etc.). When offline, by extracting a large amount of similar interaction data, the clarification query information generation strategy of the small model is retrained, so that when the small model encounters such a vague question next time, it can generate better clarification query information. In this way, when another user asks a similar unclear question, the small model can generate better clarification query information, clarify the user's needs in fewer rounds, and thus generate the response content more quickly. In this process, each interaction between the user and the system is no longer an isolated transaction, but a contribution to the evolution of the system, quietly enhancing the overall capabilities of the system. Compared with traditional models with fixed performance after static deployment, the system can dynamically adapt to new requirements and new scenarios, and continuously improve the response quality.

[0152] In an alternative embodiment, the small model or the large model may be used to generate the response content based on the updated question. Optionally, the small model may be used to generate candidate response content based on the updated question; an evaluation model is used to evaluate the candidate response content to obtain the confidence level of the candidate response content; when the confidence level is greater than the set threshold, the candidate response content is determined as the response content; when the confidence level is not greater than the set threshold, the large model is used to generate the response content based on the updated question. Thus, when the small model can generate a response content with a confidence level greater than the set threshold, the large model may not be called, reducing the call frequency and processing load of the large model and avoiding waste of computing resources.

[0153] In an alternative embodiment, if the satisfaction feedback corresponding to the response content generated by the large model is satisfaction, the small model may be trained based on the updated question and the response content generated by the large model, so that the small model learns the response thinking of the large model, improves the performance of the small model, and enables the small model to achieve a task performance similar to that of the large model while maintaining a small model size.

[0154] Thus, on the one hand, the small model updates the question at least once, making the question input to the large model of higher quality. As a result, the response content generated by the large model better meets the user's expectations, enhancing the output quality of the large model. On the other hand, some high-quality response examples generated by the large model during operation are added to the training set of the small model, enabling the small model to learn the response thinking of the large model and improving the output ability of the small model. This co-evolution can continuously improve the system performance.

[0155] As an example but not a limitation, in this system, the question input by the user first enters the "small model layer" and undergoes at least one round of question updates through a series of small model units. During each round of question update, the small model has at least one clarification interaction with the user. If the updated question is considered to be fully understood and the confidence level of the response content generated by the small model is greater than the set threshold, there is no need to call the large model, and the response content generated by the small model is directly presented to the user; if the updated question is considered to be fully understood but the confidence level of the response content generated by the small model is not greater than the set threshold, it enters the "large model layer" and the large model is called to generate the response content. In addition, some high-quality response examples generated by the large model during operation are used to train and update the small model. Thus, both the response quality and the computational efficiency are ensured, and a more cost-effective processing than a single large model can be achieved.

[0156] In an optional embodiment, the user interface (UI) of this system can adopt a conversational interaction design, guiding the user to provide information step by step in the form of a chat. Thus, the user does not need to understand complex technical details and can improve the question they input step by step according to the interface prompts. For example, when it is detected that the question input by the user is unclear, the interface will automatically pop up clarification query information (such as clarification questions or options) for the user to supplement (like an intelligent questionnaire). In the parameter template scenario, the UI asks for key parameters one by one in the form of a form or a dialogue, allowing the user to fill in the blanks or make selections. The whole process visually shows which information has been obtained and which information is still missing, thus reducing the user's thinking burden. The user can feel that the question they input is gradually refined and improved during the interaction process. This guided interface hides professional prompting engineering in the interaction process, enabling any user to ask high-quality questions.

[0157] In an alternative embodiment, a visual feedback module for input quality assessment can also be integrated into the system UI. After the user enters the corresponding clarification feedback information for each clarification query, the small model scores the clarity, completeness, etc. of the currently known information and provides the user with the improvement status of the question in the form of simple icons or colors. For example, when the input is obviously ambiguous, a prompt of "Needs clarification" can be lit up beside the interface to inform the user that the current question needs to be improved. When the input information is insufficient, the UI can display the "Question information completeness" in the form of a progress bar to inform the user of the improvement status of the question. These intuitive feedbacks can help the user understand the thinking process of the system and guide the user to participate in the optimization cycle. In this process, the user is no longer passively providing questions, but collaborating with the system to jointly improve the problem statement.

[0158] In an alternative embodiment, the system can also provide flexible input methods and template reuse. For example, for structured information requirements, the UI can automatically switch to the form mode (such as filling in item by item in the above parameter list). For questions in common domains, the system can provide preset question templates. After the user selects the corresponding scenario, the interface will display a template example, and the user only needs to fill in their specific content. For example, for inquiries about the weather, reservation services, etc., corresponding standard templates can be set to reduce the user's thinking burden. At the same time, the interface allows the user to return and modify the information in the previous steps at any time, and the system will update the prompt suggestions accordingly, making the whole process smooth and flexible.

[0159] By way of example and not limitation, assume that the user wants to obtain the installation guide for a certain software, and the initially entered question is "How to install Software X?". The system UI detects the lack of environmental information, displays the clarification query information: "Please select your operating system", provides Windows / Linux options, and at the same time displays the progress bar of "Question information completeness". After the user selects Windows, the progress bar of "Question information completeness" increases. Then the UI displays the clarification query information again: "Which version do you want to install? (such as the latest version or a specific version number)". The user enters the version number. At this time, the system has collected elements such as the software name, operating system, and version, and the information has been supplemented completely. At this time, the progress bar of "Question information completeness" is full, or the progress bar of "Question information completeness" is replaced with "Information completeness: √". Subsequently, the small model generates an updated question and submits it to the large model, and the large model gives the detailed installation steps for Software X under Windows. Thus, the UI gradually guides the user to complete the question details, avoiding the situation where the answer from the large model is general or inapplicable due to a one-time ambiguous question. Moreover, the whole interaction is very intuitive for the user, like filling out a questionnaire, enhancing the user's participation and trust while improving the input quality.

[0160] In an optional embodiment, the process from when the system receives a question input by the user to when it generates a response content may include the following implementation steps:

[0161] 1. Obtain the question input by the user and identify whether the question is abnormal

[0162] The user inputs a question on the interface. The small model obtains the question input by the user and identifies whether there is any abnormality in the question input by the user.

[0163] 2. In the case where the question input by the user is abnormal, based on the type of abnormality identified by the small model, perform at least one round of question update

[0164] It should be noted that if the small model believes that the question input by the user is not abnormal, it can directly generate candidate response content based on the question input by the user, and determine whether to input the question input by the user into the large model according to the confidence level of the candidate response content, so that the large model generates the response content.

[0165] 3. During each round of question update, the small model generates and displays at least one clarification query message to the user, and obtains at least one piece of clarification feedback information corresponding to the clarification query message

[0166] During each round of question update, the small model generates and displays at least one clarification query message to the user. For example, when it detects that the time or object is not clear, it generates and displays the following clarification query message to the user: "Which one do you specifically refer to...?". In this step, the small model can use pre-trained templates and context to automatically generate clarification query messages, aiming to obtain additional information input from the user.

[0167] The user views the clarification query message and answers the clarification query message in the form of text or multiple-choice questions. The system receives the clarification feedback information input by the user and merges it into the original question to form an updated question. Subsequently, the small model identifies whether there is any abnormality in the updated question. If so, it enters the next round of question update. This process repeats, forming a multi-round interaction: the small model continuously enriches the question details, and the user cooperates to provide feedback until the updated question is not abnormal.

[0168] Optionally, during the first round of question update, based on the type of abnormality identified by the small model (i.e., the type of abnormality to which the question to be updated in the first round belongs, that is, the type of abnormality to which the input question belongs), display the corresponding clarification query message and obtain the corresponding clarification feedback information. Then, based on the clarification query message and the clarification feedback information, perform question update to obtain the question updated in the first round.

[0169] Then, use a small model to identify whether there are any anomalies in the questions updated in the first round. If there are anomalies in the questions updated in the first round, based on the anomaly type identified by the small model (i.e., the anomaly type to which the questions to be updated in the second round belong, and also the anomaly type to which the questions updated in the first round belong), perform the second-round question update.

[0170] During the second-round question update process, based on the anomaly type identified by the small model (i.e., the anomaly type to which the questions to be updated in the first round belong, and also the anomaly type to which the input questions belong), display the corresponding clarification inquiry information and obtain the corresponding clarification feedback information. Then, based on the clarification inquiry information and the clarification feedback information, perform question update to obtain the questions updated in the second round.

[0171] Then, use a small model to identify whether there are any anomalies in the questions updated in the second round, and so on, until in a certain round of question update process, it is identified by the small model that there are no anomalies in the questions updated in that round, and then stop the question update.

[0172] 4. The small model generates candidate response content based on the updated questions, and based on the confidence level of the candidate response content, determines whether to input the updated questions into the large model.

[0173] If the confidence level of the candidate response content is greater than the set threshold, the system can directly jump to step 5 and display the candidate response content generated by the small model to the user without calling the large model, saving resources.

[0174] 5. Input the updated questions into the large model, and the large model generates response content.

[0175] Because the previous steps have ensured that there are no anomalies in the updated questions, the large model can now focus on leveraging its powerful generation and reasoning capabilities to generate response content with higher quality and more pertinence based on the updated questions.

[0176] 6. Display the response content and obtain the satisfaction feedback corresponding to the response content.

[0177] The system displays the generated response content to the user. At the same time, the system requests the user to give a simple evaluation of the response content (such as a star rating or a binary choice of whether it is helpful) to obtain the satisfaction feedback corresponding to the response content. If the user is satisfied, this process ends; if the user is not satisfied or has new supplementary information, the small model can be used to trigger a new round of optimization based on the user's feedback. For example, when the user points out that there are omissions in the response content, the small model can be used to ask the user or the context about the omissions, and then update the questions for re-answering. This mechanism allows for re-optimization after answering and can form a closed loop to improve the response integrity and accuracy.

[0178] It should be noted that in most cases, the updated question obtained after at least one round of question updates is already high-quality input, and the response content generated based on the updated question is also relatively complete, and it rarely triggers a new round of optimization.

[0179] 7. Interactive Data Saving

[0180] Save each piece of interactive data (including the input question, at least one clarification query message, the clarification feedback message corresponding to at least one clarification query message, the updated question, the response content, the satisfaction feedback corresponding to the response content, etc.) for training the small model and / or the large model.

[0181] As an example rather than a limitation, the user inputs the question: "I want to buy a new mobile phone. Do you have any suggestions?" The small model identifies that there are abnormalities in this question (the question lacks necessary information: purchase budget, purchase requirements, brand preference), so it generates and displays a clarification query message: "What is your budget range for purchasing a mobile phone? Do you value performance or photography more?" Obtain the clarification feedback message corresponding to the clarification query message input by the user: "The budget is approximately within 3000 yuan, and I hope the photography effect is good." After receiving this clarification feedback message, the small model continues to generate and display a clarification query message: "Do you have a brand preference? (Such as brand A, brand B, etc.)" Obtain the clarification feedback message corresponding to the clarification query message input by the user: "I hope it is brand A." After receiving this clarification feedback message, the small model determines that all the information to be supplemented has been supplemented, and then updates the user input question based on the clarification query message and the clarification feedback message to obtain an updated question, such as: "Recommend a new mobile phone for users with a budget of less than 3000 yuan, preferring brand A, and mainly focusing on the photography effect." Then the small model identifies whether there are abnormalities in the updated question. If it is determined that there are no abnormalities in the updated question, it generates a candidate response content based on the updated question. If the confidence level of the candidate response content is greater than the set threshold, it displays the candidate response content generated by the small model to the user. If the confidence level of the candidate response content is not greater than the set threshold, it calls the large model, and the large model generates a response content based on the updated question: "It is recommended that you consider the new mobile phone of XX brand. The price of this model is about 2800 yuan, and the photography hardware is excellent... (detailed reasons are introduced).", and the user is satisfied with the response content and gives a positive evaluation. The process ends.

[0182] In addition, the interactive design of the user interface is also crucial for the smooth operation of the system. The following is an interface interaction example to illustrate how users can experience the input optimization of this system on the front end.

[0183] Suppose the user enters a vague question in the dialog box. The UI example process is as follows:

[0184] 1. User enters a question

[0185] User input question: "I want to plan a trip." (The interface displays the user's question).

[0186] 2. The system identifies whether there is an abnormality in the question:

[0187] The small model identifies that there is an abnormality in this question (this question lacks necessary information: travel destination, budget, travel time). The UI immediately displays a clarification question: "Can you provide the travel destination or budget?" [A hint bubble appears, possibly accompanied by a question mark emoji or highlighting] (The user sees the system's question, and an input box or options are provided below)

[0188] 3. User answer

[0189] The user inputs the clarification feedback information corresponding to the above clarification question: "The approximate budget is 10,000 yuan, and I want to go to an island." (The interface updates the conversation and displays that new information has been obtained, such as listing "Budget = 10,000 yuan" and "Preference = island" on the side)

[0190] 4. The system continues to display the clarification question

[0191] The small model finds that the travel time has still not been specified, and the UI displays the clarification question again: "Okay. When do you plan to travel?"

[0192] 5. User answer

[0193] The user inputs the clarification feedback information corresponding to the above clarification question: "Next Spring Festival." (The interface is updated, and the side information adds "Time = Spring Festival")

[0194] 6. The system updates the question

[0195] The small model determines that all the information to be supplemented has been supplemented, and then updates the user input question based on the clarification question and the clarification feedback information, obtaining the updated question: "The user plans a trip with a budget of about 10,000 yuan, a preference for an island, and the travel time is the Spring Festival of xx year. Please provide the best travel plan."

[0196] The small model identifies whether there is an abnormality in the updated question. If it is determined that there is no abnormality in the information of the updated question, candidate reply content is generated based on the updated question. If the confidence level of the candidate reply content is greater than the set threshold, the candidate reply content generated by the small model is displayed to the user. If the confidence level of the candidate reply content is not greater than the set threshold, the large model is called, and the large model generates the reply content based on the updated question. At this time, the UI can display a prompt based on the updated question: "Your needs have been understood: budget about 10,000 yuan, preference for an island, travel time Spring Festival. Now querying the best travel plan for you..." (At this time, the UI may display a loading animation indicating that the large model is generating an answer)

[0197] 7. Display the reply content

[0198] Show the reply content: "It is recommended that you go to Place B. The climate in Place B is warm during the Spring Festival and is very suitable for a seaside holiday... (The following omits the detailed itinerary and reasons)".

[0199] 8. Obtain the satisfactory feedback corresponding to the reply content

[0200] After the user reads the reply content, click the "Satisfied" button. The interface prompts "Thank you for your feedback!" and the conversation ends.

[0201] In an optional embodiment, to quantify the advantages of the system, the following evaluation plan can be formulated, including offline metric evaluation and online user testing:

[0202] Input quality evaluation: Establish a set of criteria (IQS, Input Quality Score) for evaluating the quality of input questions, and score them from dimensions such as clarity, integrity, and context relevance. For example, for a set of representative user query cases, let the traditional single-round system and this system respectively generate questions submitted to the large model, and submit them to experts or an independent small model to review their quality. Among them, the review can be carried out by counting whether ambiguous words are eliminated, whether necessary parameters are complete, etc. The expected result is that for complex query scenarios, the average IQS of the questions generated by this system and submitted to the large model is more than 30% higher than the average IQS of the questions generated by the traditional single-round system and submitted to the large model (i.e., the original input questions), which proves that the input quality has been significantly improved.

[0203] Answer accuracy and integrity evaluation: For several standard Q&A tasks (which can be selected from public QA datasets or actual business problems), compare the differences in accuracy and integrity between the answers of this system and the baseline large model's direct answers. Optionally, the two methods can be allowed to generate answers respectively, and then the accuracy, recall rate, etc. can be calculated by manual annotation or by referring to the standard answers. For example, in a comprehensive query containing multiple sub-questions, check whether the answer covers all sub-question points. It is expected that this system can improve both in terms of correctness and answer coverage rate. For example, in a question containing 5 key points, the baseline large model's direct answer on average only answers and covers 3 key points correctly, while this system can reach more than 4. In addition, the proportion of irrelevant content or incorrect "hallucination" information in the answer can also be counted to verify that this system can reduce such errors to the minimum.

[0204] Inference Efficiency Evaluation: Record the computing resources and time consumed by this system and the system using only the large model to process a certain number of queries under the same conditions. Among them, the relevant indicators include: the number of times the large model is called per query on average, the number of tokens processed by the large model per query on average, the average response latency, etc. Ideally, the average number of calls to the large model in this system is close to 1 (most simple Q&A do not call or only call the large model once), while the large model of the system using only the large model often requires multiple attempts. In addition, the number of tokens processed per query should also be reduced due to input optimization. The expected result is that, on the premise of ensuring quality, this system can save at least 30%-50% of the computing power consumption of the large model and reduce the response latency perceived by users by more than 20%. For example, if the system using only the large model consumes 1000 large model tokens per request on average, this system may only need 600-700 to complete the same or even better answers.

[0205] User Satisfaction Test: Conduct an A / B test or user research, alternately assign this system and the traditional system to real users for use, and collect satisfaction scores and subjective feedback after a period of time. Compare the satisfaction differences of the two groups of users in terms of answer relevance, time taken to obtain answers, interaction experience, etc. The indicators can be the average satisfaction score (such as on a ten-point scale) and the proportion of "willing to use / recommend again", etc. The expected result is that the user satisfaction of using this system is significantly higher. For example, in a controlled experiment with 50 people, 90% of the users said that the questioning process of this system is more considerate and the answers are more in line with expectations, while this proportion of the traditional system may be less than 70%. In addition, user comments can also be collected to analyze in which aspects this system wins user recognition (such as "the question is clarified before answering, feeling more accurate", etc.).

[0206] Robustness and Generalization Evaluation: To test the performance of the system on different domains and different types of questions, a cross-domain test set can be used to evaluate this system. Evaluate whether the system is still robust in processing novel questions not covered by the training data. For example, for some questions with unconventional wording, observe whether the small model can still propose meaningful clarifications; for questions with dense domain terms, whether the system can still obtain the correct meaning through multiple rounds of interaction. This part of the evaluation can help discover the weaknesses of the system and continuously improve it. At the same time, it can also demonstrate the strong generalization ability of this system, proving that it is not only effective for familiar questions, but also can gradually figure out unfamiliar questions through interaction, which is better than models without interaction.

[0207] Comparative experiment: The system can also be compared with the effects of other existing optimization strategies (such as optimizing only with the prompt template or relying solely on the self-reflection of the large model). This will further highlight the unique advantages of the multi-round small-large model collaboration solution. For example, the comparison shows that on the same test set, the accuracy of the system's answers is several percentage points higher than that of a single-round optimization algorithm, while the inference cost is much lower. Optionally, the comparison data can be presented in the form of tables and charts, thereby scientifically quantifying the advanced nature and innovation of the system.

[0208] Through the above multi-level evaluations, the improvement extent of the system in terms of input quality, response quality, efficiency, and user experience can be comprehensively verified.

[0209] The present disclosure also provides a model training method. Figure 5 It is a schematic flowchart of the model training method provided according to the fifth embodiment of the present disclosure.

[0210] As Figure 5 shown, the model training method includes:

[0211] Step 501, obtaining multiple groups of training samples, where the training samples include the input question, at least one clarification query information for any round of question update and the corresponding clarification feedback information, the updated question, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content.

[0212] In the embodiments of the present disclosure, each group of training samples is real interaction data, rather than an artificially constructed data set.

[0213] In an optional embodiment, the interaction data (including the input question, at least one clarification query information, the clarification feedback information corresponding to at least one clarification query information, the updated question, the response content, the satisfaction feedback corresponding to the response content, etc.) can be structurally stored and regularly added to the training data pool.

[0214] It should be noted that during the storage process, sensitive information needs to be desensitized to ensure user privacy security, while retaining the necessary fields for analysis and training.

[0215] In an optional embodiment, the quality of the interaction data may vary, and it can be screened and cleaned by a small model and / or simple rules. For example, filtering out user-entered random characters or obviously invalid sessions. At the same time, manual annotation can be introduced to evaluate whether some model answers are correct and whether the clarification questions are appropriate. These annotation information will be used as training supervision signals to improve the quality of the training data.

[0216] Step 502, training a first model based on multiple groups of training samples.

[0217] In an alternative embodiment, for any set of training samples, when the response content in the training samples is generated by the first model, the first model can be trained based on the questions input in the training samples, at least one clarification query information for any round of question update and the corresponding clarification feedback information, the updated questions, the response content generated based on the updated questions, and the satisfaction feedback corresponding to the response content. Thus, by training the first model with the input questions, at least one clarification query information for any round of question update and the corresponding clarification feedback information, and the satisfaction feedback corresponding to the response content, the ability of the first model to generate clarification query information can be improved, and the adaptability of the first model in dealing with complex and variable user requirements can be enhanced. By training the first model with the updated questions, the response content, and the satisfaction feedback corresponding to the response content, the accuracy of the first model in generating the response content can be effectively improved.

[0218] In an alternative embodiment, for any set of training samples, when the response content in the training samples is generated by the second model, the second model can be trained based on the updated questions in the training samples, the response content generated based on the updated questions, and the satisfaction feedback corresponding to the response content, and the first model can be trained based on the questions input in the training samples, at least one clarification query information for any round of question update and the corresponding clarification feedback information, and the satisfaction feedback corresponding to the response content. Thus, by training the second model with the updated questions, the response content, and the satisfaction feedback corresponding to the response content, the accuracy of the second model in generating the response content can be effectively improved. By training the first model with the input questions, at least one clarification query information for any round of question update and the corresponding clarification feedback information, and the satisfaction feedback corresponding to the response content, the ability of the first model to generate clarification query information can be improved, and the adaptability of the first model in dealing with complex and variable user requirements can be enhanced.

[0219] By way of example and not limitation, a continuous training process can be established for the first model. For example, newly collected interaction data can be incrementally added to the original training set of the first model, and the first model can be fine-tuned and updated at regular intervals (such as weekly or when a certain amount of data is reached). The training objectives can include: enabling the first model to better predict when clarification is needed, generating more appropriate clarification query information, and correctly parsing user answers, etc. Since the training samples include the question-answer pairs actually feedback by users, the first model can learn patterns that are closer to real human query habits. For example, if it is statistically found that many users omit reporting the budget when asking "recommend a mobile phone", the first model will learn the importance of always asking about the budget first, and thus be more proactive in asking relevant clarification query information next time. This kind of training that feeds the model with real data can continuously improve the question guiding ability of the first model.

[0220] In an alternative embodiment, for any set of training samples, when the response content in the training samples is generated by the second model and the satisfaction feedback in the training samples is satisfied, the first model can be trained based on the updated questions in the training samples and the response content generated based on the updated questions. Thereby, the first model can learn the response thinking of the second model, thus significantly improving the response ability of the first model.

[0221] In an alternative embodiment, the first model and / or the second model can be trained in an automated manner. For example, scripts can be constructed to regularly extract training samples from the log database, automatically complete preprocessing and training set updates, use machine learning operations (MLOps) technology pipelines to automatically train and deploy model versions, and set up automatic generation of evaluation reports to alert developers to pay attention. Thereby, the maintenance cost can be reduced through a high degree of automation, enabling the model to improve itself with less manual intervention (developers mainly observe indicators and adjust a small number of optimization strategies rather than manually collate data to train the model).

[0222] It should be noted that the above training process needs to be carried out offline to avoid affecting the stability of the online service. In addition, the updated model needs to undergo strict testing before being deployed online to ensure that the improvement is indeed effective.

[0223] In an alternative embodiment, after each round of training, the performance of the model can be evaluated using a validation set (including historical data and a portion of new data). For example, evaluate the accuracy rate, recall rate (the proportion of successful inquiries when there is ambiguity), etc. of the first model in generating clarification inquiry questions, and evaluate whether the user satisfaction has increased compared to before, and so on. Optionally, the system performance can be tracked by establishing input quality scoring metrics and answer satisfaction metrics. Among them, the input quality scoring metric is used to comprehensively measure the scores of the questions finally submitted to the large model in terms of clarity and integrity (which can be scored by humans or the large model), and the answer satisfaction metric is used to determine the quality of the model output in combination with user evaluations. For example, by comparing these metrics before and after the update, it can be determined whether the new model version is better than the old version. If a certain training reduces the performance, roll back the model or adjust the training hyperparameters. Thus, a continuous iterative optimization engineering cycle is formed: data → training → evaluation → online → collect data again, ensuring that the model evolves in the correct direction.

[0224] In the embodiments of the present disclosure, multiple sets of training samples are obtained, where each training sample includes an input question, at least one clarification inquiry message for any round of question update and the corresponding clarification feedback message, the updated question, a response content generated based on the updated question, and a satisfaction feedback corresponding to the response content. Based on the multiple sets of training samples, a first model is trained. In the present disclosure, each set of training samples is real interaction data. By training the first model based on the multiple sets of training samples, the first model can learn the user's real query intention and expected result, thereby significantly improving the model performance.

[0225] Corresponding to the data processing method provided in the above Figures 1 to 4 embodiment, the present disclosure further provides a data processing device. Since the data processing device provided in the embodiments of the present disclosure corresponds to the data processing method provided in the above Figures 1 to 4 embodiment, the implementation manner of the data processing method is also applicable to the data processing device provided in the embodiments of the present disclosure, and will not be described in detail in the embodiments of the present disclosure.

[0226] Figure 6 It is a schematic structural diagram of a data processing device provided in the sixth embodiment of the present disclosure.

[0227] As Figure 6 shown, the data processing device includes:

[0228] A first processing module 601, configured to perform at least one round of question update on the question by using a first model according to the input question;

[0229] A second processing module 602, configured to display at least one clarification inquiry message and obtain the corresponding clarification feedback message during any round of question update;

[0230] An update module 603, configured to update the question according to at least one clarification inquiry message and the clarification feedback message corresponding to at least one clarification inquiry message;

[0231] A generation module 604, configured to generate a response content based on the updated question.

[0232] As a possible implementation manner of the embodiments of the present disclosure, the first processing module 601 includes:

[0233] An identification unit, configured to identify whether there is an abnormality in the question by using a first model according to the input question;

[0234] An update unit, configured to perform at least one round of question update according to the abnormality type identified by the first model when there is an abnormality in the question.

[0235] As a possible implementation manner of the embodiments of the present disclosure, the abnormality type includes at least one of the following:

[0236] The information of the problem does not meet the corresponding information sufficiency criterion;

[0237] The problem lacks context background information;

[0238] The problem has ambiguous wording;

[0239] The problem has redundant information.

[0240] As a possible implementation manner of an embodiment of the present disclosure, the first model includes multiple types of agents; the recognition unit is further configured to:

[0241] According to the input problem, use the first agent to determine the corresponding second agent for the problem, and distribute the problem to the corresponding second agent;

[0242] Use the second agent to identify whether there is an abnormality in the problem, and in the case where there is an abnormality in the problem, determine the type of abnormality to which the problem belongs.

[0243] As a possible implementation manner of an embodiment of the present disclosure, the second agent includes at least one of a clarity agent, a relevance agent, a precision agent, and a conciseness agent; the recognition unit is further configured to perform at least one of the following:

[0244] Use the clarity agent to determine whether the information of the problem meets the corresponding information sufficiency criterion, so as to determine that the type of abnormality to which the problem belongs is that the information of the problem does not meet the corresponding information sufficiency criterion when the information of the problem does not meet the corresponding information sufficiency criterion;

[0245] Use the relevance agent to determine whether the problem lacks context background information, so as to determine that the type of abnormality to which the problem belongs is that the problem lacks context background information when the problem lacks context background information;

[0246] Use the precision agent to determine whether the problem has ambiguous wording, so as to determine that the type of abnormality to which the problem belongs is that the problem has ambiguous wording when the problem has ambiguous wording;

[0247] Use the conciseness agent to determine whether the problem has redundant information, so as to determine that the type of abnormality to which the problem belongs is that the problem has redundant information when the problem has redundant information.

[0248] As a possible implementation manner of an embodiment of the present disclosure, the second processing module 602 is further configured to:

[0249] In any round of problem update process, display the corresponding clarification inquiry information according to the type of abnormality to which the problem to be updated in this round belongs.

[0250] As a possible implementation manner of an embodiment of the present disclosure, the second processing module 602 is further configured to perform at least one of the following:

[0251] During any round of question update process, when the exception type of the question to be updated in this round is that the information of the question does not meet the corresponding information sufficiency standard, display the clarification inquiry information about the information to be supplemented;

[0252] During any round of question update process, when the exception type of the question to be updated in this round is that the question lacks context background information, display the clarification inquiry information about the context background information to be added;

[0253] During any round of question update process, when the exception type of the question to be updated in this round is that the question has ambiguous wording, display the clarification inquiry information about the ambiguous wording to be corrected;

[0254] During any round of question update process, when the exception type of the question to be updated in this round is that the question has redundant information, display the clarification inquiry information about the redundant information to be compressed.

[0255] As a possible implementation manner of the embodiment of the present disclosure, the update module 603 is further configured to:

[0256] Semantically fuse at least one clarification inquiry information, the clarification feedback information corresponding to at least one clarification inquiry information, and the question to be updated in this round to obtain the updated question.

[0257] As a possible implementation manner of the embodiment of the present disclosure, the generation module 604 is further configured to:

[0258] When there is no exception in the updated question, adopt one of the first model and the second model to generate a reply content according to the updated question; wherein, the second model includes a large model, and the first model includes a small model with a model hierarchy depth less than that of the large model.

[0259] As a possible implementation manner of the embodiment of the present disclosure, the generation module 604 is further configured to:

[0260] Generate a candidate reply content by using the first model based on the updated question;

[0261] Use an evaluation model to evaluate the candidate reply content to obtain the confidence level of the candidate reply content;

[0262] When the confidence level is greater than the set threshold, determine the candidate reply content as the reply content;

[0263] When the confidence level is not greater than the set threshold, generate a reply content by using the second model based on the updated question.

[0264] As a possible implementation manner of the embodiment of the present disclosure, the above device further includes:

[0265] The first acquisition module is configured to acquire a satisfaction feedback corresponding to a reply content.

[0266] In an embodiment of the present disclosure, according to an input question, a first model is adopted to perform at least one round of question update on the question; during any round of question update, at least one clarification inquiry message is displayed, and a corresponding clarification feedback message is acquired; the question is updated according to at least one clarification inquiry message and the clarification feedback message corresponding to at least one clarification inquiry message; a reply content is generated based on the updated question. In the present disclosure, through multiple clarification interactions on the input question by adopting the first model, the updated question is more accurate and complete, and thus the reply content generated based on the updated question is significantly improved in terms of relevance and integrity, solving the problem in the related art that the answer is off-topic or incomplete due to unclear questioning, and improving user satisfaction.

[0267] As above Figure 5 Corresponding to the model training method provided in the above Figure 5 embodiment, the present disclosure further provides a model training apparatus. Since the model training apparatus provided in the embodiment of the present disclosure corresponds to the model training method provided in the above

[0268] Figure 7 is a schematic structural diagram of a data processing apparatus according to the seventh embodiment of the present disclosure.

[0269] As Figure 7 shown, the model training apparatus includes:

[0270] A second acquisition module 701 is configured to acquire multiple groups of training samples, where the training samples include an input question, at least one clarification inquiry message and a corresponding clarification feedback message for any round of question update, an updated question, a reply content generated based on the updated question, and a satisfaction feedback corresponding to the reply content;

[0271] A first training module 702 is configured to train a first model based on multiple groups of training samples.

[0272] As a possible implementation manner of the embodiment of the present disclosure, the first training module 702 is further configured to:

[0273] For any group of training samples, when the reply content in the training samples is generated by the first model, the first model is trained based on the input question, at least one clarification inquiry message and a corresponding clarification feedback message for any round of question update, the updated question, the reply content generated based on the updated question, and the satisfaction feedback corresponding to the reply content in the training samples.

[0274] As a possible implementation manner of the embodiments of the present disclosure, the first training module 702 is further configured to:

[0275] For any set of training samples, when the response content in the training samples is generated by the second model, based on the updated question in the training samples, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content, train the second model, and based on the question input in the training samples, at least one clarification query information for any round of question update, the corresponding clarification feedback information, and the satisfaction feedback corresponding to the response content, train the first model.

[0276] As a possible implementation manner of the embodiments of the present disclosure, the above device further includes:

[0277] A second training module, configured to, for any set of training samples, when the response content in the training samples is generated by the second model and the satisfaction feedback in the training samples is satisfied, based on the updated question in the training samples and the response content generated based on the updated question, train the first model.

[0278] In the embodiments of the present disclosure, multiple sets of training samples are obtained, where the training samples include the input question, at least one clarification query information for any round of question update and the corresponding clarification feedback information, the updated question, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content; based on the multiple sets of training samples, train the first model. In the present disclosure, each set of training samples is real interaction data. By training the first model based on the multiple sets of training samples, the first model can learn the real query intention and expected result of the user, thereby significantly improving the model performance.

[0279] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0280] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device 800 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0281] As Figure 8As shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to computer programs stored in the ROM (Read-Only Memory) 802 or computer programs loaded from the storage unit 808 into the RAM (Random Access Memory) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An I / O (Input / Output) interface 805 is also connected to the bus 804.

[0282] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0283] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the data processing method or the model training method. For example, in some embodiments, the data processing method or the model training method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above, or one or more steps of the model training method described above, can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method or the model training method by any other suitable means (e.g., by means of firmware).

[0284] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0285] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0286] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0287] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0288] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0289] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system or a server combined with a blockchain.

[0290] Among them, it should be noted that artificial intelligence is a discipline that studies enabling a computer to simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0291] It should be understood that various forms of processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is made herein.

[0292] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data processing method, comprising: According to the input question, using a first model to perform at least one round of question update on the question; During any round of question update, displaying at least once the clarification inquiry information and obtaining the corresponding clarification feedback information; Updating the question according to the at least once clarification inquiry information and the clarification feedback information corresponding to the at least once clarification inquiry information; Generating a response content based on the updated question.

2. The method according to claim 1, wherein, The step of using a first model to perform at least one round of question update on the question according to the input question includes: According to the input question, using the first model to identify whether there is an abnormality in the question; In the case where there is an abnormality in the question, performing at least one round of question update according to the identified abnormality type by the first model.

3. The method according to claim 2, wherein, The abnormality type includes at least one of the following: The information of the question does not meet the corresponding information sufficiency standard; The question lacks context background information; The question has ambiguous wording; The question has redundant information.

4. The method according to claim 2, wherein, The first model includes multiple types of agents; the step of using the first model to identify whether there is an abnormality in the question according to the input question includes: According to the input question, using a first agent to determine the corresponding second agent for the question and distributing the question to the corresponding second agent; Using the second agent to identify whether there is an abnormality in the question and, in the case where there is an abnormality in the question, determining the abnormality type to which the question belongs.

5. The method according to claim 4, wherein, The second agent includes at least one of a clarity agent, a relevance agent, an accuracy agent, and a conciseness agent; the step of using the second agent to identify whether there is an abnormality in the question and, in the case where there is an abnormality in the question, determining the abnormality type to which the question belongs includes at least one of the following: Using the clarity agent to determine whether the information of the question meets the corresponding information sufficiency standard, so as to determine that the abnormality type to which the question belongs is that the information of the question does not meet the corresponding information sufficiency standard in the case where the information of the question does not meet the corresponding information sufficiency standard; Using the relevance agent to determine whether the question lacks context background information, so as to determine that the abnormality type to which the question belongs is that the question lacks context background information in the case where the question lacks context background information; Using the accuracy agent to determine whether the question has ambiguous wording, so as to determine that the abnormality type to which the question belongs is that the question has ambiguous wording in the case where the question has ambiguous wording; Using the conciseness agent to determine whether the question has redundant information, so as to determine that the abnormality type to which the question belongs is that the question has redundant information in the case where the question has redundant information.

6. The method according to claim 2, wherein The step of displaying at least once the clarification inquiry information during any round of question update includes: During any round of question update, displaying the corresponding clarification inquiry information according to the abnormality type to which the question to be updated in this round belongs.

7. The method according to claim 6, wherein The step of displaying the corresponding clarification inquiry information according to the abnormality type to which the question to be updated in this round belongs during any round of question update includes at least one of the following: During any round of question update process, when the exception type of the question to be updated in this round is that the information of the question does not meet the corresponding information sufficiency standard, display the clarification inquiry information about the information to be supplemented; When the exception type of the question to be updated in this round is that the question lacks context background information, display the clarification inquiry information about the context background information to be added; When the exception type of the question to be updated in this round is that the question has ambiguous wording, display the clarification inquiry information about the ambiguous wording to be corrected; When the exception type of the question to be updated in this round is that the question has redundant information, display the clarification inquiry information about the redundant information to be compressed.

8. The method according to claim 1, wherein, The question update according to the at least one clarification inquiry information and the clarification feedback information corresponding to the at least one clarification inquiry information includes: Semantically fuse the at least one clarification inquiry information, the clarification feedback information corresponding to the at least one clarification inquiry information, and the question to be updated in this round to obtain the updated question.

9. The method according to claim 1, wherein The generation of the response content based on the updated question includes: When there is no exception in the updated question, use one of the first model and the second model to generate the response content based on the updated question; wherein, the second model includes a large model, and the first model includes a small model with a model layer depth less than that of the large model.

10. The method according to claim 9, wherein, The use of one of the first model and the second model to generate the response content based on the updated question includes: Use the first model to generate candidate response content based on the updated question; Use an evaluation model to evaluate the candidate response content to obtain the confidence of the candidate response content; When the confidence is greater than the set threshold, determine the candidate response content as the response content; When the confidence is not greater than the set threshold, use the second model to generate the response content based on the updated question.

11. The method according to any one of claims 1-10, wherein, The method further includes: Obtain the satisfaction feedback corresponding to the response content.

12. A model training method, including: Obtain multiple groups of training samples, where the training samples include the input question, at least one clarification inquiry information and the corresponding clarification feedback information for any round of question update, the updated question, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content; Train the first model based on the multiple groups of training samples.

13. The method according to claim 12, wherein The training of the first model based on the multiple groups of training samples includes: For any group of training samples, when the response content in the training samples is generated by the first model, train the first model based on the input question, at least one clarification inquiry information and the corresponding clarification feedback information, the updated question, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content in the training samples.

14. The method according to claim 12, wherein, The training of the first model based on the multiple groups of training samples includes: For any set of training samples, when the response content in the training samples is generated by the second model, based on the updated question in the training samples, the response content generated based on the updated question, and the satisfaction feedback corresponding to the response content, train the second model, and based on the input question in the training samples, at least one clarification query information for any round of question update, the corresponding clarification feedback information, and the satisfaction feedback corresponding to the response content, train the first model.

15. The method according to any one of claims 12 - 14, wherein, The method further includes: For any set of training samples, when the response content in the training samples is generated by the second model and the satisfaction feedback in the training samples is satisfied, based on the updated question in the training samples and the response content generated based on the updated question, train the first model.

16. A data processing device, comprising: A first processing module, configured to perform at least one round of question update on the question by using a first model according to the input question; A second processing module, configured to display at least one clarification query information and obtain corresponding clarification feedback information during any round of question update; An update module, configured to perform question update according to the at least one clarification query information and the clarification feedback information corresponding to the at least one clarification query information; A generation module, configured to generate a response content according to the updated question.

17. A model training device, comprising: A second acquisition module, configured to acquire multiple sets of training samples, where the training samples include an input question, at least one clarification query information for any round of question update, the corresponding clarification feedback information, a response content generated based on the updated question, and the satisfaction feedback corresponding to the response content; A first training module, configured to train a first model based on the multiple sets of training samples.

18. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-11, or execute the method according to any one of claims 12-15.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11, or execute the method according to any one of claims 12-15.

20. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-11 when executed by a processor, or implements the method according to any one of claims 12-15.