Question and answer method
By introducing business guidance tools and dialogue guidance models into the large model, analyzing customer text information and generating guidance suggestions, the problem of insufficient accuracy and standardization of Q&A in specific domains of the large model is solved, and more efficient business process standardization and risk control are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies show that large models have low accuracy and standardization in question answering within specific domains, making it difficult to meet the banking industry's stringent requirements for standardized service processes.
By utilizing the first major model to invoke business guidance tools, analyzing the text information input by customers, generating guidance suggestions based on the trained dialogue guidance model, including recommended intent and recommended response information, and combining historical dialogue information and business intent transfer processes, the question-and-answer process is optimized.
It improves the accuracy and standardization of large-scale models in specific domain question answering, ensures the standardization of business processes, and reduces the risk of business violations.
Smart Images

Figure CN121809680A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a question and answer method. BACKGROUND
[0002] With the rise of large model technology, large models are usually applied in business scenarios such as bank customer service, sales, and collection to interact with customers.
[0003] In related technologies, a general large model is usually used to interact with customers in combination with prompt words. Business knowledge and script specifications are input to the general large language model through long prompt words, expecting the general large language model to play a professional role according to the instructions. For example, an insurance company describes the script specifications and processes of insurance sales in detail in the prompt words, and expects the general large language model to play the role of a sales consultant.
[0004] However, this approach has weak business process control and is difficult to achieve standardized services. The general large model lacks a deep understanding of the standard processes in specific fields, and even with long prompt words, the general large model cannot maintain a clear business path in complex conversations. For example, in a credit card sales scenario, the general large model may be eager to introduce product benefits without completing the preliminary assessment of the customer's qualifications, or may omit the necessary risk tolerance assessment link when the customer consults about financial planning. The randomness of this conversation path not only reduces the professionalism of the service, but also may lead to business violations, which cannot meet the strict requirements of the banking industry for standardized service processes.
[0005] Therefore, how to improve the accuracy and standardization of large models in specific field question and answer has become a problem to be solved. SUMMARY
[0006] Embodiments of the present application provide a question and answer method to solve the problem of low accuracy and standardization of large models in specific field question and answer in the prior art.
[0007] The present application provides a question and answer method, which comprises: If the text information input by the customer is received, a first large model is used to call a business guide tool to analyze the text information, and obtain a guide suggestion output by the business guide tool, the guide suggestion comprising a recommended intention and recommended reply information corresponding to the recommended intention, wherein the recommended intention and the recommended reply information are a script guide model deployed in the business guide tool and trained to determine the content included in the text information and the target intention of the text information. The first large model generates target reply information corresponding to the text information based on the guide suggestion.
[0008] Furthermore, before analyzing the text information input by the customer using the first major model to invoke the business guidance tool after receiving the text information, the method further includes: Obtain the historical dialogue information prior to the text information; The text information and the historical dialogue information are used to construct dialogue information, and the dialogue information is used to update the text information.
[0009] Furthermore, the business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines. The connecting lines are used to describe the flow order between the corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order. The process by which the business guidance tool determines the guidance suggestion also includes: Determine the target intent flow sequence of the target intent and the recommended intent; Among the multiple business intent transfer processes, find the first business intent transfer process that contains the target intent flow sequence; Determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; Based on the weight of each connector identifier, a target weight is determined and added to the output of the guidance suggestion, wherein the target weight is used to describe the probability that responding in accordance with the recommendation intent will lead to business success.
[0010] Furthermore, the fine-tuning training process of the dialogue guidance model includes: Obtain target training data from the sample set, and the labels corresponding to the target training data; the target training data includes dialogue information, and a first intent corresponding to at least one sentence in the dialogue information; the labels are used to identify standard response information made in response to the last sentence in the dialogue information, and the standard intent of the standard response information; The target training data is input into the basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data. A target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight. The preset weight includes a first weight corresponding to the first deviation and a second weight corresponding to the second deviation. The basic large model is fine-tuned and trained based on the target reward value to obtain the speech guidance model.
[0011] Furthermore, the sample set includes positive samples, and the process of obtaining the positive samples and labels includes: Based on the performance data of each salesperson, outstanding salespersons are identified. Obtain a preset number of business dialogue messages from the historical dialogue information corresponding to the outstanding business personnel; For each of the aforementioned business dialogue messages, a portion of the dialogue information is extracted from the business dialogue message. The extracted portion of the dialogue information is determined as training data for positive samples. The next sentence at the extracted position in the business dialogue message is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined.
[0012] Furthermore, the sample set also includes negative samples, and the process of obtaining the negative samples and labels includes at least one of the following: Acquire sample dialogue information and use a general large model to generate standard response information and standard intent for the sample dialogue information. Use the sample dialogue information as training data for negative samples, and save the standard response information and standard intent as labels corresponding to the sample dialogue information; or Determine the target business domain for which the dialogue guidance model is applied, acquire business dialogue information from non-target business domains, extract a portion of the dialogue information from the non-target business domains, use this extracted portion as training data for negative samples, take the next sentence at the extracted position in the non-target business domains as the standard response information for this training data, determine the standard intent of the next sentence at the extracted position, and save the standard response information and the standard intent as labels corresponding to the training data; or Based on the performance data of each salesperson, non-performing salespersons are identified; sales dialogue information of the non-performing salespersons is obtained, a portion of the dialogue information is extracted, and the extracted portion of the dialogue information is used as training data for negative samples. The next sentence after the extracted position in the dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence after the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data.
[0013] Furthermore, if the target training data is negative samples, after determining the target reward value and before fine-tuning the basic large model based on the target reward value to obtain the dialogue guidance model, the method further includes: Determine the opposite of the target reward value, and update the target reward value using the opposite.
[0014] Furthermore, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Obtain the target business conversion result corresponding to the target training data, and the business performance reward value pre-saved for the target business conversion result; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the business performance reward value, wherein the preset weight includes the first weight, the second weight, and the third weight corresponding to the business performance reward value.
[0015] Furthermore, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Based on the first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent, determine the predicted intent flow sequence corresponding to the target training data; Among the multiple pre-saved intent transfer processes, find the second business intent transfer process that contains the predicted intent transfer sequence; Determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; The intention flow reward value is determined based on the weight of each connection identifier; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and the fourth weight corresponding to the intent flow reward value.
[0016] Furthermore, the process of determining the preset weights includes: If the current round of fine-tuning training of the dialogue guidance model is in the first round interval, then the basic ability training weight is used as the preset weight, and the maximum weight in the basic ability training weight is the first weight. If the current round of fine-tuning training of the dialogue guidance model is in the second round interval, then the balance ability training weight is used as the preset weight, and the maximum weight in the balance ability training weight is the second weight. If the current training round for fine-tuning the dialogue guidance model is in the third round interval, then the process capability training weight is used as the preset weight, and the maximum weight in the process capability training weight is the fourth weight. The minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
[0017] Furthermore, the process of determining the business intent transfer flow includes: Obtain any historical call text and the second intent corresponding to each sentence in that historical call text; The historical call text and the second intent are processed using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; The weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence are determined based on a preset algorithm, and the business intent transfer process is obtained.
[0018] Furthermore, the process of constructing the historical call text includes: Retrieve a preset amount of basic call text; Based on the role, sentences in the preset number of basic call texts are classified to obtain a sentence set corresponding to each role; For each set of sentences, clustering is performed based on the semantics of each sentence in the set to obtain sentence groups; Obtain a business dictionary pre-saved for the target business domain applied to the script guidance model, and a suggested intent saved for each keyword in the business dictionary; For each sentence group, the business dictionary, the suggested intent corresponding to each keyword, and the sentence group are input into the second large model, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword. The target base call text to which each sentence in each sentence group belongs is determined, and the corresponding second intent is identified for the corresponding sentence in the target base call text to obtain the historical call text.
[0019] Furthermore, after obtaining the sentence group and before acquiring the business dictionary pre-saved for the target business domain applied to the dialogue guidance model, the method further includes: Count the number of sentences in each sentence group, and delete sentence groups whose number is less than a threshold.
[0020] This application embodiment also provides a question-answering device, the device comprising: The guidance module is used to, upon receiving text information input by a customer, use the first major model to call the business guidance tool to analyze the text information and obtain guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommendation response information corresponding to the recommendation intent. The recommendation intent and the recommendation response information are determined by the trained dialogue guidance model deployed in the business guidance tool, based on the content of the text information and the target intent of the text information. The response module is used by the first large model to generate target response information corresponding to the text information based on the guidance suggestions.
[0021] Furthermore, the device also includes: The acquisition module is used to acquire historical dialogue information prior to the text information; to construct dialogue information using the text information and the historical dialogue information; and to update the text information using the dialogue information.
[0022] Furthermore, the business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines. The connecting lines are used to describe the flow order between the corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order. The guidance module is specifically used to determine the target intent flow sequence of the target intent and the recommended intent; find a first business intent transfer process containing the target intent flow sequence among the multiple business intent transfer processes; determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; determine the target weight according to the weight of each connection line identifier and add it to the guidance suggestion for output, wherein the target weight is used to describe the probability that responding according to the recommended intent can promote business success.
[0023] Furthermore, the device also includes: A training module is used to acquire target training data in a sample set and the corresponding labels for the target training data. The target training data includes dialogue information and a first intent corresponding to at least one sentence in the dialogue information. The labels are used to identify standard response information made in response to the last sentence in the dialogue information and the standard intent of the standard response information. The target training data is input into a basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data. A target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and preset weights. The preset weights include a first weight corresponding to the first deviation and a second weight corresponding to the second deviation. The basic large model is fine-tuned based on the target reward value to obtain the dialogue guidance model.
[0024] Furthermore, the sample set includes positive samples, and the acquisition module is also used to determine outstanding sales personnel based on the performance data of each sales personnel; to acquire a preset number of business dialogue information from the historical dialogue information corresponding to the outstanding sales personnel; for each business dialogue information, to extract a portion of the dialogue information, to determine the extracted portion of the dialogue information as training data for the positive samples, to use the next sentence at the extracted position in the business dialogue information as the standard response information for the training data, and to determine the standard intent of the next sentence at the extracted position.
[0025] Furthermore, the sample set also includes negative samples. The acquisition module is further configured to acquire sample dialogue information and generate standard response information and standard intent of the sample dialogue information using a general large model. The sample dialogue information is used as training data for negative samples, and the standard response information and standard intent are saved as labels corresponding to the sample dialogue information. Alternatively, the target business domain of the dialogue guidance model is determined, business dialogue information in non-target business domains is acquired, a portion of the dialogue information in the non-target business domains is extracted, the extracted portion of the dialogue information is used as training data for negative samples, and the next sentence at the extracted position in the non-target business domains is selected. As the standard response information for the training data, and determining the standard intent of the next sentence at the truncated position, the standard response information and the standard intent are saved as labels corresponding to the training data; or, based on the performance data of each salesperson, non-excellent salespersons are identified; the sales dialogue information of the non-excellent salespersons is obtained, a portion of the dialogue information is truncated from the sales dialogue information, the truncated portion of the dialogue information is determined as the training data for negative samples, the next sentence at the truncated position in the sales dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence at the truncated position is determined, the standard response information and the standard intent are saved as labels corresponding to the training data.
[0026] Furthermore, if the target training data is a negative sample, the training module is also used to determine the opposite of the target reward value and update the target reward value using the opposite.
[0027] Furthermore, the training module is also used to obtain the target business conversion result corresponding to the target training data, and the business performance reward value saved in advance for the target business conversion result; and to determine the target reward value according to the first deviation, the second deviation, the preset weight and the business performance reward value, wherein the preset weight includes the first weight, the second weight and the third weight corresponding to the business performance reward value.
[0028] Furthermore, the training module is also configured to: determine a predicted intent flow sequence corresponding to the target training data based on a first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent; search for a second business intent transfer process containing the predicted intent flow sequence among a plurality of pre-saved intent transfer processes; determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; determine an intent flow reward value based on the weight identified by each connection line; and determine a target reward value based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and a fourth weight corresponding to the intent flow reward value.
[0029] Furthermore, the training module is specifically configured to: if the current round of fine-tuning training of the dialogue guidance model is in the first round interval, then use the basic ability training weight as the preset weight, where the maximum weight in the basic ability training weight is the first weight; if the current round of fine-tuning training of the dialogue guidance model is in the second round interval, then use the balanced ability training weight as the preset weight, where the maximum weight in the balanced ability training weight is the second weight; if the current round of fine-tuning training of the dialogue guidance model is in the third round interval, then use the process ability training weight as the preset weight, where the maximum weight in the process ability training weight is the fourth weight, wherein the minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
[0030] Furthermore, the device also includes: The determination module is used to obtain any historical call text and the second intent corresponding to each sentence in the historical call text; process the historical call text and the second intent using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; determine the weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence based on a preset algorithm to obtain the business intent transfer process.
[0031] Furthermore, the determining module is also used to acquire a preset number of basic call texts; classify the sentences in the preset number of basic call texts based on roles to obtain a sentence set corresponding to each role; for each sentence set, perform clustering processing based on the semantics of each sentence in the sentence set to obtain sentence groups; acquire a business dictionary pre-saved for the target business domain applied to the dialogue guidance model, and a suggested intent saved for each keyword in the business dictionary; for each sentence group, input the business dictionary, the suggested intent corresponding to each keyword, and the sentence group into the second large model, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword; determine the target basic call text to which each sentence in each sentence group belongs, and identify the corresponding second intent for the corresponding sentence in the target basic call text to obtain historical call text.
[0032] Furthermore, the determining module is also used to count the number of sentences included in each sentence group and delete sentence groups whose number is less than a number threshold.
[0033] This application also provides an electronic device, which includes a processor for executing a computer program stored in a memory to implement the steps of any of the question-and-answer methods described above.
[0034] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the question-and-answer methods described above.
[0035] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the steps of any of the question-and-answer methods described above.
[0036] In this embodiment, after receiving text information input by the customer, the first large model calls the business guidance tool to analyze the text information and obtain guidance suggestions output by the business guidance tool. Since the guidance suggestions include the recommendation intent and the corresponding recommended response information, the first large model can refer to the guidance suggestions to generate the target response information corresponding to the text information. Because the business guidance tool is equipped with a trained dialogue guidance model, this dialogue guidance model can accurately determine the guidance suggestions based on the content included in the text information and the target intent of the text information. Even if the first large model has not been trained on data from a specific domain, it can still accurately respond based on the guidance suggestions, improving the accuracy and standardization of the large model's question-and-answer performance in a specific domain. Attached Figure Description
[0037] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 A flowchart illustrating a question-and-answer process provided in an embodiment of this application; Figure 2 A schematic diagram of a question-and-answer process provided in an embodiment of this application; Figure 3 This is a schematic diagram of a question-and-answer device structure provided in an embodiment of this application; Figure 4 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art are within the scope of protection of this application.
[0040] The acquisition, transmission, storage, and use of data in this application all comply with the requirements of relevant national laws and regulations.
[0041] This application provides a question-and-answer method. In this method, if text information input by a customer is received, a first model is used to call a business guidance tool to analyze the text information and obtain guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommended response information corresponding to the recommendation intent. The recommendation intent and the recommended response information are determined by a trained dialogue guidance model deployed in the business guidance tool, based on the content of the text information and the target intent of the text information. Based on the guidance suggestions, the first model generates the target response information corresponding to the text information.
[0042] To facilitate subsequent understanding, before introducing specific implementation methods, the definitions of some abbreviations and key terms involved in this application will be explained.
[0043] Direct Preference Optimization (DPO) is a machine learning method that trains and optimizes model preferences by comparing positive and negative samples.
[0044] Model Context Protocol (MCP): A standardized model service interface protocol that supports tool calls and resource access, enabling convenient integration of models with external systems.
[0045] Workflow Decision Tree: A business process diagram automatically extracted from actual dialogue data. Its nodes represent dialogue intentions, and edges represent the transition relationships between intentions and their statistically derived success probability weights.
[0046] Dialogue intent refers to the business objective desired to be achieved in a particular round of dialogue. For example, in a sales scenario, this includes "opening greetings," "demand assessment," "handling price objections," and "closing the deal."
[0047] Positive samples: refer to real and effective dialogue response data from high-performing agents.
[0048] Negative samples: refer to low-quality dialogue responses used for comparative training, typically including responses that lack business relevance generated by general large models, or suboptimal responses from ordinary agents.
[0049] Script guidance: refers to intelligently guiding the dialogue model to generate response content that conforms to the optimal business path based on the preset business process and the current dialogue status.
[0050] Example 1: Figure 1 This application provides a flowchart illustrating a question-and-answer process, as shown in the embodiments below. Figure 1 As shown, the process includes the following steps: S101: If text information input by a customer is received, the first major model is used to call the business guidance tool to analyze the text information and obtain the guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommendation response information corresponding to the recommendation intent. The recommendation intent and the recommendation response information are determined by the trained dialogue guidance model deployed in the business guidance tool based on the content of the text information and the target intent of the text information.
[0051] The question-and-answer method provided in this application can be applied to electronic devices, such as mobile terminals, servers, and PCs.
[0052] In this embodiment, text information input by the customer can be received in real time. This text information can be input by the customer during a conversation with the electronic device. The text information can be directly input by the customer or obtained through voice recognition based on the customer's input. For example, the text information could be: "The product you recommended is too expensive," "Are there any other similar products?", "I'll think about it," etc.
[0053] If it is confirmed that text information input by the customer has been received, in order to generate the target response information corresponding to that text information, in this embodiment of the application, the first major model can be used to call the business guidance tool to analyze the received text information and obtain the guidance suggestions output by the business guidance tool.
[0054] In this embodiment, after receiving a request to analyze the text information, the business guidance tool can analyze the content and target intent of the text information based on the deployed and trained dialogue guidance model, thereby providing the intent corresponding to the next response. For ease of description, this intent can be referred to as the recommended intent. To enable the primary model to make more accurate references, in this embodiment, the dialogue guidance model can also provide recommended response information corresponding to the recommended intent. For example, the guidance information could be: "Recommended next intent: Acknowledge feelings and explore budget; Recommended response information: I understand your concerns; insurance is indeed a long-term investment. To recommend the most suitable plan for you, would you like to know your approximate premium budget?"
[0055] S102: The first large model generates target response information corresponding to the text information based on the guidance suggestion.
[0056] After receiving guidance suggestions from the business guidance tool, the first major model can generate the target response information corresponding to the text information based on these suggestions. This first major model can be any large model with semantic recognition capabilities, such as a large language model; those skilled in the art can choose a suitable large model for question answering as needed. Because the first major model possesses extremely strong semantic recognition capabilities, after receiving guidance suggestions, it can refer to the recommendation intent and recommended response information included in the guidance suggestions to determine the final target response information.
[0057] It should be noted that the target response information determined by the first major model can be the same as or different from the recommended response information.
[0058] In this embodiment, after receiving text information input by the customer, the first large model calls the business guidance tool to analyze the text information and obtain guidance suggestions output by the business guidance tool. Since the guidance suggestions include the recommendation intent and the corresponding recommended response information, the first large model can refer to the guidance suggestions to generate the target response information corresponding to the text information. Because the business guidance tool is equipped with a trained dialogue guidance model, this dialogue guidance model can accurately determine the guidance suggestions based on the content included in the text information and the target intent of the text information. Even if the first large model has not been trained on data from a specific domain, it can still accurately respond based on the guidance suggestions, improving the accuracy and standardization of the large model's question-and-answer performance in a specific domain.
[0059] Example 2: To further improve the accuracy of question and answer, based on the above embodiments, in this embodiment, before analyzing the text information by calling the business guidance tool using the first model after receiving text information input from the customer, the method further includes: Obtain the historical dialogue information prior to the text information; The text information and the historical dialogue information are used to construct dialogue information, and the dialogue information is used to update the text information.
[0060] Since the question-and-answer process typically involves multiple rounds of dialogue, and in most scenarios, such as telephone sales, the customer passively initiates the interaction, historical dialogue information often precedes the customer's text message. Therefore, in this embodiment, after confirming the receipt of the customer's input text message, and before using the first major model to determine the target response, historical dialogue information preceding the text message can be obtained. This text message and historical dialogue information are then used to construct dialogue information, which is then used to update the text message. In other words, the target response is subsequently determined based on the completed dialogue record constructed from this text message and historical dialogue information. This allows the business guidance tool to make more accurate predictions based on more information when determining the recommendation intent and recommended response information.
[0061] In one possible implementation, in addition to obtaining historical dialogue information prior to obtaining the text information, historical guidance suggestions output by the business guidance tool can also be obtained. That is, the target response information is subsequently determined based on the text information, historical dialogue information, and historical guidance suggestions.
[0062] Example 3: To further improve the accuracy of question and answer, based on the above embodiments, in this embodiment of the application, the business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines. The connecting lines are used to describe the flow order between corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question and answer dialogue according to the corresponding flow order. The process by which the business guidance tool determines the guidance suggestion also includes: Determine the target intent flow sequence of the target intent and the recommended intent; Among the multiple business intent transfer processes, find the first business intent transfer process that contains the target intent flow sequence; Determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; Based on the weight of each connector identifier, a target weight is determined and added to the output of the guidance suggestion, wherein the target weight is used to describe the probability that responding in accordance with the recommendation intent will lead to business success.
[0063] If the guidance suggestions provided by the business guidance tool include more information, then the first major model will have more references when determining the target response information corresponding to the text information, thus providing a more accurate response. Therefore, in this embodiment, the guidance suggestions may also include a target weight, which describes the probability that responding according to the recommended intent will lead to business success.
[0064] To determine the target weight, in this embodiment, multiple business intent transfer processes can be encapsulated within the business guidance tool. These business intent transfer processes can be used to describe the change in intent of the content communicated by the participants in any given business dialogue. These business intent transfer processes can be understood as a workflow or workflow decision tree during communication.
[0065] In this embodiment of the application, any business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines, which are used to describe the flow order between the corresponding business intents. Each connecting line is assigned a weight, which is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order.
[0066] Suppose a business dialogue involves two roles: a customer and a customer service representative. The customer service representative initiates the dialogue. The customer service representative's first statement has intent A. The customer's response to the representative's statement has intent B. The customer service representative then continues the dialogue based on the customer's response, with intent C, and so on. The business transition flow for this dialogue is: Intent A -> Intent B -> Intent C. Here, "Intent A," "Intent B," and "Intent C" represent nodes, "->" represents a connecting line, and "Intent A -> Intent B" indicates a transition from Intent A to Intent B. The weight of the connecting line between "Intent A -> Intent B" can be 0.3, and the weight of the connecting line between "Intent B -> Intent C" can be 0.6. That is, if the current intent of the customer service representative's dialogue with the customer is Intent B, and the intent of the subsequent dialogue shifts to Intent C, the probability of successful business communication is 0.6. Therefore, the customer service representative can then guide the conversation towards Intent C.
[0067] In this embodiment of the application, when using the business guidance tool to determine guidance information, in addition to determining the recommendation intent and recommendation response information, it is also possible to determine the target intent corresponding to the text information entered by the customer and the target intent flow sequence of the recommendation intent. For example, if the target intent is intent A and the recommendation intent is intent B, then the target intent flow sequence is intent A -> intent B.
[0068] In one possible implementation, if historical dialogue information is also input into the business guidance tool, the historical intent corresponding to each sentence in the historical dialogue information can be determined, and the target intent flow sequence can be determined based on the historical intent, the target intent, and the recommended intent. For example, if the historical meaning graph consists of intents C, M, and D, the target intent is intent A, and the recommended intent is intent B, then the target intent flow sequence is intent C -> intent M -> intent D -> A -> intent B.
[0069] After determining the target intent flow sequence, a first business intent transfer process containing the target intent flow sequence can be found among multiple pre-saved business intent transfer processes. Within this first business intent transfer process, each connecting line corresponding to the target intent flow sequence is identified, and a target weight is determined based on the weight identifier of each connecting line and added to the guidance suggestion output. In this embodiment, the sum of the weights corresponding to each connecting line can be determined as the target weight. Alternatively, the weight corresponding to the last connecting line can be determined as the target weight.
[0070] Example 4: To obtain a more accurate script guidance model, based on the above embodiments, the fine-tuning training process of the script guidance model in this embodiment includes: Obtain target training data from the sample set, and the labels corresponding to the target training data; the target training data includes dialogue information, and a first intent corresponding to at least one sentence in the dialogue information; the labels are used to identify standard response information made in response to the last sentence in the dialogue information, and the standard intent of the standard response information; The target training data is input into the basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data. A target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight. The preset weight includes a first weight corresponding to the first deviation and a second weight corresponding to the second deviation. The basic large model is fine-tuned and trained based on the target reward value to obtain the speech guidance model.
[0071] In related technologies, the common approach is to collect dialogue data from business scenarios and then fine-tune a large-parameter base model, either fully or efficiently, to allow the model to learn professional business communication styles from the data. For example, a bank might collect 100,000 customer service conversations and use the LoRA method to fine-tune a model with 13 billion parameters. However, this fine-tuning method is costly and difficult to adapt to multiple scenarios. For instance, banking business scenarios are diverse and highly specialized, covering multiple independent areas such as credit card marketing, loan consultation, wealth management recommendations, and complaint handling. Fine-tuning a dedicated large model for each scenario requires significant investment in computing resources, storage costs, and time. Furthermore, the communication standards, process requirements, and risk control points differ significantly across business scenarios, and isolated deployment between models leads to low resource utilization, difficulty in scenario switching, and an inability to adapt to the rapid iteration and cross-departmental collaboration needs of banking operations. Therefore, in this embodiment, a lightweight base model can be fine-tuned to obtain a lightweight communication guidance model applicable to a specific domain. For example, this base model could be an 800 million parameter Qwen3 model.
[0072] To train the basic large-scale model, in this embodiment, a sample set can be pre-configured. This sample set stores multiple training data sets and a label corresponding to each training data set. The basic large-scale model can then be fine-tuned based on each training data set in this sample set. In this embodiment, any training data set includes dialogue information, which includes at least one sentence. The training data also includes a first intent corresponding to at least one sentence in the dialogue information, so that a predicted intent can be obtained based on the first intent during subsequent fine-tuning training. In this embodiment, the label corresponding to any training data set represents the standard response information to the last sentence in the corresponding dialogue information, and the standard intent of the standard response information.
[0073] In this embodiment of the application, during fine-tuning training, any training data can be obtained from the sample set. For ease of description, this obtained training data is referred to as the target training data. Furthermore, to facilitate the adjustment of model parameters, the labels corresponding to the target training data can also be obtained.
[0074] The target training data is input into a base model, which analyzes the dialogue information in the target training data and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent. The predicted response information is the response made by the base model based on the last sentence of the dialogue information, and the predicted intent is the intent corresponding to the predicted response information.
[0075] After obtaining the predicted intent and predicted response information, a first deviation between the predicted response information and the standard response information, and a second deviation between the predicted intent and the standard intent, can be determined. For example, the semantic similarity between the predicted response information and the standard response information can be determined, and the difference between a set value and this semantic similarity is determined as the first deviation, where the set value can be 1. Similarly, the semantic similarity between the predicted intent and the standard intent can be determined, and the difference between the set value and this semantic similarity is determined as the second deviation.
[0076] In related technologies, simple weighted summation of multiple training objectives easily leads to conflicts and instability, lacking a phased and hierarchical progressive optimization strategy. Therefore, in this embodiment, more weights can be allocated to the content that the model is expected to focus on learning, and preset weights can be obtained during model training. These preset weights include a first weight corresponding to a first deviation and a second weight corresponding to a second deviation. The first weight can be understood as a weight configured for speech quality, and the second weight can be understood as a weight configured for intent accuracy.
[0077] After obtaining the preset weights, the target reward value can be determined based on the first deviation, the second deviation, and the preset weights. Specifically, the target reward value = first deviation × first weight + second deviation × second weight.
[0078] After obtaining the target reward value, the basic large model can be fine-tuned and trained based on this target reward value to obtain the dialogue guidance model. How to fine-tune and train the basic large model based on the target reward value is existing technology, and this application embodiment will not elaborate on this process.
[0079] In this embodiment, a convergence condition is preset. This convergence condition may be that the number of times the semantic similarity between the predicted response information and the standard response information of the target training data in the sample set exceeds a preset threshold is greater than a set number, and / or the number of times the semantic similarity between the predicted intent and the standard intent exceeds a preset threshold is greater than a set number. Alternatively, it may be that the number of iterations of model training reaches a set maximum number of iterations, etc. This embodiment does not limit the specifics. When the convergence condition is met, the fine-tuning training of the dialogue guidance model can be considered complete, resulting in a trained dialogue guidance model. This dialogue guidance model can be encapsulated in a business guidance tool to assist the first main model in completing question-and-answer processing.
[0080] If the model fine-tuning training does not meet the convergence condition, the currently obtained dialogue guidance model can be determined as the base model. In other words, when the convergence condition is not met, each fine-tuning training is based on the dialogue guidance model obtained in the previous round of training.
[0081] Example 5: To further improve the accuracy of model fine-tuning training, based on the above embodiments, in this embodiment, the sample set includes positive samples, and the process of obtaining the positive samples and labels includes: Based on the performance data of each salesperson, outstanding salespersons are identified. Obtain a preset number of business dialogue messages from the historical dialogue information corresponding to the outstanding business personnel; For each of the aforementioned business dialogue messages, a portion of the dialogue information is extracted from the business dialogue message. The extracted portion of the dialogue information is determined as training data for positive samples. The next sentence at the extracted position in the business dialogue message is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined.
[0082] Taking the banking industry as an example, the banking sector has accumulated massive amounts of high-quality customer service dialogue data, which contains successful marketing strategies and risk avoidance experience from top account managers. Related technologies rely on manual extraction of business rules or simple text imitation, lacking a mechanism to automatically extract "high-conversion-rate dialogue paths" and "effective risk control nodes" from successful cases. This prevents the systematic accumulation and replication of valuable business knowledge assets, impacting the overall improvement of service levels. In this embodiment, historical data can be automatically analyzed to construct a training set.
[0083] In this embodiment, to improve the quality of model fine-tuning training, a preset number of historical dialogues can be selected from the historical dialogue information of high-performing agents as training data in the sample set. For example, 8000 high-quality interactions can be selected from the dialogues of the top 20% of performing agents. In this embodiment, this acquired training data can be used as positive samples.
[0084] In this embodiment, outstanding sales personnel can be identified based on their performance data. It should be noted that those skilled in the art can configure methods for determining outstanding sales personnel based on performance data as needed. The number of outstanding sales personnel can be one or more. The decision-making logic for outstanding agents is implicit, and traditional statistical methods or manual summarization are difficult to extract systematically and quantitatively.
[0085] In this embodiment of the application, after identifying outstanding sales personnel, historical dialogue information corresponding to the outstanding sales personnel can be obtained, and a preset number of dialogue messages can be obtained from the historical dialogue information.
[0086] Since each training data in the sample set also corresponds to a label, in this embodiment of the application, for each business dialogue information, a portion of the dialogue information can be extracted from the business dialogue information, the extracted portion of the dialogue information can be determined as the training data of the positive sample, the next sentence of the extracted position in the business dialogue information can be used as the standard response information of the training data, and the standard intent of the next sentence of the extracted position can be determined.
[0087] In this embodiment, the segment can be randomly selected from any position within the business dialogue information. Since the dialogue guidance model aims to predict the customer service representative's response based on the customer's statements, in this embodiment, the segment can be selected after any sentence spoken by the customer.
[0088] When determining the standard intent, it can be based on a trained intent recognition model or determined by keyword comparison. This application does not limit this, and those skilled in the art can configure it as needed.
[0089] Specifically, suppose a certain business dialogue includes: sentence 1, sentence 2, sentence 3, and sentence 4, and the extracted dialogue information can be: sentence 1 and sentence 2. Then, "sentence 1 and sentence 2" can be determined as the training data of positive samples, "sentence 3" can be determined as the standard response information of the training data, and the intent corresponding to "sentence 3" can be determined as the standard intent.
[0090] To further improve the accuracy of model fine-tuning training, based on the above embodiments, in this embodiment, the sample set also includes negative samples, and the process of obtaining the negative samples and labels includes at least one of the following: Acquire sample dialogue information and use a general large model to generate standard response information and standard intent for the sample dialogue information. Use the sample dialogue information as training data for negative samples, and save the standard response information and standard intent as labels corresponding to the sample dialogue information; or Determine the target business domain for which the dialogue guidance model is applied, acquire business dialogue information from non-target business domains, extract a portion of the dialogue information from the non-target business domains, use this extracted portion as training data for negative samples, take the next sentence at the extracted position in the non-target business domains as the standard response information for this training data, determine the standard intent of the next sentence at the extracted position, and save the standard response information and the standard intent as labels corresponding to the training data; or Based on the performance data of each salesperson, non-performing salespersons are identified; sales dialogue information of the non-performing salespersons is obtained, a portion of the dialogue information is extracted, and the extracted portion of the dialogue information is used as training data for negative samples. The next sentence after the extracted position in the dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence after the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data.
[0091] To further improve the accuracy of model training, in this embodiment of the application, the sample set may also include negative samples, so that the base model can learn how to make positive and effective responses during the training process.
[0092] In one possible implementation, when acquiring negative samples, sample dialogue information can be acquired, and standard response information and standard intent of the sample dialogue information can be generated using a general large model. The sample dialogue information can be used as training data for negative samples, and the standard response information and standard intent can be saved as labels corresponding to the sample dialogue information.
[0093] Specifically, Deepseek-R1 can be used to generate generic responses for the same context. For example, Deepseek-R1 can be used to generate standard response information and standard intent from training data of positive samples. In this embodiment, the negative sample generated based on this method lacks business knowledge.
[0094] In one possible implementation, when acquiring negative samples, the target business domain for which the trained dialogue guidance model is applied can be determined, and business dialogue information from non-target business domains can be acquired. A portion of the dialogue information from the non-target business domain is then extracted and used as training data for the negative samples. The next sentence following the extracted position in the non-target business domain dialogue information is used as the standard response information for this training data, and the standard intent of the next sentence is determined. The standard response information and the standard intent are then saved as labels corresponding to the training data.
[0095] Specifically, business dialogue information from other financial industries (such as bank wealth management) can be obtained. For each business dialogue, a portion of the dialogue information is extracted, and the next sentence at the extracted position is determined as the standard response information, along with the standard intent of the next sentence. In this embodiment, this method can generate negative samples from other business domains that do not match the target business domain.
[0096] In one possible implementation, when acquiring negative samples, non-performing sales personnel can be identified based on the performance data of each salesperson; the sales dialogue information of the non-performing sales personnel can be acquired, a portion of the dialogue information can be extracted from the dialogue information, the extracted portion of the dialogue information can be used as training data for negative samples, the next sentence after the extracted position in the dialogue information can be used as the standard response information for the training data, and the standard intent of the next sentence after the extracted position can be determined. The standard response information and the standard intent can be saved as labels corresponding to the training data.
[0097] Specifically, it can obtain business dialogue information from the company's regular agents and determine standard intent and standard response information based on this information. In this embodiment, this method can generate negative samples of "non-optimal processes".
[0098] To further improve the accuracy of model training, based on the above embodiments, in this embodiment, if the target training data is a negative sample, before fine-tuning the basic large model based on the target reward value to obtain the dialogue guidance model after determining the target reward value, the method further includes: Determine the opposite of the target reward value, and update the target reward value using the opposite.
[0099] Since the sample set includes both positive and negative samples, the target training data obtained from the sample set may be negative samples.
[0100] If the target training data obtained is a negative sample, in order to make the model training more deviate from the negative sample, in this embodiment of the application, after determining the target reward value, before fine-tuning the basic large model based on the target reward value, the opposite number of the target reward value can be determined, and the target reward value can be updated using the opposite number.
[0101] Example 6: To further improve the effectiveness of model fine-tuning training, based on the above embodiments, in this embodiment, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Obtain the target business conversion result corresponding to the target training data, and the business performance reward value pre-saved for the target business conversion result; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the business performance reward value, wherein the preset weight includes the first weight, the second weight, and the third weight corresponding to the business performance reward value.
[0102] Since customer conversations in related technologies are generally purposeful, such as product promotion or requesting feedback, this application can save the corresponding business conversion result for each conversation. For example, if customer A purchases product A within three days of their conversation with customer service, the business conversion result can be considered "successful." If customer A provides feedback within 24 hours of their conversation, the business conversion result can be determined as "successful." If customer A does not take any follow-up action within three days of their conversation, the business conversion result can be considered "failed."
[0103] To further improve the effectiveness of model fine-tuning training, after determining the predicted response information and prediction intent, and before determining the target reward value, it is also possible to obtain the target business conversion result corresponding to the target training data, as well as the business performance reward value saved for that target business conversion result. For example, if the business conversion result is "failure," then the business performance reward value can be -1; if the business conversion result is "success," then the business performance reward value can be 1.
[0104] After obtaining the business performance reward value, when determining the target reward value based on the first deviation, the second deviation, and the preset weight, the target reward value can be determined according to the first deviation, the second deviation, the preset weight, and the business performance reward value. Since the determination of the target reward value currently also considers business conversion results, in this embodiment, the preset weight may further include the first weight, the second weight, and a third weight corresponding to the business performance reward value. Specifically, the target reward value = first deviation × first weight + second deviation × second weight + business performance reward value × third weight.
[0105] Example 7: To further improve the effectiveness of model fine-tuning training, based on the above embodiments, in this embodiment, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Based on the first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent, determine the predicted intent flow sequence corresponding to the target training data; Among the multiple pre-saved intent transfer processes, find the second business intent transfer process that contains the predicted intent transfer sequence; Determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; The intention flow reward value is determined based on the weight of each connection identifier; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and the fourth weight corresponding to the intent flow reward value.
[0106] In the highly regulated banking industry, the unpredictability of large-scale model outputs poses a substantial risk. The "black box" nature of large-scale models makes it difficult to guarantee strict adherence to critical business processes on every occasion, for example: a) A risk assessment must be completed before recommending investment or financial products; b) Interest rates and fees must be clearly disclosed when introducing loan products; c) Annual fee policies and repayment responsibilities must be fully explained in credit card marketing; d) Existing technology cannot ensure that these mandatory steps are not overlooked. If the output of the large model deviates from compliance requirements, it may lead to customer complaints, regulatory penalties, or even legal disputes, which may seriously affect the bank's reputation and operational stability.
[0107] In particular, the impact of these technological deficiencies is especially pronounced in banking operations. The banking industry has far higher requirements for process standardization, risk control, and compliance than other industries; any deviation from the established procedures could potentially cross regulatory red lines.
[0108] Therefore, during the model training process, the model can be trained to have the ability to analyze the order of intent flow.
[0109] In this embodiment, after determining the predicted response information and the predicted intent, but before determining the target reward value, the predicted intent flow sequence corresponding to the target training data can be determined based on the first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent. Then, a second business intent transfer process containing the predicted intent flow sequence is searched among a pre-saved plurality of intent transfer processes. Next, each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process is determined, and the intent flow reward value is determined based on the weight identified by each connection line. The process of determining the intent flow reward value is consistent with the process of determining the target weight described in the above embodiments, and this embodiment will not elaborate on the process of determining the intent flow reward value.
[0110] After determining the intent flow reward value, the target reward value can be determined based on the first deviation, the second deviation, the preset weight, and the intent flow reward value. Since the determination of the target reward value also considers the correctness of intent flow, in this embodiment, the preset weight may further include the first weight, the second weight, and the fourth weight corresponding to the intent flow reward value. Specifically, the target reward value = first deviation × first weight + second deviation × second weight + intent flow reward value × fourth weight.
[0111] Example 8: To further improve the effectiveness of model fine-tuning training, based on the above embodiments, in this embodiment, the process of determining the preset weights includes: If the current round of fine-tuning training of the dialogue guidance model is in the first round interval, then the basic ability training weight is used as the preset weight, and the maximum weight in the basic ability training weight is the first weight. If the current round of fine-tuning training of the dialogue guidance model is in the second round interval, then the balance ability training weight is used as the preset weight, and the maximum weight in the balance ability training weight is the second weight. If the current training round for fine-tuning the dialogue guidance model is in the third round interval, then the process capability training weight is used as the preset weight, and the maximum weight in the process capability training weight is the fourth weight. The minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
[0112] To further improve the effect of model fine-tuning training, different training focuses can be set at different training stages in this embodiment. Therefore, different weight combinations can be set at different training stages in this embodiment.
[0113] In this embodiment of the application, multiple round intervals can be pre-divided, and different round intervals correspond to different weight combinations.
[0114] In one possible implementation, the training rounds can be divided into three round intervals. The first round interval corresponds to the basic ability training weights, which are configured with a first weight, a second weight, a third weight, and a fourth weight, where the first weight is the maximum weight. For ease of description, this first round interval can be referred to as the first round interval. The second round interval corresponds to the balance ability training weights, which are configured with a first weight, a second weight, a third weight, and a fourth weight, where the second weight is the maximum weight. For ease of description, this second round interval can be referred to as the second round interval. The third round interval corresponds to the process ability training weights, which are configured with a first weight, a second weight, a third weight, and a fourth weight, where the fourth weight is the maximum weight. For ease of description, this third round interval can be referred to as the third round interval.
[0115] In this embodiment, the minimum value between the two endpoints of the third round interval is greater than the maximum value between the two endpoints of the second round interval. The minimum value between the two endpoints of the second round interval is greater than the maximum value between the two endpoints of the first round interval.
[0116] In one possible implementation, if the current round of fine-tuning training of the dialogue guidance model is in the first round interval, the basic ability training weights can be used as preset weights.
[0117] In one possible implementation, if the current round of fine-tuning training of the dialogue guidance model is in the second round interval, the balance ability training weights can be used as preset weights.
[0118] In one possible implementation, if the current round of fine-tuning training of the dialogue guidance model is in the third round range, the process capability training weights can be used as preset weights.
[0119] Specifically, in the early stages of fine-tuning training (e.g., rounds 1-2), the preset weights are: α=0.5, β=0.3, γ=0.2, δ=0.0. Here, α represents the first weight, β the second weight, γ the fourth weight, and δ the third weight. This stage focuses on training the model's fundamental capabilities, such as learning the conversational styles of excellent sales representatives.
[0120] During the mid-stage of fine-tuning training (e.g., rounds 3-4), the preset weights are: α=0.3, β=0.3, γ=0.3, δ=0.1. This stage focuses on the balanced development of the model, such as enhancing the accuracy of intent recognition.
[0121] In the later stages of fine-tuning training (e.g., rounds 5-6), preset weights are used: α=0.2, β=0.2, γ=0.4, δ=0.2. This stage focuses on training the model to focus on business processes, such as deeply optimizing business process decision-making capabilities.
[0122] It should be noted that the above rounds are set only for the convenience of describing the solution, and those skilled in the art can configure them as needed.
[0123] To address the issue of single-objective optimization in traditional fine-tuning methods, this application proposes a multi-objective optimization framework that incorporates Workflow constraints: 1. Layered comparative learning mechanism.
[0124] Construct a three-level negative sample system: responses generated by general models (lacking professionalism), high-quality responses from other business scenarios (scenario mismatch), and non-optimal responses within the same scenario (inappropriate process).
[0125] Through tiered comparison, the model gradually learns from basic language skills to professional business processes.
[0126] 2. Dynamic multi-objective reward function.
[0127] Establish a four-dimensional reward evaluation system: Target reward value = α × Script quality reward + β × Intent accuracy reward + γ × Workflow compliance reward + δ × Business performance reward, where the coefficients adopt a dynamic adjustment strategy: Early training stage: α=0.5, β=0.3, γ=0.2, δ=0.0 (focusing on basic capabilities), Mid-training stage: α=0.3, β=0.3, γ=0.3, δ=0.1 (balanced development), Late training stage: α=0.2, β=0.2, γ=0.4, δ=0.2 (focusing on business processes).
[0128] Sequence-level consistency constraints: Design a dialogue trajectory evaluation algorithm to score the overall path of multi-turn dialogues based on Workflow conformity; introduce a forward-looking reward mechanism to not only evaluate the quality of the current response, but also predict its impact on subsequent dialogue paths.
[0129] Example: When training a bank wealth management recommendation model, when a customer asks about "principal-protected products", the model only receives a Workflow reward of 0.3 for the response "directly recommend money market funds" (because it skips the risk assessment step), while the response "conduct a risk assessment first and then recommend suitable products" receives a high reward of 0.9, effectively guiding the model to follow compliance procedures.
[0130] Example 9: To further improve the accuracy of question and answer, based on the above embodiments, in this embodiment, the process of determining the business intent transfer flow includes: Obtain any historical call text and the second intent corresponding to each sentence in that historical call text; The historical call text and the second intent are processed using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; The weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence are determined based on a preset algorithm, and the business intent transfer process is obtained.
[0131] To construct a business intent transfer process that conforms to actual business conditions, this embodiment of the application can obtain any historical call text and the second intent corresponding to each sentence in the historical call text. The historical call text can be written by someone skilled in the art, or it can be text corresponding to call content authorized and agreed upon by the customer and saved from historical work. The second intent can be determined by a pre-trained model or it can be manually annotated.
[0132] In this embodiment, a sequence pattern mining algorithm can be used to process the historical call text and the second intent to obtain the intent transition sequence corresponding to the historical call text. The sequence pattern mining algorithm can be a prefix-based sequence pattern mining algorithm (PrefixSpan algorithm).
[0133] After obtaining the intent transfer sequence corresponding to the historical call text output by the sequence pattern mining algorithm, the weights of the connecting lines between adjacent nodes in the intent transfer sequence can be determined based on a preset algorithm to obtain the business intent transfer process.
[0134] Specifically, the weight of the connection line between any two adjacent nodes can be determined based on the following formula:
[0135] 1. Indicates from node To the node The weights corresponding to the connecting lines between them.
[0136] 2. : Represents the success rate of transactions, i.e., the percentage of transactions completed after a certain point in the historical data. The empirical probability of a final transaction. Definition: .
[0137] Calculation method: Let For Given the set of all leaf nodes starting from the origin, then:
[0138] in: The leaf nodes are marked with transaction indicators (1 = completed, 0 = not completed). Therefore, the success rate gain can be calculated as follows: .
[0139] The non-negativity constraint in the above formula: If positive, the transition is a positive decision. If zero, the transition is a deteriorating path (which can be used for backpropagation). In the above formula, The value is or In other words, This indicates that in historical data, the nodes passed through... The empirical probability of the final transaction; This indicates that in historical data, the nodes passed through... The empirical probability of the final transaction.
[0140] 3. : Indicates the shift in confidence level. Reflected in historical dialogues, from Transferred to Statistical stability and confidence level.
[0141] definition: in, Indicates the number of times this transition has occurred in history; the denominator is from The total number of all transfers starting from the origin.
[0142] 4. : Represents the importance weight of a node. It characterizes the business importance of a node in the entire dialogue; the higher the frequency of occurrence, the more critical it is. Therefore, in the above formula... This indicates a node. Importance weights; This indicates a node. Importance weights.
[0143] Normalization determination method:
[0144] Value range: .
[0145] In the above formula This indicates the joint importance term. In this embodiment, the geometric mean is used to prevent edge weight distortion caused by an excessively high frequency of a single end node.
[0146] 5. : Indicates the shortest distance to the transaction. Node The shortest hop count from the most recent successfully traded leaf node. Then, in the above formula... This indicates a node. The shortest number of hops from the leaf node to the most recent successful transaction; This indicates a node. The shortest number of hops from the leaf node to the most recent successful transaction.
[0147] definition: .
[0148] In the above formula Then, it represents the time-series approximation factor, used to measure the approximation from... Transferred to Is it closer to a final deal in terms of time and steps? Definition: ,in, Attenuation coefficient (empirically taken as 0.3–0.7); if If the value of this item is greater than 1, it indicates that the transaction is clearly approaching; if If the value is less than 1, then the value of that item is less than 1.
[0149] To facilitate understanding of the above weight determination formula, Table 1 below explains the physical meaning of the relevant parameters.
[0150] Table 1
[0151] To further improve the accuracy of question and answer, based on the above embodiments, in this application embodiment, the process of constructing the historical call text includes: Retrieve a preset amount of basic call text; Based on the role, sentences in the preset number of basic call texts are classified to obtain a sentence set corresponding to each role; For each set of sentences, clustering is performed based on the semantics of each sentence in the set to obtain sentence groups; Obtain a business dictionary pre-saved for the target business domain applied to the script guidance model, and a suggested intent saved for each keyword in the business dictionary; For each sentence group, the business dictionary, the suggested intent corresponding to each keyword, and the sentence group are input into the second large model, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword. The target base call text to which each sentence in each sentence group belongs is determined, and the corresponding second intent is identified for the corresponding sentence in the target base call text to obtain the historical call text.
[0152] To construct historical call transcripts, a preset number of basic call transcripts can be obtained in this embodiment. For example, dialogue transcripts can be obtained from 50,000 historical call recordings accumulated by an insurance company using speech-to-text technology, and these transcripts can then be used as the basic call transcripts.
[0153] In one possible implementation, after obtaining the basic call text, dialogues with clear structure and complete interaction rounds can be automatically filtered out, and the final result tags (sold / unsold, customer satisfaction rating) can be obtained by associating with the business system.
[0154] Specifically, when filtering dialogues with clear structure and complete interaction rounds, at least one of the following strategies can be used: The basic call text includes the dialogue between the two characters; The number of dialogue rounds between the two characters in the basic call text is no less than a preset threshold; The length of the content expressed by any character in the basic call text is no less than the threshold.
[0155] After obtaining the final basic call transcripts, each sentence in each basic call transcript can be categorized based on the role, resulting in a sentence set corresponding to each role. In other words, sentences expressed by the same role in all basic call transcripts are stored in the same sentence set.
[0156] After obtaining the sentence set corresponding to each role, we can perform clustering based on the semantics of each sentence in each sentence set to obtain sentence groups. In other words, we analyze the semantics of each sentence in the sentence set corresponding to a certain role and group sentences with similar semantics into a sentence group.
[0157] For example, unsupervised clustering methods can be used to automatically discover the core intent in a dialogue. First, the BAAI General Embedding - Multi-Functionality, Multi-Linguality, Multi-Granularity (BGE-M3) model is used to convert each round of dialogue into a semantic vector, and then the DBSCAN clustering algorithm is used to identify semantically similar sentences.
[0158] To determine the secondary intent corresponding to each sentence in each sentence group, after obtaining each sentence group, we can retrieve a business dictionary pre-stored for the target business domain applied to the dialogue guidance model, as well as suggested intents stored for each keyword in that business dictionary. In other words, a pre-stored correspondence between different keywords and suggested intents is established. Subsequently, the secondary intent corresponding to each sentence can be determined based on this correspondence.
[0159] In this embodiment, for each sentence group, the business dictionary, the suggested intent corresponding to each keyword, and each sentence included in the sentence group can be input into the second large model. When the second large model determines that the sentence group contains a target keyword from the business dictionary, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword.
[0160] Specifically, by combining the business dictionary of the insurance industry (including professional terms such as "premium," "coverage," and "claims"), the secondary intent corresponding to each sentence group can be determined, and intent tags with business meaning can be automatically generated, such as "demand exploration - price," "product recommendation - critical illness insurance," and "objection handling - insufficient budget." It should be noted that those skilled in the art can configure the intents that can be included as needed.
[0161] For example, when a customer says "How much does this insurance cost?", it is clustered into the "Needs Exploration - Price" intent, and the agent's reply "The annual premium is 2,000 yuan. Are you more concerned about the coverage or your budget?" is identified as "Product Recommendation - Price Inquiry and Budget Inquiry".
[0162] After obtaining the second intent corresponding to each sentence group, the target base call text to which each sentence in each sentence group belongs can be determined. The corresponding second intent is then identified for the corresponding sentence in the target base call text to obtain the historical call text.
[0163] In other words, when determining intent, all sentences in the basic call text are broken down and clustered. Sentences with similar semantics are grouped into the same sentence group, and a secondary intent corresponding to each sentence group is determined. Then, all the broken-down sentences are regressed and mapped back to the original basic call text, thereby identifying the basic call text with the secondary intent as the historical call text.
[0164] To improve the efficiency of constructing historical call texts, based on the above embodiments, in this embodiment, after obtaining the sentence groups and before obtaining the business dictionary pre-saved for the target business domain applied to the dialogue guidance model, the method further includes: Count the number of sentences in each sentence group, and delete sentence groups whose number is less than a threshold.
[0165] Since the content expressed by each person is uncontrollable, some people may express niche content that is not relevant to model training. Therefore, in this embodiment, after obtaining each sentence group, before acquiring the business dictionary pre-saved for the target business domain of the dialogue guidance model, the number of sentences included in each sentence can be counted, and sentence groups with a number less than a threshold can be deleted. For example, this threshold can be 10, 5, etc., and those skilled in the art can set it based on the amount of basic call text obtained.
[0166] In this application embodiment, addressing the issue of traditional methods relying on large amounts of labeled data, a workflow automatic mining method integrating weakly supervised learning and business rule guidance is proposed. This method achieves a seamless transformation from raw dialogue to a business process decision tree through three levels of automated processing: 1. Automatic intent discovery and labeling.
[0167] An unsupervised method based on semantic clustering is employed to automatically discover the core intent categories in dialogues. First, dialogue turns are converted into vector representations using a text embedding model. Then, a clustering algorithm is used to identify semantically similar dialogue segments, automatically forming initial intent categories.
[0168] By combining business dictionaries and existing communication standards, weakly supervised signals are constructed to calibrate and name the clustering results, forming an intent labeling system with business meaning, which significantly reduces the reliance on manual annotation.
[0169] 2. Automated abstraction of dialogue sequences.
[0170] An algorithm based on sequence pattern mining is designed to automatically identify key state transition points in a dialogue. By analyzing the semantic coherence and business logic connections between adjacent dialogue rounds, a complete intent transition sequence is constructed.
[0171] Introduce a business goal-oriented sequence filtering mechanism to retain only intent transfer paths that have a significant impact on the final business outcome, ensuring the simplicity and usability of the Workflow.
[0172] 3. Intelligent weighting of success paths.
[0173] Establish a multi-dimensional business performance evaluation system that comprehensively considers multiple indicators such as transaction completion rate, customer satisfaction, and process compliance, and assigns differentiated weights to different success paths.
[0174] An innovative time-decay weight allocation mechanism is designed, where decision steps closer to business objectives receive higher weights, ensuring that the workflow accurately reflects the importance of key decision points.
[0175] Example: In a bank's credit card marketing scenario, the system automatically extracted core business paths from 30,000 historical recordings. Analysis revealed that when customers mentioned "annual fees," the success weight of the path taken by top-performing account managers—"explaining the annual fee waiver policy → emphasizing card benefits → guiding immediate application"—reached 0.88, while the path directly leading to the "guiding application" stage had a weight of only 0.45. This finding provided data support for model optimization, significantly improving marketing conversion rates.
[0176] Example 10: The question-and-answer process will be explained below with reference to a specific example. Figure 2 A question-and-answer process diagram provided for an embodiment of this application, such as Figure 2 As shown, the process comprises an end-to-end workflow with three core stages. The following explanation uses an insurance sales case study that runs throughout the process to illustrate the complete technological loop from raw data to intelligent applications.
[0177] Case Background: A large insurance company wanted to improve the sales conversion rate and professionalism of its intelligent customer service system. The company had accumulated nearly a year's worth of historical call data, but its existing dialogue system, based on a general large model, had a conversion rate of only 9.2% and suffered from problems such as non-standard scripts and chaotic processes.
[0178] Phase 1: Automatic Workflow Mining Based on Weakly Supervised Learning.
[0179] The core objective of this stage is to automatically extract the optimal business process from the raw dialogue data, without relying on a large amount of manual annotation.
[0180] First, a preset amount of historical dialogue data is acquired, and the intent of each sentence is labeled based on unsupervised clustering and weakly supervised intent recognition.
[0181] Next, automatic abstraction of intent transfer is achieved based on sequence pattern mining algorithms: high-frequency intent transfer paths are automatically extracted using sequence pattern mining algorithms (such as PrefixSpan). For example, a typical path for a top agent is: "Opening greeting → Needs assessment → Product recommendation → Objection handling → Conversion facilitation".
[0182] Next, intelligent weighting of the success path is implemented based on time-series decay and multi-dimensional features. An intelligent weighting mechanism based on business outcomes is introduced: for dialogues leading to a final transaction, the path receives a base score, and a time-series decay factor is applied, resulting in higher weights for key steps closer to the transaction.
[0183] Finally, a structured Workflow decision tree is obtained.
[0184] Example Output: The generated Workflow graph clearly shows that the edge weight from the "Customer raises price objection" node to "Agree on feelings and explore budget" is as high as 0.87, while the path weight directly to "Offer discount" is only 0.52. This quantitative result reveals the secret to the success of excellent agents.
[0185] Phase 2: Multi-objective DPO training guided by a defined workflow.
[0186] The goal of this phase is to train a lightweight bootstrapping model that can both understand the business and make optimal decisions.
[0187] In this stage, a sample set including positive and negative samples is constructed by using a hierarchical sample construction method.
[0188] The model is then trained based on the training data included in the sample set. The target reward value is determined through a multi-objective reward function with designed dynamic weights: Target reward value = α × Style reward + β × Intent reward + γ × Workflow reward + δ × Result reward. Here, the style reward is the first bias described in the above embodiments, the intent reward is the second bias described in the above embodiments, the Workflow reward is the intent flow reward value described in the above embodiments, and the result reward is the business performance reward value described in the above embodiments.
[0189] Once the target reward value is determined, the basic large model can be fine-tuned based on this target reward value. If the trained model does not meet the convergence condition, a model training step based on the training data included in the sample set can be performed. If the trained model meets the convergence condition, a trained lightweight dialogue guidance model can be obtained.
[0190] Performance verification: After training, the workflow compliance rate of the dialogue guidance model on the test set increased significantly from 0.42 to 0.78, and the estimated conversion rate of simulated dialogues increased from 35% to 61%, demonstrating the significant improvement of the method in business decision-making capabilities.
[0191] The third stage involves lightweight service-oriented deployment and collaborative applications.
[0192] This stage encapsulates the trained guidance capabilities into standardized services, enabling collaborative work with large models.
[0193] Service-oriented deployment. The finely tuned 800 million parameter Qwen3 model and Workflow graph are packaged together as an MCP service. This business guidance tool is deployed on a dedicated inference engine, providing business guidance capabilities with millisecond-level response.
[0194] Collaborative workflow: Step 1: The customer sends a message to the cloud-based big data model connected to this system: "This critical illness insurance is too expensive." Step 2: The large language model (the first large model mentioned above) immediately calls the guide_next_response tool of the MCP service, taking the current dialogue history and the latest customer message as input.
[0195] Step 3: The dialogue guidance model in the MCP service completes inference within 100 milliseconds: it identifies the current business node as "price objection", queries the Workflow graph, and returns guidance suggestions: "Recommended next step intention: acknowledge feelings and explore budget. Weight: 0.87."
[0196] Step 4: Based on this precise business guidance and its superior language generation capabilities, the large language model outputs the final response: "I understand your concerns; insurance is indeed a long-term investment. That's why we should ensure that every penny of premium is spent wisely. To recommend the most suitable plan for you, would you like to know your approximate premium budget?"
[0197] The implementation of this complete solution enabled the insurance company to achieve a qualitative leap in the business capabilities of its intelligent customer service system without significantly increasing computing costs.
[0198] To adapt to different resources and conditions, three alternative solutions are also proposed in the embodiments of this application: Option 1: Rule-enhanced Workflow Solution. Suitable for scenarios with small amounts of data. The core workflow skeleton is manually drawn by business experts, and rule violation penalties are added during DPO training. Advantages include fast startup and strong compliance; disadvantages include poor flexibility and reliance on expert experience.
[0199] Option 2: Lightweight Intent Routing Solution. Suitable for edge environments with extremely limited computing resources. It uses only a very small model for intent recognition and subsequent intent recommendation, while a larger model handles the dialogue generation. Advantages include extremely low latency and low resource consumption; disadvantages include weaker business control capabilities.
[0200] Option 3: Enhanced Example-Based Solution. Suitable for teams without model training capabilities. This involves building a high-quality dialogue example library, matching similar scenarios through vector retrieval, and using high-weighted historical excellent responses as examples to fill in the prompt words. The advantages are zero training and ease of implementation; the disadvantages are limited generalization ability and constraint by the length of the prompt words.
[0201] Recommendations for choosing a solution: Use the optimal solution when data is plentiful; use solution one when data is limited or in scenarios with strong compliance requirements; use solution two in environments with demanding computing power; and use solution three when rapidly validating concepts.
[0202] The question-and-answer solution provided in this application has yielded the following significant results in real-world applications: Result 1: Significant improvement in business processes and conversion rates. Workflow compliance rate increased from 35% to 81%, sales conversion rate increased from 9.2% to 21.8%, and accuracy in handling key milestones improved by over 100%.
[0203] Effect 2: Significantly reduced resource costs. Model storage usage is reduced by approximately 95%, from 78GB in traditional multi-model deployments to 3.5GB; training costs are reduced by over 80%; and scene switching time is reduced from minutes to seconds.
[0204] Effect 3: Enhanced deployment agility and flexibility. The launch cycle for new business scenarios is shortened by approximately 40%, and rapid switching between multiple scenarios and A / B testing can be achieved.
[0205] Benefit 4: Provides explainability and compliance assurance. Each conversation's workflow path is traceable, facilitating business review and optimization. By setting mandatory nodes (such as "risk warnings") in the workflow graph, the script violation rate decreased from 12% to 2%.
[0206] Effect 5: Achieving data-driven continuous optimization. The system automatically updates the Workflow weekly from new data, discovering tacit knowledge such as "exploring the budget first and then discussing value" having a 31% higher conversion rate than "directly emphasizing value," driving continuous business growth.
[0207] The question-answering method provided in this application primarily enhances the professional communication skills and business process control capabilities of large language models in dialogue scenarios such as customer service and sales by automatically mining business process decision trees from historical high-performing dialogue data and training small-scale models using an innovative, workflow-guided direct preference optimization method. This is achieved through a system deployed as a model context protocol service, enabling general-purpose large models to quickly acquire topic guidance and process control capabilities for specific business scenarios without the need for costly fine-tuning of the massive models with hundreds of billions of parameters. This application addresses the issues of unprofessional communication skills and weak process control in large models within vertical business scenarios, avoiding the high costs and deployment inconvenience associated with fine-tuning massive models with hundreds of billions of parameters.
[0208] In the question-and-answer process provided in this application embodiment, the accuracy of Workflow mining depends on data quality. To address this issue, multi-dimensional success metrics (conversion rate, satisfaction, etc.) can be used to comprehensively evaluate dialogue quality; frequency thresholds can be set to filter noisy paths; a manual verification mechanism by business experts can be established; and continuous iterative updates of the Workflow can be achieved.
[0209] In the question-and-answer process provided in this application embodiment, an uncertainty detection mechanism can be introduced to inform the first major model to refer to it with caution when the confidence level is low; a fallback strategy for the first major model can be designed; and differentiated guiding weights can be adopted for high-frequency and low-frequency scenarios.
[0210] In the question-and-answer process provided in this application embodiment, there may be a problem where the Workflow is too rigid, leading to an unnatural dialogue. To solve this problem, in this application embodiment, the Workflow can be treated as a "soft constraint" rather than a "hard rule"; the Workflow weight is dynamically adjusted according to the dialogue context; multiple high-quality paths are reserved in the graph for the model to choose from; and a real-time human intervention interface is provided.
[0211] Example 11: Based on the same inventive concept, embodiments of this application provide a question-and-answer device. Figure 3 This application provides a schematic diagram of a question-and-answer device structure, which includes: The guidance module 301 is used to, upon receiving text information input by a customer, use the first model to call the business guidance tool to analyze the text information and obtain guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommendation response information corresponding to the recommendation intent. The recommendation intent and the recommendation response information are determined by the trained dialogue guidance model deployed in the business guidance tool, based on the content included in the text information and the target intent of the text information. The response module 302 is used by the first large model to generate target response information corresponding to the text information based on the guidance suggestion.
[0212] In one possible implementation, the device further includes: The acquisition module 303 is used to acquire historical dialogue information prior to the text information; to construct dialogue information using the text information and the historical dialogue information; and to update the text information using the dialogue information.
[0213] In one possible implementation, the business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines. The connecting lines are used to describe the flow order between the corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order. The guidance module 301 is specifically used to determine the target intent flow sequence of the target intent and the recommended intent; find a first business intent transfer process containing the target intent flow sequence among the multiple business intent transfer processes; determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; determine the target weight according to the weight of each connection line identifier and add it to the guidance suggestion for output, wherein the target weight is used to describe the probability that responding according to the recommended intent can promote business success.
[0214] In one possible implementation, the device further includes: Training module 304 is used to acquire target training data in a sample set and labels corresponding to the target training data; the target training data includes dialogue information and a first intent corresponding to at least one sentence in the dialogue information; the labels are used to identify standard response information made in response to the last sentence in the dialogue information and the standard intent of the standard response information; the target training data is input into a basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data; a target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and preset weights, wherein the preset weights include a first weight corresponding to the first deviation and a second weight corresponding to the second deviation; the basic large model is fine-tuned based on the target reward value to obtain the dialogue guidance model.
[0215] In one possible implementation, the sample set includes positive samples. The acquisition module 303 is further configured to: identify outstanding sales personnel based on the performance data of each sales personnel; acquire a preset number of business dialogue messages from the historical dialogue information corresponding to the outstanding sales personnel; for each business dialogue message, extract a portion of the dialogue information from the business dialogue message, determine the extracted portion of the dialogue information as training data for the positive samples, use the next sentence at the extracted position in the business dialogue message as the standard response information for the training data, and determine the standard intent of the next sentence at the extracted position.
[0216] In one possible implementation, the sample set further includes negative samples. The acquisition module 303 is further configured to acquire sample dialogue information and generate standard response information and standard intent of the sample dialogue information using a general large model. The sample dialogue information is used as training data for negative samples, and the standard response information and standard intent are saved as labels corresponding to the sample dialogue information. Alternatively, the target business domain of the dialogue guidance model is determined, business dialogue information in non-target business domains is acquired, a portion of the dialogue information in the non-target business domains is extracted, and the extracted portion of the dialogue information is used as training data for negative samples. The extracted position in the business dialogue information in the non-target business domains is further defined. The next sentence of the extracted position is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data. Alternatively, based on the performance data of each business person, non-excellent business people are identified. Business dialogue information of the non-excellent business people is obtained, a portion of the dialogue information is extracted from the business dialogue information, the extracted portion of the dialogue information is determined as training data for negative samples, the next sentence of the extracted position in the business dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data.
[0217] In one possible implementation, if the target training data is a negative sample, the training module 304 is further configured to determine the opposite of the target reward value and update the target reward value using the opposite.
[0218] In one possible implementation, the training module 304 is further configured to acquire the target business conversion result corresponding to the target training data, and the business performance reward value pre-saved for the target business conversion result; and determine the target reward value based on the first deviation, the second deviation, the preset weight, and the business performance reward value, wherein the preset weight includes the first weight, the second weight, and the third weight corresponding to the business performance reward value.
[0219] In one possible implementation, the training module 304 is further configured to: determine a predicted intent flow sequence corresponding to the target training data based on a first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent; search for a second business intent transfer process containing the predicted intent flow sequence among a plurality of pre-saved intent transfer processes; determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; determine an intent flow reward value based on the weight identified by each connection line; and determine a target reward value based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and a fourth weight corresponding to the intent flow reward value.
[0220] In one possible implementation, the training module 304 is specifically configured to: if the current round of fine-tuning training of the dialogue guidance model is in the first round interval, use the basic ability training weight as the preset weight, where the maximum weight in the basic ability training weight is the first weight; if the current round of fine-tuning training of the dialogue guidance model is in the second round interval, use the balanced ability training weight as the preset weight, where the maximum weight in the balanced ability training weight is the second weight; if the current round of fine-tuning training of the dialogue guidance model is in the third round interval, use the process ability training weight as the preset weight, where the maximum weight in the process ability training weight is the fourth weight, wherein the minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
[0221] In one possible implementation, the device further includes: The determination module 305 is used to obtain any historical call text and the second intent corresponding to each sentence in the historical call text; process the historical call text and the second intent using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; determine the weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence based on a preset algorithm to obtain the business intent transfer process.
[0222] In one possible implementation, the determining module 305 is further configured to: acquire a preset number of basic call texts; classify sentences in the preset number of basic call texts based on roles to obtain a sentence set corresponding to each role; perform clustering processing on the semantics of each sentence in each sentence set to obtain sentence groups; acquire a business dictionary pre-saved for the target business domain applied to the dialogue guidance model, and a suggested intent saved for each keyword in the business dictionary; input the business dictionary, the suggested intent corresponding to each keyword, and the sentence group into a second large model for each sentence group, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword; determine the target basic call text to which each sentence in each sentence group belongs, and identify the corresponding second intent for the corresponding sentence in the target basic call text to obtain historical call texts.
[0223] In one possible implementation, the determining module 305 is further configured to count the number of sentences included in each sentence group and delete sentence groups whose number is less than a threshold.
[0224] Example 12: Based on the same inventive concept, embodiments of this application provide an electronic device that can implement the steps of the question-and-answer method described above. Figure 4 This application provides a schematic diagram of an electronic device structure, such as... Figure 4 As shown, it includes: processor 401, communication interface 402, memory 403 and communication bus 404, wherein processor 401, communication interface 402 and memory 403 communicate with each other through communication bus 404. The memory 403 stores a computer program. When the program is executed by the processor 401, the processor 401 performs the following steps: If a text message is received from a customer, the first major model is used to call the business guidance tool to analyze the text message and obtain the guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommendation response information corresponding to the recommendation intent. The recommendation intent and the recommendation response information are determined by the trained dialogue guidance model deployed in the business guidance tool, based on the content of the text message and the target intent of the text message. Based on the guidance suggestions, the first large model generates target response information corresponding to the text information.
[0225] In one possible implementation, before analyzing the text information input by the customer using the first model to invoke the business guidance tool after receiving the text information, the method further includes: Obtain the historical dialogue information prior to the text information; The text information and the historical dialogue information are used to construct dialogue information, and the dialogue information is used to update the text information.
[0226] In one possible implementation, the business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and the nodes are connected by connecting lines. The connecting lines are used to describe the flow order between the corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order. The process by which the business guidance tool determines the guidance suggestion also includes: Determine the target intent flow sequence of the target intent and the recommended intent; Among the multiple business intent transfer processes, find the first business intent transfer process that contains the target intent flow sequence; Determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; Based on the weight of each connector identifier, a target weight is determined and added to the output of the guidance suggestion, wherein the target weight is used to describe the probability that responding in accordance with the recommendation intent will lead to business success.
[0227] In one possible implementation, the fine-tuning training process of the dialogue guidance model includes: Obtain target training data from the sample set, and the labels corresponding to the target training data; the target training data includes dialogue information, and a first intent corresponding to at least one sentence in the dialogue information; the labels are used to identify standard response information made in response to the last sentence in the dialogue information, and the standard intent of the standard response information; The target training data is input into the basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data. A target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight. The preset weight includes a first weight corresponding to the first deviation and a second weight corresponding to the second deviation. The basic large model is fine-tuned and trained based on the target reward value to obtain the speech guidance model.
[0228] In one possible implementation, the sample set includes positive samples, and the process of obtaining the positive samples and labels includes: Based on the performance data of each salesperson, outstanding salespersons are identified. Obtain a preset number of business dialogue messages from the historical dialogue information corresponding to the outstanding business personnel; For each of the aforementioned business dialogue messages, a portion of the dialogue information is extracted from the business dialogue message. The extracted portion of the dialogue information is determined as training data for positive samples. The next sentence at the extracted position in the business dialogue message is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined.
[0229] In one possible implementation, the sample set further includes negative samples, and the process of obtaining the negative samples and labels includes at least one of the following: Acquire sample dialogue information and use a general large model to generate standard response information and standard intent for the sample dialogue information. Use the sample dialogue information as training data for negative samples, and save the standard response information and standard intent as labels corresponding to the sample dialogue information; or Determine the target business domain for which the dialogue guidance model is applied, acquire business dialogue information from non-target business domains, extract a portion of the dialogue information from the non-target business domains, use this extracted portion as training data for negative samples, take the next sentence at the extracted position in the non-target business domains as the standard response information for this training data, determine the standard intent of the next sentence at the extracted position, and save the standard response information and the standard intent as labels corresponding to the training data; or Based on the performance data of each salesperson, non-performing salespersons are identified; sales dialogue information of the non-performing salespersons is obtained, a portion of the dialogue information is extracted, and the extracted portion of the dialogue information is used as training data for negative samples. The next sentence after the extracted position in the dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence after the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data.
[0230] In one possible implementation, if the target training data is negative samples, after determining the target reward value and before fine-tuning the basic large model based on the target reward value to obtain the dialogue guidance model, the method further includes: Determine the opposite of the target reward value, and update the target reward value using the opposite.
[0231] In one possible implementation, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight, the method further includes: Obtain the target business conversion result corresponding to the target training data, and the business performance reward value pre-saved for the target business conversion result; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the business performance reward value, wherein the preset weight includes the first weight, the second weight, and the third weight corresponding to the business performance reward value.
[0232] In one possible implementation, after determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight, the method further includes: Based on the first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent, determine the predicted intent flow sequence corresponding to the target training data; Among the multiple pre-saved intent transfer processes, find the second business intent transfer process that contains the predicted intent transfer sequence; Determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; The intention flow reward value is determined based on the weight of each connection identifier; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and the fourth weight corresponding to the intent flow reward value.
[0233] In one possible implementation, the process of determining the preset weights includes: If the current round of fine-tuning training of the dialogue guidance model is in the first round interval, then the basic ability training weight is used as the preset weight, and the maximum weight in the basic ability training weight is the first weight. If the current round of fine-tuning training of the dialogue guidance model is in the second round interval, then the balance ability training weight is used as the preset weight, and the maximum weight in the balance ability training weight is the second weight. If the current training round for fine-tuning the dialogue guidance model is in the third round interval, then the process capability training weight is used as the preset weight, and the maximum weight in the process capability training weight is the fourth weight. The minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
[0234] In one possible implementation, the process of determining the business intent transfer flow includes: Obtain any historical call text and the second intent corresponding to each sentence in that historical call text; The historical call text and the second intent are processed using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; The weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence are determined based on a preset algorithm, and the business intent transfer process is obtained.
[0235] In one possible implementation, the process of constructing the historical call text includes: Retrieve a preset amount of basic call text; Based on the role, sentences in the preset number of basic call texts are classified to obtain a sentence set corresponding to each role; For each set of sentences, clustering is performed based on the semantics of each sentence in the set to obtain sentence groups; Obtain a business dictionary pre-saved for the target business domain applied to the script guidance model, and a suggested intent saved for each keyword in the business dictionary; For each sentence group, the business dictionary, the suggested intent corresponding to each keyword, and the sentence group are input into the second large model, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword. The target base call text to which each sentence in each sentence group belongs is determined, and the corresponding second intent is identified for the corresponding sentence in the target base call text to obtain the historical call text.
[0236] In one possible implementation, after obtaining the sentence group and before acquiring the business dictionary pre-saved for the target business domain applied to the speech guidance model, the method further includes: Count the number of sentences in each sentence group, and delete sentence groups whose number is less than a threshold.
[0237] Since the principle of the above-mentioned electronic device in solving problems is similar to that of the question-and-answer method, the implementation of the above-mentioned electronic device can be found in the embodiments of the method, and repeated parts will not be described again.
[0238] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 402 is used for communication between the above-mentioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.
[0239] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0240] Example 13: Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to execute the steps of the question-and-answer method described above.
[0241] Example 14: Based on the same inventive concept, this application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute any of the target detection methods discussed above. Since the principle of the above-described computer program product in solving the problem is similar to that of the question-and-answer method, the implementation of the above-described computer program product can be referred to the implementation of the method, and repeated details will not be repeated.
[0242] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0243] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0244] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0245] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of user-operated steps to be executed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0246] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A question-and-answer method, characterized in that, The method includes: If a text message is received from a customer, the first major model is used to call the business guidance tool to analyze the text message and obtain the guidance suggestions output by the business guidance tool. The guidance suggestions include the recommendation intent and the recommendation response information corresponding to the recommendation intent. The recommendation intent and the recommendation response information are determined by the trained dialogue guidance model deployed in the business guidance tool, based on the content of the text message and the target intent of the text message. Based on the guidance suggestions, the first large model generates target response information corresponding to the text information.
2. The method according to claim 1, characterized in that, Before analyzing the text information input by the customer using the first major model and the business guidance tool after receiving the text information, the method further includes: Obtain the historical dialogue information prior to the text information; The text information and the historical dialogue information are used to construct dialogue information, and the dialogue information is used to update the text information.
3. The method according to claim 1, characterized in that, The business guidance tool also encapsulates multiple business intent transfer processes. Each business intent transfer process includes multiple nodes, each node is used to describe a business intent, and nodes are connected by connecting lines. The connecting lines are used to describe the flow order between corresponding business intents, and each connecting line is marked with a weight. The weight is used to describe the probability of successfully conducting a question-and-answer dialogue according to the corresponding flow order. The process by which the business guidance tool determines the guidance suggestion also includes: Determine the target intent flow sequence of the target intent and the recommended intent; Among the multiple business intent transfer processes, find the first business intent transfer process that contains the target intent flow sequence; Determine each connection line corresponding to the target intent flow sequence in the first business intent transfer process; Based on the weight of each connector identifier, a target weight is determined and added to the output of the guidance suggestion, wherein the target weight is used to describe the probability that responding in accordance with the recommendation intent will lead to business success.
4. The method according to claim 1, characterized in that, The fine-tuning training process of the dialogue guidance model includes: Obtain target training data from the sample set, and the labels corresponding to the target training data; the target training data includes dialogue information, and a first intent corresponding to at least one sentence in the dialogue information; the labels are used to identify standard response information made in response to the last sentence in the dialogue information, and the standard intent of the standard response information; The target training data is input into the basic large model, which analyzes the dialogue information and the first intent corresponding to at least one sentence in the dialogue information to determine the predicted response information and the predicted intent for responding to the target training data. A target reward value is determined based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight. The preset weight includes a first weight corresponding to the first deviation and a second weight corresponding to the second deviation. The basic large model is fine-tuned and trained based on the target reward value to obtain the speech guidance model.
5. The method according to claim 4, characterized in that, The sample set includes positive samples, and the process of obtaining the positive samples and labels includes: Based on the performance data of each salesperson, outstanding salespersons are identified; A preset number of business dialogue messages are obtained from the historical dialogue information corresponding to the outstanding business personnel. For each of the aforementioned business dialogue messages, a portion of the dialogue information is extracted from the business dialogue message. The extracted portion of the dialogue information is determined as training data for positive samples. The next sentence at the extracted position in the business dialogue message is used as the standard response information for the training data, and the standard intent of the next sentence at the extracted position is determined.
6. The method according to claim 5, characterized in that, The sample set also includes negative samples, and the process of obtaining the negative samples and labels includes at least one of the following: Obtain sample dialogue information, and use a general large model to generate standard response information and standard intent for the sample dialogue information. Use the sample dialogue information as training data for negative samples, and save the standard response information and standard intent as labels corresponding to the sample dialogue information. or Determine the target business domain for which the dialogue guidance model is applied, acquire business dialogue information from non-target business domains, extract a portion of the dialogue information from the non-target business domains, use this extracted portion as training data for negative samples, take the next sentence at the extracted position in the non-target business domains as the standard response information for this training data, determine the standard intent of the next sentence at the extracted position, and save the standard response information and the standard intent as labels corresponding to the training data; or Based on the performance data of each salesperson, non-performing salespersons are identified; sales dialogue information of the non-performing salespersons is obtained, a portion of the dialogue information is extracted, and the extracted portion of the dialogue information is used as training data for negative samples. The next sentence after the extracted position in the dialogue information is used as the standard response information for the training data, and the standard intent of the next sentence after the extracted position is determined. The standard response information and the standard intent are saved as labels corresponding to the training data.
7. The method according to claim 6, characterized in that, If the target training data is negative samples, after determining the target reward value and before fine-tuning the basic large model based on the target reward value to obtain the dialogue guidance model, the method further includes: Determine the opposite of the target reward value, and update the target reward value using the opposite.
8. The method according to claim 4, characterized in that, After determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Obtain the target business conversion result corresponding to the target training data, and the business performance reward value pre-saved for the target business conversion result; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the business performance reward value, wherein the preset weight includes the first weight, the second weight, and the third weight corresponding to the business performance reward value.
9. The method according to claim 4, characterized in that, After determining the predicted response information and predicted intent for responding to the target training data, and before determining the target reward value based on the first deviation between the predicted response information and the standard response information, the second deviation between the predicted intent and the standard intent, and the preset weights, the method further includes: Based on the first intent corresponding to at least one dialogue in the dialogue information identified in the target training data and the predicted intent, determine the predicted intent flow sequence corresponding to the target training data; Among the multiple pre-saved intent transfer processes, find the second business intent transfer process that contains the predicted intent transfer sequence; Determine each connection line corresponding to the predicted intent flow sequence in the second business intent transfer process; The intention flow reward value is determined based on the weight of each connection identifier; The step of determining the target reward value based on a first deviation between the predicted response information and the standard response information, a second deviation between the predicted intent and the standard intent, and a preset weight includes: A target reward value is determined based on the first deviation, the second deviation, the preset weight, and the intent flow reward value, wherein the preset weight includes the first weight, the second weight, and the fourth weight corresponding to the intent flow reward value.
10. The method according to any one of claims 4, 8, and 9, characterized in that, The process of determining the preset weights includes: If the current round of fine-tuning training of the dialogue guidance model is in the first round interval, then the basic ability training weight is used as the preset weight, and the maximum weight in the basic ability training weight is the first weight. If the current round of fine-tuning training of the dialogue guidance model is in the second round interval, then the balance ability training weight is used as the preset weight, and the maximum weight in the balance ability training weight is the second weight. If the current training round for fine-tuning the dialogue guidance model is in the third round interval, then the process capability training weight is used as the preset weight, and the maximum weight in the process capability training weight is the fourth weight. The minimum value of the third round interval is greater than the maximum value of the second round interval, and the minimum value of the second round interval is greater than the maximum value of the first round interval.
11. The method according to claim 3 or 8, characterized in that, The process of determining the business intent transfer flow includes: Obtain any historical call text and the second intent corresponding to each sentence in that historical call text; The historical call text and the second intent are processed using a sequence pattern mining algorithm to obtain the intent transfer sequence corresponding to the historical call text; The weights corresponding to the connecting lines between adjacent nodes in the intent transfer sequence are determined based on a preset algorithm, and the business intent transfer process is obtained.
12. The method according to claim 11, characterized in that, The process of constructing the historical call text includes: Retrieve a preset amount of basic call text; Based on the role, sentences in the preset number of basic call texts are classified to obtain a sentence set corresponding to each role; For each set of sentences, clustering is performed based on the semantics of each sentence in the set to obtain sentence groups; Obtain a business dictionary pre-saved for the target business domain applied to the script guidance model, and a suggested intent saved for each keyword in the business dictionary; For each sentence group, the business dictionary, the suggested intent corresponding to each keyword, and the sentence group are input into the second large model, so that when the second large model determines that the target keyword in the business dictionary exists in the sentence group, it determines the second intent corresponding to the sentence group based on the suggested intent corresponding to the target keyword. The target base call text to which each sentence in each sentence group belongs is determined, and the corresponding second intent is identified for the corresponding sentence in the target base call text to obtain the historical call text.
13. The method according to claim 12, characterized in that, After obtaining the sentence group and before acquiring the business dictionary pre-saved for the target business domain applied to the speech guidance model, the method further includes: Count the number of sentences in each sentence group, and delete sentence groups whose number is less than a threshold.