Dialogue generation method and device

By clustering and evaluating customer service scripts, standard operating procedures are generated, solving the inefficiency problem caused by the diversity of customer service scripts and achieving efficient and low-cost generation of standard operating procedures.

CN115705362BActive Publication Date: 2025-11-11ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110939632.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-16
Publication Date
2025-11-11
Estimated Expiration
2041-08-16

AI Technical Summary

Technical Problem

In existing technologies, the types of dialogues between users and customer service are diverse and numerous, resulting in low efficiency of standard operating procedures for customer service personnel to provide customer service, and increasing the labor costs of enterprises.

Method used

By clustering semantically similar customer service scripts into script nodes, and selecting typical scripts for each node, a script node identifier sequence is generated using a pre-trained deep neural network and clustering algorithm. Frequent subsequences and subsequences whose quality assessment exceeds a threshold are then selected to generate a standard operating procedure.

Benefits of technology

It improved the efficiency of generating standard operating procedures, reduced labor costs for enterprises, and ensured the quality and efficiency of the generated procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705362B_ABST
    Figure CN115705362B_ABST
Patent Text Reader

Abstract

This specification provides a dialogue generation method and apparatus. The dialogue generation method includes: clustering semantically similar dialogues into at least one dialogue node, and selecting one or more typical dialogues for each dialogue node, wherein each of the at least one dialogue node has a corresponding node identifier; obtaining a dialogue node identifier sequence corresponding to the dialogue based on the dialogue node to which each dialogue belongs; obtaining frequent subsequences in the dialogue node identifier sequence; obtaining one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtaining a standard operating procedure for the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method for generating dialogues. This specification also relates to a dialogue generation apparatus, a computing device, and a computer-readable storage medium. Background Technology

[0002] With the booming development of the internet and service industry, more and more companies are providing online consultation services to their users through human customer service. However, as the complexity of online consultations increases and the training costs for customer service personnel continue to rise, many companies are trying to reduce the complexity of customer service and save on customer service personnel training costs by selecting standard operating procedures for customer service from the dialogues between users and customer service representatives.

[0003] In existing technologies, the main approach relies on experienced customer service personnel to analyze and summarize the dialogues between users and customer service representatives to obtain standard operating procedures. However, due to the large number of dialogues between users and customer service representatives, and the variety of user questions, the types of dialogues are diverse, which further reduces the efficiency of customer service personnel in obtaining standard operating procedures for customer service and increases the company's labor costs. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a dialogue generation method. This specification also relates to a dialogue generation apparatus, a computing device, and a computer-readable storage medium to address the technical deficiencies existing in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a dialogue generation method is provided, comprising:

[0006] Sentences with similar semantics are clustered into at least one speech node, and one or more typical speech words are selected for each speech node, wherein each speech node in the at least one speech node has a corresponding node identifier;

[0007] Based on the speech node to which each speech in the dialogue belongs, obtain the speech node identifier sequence corresponding to the dialogue;

[0008] Frequent subsequences are obtained from the speech node identifier sequence;

[0009] Obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

[0010] According to a second aspect of the embodiments of this specification, a dialogue generation apparatus is provided, comprising:

[0011] The clustering module is configured to cluster semantically similar statements into at least one statement node, and select one or more typical statements for each statement node, wherein each statement node in the at least one statement node has a corresponding node identifier.

[0012] The first acquisition module is configured to obtain the sequence of speech node identifiers corresponding to the dialogue based on the speech node to which each speech in the dialogue belongs;

[0013] The second acquisition module is configured to acquire frequent subsequences from the speech node identifier sequence;

[0014] The evaluation module is configured to obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and to obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

[0015] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:

[0016] Memory and processor;

[0017] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of any of the dialogue generation methods.

[0018] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of any of the dialogue generation methods.

[0019] The dialogue generation method provided in this specification clusters semantically similar dialogues into at least one dialogue node, and selects one or more typical dialogues for each dialogue node, wherein each of the at least one dialogue node has a corresponding node identifier; obtains a dialogue node identifier sequence corresponding to the dialogue based on the dialogue node to which each dialogue belongs; obtains frequent subsequences from the dialogue node identifier sequence; obtains one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtains the standard operation flow of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

[0020] Specifically, by selecting one or more typical dialogues for each dialogue node generated after clustering, and performing quality assessment on the frequent subsequences obtained from the dialogue node identification sequence, one or more frequent subsequences with quality assessment exceeding a preset quality threshold are obtained. Based on one or more frequent subsequences and one or more typical dialogues, the standard operating procedure for obtaining the dialogue is quickly obtained, which effectively improves the efficiency of obtaining the standard operating procedure for the dialogue and avoids the problem of high labor costs for enterprises due to the large number and variety of dialogues, thus effectively reducing the labor costs of enterprises. Attached Figure Description

[0021] Figure 1 This is a flowchart of a dialogue generation method provided in one embodiment of this specification;

[0022] Figure 2 This is a schematic diagram illustrating the acquisition of a partial order subsequence according to an embodiment of this specification;

[0023] Figure 3 This is a flowchart illustrating a dialogue generation method for generating standard operating procedures for customer service, as provided in one embodiment of this specification.

[0024] Figure 4 This is a schematic diagram of the structure of a dialogue generation device provided in one embodiment of this specification;

[0025] Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0026] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0027] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0028] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0029] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0030] Customer service dialogue: The company's customer service personnel communicate with customers by phone or text and provide support.

[0031] SOP: Standard Operating Procedure, or SOP for short, refers to the operating procedures established to achieve specific goals. New employees can achieve these goals by following the standard operating procedures.

[0032] Statistical language model: This is a technique for modeling the probability of text generation. It decomposes the probability of text generation into an autoregressive generation probability of each word in a chain and introduces the Markov property assumption to fit the autoregressive generation probability based on the first few words of each word.

[0033] Edit distance: a measure of the minimum number of editing operations required to transform one piece of text into another.

[0034] KMeans model: k-means clustering algorithm, or simply KMeans model, is an iterative clustering analysis algorithm model.

[0035] PrefixSpan algorithm: A sequence pattern algorithm that extracts frequently occurring partial order patterns from a large number of partial order phenomena.

[0036] ID: Identity document, abbreviated as ID.

[0037] Standard Operating Procedures (SOPs) are operational processes designed to achieve specific goals. They are typically developed and refined by experienced senior personnel, helping new employees quickly get up to speed and achieve objectives. However, SOPs in customer service have the following problems:

[0038] (1) Customers have a variety of problems and need a lot of SOPs to cover most scenarios, resulting in high construction costs.

[0039] (2) Customer issues are dynamic, and the SOPs that are developed often lag behind the actual customer service work.

[0040] Based on this, the embodiments in this specification are based on the large customer service teams of large enterprises, generating hundreds of thousands of communication records every month. This allows for the extraction of diverse and timely dialogue patterns from real, large-scale customer service conversations, assisting customer service managers in developing SOPs for more scenarios and promptly building SOPs for emerging scenarios.

[0041] This specification provides a dialogue generation method, and also relates to a dialogue generation apparatus, a computing device, and a computer-readable storage medium, which will be described in detail in the following embodiments.

[0042] See Figure 1 , Figure 1 A flowchart of a dialogue generation method according to an embodiment of this specification is shown, which specifically includes the following steps:

[0043] Step 102: Cluster semantically similar statements into at least one statement node, and select one or more typical statements for each statement node.

[0044] Each of the at least one speech node has a corresponding node identifier.

[0045] This script can be understood as a script used to generate standard operating procedures (SOPs); and the script used to generate SOPs will differ depending on the specific SOP. For example, in the case of a SOP for customer service, the script could be a customer service script selected from a large amount of dialogue data between users and customer service representatives; specifically, it could be: "User: What discount is available for this cup? Customer Service: This cup is currently 20% off." In the case of a SOP for answering questions about classical Chinese poems, the script could be a script selected from a large amount of dialogue data between students and teachers; specifically, it could be: "Student: Teacher, who wrote 'Quiet Night Thoughts'? Teacher: 'Quiet Night Thoughts' was written by Li Bai."

[0046] A script node can be understood as a node containing multiple scripts. Each script node contains multiple scripts with similar semantics. When the script is a customer service script, each script node will contain multiple customer service scripts with similar semantics. When the script is a script for answering classical Chinese poems, each script node will contain multiple scripts for answering classical Chinese poems with similar content.

[0047] When the script is a customer service script, at least one script node contains multiple semantically similar scripts. For example, a greeting-type script node stores multiple semantically similar greeting-type customer service scripts, and a product discount-type script node stores multiple semantically similar product discount-type customer service scripts. Multiple greeting-type customer service scripts could include "User: Hello; Customer Service Representative: Hello, customer!" or "User: Hello; Customer Service Representative: Hello!". Multiple product discount-type customer service scripts could include "User: Is the cup on sale?; Customer Service Representative: This cup is 20% off" or "User: Is the cup on sale?; Customer Service Representative: This cup is currently on promotion, 20% off during the promotion period."

[0048] Typical customer service scripts can be understood as the most representative scripts within a script node. In the case of customer service scripts, typical scripts are those with low confusion values ​​and the highest representativeness. A script with low confusion value can be understood as the one with clear and unambiguous meaning among all customer service scripts included in a script node. For example, customer service scripts in a script node may include, but are not limited to, Customer Service Script 1 and Customer Service Script 2. Customer Service Script 1 is "User: Is there a discount on this cup? Customer Service: All items in this store are 20% off." Customer Service Script 2 is "User: Is there a discount on this cup? Customer Service: This cup is currently 20% off." Compared to Customer Service Script 1, Customer Service Script 2 provides a clearer and more unambiguous answer to the question, therefore, Customer Service Script 2 has a lower confusion value. The most representative customer service scripts can be understood as the most representative customer service scripts included in a process script node. For example, customer service scripts in a process script node may include, but are not limited to, Customer Service Script 1, Customer Service 2, and Customer Service 3. Customer service script 1 is "User: Hello; Customer Service: Dear customer, hello!"; Customer service script 2 is "User: Hello; Customer Service: Hello, customer!"; Customer service script 3 is "User: Hello; Customer Service: Hello!". Compared to customer service scripts 1 and 2, customer service script 1 is more representative of user responses, therefore customer service script 1 is more representative.

[0049] A node identifier can be understood as an identifier representing a dialogue node. Each node identifier can uniquely represent a dialogue node. The node identifier can be a character or string composed of numbers, letters, or some special letters that can uniquely represent a dialogue node. The specific configuration can be set according to actual needs. This manual does not impose any restrictions on the composition of the node identifier. For example, the node identifier can be the dialogue node ID of the dialogue node.

[0050] Specifically, after obtaining multiple scripts, scripts with similar semantics are clustered to obtain at least one script node with a node identifier, and one or more corresponding typical scripts are selected for each script node.

[0051] For example, in the scenario of applying the dialogue generation method to generate standard operating procedures for customer service, we will further explain the generation of script nodes and the selection of typical scripts for script nodes. Here, the standard operating procedure can be the standard operating procedure for customer service, the script can be the customer service script, and the script node can be the customer service script node.

[0052] The process building platform pre-collects all dialogue data generated between users and customer service representatives, obtains multiple customer service scripts from these dialogues, and clusters semantically similar customer service scripts to obtain at least one customer service script node. Each customer service script node has a script identifier corresponding to itself, and the customer service script node contains multiple semantically similar customer service scripts. Furthermore, from the multiple customer service script nodes contained in each customer service script node, one or more typical customer service scripts corresponding to that customer service script node are selected.

[0053] In the process of clustering semantically similar utterances into at least one utterance node, the efficiency of clustering is low due to the diverse types and large number of utterances. Therefore, in this embodiment, the utterances are clustered based on vectorized semantic representations to solve the above problem. The specific implementation is as follows:

[0054] The process of clustering semantically similar utterances into at least one utterance node includes:

[0055] The utterances from multiple dialogues are input into a pre-trained deep neural network to obtain vectorized semantic representations of utterances.

[0056] The vectorized discourse semantic representation is clustered according to the clustering algorithm to obtain at least one cluster, and the at least one cluster is used as at least one discourse node, wherein the discourses in the same cluster are semantically similar discourses.

[0057] The pre-trained deep neural network can be any type of neural network used to convert utterances into vectorized semantic representations. In the embodiments of this specification, any pre-trained deep neural network can be used to convert utterances into vectorized semantic representations according to the needs of actual applications, and the embodiments of this specification do not impose any limitations on this. For example, the pre-trained deep neural network can be the pre-trained deep neural network BERT.

[0058] The clustering algorithm can be any clustering algorithm used for clustering dialogues. In the embodiments of this specification, any clustering algorithm can be used to cluster dialogues according to the needs of actual applications, and this specification does not impose any limitations on this. For example, the clustering algorithm can be the KMeans clustering algorithm.

[0059] Specifically, after obtaining the dialogues from multiple dialogues, the dialogues are input into a pre-trained deep neural network to vectorize the dialogues, thereby obtaining a vectorized semantic representation of the dialogues corresponding to the dialogues; and the vectorized semantic representation of the dialogues is clustered by a clustering algorithm to obtain at least one cluster, wherein the dialogues in the same cluster are semantically similar, and the at least one cluster is regarded as at least one dialogue node.

[0060] Continuing with the previous example, we will further explain how to obtain at least one utterance node based on vectorized utterance semantic representation. Clustering can be achieved by using the KMeans clustering algorithm to cluster the vectorized utterance semantic representation, resulting in clusters of utterance nodes.

[0061] The process building platform pre-collects all dialogue data generated between users and customer service representatives, obtains multiple customer service scripts from this dialogue data, and inputs each customer service script into a pre-trained deep neural network BERT. BERT then vectorizes each customer service script to obtain a vectorized semantic representation of the customer service script, which is the semantic vector of the customer service script.

[0062] After obtaining the semantic vector of the customer service script, the KMeans clustering algorithm is used to cluster the semantic vector of the customer service script to obtain at least one cluster of customer service scripts. Each cluster contains customer service scripts with similar semantics, and the obtained cluster of at least one customer service script is used as a customer service script node.

[0063] In the embodiments of this specification, a clustering algorithm is used to cluster the vectorized semantic representation of discourse obtained based on a pre-trained deep neural network, thereby obtaining at least one cluster and using the at least one cluster as at least one discourse node. This avoids the problem of low clustering efficiency caused by the diversity and large number of discourses, effectively improves the efficiency of clustering discourses, further improves the efficiency of obtaining standard operating procedures, and saves time costs for enterprises.

[0064] Furthermore, after obtaining at least one script node, the specific method for selecting one or more typical scripts for each script node is as follows:

[0065] The selection of one or more typical dialogues for each dialogue node includes:

[0066] Obtain the cluster center of each utterance node and the perplexity value of each utterance in each utterance node;

[0067] Based on the cluster center of each dialogue node and the confusion value of each dialogue in each dialogue node, one or more typical dialogues corresponding to each dialogue node are determined.

[0068] The confusion value can be understood as a numerical value that represents the degree of confusion of the script. When the script is a customer service script, the confusion value is used to indicate whether the customer service script is clear and unambiguous. The smaller the confusion value, the clearer the customer service script is.

[0069] Cluster center can be understood as the center of each cluster.

[0070] Specifically, after clustering semantically similar dialogues into at least one dialogue node, the cluster center of each dialogue node and the confusion value corresponding to the dialogue in each dialogue node are obtained, and one or more typical dialogues are selected for each dialogue node based on the confusion value and the cluster center.

[0071] Following the previous example, we will provide a more detailed explanation of how to determine one or more typical dialogues corresponding to each dialogue node based on cluster centers and confusion values.

[0072] The acquired customer service scripts are clustered using a clustering model to obtain clusters. Each cluster can then have a cluster center calculated. Simultaneously, by evaluating the customer service scripts, the confusion value corresponding to each customer service script in each script node is obtained.

[0073] After obtaining the cluster center and the confusion value, one or more typical dialogues corresponding to each dialogue node can be determined based on the cluster center and the confusion value.

[0074] In the embodiments of this specification, by obtaining the cluster center of each dialogue node and the confusion value of the dialogue, the typical dialogue of each dialogue node is determined, making the obtained typical dialogue more accurate, effectively improving the quality of the typical dialogue, and realizing the subsequent generation of high-quality standard operating procedures through typical dialogue.

[0075] In obtaining the perplexity value of each dialogue node, due to the diverse types and large number of dialogues, relying solely on manual perplexity analysis would result in low efficiency in acquiring typical dialogues, further reducing the efficiency of generating standard operating procedures and increasing the company's labor costs. Therefore, this embodiment uses a pre-trained statistical language model to evaluate dialogues and obtain their perplexity values ​​to address the aforementioned problems. The specific implementation is as follows:

[0076] Obtaining the confusion value of the dialogue in each dialogue node includes:

[0077] Each utterance in each utterance node is input into a pre-trained statistical language model to obtain the perplexity value corresponding to each utterance in each utterance node.

[0078] The pre-trained statistical language model can be any type of statistical language model, and the embodiments in this specification do not impose any limitations on it.

[0079] Using the previous example, inputting each dialogue in each dialogue node into a pre-trained statistical language model to obtain the perplexity value corresponding to each dialogue in each dialogue node can be understood as inputting multiple semantically similar customer service dialogues contained in each customer service dialogue node into a pre-trained statistical language model to obtain the perplexity value corresponding to each customer service dialogue.

[0080] In another scenario, this manual can also collect customer service scripts from multiple customer conversations, build a corpus based on all the scripts, and then build a statistical language model on the corpus. For each customer service script in the corpus, the statistical language model will assign a confusion score, thereby obtaining the confusion value corresponding to each customer service script.

[0081] In the embodiments of this specification, accurate and high-quality typical dialogues can be obtained through the cluster center and the perplexity value of each dialogue node. However, the process of obtaining typical dialogues from these two perspectives cannot be intuitively displayed, making it impossible to monitor the process of generating standard operating procedures in real time. Therefore, in the embodiments of this specification, the dialogues are sorted based on the cluster center and perplexity value, and typical dialogues are generated based on the target sorting result to solve the above problems. The specific method for generating standard operating procedures is as follows:

[0082] The step of determining one or more typical dialogues corresponding to each dialogue node based on the cluster center of each dialogue node and the confusion value of each dialogue in each dialogue node further includes steps A1 to A2:

[0083] Step A1: Based on the confusion value of each dialogue in each dialogue node and the cluster center of each dialogue node, obtain the target ranking result of each dialogue in each dialogue node.

[0084] The target ranking result can be understood as the ranking result obtained by sorting the utterances in the utterance nodes based on the cluster centers and the perplexity values.

[0085] In practice, after obtaining the perplexity value of the dialogue in each dialogue node and the cluster center of each dialogue node, the dialogue in each dialogue node is sorted based on the perplexity value and the cluster center to obtain the target sorting result of the dialogue in each dialogue node.

[0086] Further, based on the perplexity value of each utterance in each utterance node and the cluster center of each dialogue node, the target ranking result for each utterance in each dialogue node is obtained, including:

[0087] Determine the distance value between each dialogue in each dialogue node and the cluster center, and sort each dialogue in each dialogue node according to the distance value to obtain the distance value sorting result;

[0088] Based on the confusion value corresponding to each dialogue in each dialogue node, sort each dialogue in each dialogue node to obtain the confusion value sorting result;

[0089] Based on the distance value sorting result and the confusion value sorting result, the target sorting result for each utterance in each dialogue node is obtained.

[0090] The confusion value sorting result can be understood as the sorting result obtained by sorting the dialogues in each dialogue node according to the confusion value corresponding to the dialogue in each dialogue node.

[0091] When the script is a customer service script, the distance value can be understood as the distance between the customer service script in each script node and the cluster center.

[0092] The distance value sorting result can be understood as the sorting of the utterances in each utterance node based on their distance values.

[0093] Specifically, after obtaining the cluster centers and perplexity values, the object generation platform first sorts the dialogues in each dialogue node from smallest to largest based on their perplexity values, obtaining a perplexity value sorting result; second, it calculates the distance between the dialogues in each dialogue node and the corresponding cluster center, and sorts the dialogues in each dialogue node from smallest to largest based on the distance values, obtaining a distance value sorting result; finally, based on the perplexity value sorting result and the distance value sorting result, it obtains the target sorting result for the dialogues in each dialogue node.

[0094] Following the previous example, we will provide a more detailed explanation of how to obtain the target ranking result of the speech in each speech node based on the obtained confusion value ranking result and distance value ranking result.

[0095] After scoring the confusion value of customer service scripts in the dialogue node using a pre-trained statistical language model, the customer service scripts contained under each script node are sorted from smallest to largest according to their confusion value scores, with the customer service scripts with smaller confusion values ​​being ranked higher, thus obtaining the confusion value ranking result.

[0096] After obtaining the cluster center of each cluster (script node), the distance between the customer service script in the cluster and the cluster center in the cluster is calculated, and the scripts are sorted according to the distance from smallest to largest, with the smaller the distance, the higher the ranking, thus obtaining the distance value sorting result.

[0097] By combining the confusion value sorting results and the distance value sorting results, the target sorting result of customer service scripts in each process script node can be obtained.

[0098] In the embodiments of this specification, the target ranking result of the dialogues in each dialogue node is obtained based on the confusion value ranking result obtained by ranking the dialogues by confusion value and the distance value ranking result obtained by ranking the dialogues by cluster center. This intuitively demonstrates the process of obtaining typical dialogues from the perspectives of cluster centers and the confusion value of the dialogues, enabling real-time monitoring of the standard operating procedure generation process and ensuring the quality of the generated standard operating procedure.

[0099] Furthermore, in the process of obtaining the target ranking result based on the confusion value ranking result and the distance value ranking result, since there are certain differences between the confusion value ranking result and the distance value ranking result, how to accurately obtain the target ranking result becomes a problem that needs to be solved. Based on this, the target ranking result of the speech in each speech node is obtained by calculating the confusion value order score, the distance value order score, and the target ranking score. The specific implementation method is as follows:

[0100] The step of obtaining the target ranking result for each utterance in each dialogue node based on the ranking results of the distance value and the confusion value includes:

[0101] Based on the distance value sorting result, determine the corresponding distance value order score for each speech in each speech node;

[0102] Based on the confusion value sorting result, determine the corresponding confusion value order score for each dialogue in each dialogue node;

[0103] The target ranking score for each dialogue in each dialogue node is obtained by taking a weighted average of the distance value ranking score and the confusion value ranking score.

[0104] The target ranking score is used to rank each speech in each speech node to obtain the target ranking result.

[0105] The confusion value order score can be understood as the score of a word obtained by changing the position of the word in the confusion value sorting result. The earlier the word appears in the sorting result, the higher its score.

[0106] Distance value order score can be understood as the score obtained by transforming the position of the words in the distance value sorting result. The earlier the word appears in the sorting result, the higher its score.

[0107] The target ranking score can be understood as a weighted average of the confusion value ranking score and the distance value ranking score to obtain the score corresponding to the rhetoric.

[0108] Specifically, after obtaining the confusion value sorting results and the distance value sorting results, the object generation platform first determines the confusion value order score of each dialogue in each dialogue node based on the confusion value sorting results; secondly, it determines the distance value order score of each dialogue in each dialogue node based on the distance value sorting results; finally, it determines the target sorting score of the dialogue in each dialogue node based on the confusion value order score and the distance value order score; and then sorts the dialogues in each dialogue node from smallest to largest based on the target sorting score to obtain the target sorting result of the dialogues in each dialogue node.

[0109] Following the previous example, we will provide a more detailed explanation of how to obtain the target ranking result of the speech in each speech node by using the confusion value order score, distance value order score, and target ranking score.

[0110] After obtaining the distance value ranking results determined by the distance between the customer service script and the cluster center, and the confusion value ranking results determined by the statistical language model, the confusion value ranking results and the distance value ranking results describe the quality and typicality of the customer service script from different perspectives. Therefore, it is necessary to integrate the two to obtain the final ranking result, thereby obtaining the most high-quality typical script.

[0111] First, the confusion value sorting result is converted into a confusion value order score. Specifically, for a confusion value sorting result of length N, the score calculated using the formula 1-i / N is used to score the i-th element in the confusion value sorting result. Thus, each customer service script in the confusion value sorting result will receive a score between 0 and 1. The earlier the customer service script appears in the confusion value sorting result, the higher its score. For example, if a script node contains customer service scripts A, B, and C, and the order of customer service scripts A, B, and C in the confusion value sorting result is [customer service script C, customer service script B, customer service script A], the corresponding confusion value order scores are 0, 0.333, and 0.666.

[0112] Secondly, the distance value sorting results are converted into distance value order scores. Specifically, for a distance value sorting result of length N, the score calculated using the formula 1-i / N is used to score the i-th element in the distance value sorting result. Thus, each customer service script in the distance value sorting result will receive a score between 0 and 1. The earlier the customer service script appears in the distance value sorting result, the higher its score. For example, if a certain process script node contains customer service scripts A, B, and C, and the order of customer service scripts A, B, and C in the distance value sorting result is [customer service script A, customer service script C, customer service script B], the corresponding distance value order scores are 0.666, 0, and 0.333, respectively.

[0113] Finally, a weighted average is calculated using the confusion value ranking score and the distance value ranking score to obtain the target ranking score for the customer service scripts at each process script node. The weighted average means averaging the scores of the customer service scripts at each process script node. For example, the final score for customer service script A is (0 + 0.666) / 2, the final score for customer service script B is (0.333 + 0) / 2, and the final score for C is (0.666 + 0.333) / 2. This yields the target ranking score for the customer service scripts at each process script node.

[0114] Based on the target ranking score of the customer service scripts in each process script node, the customer service scripts in the process script node are sorted from smallest to largest to obtain the target ranking result of the customer service scripts in each process script node.

[0115] In this embodiment of the specification, the speech in each speech node is sorted by the determined confusion value order score, distance value order score and target sort score, thereby obtaining the target sorting result of the speech in each speech node. This solves the problem that it is difficult to accurately obtain the target sorting result due to the certain difference between the confusion value sorting result and the distance value sorting result.

[0116] Step A2: Determine one or more typical dialogues corresponding to each dialogue node based on the target sorting result.

[0117] Further, determining one or more typical dialogues corresponding to each dialogue node based on the target sorting result includes:

[0118] The ranked dialogue scripts are obtained from the target ranking results, and the ranked dialogue scripts are compared with the typical dialogue scripts in the typical dialogue script set for similarity. The ranked dialogue scripts are the ranked dialogue scripts in each dialogue node.

[0119] The sorted dialogues with a similarity less than or equal to a preset similarity threshold are stored in the typical dialogue set as typical dialogues;

[0120] If the similarity is greater than the preset similarity threshold, the next ranked script will be compared with the typical scripts in the typical script set.

[0121] When the sorting of dialogues is completed or the number of typical dialogues in the typical dialogue set is greater than or equal to a preset threshold, the typical dialogues in the typical dialogue set are used as one or more typical dialogues corresponding to each dialogue node.

[0122] Among them, the typical dialogue set can be understood as a set that stores typical dialogues.

[0123] Similarity can be understood as a numerical value representing the degree of similarity between a dialogue in a dialogue node and a typical dialogue in a set of typical dialogues. In practical applications, similarity is determined as follows: First, calculate the edit distance d between the dialogue q in the dialogue node and the typical dialogue q' in the set of typical dialogues; second, obtain the similarity between dialogue q and typical dialogue q' using the formula sim = (max_len - d) / min_len, where max_len is the length of the longer of dialogue q and typical dialogue q', min_len is the length of the shorter of q and q', and sim is the similarity between dialogue q and typical dialogue q'.

[0124] The preset similarity threshold can be set according to the actual application, and this specification does not limit it in any way. For example, the preset similarity threshold can be 0.7.

[0125] The preset quantity threshold can be set according to the actual application, and this specification does not impose any limitations on it. For example, the preset quantity threshold can be 5.

[0126] Continuing with the previous example, we will further explain how to use typical scripts from the typical script set as one or more typical scripts corresponding to a script node. Here, the typical scripts are selected from the customer service scripts in the script node; the typical script set can be a set that stores typical scripts.

[0127] Based on the target ranking results, the edit distance between the customer service scripts in each script node and the typical scripts in the typical script set is combined to obtain the final typical scripts.

[0128] Specifically, the sorted customer service scripts are obtained from the target sorting results corresponding to the customer service scripts of each script node; and according to the sorting order, each customer service script in each script node is compared with each typical script in the typical script set corresponding to each script node to obtain the similarity between the customer service script and each typical script.

[0129] If the similarity between the customer service script and each typical script in the typical script set is less than or equal to the preset similarity threshold of 0.7, the customer service script is stored in the typical script set, thus identifying it as a typical script.

[0130] If the similarity between the current customer service script and any typical script in the typical script set is greater than the preset similarity threshold of 0.7, then the current customer service script is skipped, and the next customer service script is compared with each typical script in the typical script set.

[0131] After comparing the sorted scripts with each typical script in the typical script set, or if the number of typical scripts in the typical script set is greater than or equal to a preset threshold of 5, the similarity comparison between the sorted customer service scripts and the typical scripts in the typical script set ends. All the typical scripts in the typical script set are taken as the typical scripts of the script node corresponding to the target sorting result, thereby obtaining one or more typical scripts corresponding to each script node.

[0132] In the embodiments of this specification, by comparing the similarity between the sorted scripts and the typical scripts in the typical script set, the typical scripts in the typical script set are determined, thereby obtaining high-quality typical scripts. Furthermore, the standard operating procedure is generated using the high-quality typical scripts, which further ensures the quality of the standard operating procedure.

[0133] Step 104: Obtain the sequence of speech node identifiers corresponding to the dialogue based on the speech node to which each speech belongs in the dialogue.

[0134] The utterance node identifier sequence can be understood as a sequence composed of the node identifiers of multiple consecutive utterance nodes, used to determine frequent subsequences.

[0135] Frequent subsequences can be understood as partially ordered subsequences of the speech node identifier sequence, which are composed of node identifiers of multiple consecutive speech nodes.

[0136] Continuing with the previous example, we will further explain how to obtain the sequence of speech node identifiers based on the speech node to which each speech belongs in the dialogue. The node identifier can be a speech node ID.

[0137] After clustering semantically similar customer service scripts into an exponential customer script node and configuring a script node ID for each script node, the complete dialogue process between the customer service representative and the user is obtained from the dialogue. Each customer service script in each complete dialogue process has a corresponding script node ID. Based on this, the complete dialogue process between the customer service representative and the user is represented by the script node ID corresponding to the customer service script, thereby obtaining the script node identifier sequence represented by the script node ID.

[0138] See Figure 2 , Figure 2 This is a schematic diagram illustrating the acquisition of a partially ordered subsequence provided in one embodiment of this specification. After configuring a dialogue node ID for each dialogue node, the complete communication process between the customer service representative and the user is obtained from the dialogue. Each customer service dialogue in each complete communication process has a corresponding dialogue node ID. Based on this, the complete communication process is represented by the dialogue node ID corresponding to that customer service dialogue, thereby obtaining a dialogue node sequence corresponding to at least two types of process dialogue nodes, such as... Figure 2 The symbols “A→B→C→D→E” and “A→X→C→Y→E” are used.

[0139] Step 106: Obtain the frequent subsequences in the speech node identifier sequence.

[0140] After representing customer service dialogues as a sequence of script node identifiers, one approach to extracting high-quality, frequent subsequence patterns is to statistically analyze the transition probabilities between adjacent script nodes and construct a state transition graph between them. Based on the state transition graph, the path with the highest generation probability can be searched as the high-quality pattern.

[0141] However, this scheme has a flaw. When constructing the state transition diagram for the dialogue nodes, it essentially assumes a Markov property between nodes, meaning the next node depends only on the preceding nodes (usually the first node). This assumption leads to the following two problems:

[0142] (1) A high-quality path is formed by a combination of several transition edges. Each edge is a relatively reasonable edge, but the entire path cannot be guaranteed to be a reasonable path.

[0143] (2) SOP is an abstraction of the real dialogue content, which only includes the key nodes. However, the real dialogue also contains a large number of non-key nodes. The state transition diagram counts the transition probabilities of all nodes, which is inconsistent with the goal of SOP.

[0144] Based on this, in one embodiment of this specification, the complete subsequence is modeled directly, rather than the state transition process. Therefore, the high-quality frequent subsequences mined are all real-world patterns. By using frequent subsequences, non-critical nodes in the dialogue can be ignored. This addresses the problems of existing technologies, and the specific implementation is as follows:

[0145] The process of obtaining frequent subsequences in the speech node identifier sequence includes:

[0146] The utterance node identifier sequence is processed according to the sequence mining algorithm to obtain the prefix of the utterance node identifier sequence and the prefix projection corresponding to the prefix, and the number of utterance nodes with the same identifier in the prefix projection is determined.

[0147] Frequent subsequences are obtained based on the prefix and the number of discourse nodes with the same identifier.

[0148] Sequence mining algorithms can be understood as algorithms that analyze and mine dialogue sequence nodes. One such algorithm is the PrefixSpan algorithm, which processes dialogue node identifier sequences to obtain frequent subsequences.

[0149] When the sequence mining algorithm is the PrefixSpan algorithm, the prefix can be understood as a subsequence consisting of a specific number of nodes preceding the utterance node identifier sequence. This specific number can be set according to the actual application, and this specification does not impose any limitation on it. Correspondingly, the prefix projection can be understood as a subsequence consisting of nodes following the prefix of the utterance node identifier sequence. The prefix of the utterance node identifier sequence plus the corresponding prefix projection constitutes the complete utterance node identifier sequence. The prefix projection contains all suffixes corresponding to the same prefix, and this prefix projection is also called the suffix and prefix corresponding projection database.

[0150] Following the previous example, we will further explain how to obtain frequent subsequences based on the number of utterance nodes with the same identifier in the prefix and the prefix projection.

[0151] The PrefixSpan algorithm is used to process the speech node identifier sequence to determine the prefix of the speech node identifier sequence, and the prefix projection corresponding to the prefix is ​​determined from the speech node identifier sequence based on the prefix.

[0152] After determining the prefix projection corresponding to the prefix, the PrefixSpan algorithm determines the number of uniformly identified utterance nodes in the prefix projection. For utterance nodes whose number exceeds the support threshold, the utterance node is added to the prefix of the utterance node identifier sequence, thus obtaining frequent subsequences of the utterance node sequence, such as... Figure 2The partially ordered subsequence (frequent subsequence) "A→C→E" is shown in the figure; the support threshold can be set according to the actual application scenario, and this specification does not impose any limitations on it. For example, the support threshold can be 5.

[0153] In the embodiments of this specification, a sequence mining algorithm is used to analyze and mine the phonology sequence identifier nodes, thereby determining the number of phonology nodes with the same identifier in the prefix and the prefix projection, and obtaining frequent subsequences based on the number of phonology nodes with the same identifier in the prefix and the prefix projection, which improves the efficiency of obtaining frequent subsequences and saves computer processing resources.

[0154] Further, obtaining frequent subsequences based on the prefix and the number of utterance nodes with the same identifier includes:

[0155] If the number is less than a preset threshold, the prefix is ​​considered a frequent subsequence.

[0156] If the number is greater than or equal to a preset number threshold, the prefix is ​​updated based on the utterance nodes corresponding to the number to obtain the updated prefix;

[0157] The updated prefix projection corresponding to the updated prefix is ​​obtained from the speech node identifier sequence until the number of speech nodes with the same identifier in the updated prefix projection is less than the preset number threshold.

[0158] Specifically, when the sequence analysis algorithm is the PrefixSpan algorithm, the preset quantity threshold can be understood as the support threshold; this preset quantity threshold is set according to the actual application, and the embodiments in this specification do not impose any limitations on it. For example, the preset quantity threshold can be 5.

[0159] Specifically, the sequence mining algorithm determines the number of utterance nodes with the same identifier in the prefix projection. If the number is less than a preset threshold, the prefix is ​​considered a frequent subsequence. If the number of utterance nodes with the same identifier in the prefix projection is greater than or equal to the preset threshold, the prefix is ​​updated based on the utterance nodes corresponding to that number, thereby obtaining the updated prefix. The updated prefix projection corresponding to the updated prefix is ​​obtained from the utterance sequence identifier node, until the number of utterance nodes with the same identifier in the updated prefix projection is less than the preset threshold.

[0160] Following the previous example, after determining the number of speech nodes with the same identifier in the prefix projection, the PrefixSpan algorithm determines the prefix of the speech node identifier sequence as a frequent subsequence if the number of speech nodes with the same identifier is less than a preset threshold.

[0161] If the number of speech nodes with the same identifier is greater than or equal to a preset threshold, the prefix of the speech node sequence is updated based on the speech nodes with the same identifier. The speech nodes with the same identifier are added to the prefix of the speech node sequence to obtain the updated prefix of the speech node sequence, i.e., the updated prefix.

[0162] The prefix projection corresponding to the updated prefix is ​​redefined, i.e., the prefix projection is updated, and the number of discourse nodes with the same identifier in the updated prefix projection is determined. It is then further determined whether the number of discourse nodes with the same identifier in the prefix projection is less than the support threshold, until the number of discourse nodes with the same identifier in the updated prefix projection is less than the support threshold. The prefix of the updated discourse node identifier sequence is then identified as a frequent subsequence. This improves the efficiency of obtaining frequent subsequences and saves computer processing resources.

[0163] Step 108: Obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

[0164] In this context, a standard operating procedure (SOP) can be understood as an object of operational processes constructed to better accomplish a specific task. Furthermore, the SOP varies depending on the application scenario of the dialogue generation method. For example, in a customer service scenario, the SOP can be a standard operating procedure for customer service, helping customer service personnel resolve user inquiries. Similarly, in an online classical Chinese poetry tutoring scenario, the SOP can be a standard operating procedure for answering classical Chinese poetry questions, helping online tutors answer students' questions about classical Chinese poetry.

[0165] Specifically, after obtaining frequent subsequences from the speech node identifier sequence, the frequent subsequences are evaluated for quality to determine the quality score corresponding to them. If the quality score is greater than a preset quality threshold, one or more frequent subsequences whose quality evaluation exceeds the preset quality threshold are obtained.

[0166] After identifying one or more high-quality frequent subsequences and the typical dialogue corresponding to each dialogue node, a standard operating procedure is generated based on the one or more frequent subsequences and the typical dialogue corresponding to each dialogue node in the one or more frequent subsequences, thereby obtaining the standard operating procedure for the dialogue.

[0167] Using the previous example, we will further explain the standard operating procedure for obtaining dialogue using one or more frequent subsequences and one or more typical phrases.

[0168] After obtaining the frequent subsequences from the speech node identifier sequence, the quality score of the frequent subsequences can be calculated and evaluated using the quality score calculation formula, thereby determining the quality score corresponding to the frequent subsequences.

[0169] The quality score corresponding to the frequent subsequence is compared with a preset quality threshold. If the quality score is greater than the preset quality threshold, one or more frequent subsequences whose quality assessment exceeds the preset quality threshold are obtained.

[0170] After identifying one or more high-quality frequent subsequences and one or more typical scripts corresponding to each customer service script node, a standard operating procedure is generated based on the one or more frequent subsequences and the typical scripts corresponding to each script node in the one or more frequent subsequences, thereby obtaining the standard operating procedure for the dialogue.

[0171] In another scenario, after identifying one or more high-quality frequent subsequences and typical scripts corresponding to each customer service script node, one or more frequent subsequences and one or more typical scripts are output to customer service personnel for organization and refinement, thereby obtaining a standard operating procedure for customer service dialogue.

[0172] The dialogue generation method provided in this manual obtains one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtains a standard operating procedure for dialogue based on one or more frequent subsequences and one or more typical dialogues. This effectively improves the efficiency of obtaining a standard operating procedure for customer service and avoids the problem of high labor costs for enterprises due to the diversity and large amount of data, thus effectively reducing the labor costs of enterprises.

[0173] In practical applications, the quality score of frequent subsequences is evaluated based on their frequency of occurrence; higher frequency generally indicates higher quality. However, this approach has a drawback: it doesn't consider the inherent cohesion of frequent subsequences. Some high-frequency nodes, such as opening remarks and closing remarks, occur extremely frequently, leading to many frequent subsequences containing these nodes also having high frequencies. However, these frequently identified subsequences lack significant information content. For example, the ternary subsequence "opening remarks → any intermediate node → closing remarks" has a high probability of occurrence, but the intermediate node can be any number. Another ternary subsequence, A→B→C, although occurring less frequently, always appears together, and its quality should be higher than the former.

[0174] Therefore, in the embodiments of this specification, the evaluation of subsequence quality combines subsequence frequency and intrinsic cohesion. If a subsequence occurs frequently but has low intrinsic cohesion, its final quality score will not be high. This effectively improves the quality of the mined subsequences and increases the efficiency of obtaining frequent subsequences, as described below:

[0175] The acquisition of one or more frequent subsequences whose quality assessment exceeds a preset quality threshold includes:

[0176] The frequent subsequences are evaluated for quality to obtain a quality score corresponding to the frequent subsequences; and the quality score is compared with a preset quality threshold to obtain one or more frequent subsequences whose quality evaluation exceeds the preset quality threshold.

[0177] The quality score can be understood as a quality score that characterizes the quality of frequent subsequences.

[0178] Specifically, after obtaining frequent subsequences, the quality of the frequent subsequences is evaluated to obtain the quality score corresponding to the frequent subsequences. The quality score is then compared with a preset quality threshold to obtain one or more frequent subsequences whose quality evaluation exceeds the preset quality threshold.

[0179] Following the previous example, we will provide a more detailed explanation of how to identify one or more frequent subsequences whose quality assessment exceeds the preset quality threshold based on the quality score and the preset quality threshold.

[0180] The quality score of each frequent subsequence is calculated using a quality assessment formula, which is:

[0181]

[0182] Among them, a i Let s be a utterance node, and s be a partially ordered subsequence (frequent subsequence), (a1, a2…a…) |s| ) represents the ID of each utterance node in s, where |s| refers to the length of s, therefore the last node is a. |s| Let score(s) be the quality score of s, P(s) be the frequency of s, which is obtained by dividing the number of times s appears by the total partial order subsequence (total number of calls), and P(a) be the frequency of a node a, which is obtained by dividing the number of times a appears by the total partial order subsequence (total number of calls). When calculating P(a), it is not required that a must be in a specific s, as long as a exists in the dialogue data between customer service and users.

[0183] After obtaining the quality score of the frequent subsequence, the quality score of the frequent subsequence is compared with a preset quality threshold. If the quality score is greater than the preset quality threshold, the frequent subsequence corresponding to the quality score is determined as a frequent subsequence whose quality assessment exceeds the preset quality threshold. The preset quality threshold can be set according to the actual application, and this specification does not limit it in any way.

[0184] In another scenario, after obtaining the quality scores of frequent subsequences, the frequent subsequences are sorted from highest to lowest quality according to their quality scores. The more frequently occurring subsequences appear at the top of the sorted list, the higher their quality. A specific number of these frequently occurring subsequences are then selected as high-quality frequent subsequences. This specific number can be set according to the actual application; this specification does not impose any limitations on this, for example, the specific number could be 100.

[0185] In the embodiments described in this specification, the evaluation of subsequence quality combines subsequence frequency and intrinsic cohesion. This effectively improves the quality of the mined subsequences and increases the efficiency of obtaining frequent subsequences.

[0186] The dialogue generation method provided in this specification selects one or more typical dialogues for each dialogue node generated after clustering, and performs quality assessment on the frequent subsequences obtained from the dialogue node identifier sequence to obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold. Based on one or more frequent subsequences and one or more typical dialogues, the standard operating procedure for obtaining the dialogue is quickly obtained, which effectively improves the efficiency of obtaining the standard operating procedure for the dialogue and avoids the problem of high labor costs for enterprises due to the large number and variety of dialogues, thus effectively reducing the labor costs of enterprises.

[0187] The following is in conjunction with the appendix Figure 3 Taking the application of the dialogue generation method provided in this specification in the scenario of generating standard operating procedures for customer service as an example, the dialogue generation method will be further explained. Figure 3 This specification illustrates a flowchart of a dialogue generation method for generating standard operating procedures in customer service scenarios, as provided in an embodiment of this specification. The method specifically includes the following steps:

[0188] Step 302: Obtain dialogue data.

[0189] Specifically, it involves acquiring a large amount of dialogue data generated by a company's customer service team during communication with its users.

[0190] Step 304: Obtain customer service scripts.

[0191] Specifically, all customer service scripts are obtained from the dialogue data.

[0192] Step 306: Determine the confusion level of each customer service script based on a statistical language model.

[0193] The perplexity level can be understood as the perplexity value in the above embodiments.

[0194] Specifically, a corpus is built based on all customer service scripts in the dialogue data, and a statistical language model is built in the corpus. The statistical language model is used to score the perplexity of each customer service script, thereby determining the perplexity of each script.

[0195] Step 308: Obtain the semantic vector of customer service scripts using the BERT model.

[0196] In this context, the semantic vector of customer service scripts can be understood as the vectorized semantic representation of the scripts in the above embodiments.

[0197] Specifically, each customer service script is input into a pre-trained deep neural network BERT to obtain the semantic vector of the customer service script.

[0198] Step 310: Cluster the semantic vectors using the KMeans model to obtain the process speech nodes (clusters).

[0199] In this context, the process script node can be understood as the script node in the above embodiment.

[0200] Specifically, the semantic vectors of customer service scripts are input into the KMeans clustering algorithm for clustering, resulting in clusters of customer service scripts. Each cluster is then identified as a process script node. This yields process script nodes, where each process script node contains multiple semantically similar customer service scripts, and each process script node has a corresponding node identifier.

[0201] Step 312: Sort the customer service scripts in the script node according to the distance between the customer service script and the cluster center.

[0202] In this context, the distance between the customer service script and the cluster center can be understood as the distance value in the above embodiment.

[0203] Specifically, for each cluster obtained by the KMeans clustering algorithm, a cluster center can be calculated.

[0204] Calculate the distance between each customer service script in a cluster and the cluster center. Then, based on the distance between the customer service scripts and the cluster center, sort the customer service scripts in ascending order, with the smaller the distance, the higher the ranking.

[0205] Step 314: Sort the customer service scripts according to their level of confusion.

[0206] Specifically, based on the confusion level of each customer service script determined in step 306, the statistical language model sorts all customer service scripts under the same process script node according to the confusion level of the customer service script from smallest to largest, with the customer service script with the lower confusion level being ranked higher.

[0207] Step 316: Merge the sorting results to obtain the final sorting result.

[0208] The final sorting result can be understood as the target sorting result in the above embodiments.

[0209] Specifically, the confusion level of customer service scripts is scored by measuring the distance between the scripts and the cluster centers and by using a statistical language model. Both methods characterize the typicality of the scripts from different perspectives. Finally, the ranking results of the two methods are merged to obtain the final ranking result.

[0210] The sorting results are merged as follows:

[0211] First, the existing sorting results are converted into ordinal scores: for a sort of length N, the score of the i-th element is 1-i / N, so each customer service script will receive a score between 0 and 1. Sorts based on cluster center distance and sorts based on statistical language models are processed in this way.

[0212] Then, a weighted average is calculated based on the two order scores to obtain the final order score. The customer service scripts in each process script node are then sorted according to this order score to obtain the final sorting result.

[0213] Step 318: Generate a set of typical scripts based on the similarity between customer service scripts in the final sorting results.

[0214] Specifically, based on the final sorting results, and combining the similarity and edit distance between the statements, a final set of typical statements is generated. The specific steps are as follows:

[0215] S1. Initialize the script set S to be empty.

[0216] S2, sorting index i = 1.

[0217] S3. Extract the i-th phrase q from the sorted list.

[0218] S4. Calculate the edit distance d between the statement q and each statement q' in the statement set S, and calculate the similarity between the statement q and each statement q' in the statement set S based on the edit distance d.

[0219] S5. If the similarity sim between the word q and any word q' in the word set S is greater than the threshold of 0.7, then skip the word q.

[0220] If the similarity sim between a word q and any word q' in the word set S is less than or equal to the threshold of 0.7, then word q is added to the word set S.

[0221] S6. If the size of the script set is greater than or equal to K, then end the loop and output the script set S.

[0222] S7, i = i + 1, switch to S3 and continue the loop.

[0223] Step 320: Represent the communication content between the user and customer service through the script node ID of the process script node, and generate a sequence composed of script node IDs.

[0224] Specifically, each process script node has a script node ID, and each customer service script in the communication dialogue between customer service and user has a corresponding script node ID. Based on this, the complete communication process is represented by the script node ID corresponding to the customer service script, thereby obtaining a script node sequence corresponding to at least two types of process script nodes.

[0225] Step 322: Determine frequent subsequences based on the PrefixSpan algorithm.

[0226] Specifically, the PrefixSpan algorithm is used to search for frequent subsequences, with a support threshold of 5 set.

[0227] The PrefixSpan algorithm is used to search for all subsequences whose occurrence frequency is greater than or equal to the support threshold. The search space is then pruned and recalled based on the occurrence frequency of frequent subsequences to determine the frequent subsequences, which are then stored in a frequent subsequence set.

[0228] Step 324: Calculate the quality score of the frequent subsequences and determine the high-quality subsequences based on the quality score.

[0229] Specifically, for each frequent subsequence in the set of frequent subsequences, a quality score is calculated, and the frequent subsequences are sorted in descending order of their scores. The more frequent subsequences are ranked higher, the higher their quality.

[0230] After sorting the frequent subsequences, the top N frequent subsequences can be obtained and identified as high-quality subsequences, where N can be 100.

[0231] Step 326: Send the typical script set and high-quality subsequences to the customer service supervisor.

[0232] Specifically, high-quality subsequences are taken as high-quality operational processes, and the high-quality operational processes and typical dialogues at each process node are taken as the final mining results and output to the customer service supervisor for sorting and accumulation.

[0233] The dialogue generation method provided in this specification is a high-quality subsequence mining method that integrates sequence frequency and intrinsic cohesion. Combined with the dialogue nodes obtained from customer service semantic clustering, it automatically mines standard operating procedures from a large number of customer service dialogues, constructing a two-stage SOP mining framework of dialogue node clustering + high-quality subsequence mining. This effectively improves the efficiency of obtaining standard operating procedures for customer service, avoids the problem of high labor costs for enterprises due to the diversity and large quantity of data, and can help enterprise customer service departments formulate service standards, accumulate customer service experience, and guide novice customer service staff in their work.

[0234] Corresponding to the above method embodiments, this specification also provides embodiments of a dialogue generation apparatus. Figure 4 A schematic diagram of a dialogue generation apparatus according to an embodiment of this specification is shown. Figure 4 As shown, the device includes:

[0235] Clustering module 402 is configured to cluster semantically similar statements into at least one statement node, and select one or more typical statements for each statement node, wherein each statement node in the at least one statement node has a corresponding node identifier.

[0236] The first acquisition module 404 is configured to obtain the speech node identifier sequence corresponding to the dialogue based on the speech node to which each speech in the dialogue belongs;

[0237] The second acquisition module 406 is configured to acquire frequent subsequences in the speech node identifier sequence;

[0238] The evaluation module 408 is configured to obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and to obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

[0239] Optionally, the clustering module 402 is further configured to:

[0240] The utterances from multiple dialogues are input into a pre-trained deep neural network to obtain vectorized semantic representations of utterances.

[0241] The vectorized discourse semantic representation is clustered according to the clustering algorithm to obtain at least one cluster, and the at least one cluster is used as at least one discourse node, wherein the discourses in the same cluster are semantically similar discourses.

[0242] Optionally, the clustering module 402 is further configured to:

[0243] Obtain the cluster center of each utterance node and the perplexity of each utterance in each utterance node;

[0244] Based on the cluster center of each dialogue node and the perplexity of each dialogue in each dialogue node, one or more typical dialogues corresponding to each dialogue node are determined.

[0245] Optionally, the clustering module 402 is further configured to:

[0246] Based on the perplexity of each dialogue in each dialogue node and the cluster center of each dialogue node, the target ranking result of each dialogue in each dialogue node is obtained;

[0247] Based on the target sorting result, determine one or more typical dialogues corresponding to each dialogue node.

[0248] Optionally, the clustering module 402 is further configured to:

[0249] Determine the distance value between each dialogue in each dialogue node and the cluster center, and sort each dialogue in each dialogue node according to the distance value to obtain the distance value sorting result;

[0250] Based on the confusion level of each dialogue in each dialogue node, sort each dialogue in each dialogue node to obtain a confusion level sorting result;

[0251] Based on the distance value sorting result and the confusion degree sorting result, the target sorting result for each utterance in each dialogue node is obtained.

[0252] Optionally, the clustering module 402 is further configured to:

[0253] Based on the distance value sorting result, a distance value order score is determined for each speech in each speech node;

[0254] Based on the confusion ranking result, a confusion order score is determined for each dialogue in each dialogue node;

[0255] The target ranking score for each dialogue in each dialogue node is obtained by taking a weighted average of the distance value ranking score and the confusion degree ranking score.

[0256] The target ranking score is used to rank each speech in each speech node to obtain the target ranking result.

[0257] Optionally, the clustering module 402 is further configured to:

[0258] The ranked dialogue scripts are obtained from the target ranking results, and the ranked dialogue scripts are compared with the typical dialogue scripts in the typical dialogue script set for similarity. The ranked dialogue scripts are the ranked dialogue scripts in each dialogue node.

[0259] The sorted dialogues with a similarity less than or equal to a preset similarity threshold are stored in the typical dialogue set as typical dialogues;

[0260] If the similarity is greater than the preset similarity threshold, the next ranked script will be compared with the typical scripts in the typical script set.

[0261] When the sorting of dialogues is completed or the number of typical dialogues in the typical dialogue set is greater than or equal to a preset threshold, the typical dialogues in the typical dialogue set are used as one or more typical dialogues corresponding to each dialogue node.

[0262] Optionally, the clustering module 402 is further configured to:

[0263] Each utterance in each utterance node is input into a pre-trained statistical language model to obtain the perplexity of each utterance in each utterance node.

[0264] Optionally, the second acquisition module 406 is further configured to:

[0265] The utterance node identifier sequence is processed according to the sequence mining algorithm to obtain the prefix of the utterance node identifier sequence and the prefix projection corresponding to the prefix, and the number of utterance nodes with the same identifier in the prefix projection is determined.

[0266] Frequent subsequences are obtained based on the prefix and the number of discourse nodes with the same identifier.

[0267] Optionally, the second acquisition module 406 is further configured to:

[0268] If the number is less than a preset threshold, the prefix is ​​considered a frequent subsequence.

[0269] If the number is greater than or equal to a preset number threshold, the prefix is ​​updated based on the utterance nodes corresponding to the number to obtain the updated prefix;

[0270] The updated prefix projection corresponding to the updated prefix is ​​obtained from the speech node identifier sequence until the number of speech nodes with the same identifier in the updated prefix projection is less than the preset number threshold.

[0271] Optionally, the evaluation module 408 is further configured to:

[0272] The frequent subsequences are evaluated for quality to obtain a quality score corresponding to the frequent subsequences; and the quality score is compared with a preset quality threshold to obtain one or more frequent subsequences whose quality evaluation exceeds the preset quality threshold.

[0273] The dialogue generation device provided in this specification selects one or more typical dialogues for each clustered dialogue node, performs quality assessment on the frequent subsequences obtained from the dialogue node identifier sequence, obtains one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and quickly obtains the standard operating procedure for dialogue based on one or more frequent subsequences and one or more typical dialogues. This effectively improves the efficiency of obtaining the standard operating procedure for dialogue and avoids the problem of high labor costs for enterprises due to the large number and variety of dialogues, thus effectively reducing the labor costs of enterprises.

[0274] The above is an illustrative scheme of a dialogue generation device according to this embodiment. It should be noted that the technical solution of this dialogue generation device and the technical solution of the dialogue generation method described above belong to the same concept. For details not described in detail in the technical solution of the dialogue generation device, please refer to the description of the technical solution of the dialogue generation method described above.

[0275] Figure 5 A structural block diagram of a computing device 500 according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0276] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0277] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0278] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 500 can also be a mobile or stationary server.

[0279] The processor 520 is used to execute computer-executable instructions, which, when executed by the processor 520, implement the steps of any of the dialogue generation methods.

[0280] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the dialogue generation method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the dialogue generation method described above.

[0281] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of any of the described dialogue generation methods.

[0282] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the dialogue generation method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the dialogue generation method described above.

[0283] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0284] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0285] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0286] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0287] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A dialogue generation method, comprising: Sentences with similar semantics are clustered into at least one speech node, and one or more typical speech words are selected for each speech node, wherein each speech node in the at least one speech node has a corresponding node identifier; Based on the speech node to which each speech in the dialogue belongs, obtain the speech node identifier sequence corresponding to the dialogue; Obtaining frequent subsequences from the speech node identifier sequence, wherein obtaining frequent subsequences from the speech node identifier sequence includes: processing the speech node identifier sequence according to a sequence mining algorithm to obtain the prefix of the speech node identifier sequence and the prefix projection corresponding to the prefix, and determining the number of speech nodes with the same identifier in the prefix projection; obtaining frequent subsequences based on the prefix and the number of speech nodes with the same identifier; Obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

2. The dialogue generation method according to claim 1, wherein clustering semantically similar utterances into at least one utterance node comprises: The utterances from multiple dialogues are input into a pre-trained deep neural network to obtain vectorized semantic representations of utterances. The vectorized discourse semantic representation is clustered according to the clustering algorithm to obtain at least one cluster, and the at least one cluster is used as at least one discourse node, wherein the discourses in the same cluster are semantically similar discourses.

3. The dialogue generation method according to claim 1, wherein selecting one or more typical dialogues for each dialogue node includes: Obtain the cluster center of each utterance node and the perplexity value of each utterance in each utterance node; Based on the cluster center of each dialogue node and the confusion value of each dialogue in each dialogue node, one or more typical dialogues corresponding to each dialogue node are determined.

4. The dialogue generation method according to claim 3, wherein determining one or more typical dialogues corresponding to each dialogue node based on the cluster center of each dialogue node and the perplexity value of each dialogue in each dialogue node includes: Based on the confusion value of each dialogue in each dialogue node and the cluster center of each dialogue node, the target ranking result of each dialogue in each dialogue node is obtained; Based on the target sorting result, determine one or more typical dialogues corresponding to each dialogue node.

5. The dialogue generation method according to claim 4, wherein obtaining the target ranking result for each dialogue in each dialogue node based on the perplexity value of each dialogue in each dialogue node and the cluster center of each dialogue node includes: Determine the distance value between each dialogue in each dialogue node and the cluster center, and sort each dialogue in each dialogue node according to the distance value to obtain the distance value sorting result; Based on the confusion value corresponding to each dialogue in each dialogue node, sort each dialogue in each dialogue node to obtain the confusion value sorting result; Based on the distance value sorting result and the confusion value sorting result, the target sorting result for each utterance in each dialogue node is obtained.

6. The dialogue generation method according to claim 5, wherein obtaining the target ranking result for each utterance in each dialogue node based on the distance value ranking result and the confusion value ranking result includes: Based on the distance value sorting result, determine the corresponding distance value order score for each speech in each speech node; Based on the confusion value sorting result, determine the corresponding confusion value order score for each dialogue in each dialogue node; The target ranking score for each dialogue in each dialogue node is obtained by taking a weighted average of the distance value ranking score and the confusion value ranking score. The target ranking score is used to rank each speech in each speech node to obtain the target ranking result.

7. The dialogue generation method according to claim 4, wherein determining one or more typical dialogues corresponding to each dialogue node based on the target sorting result includes: The ranked dialogue scripts are obtained from the target ranking results, and the ranked dialogue scripts are compared with the typical dialogue scripts in the typical dialogue script set for similarity. The ranked dialogue scripts are the ranked dialogue scripts in each dialogue node. The sorted dialogues with a similarity less than or equal to a preset similarity threshold are stored in the typical dialogue set as typical dialogues; If the similarity is greater than the preset similarity threshold, the next ranked script will be compared with the typical scripts in the typical script set. When the sorting of dialogues is completed or the number of typical dialogues in the typical dialogue set is greater than or equal to a preset threshold, the typical dialogues in the typical dialogue set are used as one or more typical dialogues corresponding to each dialogue node.

8. The dialogue generation method according to claim 4, before obtaining the target ranking result of each dialogue in each dialogue node based on the perplexity value of each dialogue in each dialogue node and the cluster center of each dialogue node, further comprising: Each utterance in each utterance node is input into a pre-trained statistical language model to obtain the perplexity value corresponding to each utterance in each utterance node.

9. The dialogue generation method according to claim 1, wherein obtaining frequent subsequences based on the prefix and the number of utterance nodes with the same identifier comprises: If the number is less than a preset threshold, the prefix is ​​considered a frequent subsequence. If the number is greater than or equal to a preset number threshold, the prefix is ​​updated based on the utterance nodes corresponding to the number to obtain the updated prefix; The updated prefix projection corresponding to the updated prefix is ​​obtained from the speech node identifier sequence until the number of speech nodes with the same identifier in the updated prefix projection is less than the preset number threshold.

10. The dialogue generation method according to claim 1, wherein obtaining one or more frequent subsequences whose quality assessment exceeds a preset quality threshold comprises: The frequent subsequences are quality evaluated to obtain the quality scores corresponding to the frequent subsequences. The quality score is then compared with a preset quality threshold to obtain one or more frequent subsequences whose quality assessment exceeds the preset quality threshold.

11. A dialogue generation apparatus, comprising: The clustering module is configured to cluster semantically similar statements into at least one statement node, and select one or more typical statements for each statement node, wherein each statement node in the at least one statement node has a corresponding node identifier. The first acquisition module is configured to obtain the sequence of speech node identifiers corresponding to the dialogue based on the speech node to which each speech in the dialogue belongs; The second acquisition module is configured to acquire frequent subsequences from the speech node identifier sequence; The second acquisition module is further configured to process the speech node identifier sequence according to a sequence mining algorithm to obtain the prefix of the speech node identifier sequence and the prefix projection corresponding to the prefix, and determine the number of speech nodes with the same identifier in the prefix projection; and obtain frequent subsequences based on the prefix and the number of speech nodes with the same identifier. The evaluation module is configured to obtain one or more frequent subsequences whose quality assessment exceeds a preset quality threshold, and to obtain the standard operating procedure of the dialogue based on the one or more frequent subsequences and the one or more typical dialogues.

12. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the dialogue generation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Speech skill recommendation method and device and electronic equipment

    CN111522937A

  • Industrial control protocol message classification method and system based on dialogue analysis

    CN113139593A