Dialogue processing method and device, electronic equipment and storage medium

CN116303951BActive Publication Date: 2026-08-07BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-03-02
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]这类场景对客服人员的应答能力具有较高的要求,如果客户人员缺乏经验,则可能无法准确判断用户诉求,从而导致客服人员无法恰当地回应用户的问题,造成商单机会流失

Benefits of technology

[0021]根据本公开的还一方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现本公开上述一方面提出的对话处理方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303951B_ABST
    Figure CN116303951B_ABST
Patent Text Reader

Abstract

The present disclosure provides a dialogue processing method and device, electronic equipment and storage medium, relating to the fields of intelligent cloud, deep learning, natural language processing, cloud computing and the like. The implementation scheme is as follows: a classification label of a target dialogue in which a target input sentence is located is acquired, the classification label being used to indicate a dialogue intention of the target dialogue; at least one candidate dialogue fragment is acquired from a dialogue library according to the classification label and the target input sentence; a target reply sentence is determined from each candidate reply sentence according to a matching degree between the classification label, the target input sentence and each candidate reply sentence in each candidate dialogue fragment, so as to reply to the target input sentence according to the target reply sentence. Thus, by combining the sentence input by a user and the dialogue intention of the dialogue in which the sentence is located, the reply sentence corresponding to the sentence is determined, which can improve the accuracy of the determination of the reply sentence. In addition, the reply sentence can also be provided to a customer service personnel to assist the customer service personnel in quickly solving the problem of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of intelligent cloud, deep learning, natural language processing, cloud computing, etc., and particularly to dialogue processing methods, devices, electronic devices, and storage media. Background Technology

[0002] With the rapid development of artificial intelligence technology and online transaction scenarios, as well as the popularization and application of internet commerce, many companies are implementing intelligent marketing and intelligent customer service solutions. These solutions aim to guide users with specific scripts to achieve precise product or service sales.

[0003] These scenarios place high demands on customer service personnel's responsiveness. If customer service personnel lack experience, they may not be able to accurately judge user needs, resulting in customer service personnel being unable to respond appropriately to user questions and causing lost business opportunities.

[0004] Therefore, it is crucial to automatically recommend response statements based on user input, in order to assist customer service personnel in quickly resolving user issues, improving service efficiency, and ultimately enhancing the user's service experience. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, and storage medium for dialogue processing.

[0006] According to one aspect of this disclosure, a dialogue processing method is provided, comprising:

[0007] Retrieve the target input statement to be replied to;

[0008] Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue;

[0009] Based on the first classification label and the target input statement, at least one candidate speech fragment is obtained from the speech database; wherein, the candidate classification label in the candidate speech fragment is similar to the first classification label, and the candidate input statement in the candidate speech fragment is similar to the target input statement;

[0010] Based on the first classification label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech segment, a target response statement is determined from each candidate response statement, so as to respond to the target input statement according to the target response statement.

[0011] According to another aspect of this disclosure, a dialogue processing apparatus is provided, comprising:

[0012] The first acquisition module is used to acquire the target input statement to be replied to;

[0013] The second acquisition module is used to acquire the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue;

[0014] The third acquisition module is used to acquire at least one candidate speech fragment from the speech database based on the first classification label and the target input statement; wherein the candidate classification label in the candidate speech fragment is similar to the first classification label, and the candidate input statement in the candidate speech fragment is similar to the target input statement;

[0015] The first determining module is used to determine a target response statement from each of the candidate response statements based on the first classification label and the target input statement, and the matching degree between the target response statement and the candidate response statements in each of the candidate speech segments, so as to respond to the target input statement according to the target response statement.

[0016] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the dialogue processing method proposed in the above aspect of this disclosure.

[0020] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided for computer instructions used to cause the computer to perform the dialogue processing method proposed in the foregoing aspect of this disclosure.

[0021] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the dialogue processing method proposed in the above aspect of this disclosure.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0024] Figure 1This is a schematic flowchart of the dialogue processing method provided in Embodiment 1 of this disclosure;

[0025] Figure 2 This is a schematic flowchart of the dialogue processing method provided in Embodiment 2 of this disclosure;

[0026] Figure 3 This is a flowchart illustrating the dialogue processing method provided in Embodiment 3 of this disclosure;

[0027] Figure 4 This is a flowchart illustrating the dialogue processing method provided in Embodiment 4 of this disclosure;

[0028] Figure 5 This is a flowchart illustrating the dialogue processing method provided in Embodiment 5 of this disclosure;

[0029] Figure 6 This is a flowchart illustrating the dialogue processing method provided in Embodiment Six of this disclosure;

[0030] Figure 7 This is a schematic flowchart of the dialogue processing method provided in Embodiment 7 of this disclosure;

[0031] Figure 8 This is a schematic diagram of the dialogue processing flow provided in the embodiments of this disclosure;

[0032] Figure 9 A schematic diagram of each dialogue segment in a marketing dialogue scenario provided in the embodiments of this disclosure;

[0033] Figure 10 This is a schematic diagram of speech extraction provided in an embodiment of the present disclosure;

[0034] Figure 11 This is a schematic diagram of the script matching process provided in an embodiment of the present disclosure;

[0035] Figure 12 This is a schematic diagram of the structure of the dialogue processing device provided in Embodiment 8 of this disclosure;

[0036] Figure 13 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0037] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0038] With the continuous development of artificial intelligence technology and the popularization and application of internet commerce, many companies are implementing intelligent marketing and intelligent customer service solutions. These solutions aim to use specific communication techniques to guide user needs and accurately promote products or services to users based on those needs.

[0039] These scenarios require experienced sales personnel to manually summarize excellent sales scripts, which places high demands on the responsiveness of customer service staff. However, the number of manually summarized scripts is limited. When users raise complex questions or statements, customer service staff, due to their lack of experience, often struggle to accurately assess user needs and respond appropriately, resulting in lost sales opportunities.

[0040] Specifically, in the script production stage, excellent scripts are manually compiled by professional call center staff. In the script recommendation stage, based on the questions or statements input by users, similar questions and corresponding candidate script sets are obtained by clustering based on semantic features. Higher-dimensional text features are constructed for each candidate script set, and the text features are classified by a classification model to obtain the classification probability of each candidate script set. This classification probability is used to indicate the probability that the corresponding candidate script set will become the recommended script. Then, the recommended script can be determined from each candidate script set based on the classification probability.

[0041] The above method has at least the following disadvantages:

[0042] First, excellent sales scripts are manually summarized by professional sales script personnel. The templates are simple and the number is small. Although they can solve common user problems, the coverage is not broad enough and cannot meet the needs of a large number of user problems that are beyond the scope of the task.

[0043] Second, in the speech recommendation stage, relying solely on semantic matching features and single classification features to generate candidate speech has a coarse feature granularity, resulting in a broad range of recommended speech content that cannot truly address the user's real needs.

[0044] Third, dialogues in marketing scenarios have distinct phases, and user feedback can be summarized as regular bottlenecks. Faced with different user bottlenecks at different stages, it is necessary to flexibly select different responses to users, but the above methods cannot meet this kind of refined response recommendation needs.

[0045] In view of at least one of the above-mentioned problems, this disclosure provides a dialogue processing method, apparatus, electronic device and storage medium.

[0046] The following description, with reference to the accompanying drawings, outlines a dialogue processing method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure.

[0047] Figure 1 This is a flowchart illustrating the dialogue processing method provided in Embodiment 1 of this disclosure.

[0048] This disclosure illustrates an example where the dialogue processing method is configured in a dialogue processing device, which can be applied to any electronic device to enable the electronic device to perform dialogue processing functions.

[0049] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0050] like Figure 1 As shown, the dialogue processing method may include the following steps:

[0051] Step 101: Obtain the target input statement to be replied to.

[0052] In this embodiment of the disclosure, the target input statement can be a dialogue statement or question statement entered by the user, and the input method includes, but is not limited to, touch input (such as swiping, clicking, etc.), keyboard input, voice input, etc. The target input statement may include at least one of text information, image information, audio information, and video information.

[0053] Optionally, when the target input statement includes image information, audio information, and video information, OCR (Optical Character Recognition) can be performed on the image information, speech recognition can be performed on the audio information, and subtitle recognition can be performed on the video information to obtain the target input statement in text form.

[0054] Step 102: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0055] In this embodiment of the disclosure, each statement entered by the user in the target dialogue where the target input statement is located can be classified to obtain a first classification label, wherein the first classification label is used to indicate the dialogue intent of the target dialogue.

[0056] As an example, in the marketing scenario of a service or product, the primary category label may include, but is not limited to, labels such as no demand, doubts about effectiveness, doubts about price, doubts about trust, doubts about business capabilities, unsubscription, emotional outburst, and other categories.

[0057] Step 103: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database; wherein the candidate category label in the candidate speech fragment is similar to the first category label, and the candidate input statement in the candidate speech fragment is similar to the target input statement.

[0058] In this embodiment of the disclosure, the dialogue library may include multiple sample dialogue fragments, wherein the sample dialogue fragments include an input statement, a category label (used to indicate the dialogue intent of the dialogue in which the input statement is located), and a response statement corresponding to the input statement. For example, the sample dialogue fragments may be in the form of: category label##input statement##response statement.

[0059] In this embodiment of the disclosure, at least one candidate speech segment can be determined from multiple sample speech segments in the speech library based on the first classification label and the target input statement. The classification label (referred to as the candidate classification label in this disclosure) in the candidate speech segment is similar to or consistent with the first classification label, and the input statement (referred to as the candidate input statement in this disclosure) in the candidate speech segment is similar to the target input statement. For example, the semantic similarity between the candidate input statement and the target input statement is higher than a set similarity threshold.

[0060] Step 104: Based on the first category label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech fragment, determine the target response statement from each candidate response statement, and respond to the target input statement according to the target response statement.

[0061] In this embodiment of the disclosure, for any candidate speech fragment, the matching degree between the first category label and the target input statement and the response statement in the candidate speech fragment (referred to as the candidate response statement in this disclosure) can be calculated. Thus, in this disclosure, the target response statement can be determined from each candidate response statement based on the matching degree of each candidate response statement.

[0062] For example, the candidate response statement with the highest matching degree can be used as the target response statement.

[0063] For example, candidate response statements with a matching degree higher than a set threshold can be used as target response statements.

[0064] For example, candidate response statements can be sorted from largest to smallest according to their matching degree, and a set number of candidate response statements at the top of the list can be used as the target response statements.

[0065] In this disclosure, a response can be given to a target input statement based on the target response statement.

[0066] In one possible implementation of this disclosure, the target input statement can be automatically replied to based on the target response statement.

[0067] In another possible implementation of this disclosure, the target response statement can be provided to customer service personnel, who can then manually respond to the target input statement in the target dialogue based on the target response statement. In other words, the function of the target response statement is to assist customer service personnel in manually responding to the target input statement entered by the user.

[0068] The dialogue processing method of this disclosure involves obtaining a first category label for the target dialogue containing the target input statement; wherein the first category label indicates the dialogue intent of the target dialogue; obtaining at least one candidate dialogue segment from a dialogue library based on the first category label and the target input statement; wherein the candidate category labels in the candidate dialogue segments are similar to the first category label, and the candidate input statements in the candidate dialogue segments are similar to the target input statement; determining the target response statement from the candidate response statements based on the matching degree between the first category label and the target input statement and the candidate response statements in each candidate dialogue segment, and responding to the target input statement based on the target response statement. This allows for the simultaneous combination of the user-input statement and the dialogue intent of the dialogue containing the statement to determine the corresponding response statement, improving the accuracy of response statement determination. Furthermore, in scenarios involving manual customer service responses, the determined response statement can be provided to customer service personnel to assist them in quickly resolving user issues, improving both service efficiency and user experience.

[0069] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein are all carried out with the consent of the user, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.

[0070] To clearly illustrate how the above embodiments determine the target response statement from each candidate response statement based on the matching degree between the first classification label and the target input statement and the candidate response statements in each candidate dialogue segment, this disclosure also proposes a dialogue processing method.

[0071] Figure 2 This is a flowchart illustrating the dialogue processing method provided in Embodiment 2 of this disclosure.

[0072] like Figure 2 As shown, the dialogue processing method may include the following steps:

[0073] Step 201: Obtain the target input statement to be replied to.

[0074] Step 202: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0075] Step 203: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database; wherein the candidate category label in the candidate speech fragment is similar to the first category label, and the candidate input statement in the candidate speech fragment is similar to the target input statement.

[0076] The explanation of steps 201 to 203 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0077] Step 204: For any candidate response statement in a candidate speech fragment, concatenate the first category label, the target input statement, and the candidate response statement to obtain the concatenated text.

[0078] In this embodiment of the disclosure, for any candidate speech fragment, the first category label and the target input statement can be concatenated with the candidate response statement in the candidate speech fragment to obtain the concatenated text.

[0079] As an example, the target input statement can be appended after the first category label, and the candidate response statement can be appended after the target input statement to obtain the concatenated text.

[0080] Step 205: Extract features from the concatenated text to obtain text features.

[0081] In this embodiment of the disclosure, features can be extracted from the spliced ​​text to obtain text features, such as high-dimensional feature vectors.

[0082] Step 206: Classify the text features to obtain the classification probability of the candidate response statement, where the classification probability is used to indicate the degree of matching between the candidate response statement and the target input statement.

[0083] In this embodiment of the disclosure, text features can be classified to obtain the classification probability of candidate response statements, wherein the classification probability is used to indicate the degree of matching between the candidate response statement and the target input statement.

[0084] As a possible approach, to improve the accuracy of classification probability calculations, deep learning techniques can be used to classify text features and obtain the classification probabilities of candidate response statements.

[0085] As an example, text features can be classified based on a binary classification model to obtain the classification probability of whether the candidate response statement matches the target input statement. When the classification probability is greater than a certain threshold, it indicates that the candidate response statement matches the target input statement, while when the classification probability is less than or equal to the certain threshold, it indicates that the candidate response statement does not match the target input statement.

[0086] Step 207: Determine the target response statement from the candidate response statements based on the classification probability of each candidate response statement, and respond to the target input statement according to the target response statement.

[0087] In this embodiment of the disclosure, the target response statement can be determined from the candidate response statements based on the classification probability of each candidate response statement.

[0088] As an example, the candidate response statement with the highest classification probability can be used as the target response statement.

[0089] As another example, candidate response statements with a classification probability higher than a set threshold can be used as target response statements.

[0090] As another example, candidate response statements can be sorted from largest to smallest according to their classification probability values, and a set number of candidate response statements at the top of the sort can be used as the target response statements.

[0091] Therefore, the target response statement can be determined based on different methods, which can improve the flexibility and applicability of the method.

[0092] In this disclosure, a response can be given to a target input statement based on the target response statement.

[0093] In one possible implementation of this disclosure, the target input statement can be automatically replied to based on the target response statement.

[0094] In another possible implementation of this disclosure, the target response statement can be provided to customer service personnel, who can then manually respond to the target input statement in the target dialogue based on the target response statement. In other words, the function of the target response statement is to assist customer service personnel in manually responding to the target input statement entered by the user.

[0095] The dialogue processing method of this disclosure can effectively determine the target response statement from each candidate response statement based on the matching degree between the candidate response statement and the target input statement, thereby improving the accuracy and reliability of the target response statement determination.

[0096] To clearly illustrate how the dialogue database is built in the above embodiments, this disclosure also proposes a dialogue processing method.

[0097] Figure 3 This is a flowchart illustrating the dialogue processing method provided in Embodiment 3 of this disclosure.

[0098] like Figure 3 As shown, the dialogue processing method may include the following steps:

[0099] Step 301: Obtain the historical dialogue statements from at least one round of dialogue.

[0100] In this embodiment of the disclosure, historical dialogue statements from at least one round of dialogue can be obtained from the dialogue log.

[0101] It should be noted that the original dialogue statements may contain meaningless interjections, repetitive segments, and typos. Directly inputting these original dialogue statements into the model for recognition or classification may cause semantic bias. Therefore, to address the above situation, in one possible implementation of this disclosure, the dialogue log can be parsed to obtain multiple original dialogue statements (referred to as initial dialogue statements in this disclosure) from at least one round of dialogue. These initial dialogue statements from each round of dialogue can then be preprocessed to obtain each historical dialogue statement from each round of dialogue. The preprocessing includes at least one of the following: interjection removal, repetition removal, typo correction, and colloquial rewriting.

[0102] This allows for the preprocessing of dialogue statements in each round of conversation, resulting in clear and fluent historical dialogue statements, which can improve the accuracy of subsequent classification results.

[0103] Step 302: Divide the historical dialogue statements in any round of dialogue to obtain at least one text fragment.

[0104] In this embodiment of the disclosure, for any round of dialogue, the historical dialogue statements in the dialogue can be divided to obtain at least one text segment.

[0105] As one possible implementation, the historical dialogue statements in the dialogue can be classified based on a classification task to obtain the classification probability of each historical dialogue statement (a probability value used to indicate whether the end of a historical dialogue statement is a segment). Based on the classification probability of each historical dialogue statement, the historical dialogue statements in the dialogue can be divided to obtain at least one text segment.

[0106] As an example, the end of each historical dialogue statement in a round of dialogue may serve as a segmentation point. By classifying the end of each historical dialogue statement in a round of dialogue using a pre-trained classification model, the classification probability of the end of each historical dialogue statement is obtained (i.e., the probability value that the end of the historical dialogue statement is a segmentation point). When the classification probability of the end of a certain historical dialogue statement is greater than a certain threshold, it means that the historical dialogue statement will be re-divided into a natural segment (referred to as a text segment in this disclosure).

[0107] For example, suppose a dialogue includes historical dialogue statement 1, historical dialogue statement 2, historical dialogue statement 3, historical dialogue statement 4, and historical dialogue statement 5. Suppose that the classification probability at the end of historical dialogue statement 2 is greater than a certain threshold, and the classification probability at the end of historical dialogue statement 4 is greater than a certain threshold, then three text segments can be obtained. One text segment contains historical dialogue statements 1 and 2, another text segment contains historical dialogue statements 3 and 4, and yet another text segment contains historical dialogue statement 5.

[0108] Step 303: Classify each text fragment to obtain a second classification label for each text fragment, wherein the second classification label is used to indicate the dialogue segment to which the text fragment belongs.

[0109] In this embodiment of the disclosure, each text fragment can be classified to obtain a second classification label for each text fragment, wherein the second classification label is used to indicate the dialogue segment to which the corresponding text fragment belongs.

[0110] As an example, in a marketing scenario for a service or product, the dialogue process may include, but is not limited to: opening remarks, product introduction, needs inquiry, background research, case introduction, casual conversation, obtaining phone numbers and contact information, and closing remarks.

[0111] In one possible implementation of this disclosure, in order to improve the accuracy of the classification results, deep learning technology can be used to classify each text segment to obtain a second classification label for each text segment.

[0112] As an example, for any text fragment, it can be input into a first classification model for classification, obtaining the classification probabilities of multiple classification labels output by the first classification model. Based on these probabilities, a second classification label can be determined from the multiple classification labels. For instance, the classification label with the highest probability can be used as the second classification label.

[0113] The first classification model is trained based on sample text fragments, which are labeled with a first annotation label to indicate the dialogue segment to which the sample text fragment belongs.

[0114] For example, a sample text fragment can be input into a first classification model for classification to obtain the classification probabilities of multiple classification labels. Based on the classification probabilities of multiple classification labels, a predicted classification label can be determined from the multiple classification labels. Thus, in this disclosure, the first classification model can be trained based on the difference between the predicted classification label and the first annotation label labeled on the sample text fragment.

[0115] In one example, the value of the loss function (hereinafter referred to as the first loss value) can be determined based on the difference between the predicted classification label and the first labeled label. The model parameters in the first classification model can then be adjusted based on the first loss value to minimize the first loss value.

[0116] It should be noted that the above example only uses minimizing the first loss value as the termination condition for model training. In actual applications, other termination conditions can be set. For example, termination conditions can also include training time reaching a set time, training times reaching a set number of times, etc. This disclosure does not impose any restrictions on this.

[0117] Therefore, it is possible to classify text fragments based on deep learning technology and obtain secondary classification labels for each text fragment, which can improve the accuracy of classification results.

[0118] Step 304: If a category label is set in each of the second category labels, generate at least one sample dialogue fragment based on each historical dialogue statement in the dialogue.

[0119] In this embodiment of the disclosure, the classification label is set to a pre-set classification label. For example, in a marketing scenario, in order to increase the chances of closing a deal, the dialogue stage indicated by the classification label can be the contact information stage.

[0120] In this embodiment of the disclosure, it can be determined whether there is a set category label in each of the second category labels. If there is no set category label in each of the second category labels, no processing is required, that is, there is no need to generate sample dialogue fragments in the dialogue library based on each historical dialogue statement in the dialogue. However, if there is a set category label in each of the second category labels, at least one sample dialogue fragment can be generated based on each historical dialogue statement in the dialogue.

[0121] Step 305: Establish a speech database based on each sample speech fragment.

[0122] In this embodiment of the disclosure, a speech library can be established based on each sample speech segment, that is, each sample speech segment can be stored in the speech library.

[0123] It should be noted that this disclosure only exemplifies the execution of steps 301 to 305 before step 306, but this disclosure is not limited to this. In practical applications, steps 301 to 305 only need to be executed before step 308. For example, steps 301 to 305 can also be executed after step 306 and before step 307; for another example, steps 301 to 305 can also be executed after step 307 and before step 308; for yet another example, steps 301 to 305 can be executed concurrently with step 306; for yet another example, steps 301 to 305 can also be executed concurrently with step 307, and so on. This disclosure does not impose any limitations in this regard.

[0124] Step 306: Obtain the target input statement to be replied to.

[0125] Step 307: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0126] Step 308: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database.

[0127] Among them, the candidate category labels in the candidate speech fragments are similar to the first category label, and the candidate input statements in the candidate speech fragments are similar to the target input statements.

[0128] Step 309: Based on the first category label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech fragment, determine the target response statement from each candidate response statement, so as to respond to the target input statement according to the target response statement.

[0129] The explanation of steps 306 to 309 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0130] The dialogue processing method of this disclosure can generate sample dialogue fragments based on historical dialogue statements, thereby improving the effectiveness of dialogue database construction. Furthermore, by automatically constructing a dialogue database based on a large number of historical dialogue statements in the dialogue log, the number of sample dialogue fragments in the database can be increased, resulting in broader coverage and solving the problems of limited quantity and limited templates in manually summarizing sample dialogue fragments.

[0131] To clearly illustrate how at least one sample dialogue fragment is generated based on each historical dialogue statement in the above embodiments, this disclosure also proposes a dialogue processing method.

[0132] Figure 4 This is a schematic flowchart of the dialogue processing method provided in Embodiment 4 of this disclosure.

[0133] like Figure 4As shown, the dialogue processing method may include the following steps:

[0134] Step 401: Obtain the historical dialogue statements from at least one round of dialogue.

[0135] Step 402: Divide the historical dialogue statements in any round of dialogue to obtain at least one text fragment.

[0136] Step 403: Classify each text fragment to obtain a second classification label, wherein the second classification label is used to indicate the dialogue segment to which the text fragment belongs.

[0137] The explanations of steps 401 to 403 can be found in the relevant descriptions in any embodiment of this disclosure, and will not be repeated here.

[0138] Step 404: If a classification label exists in each of the second category labels, group the historical dialogue statements in the dialogue to obtain at least one dialogue pair.

[0139] The dialogue pair includes historical input statements and corresponding historical response statements.

[0140] The explanation of setting category labels can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0141] In this embodiment of the disclosure, when there are set category labels in each of the second category labels, the historical dialogue statements in the above-mentioned dialogue can be divided or grouped to obtain at least one dialogue pair. Each dialogue pair includes a user-input statement (referred to as a historical input statement in this disclosure) and a corresponding response statement (referred to as a historical response statement in this disclosure).

[0142] In other words, each historical dialogue statement includes historical input data and historical response statements. A historical input statement and its corresponding historical response statement can be grouped together to obtain a dialogue pair.

[0143] Step 405: For any dialogue pair, obtain the third category label of the dialogue containing the historical input statements in the dialogue pair.

[0144] The third category label is used to indicate the dialogue intent of the dialogue in which the historical input statement is located.

[0145] In this embodiment of the disclosure, for any given dialogue pair, a third category label can be obtained for the dialogue containing the historical input statements in that dialogue pair. For example, each statement entered by the user in the dialogue containing the historical input statements can be classified to obtain a third category label, wherein the third category label is used to indicate the dialogue intent of the dialogue containing the historical input statements.

[0146] As an example, in the marketing scenario of services or products, the third category label may include, but is not limited to, labels such as no demand, doubts about effectiveness, doubts about price, doubts about trust, doubts about business capabilities, unsubscription, emotional outburst, and other categories.

[0147] Step 406: Generate sample dialogue fragments based on the third category label, the historical input statements and historical response statements in the dialogue pair.

[0148] In this embodiment of the disclosure, sample dialogue fragments can be generated based on the third category label, historical input statements, and historical response statements in the dialogue pair. For example, the sample dialogue fragment can be in the form of: third category label ##historical input statement ##historical response statement.

[0149] Step 407: Establish a speech database based on each sample speech fragment.

[0150] Step 408: Obtain the target input statement to be replied to.

[0151] Step 409: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0152] Step 410: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database.

[0153] Among them, the candidate category labels in the candidate speech fragments are similar to the first category label, and the candidate input statements in the candidate speech fragments are similar to the target input statements.

[0154] Step 411: Based on the first category label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech fragment, determine the target response statement from each candidate response statement, so as to respond to the target input statement according to the target response statement.

[0155] The explanation of steps 407 to 411 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0156] The dialogue processing method of this disclosure can effectively generate sample dialogue fragments based on classification tags, historical input statements, and historical response statements, thereby improving the effectiveness of dialogue database establishment.

[0157] To clearly illustrate the above embodiments of this disclosure, this disclosure also proposes a dialogue processing method.

[0158] Figure 5 This is a flowchart illustrating the dialogue processing method provided in Embodiment 5 of this disclosure.

[0159] like Figure 5 As shown, the dialogue processing method may include the following steps:

[0160] Step 501: Obtain the target input statement to be replied to.

[0161] The explanation of step 501 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0162] Step 502: Determine the target dialog containing the target input statement.

[0163] In this embodiment of the disclosure, the dialogue where the target input data is located can be obtained, referred to as the target dialogue in this disclosure.

[0164] Step 503: Obtain at least one first input statement from the target dialogue.

[0165] In this embodiment of the disclosure, at least one statement input by the user (referred to as the first input statement in this disclosure) can be obtained from the target dialogue.

[0166] Step 504: Use a second classification model to classify the target input statement and at least one first input statement to obtain the classification probabilities of multiple predicted labels output by the second classification model.

[0167] In this embodiment of the disclosure, the target input statement and at least one first input statement can be input into a second classification model for classification, thereby obtaining the classification probabilities of multiple predicted labels output by the second classification model.

[0168] As an example, in the marketing scenario of services or products, predictive tags may include, but are not limited to, tags for no demand, questioning effectiveness, questioning price, questioning trust, questioning business capabilities, unsubscribing, emotional outburst, and other categories.

[0169] The second classification model is trained based on sample statements, which are labeled with a second annotation label to indicate the dialogue intent of the dialogue in which the sample statement is located.

[0170] For example, a sample statement can be input into a second classification model for classification to obtain the classification probabilities of multiple predicted labels. Based on the classification probabilities of multiple predicted labels, a target classification label can be determined from the multiple predicted labels. Thus, in this disclosure, the second classification model can be trained based on the difference between the target classification label and the second annotation label labeled on the sample statement.

[0171] In one example, the value of the loss function (hereinafter referred to as the second loss value) can be determined based on the difference between the target classification label and the second labeled label. The model parameters in the second classification model can then be adjusted based on the second loss value to minimize the second loss value.

[0172] It should be noted that the above example only uses minimizing the second loss value as the termination condition for model training. In actual applications, other termination conditions can be set, such as the training time reaching a set time or the number of training iterations reaching a set number, etc. This disclosure does not impose any restrictions on this.

[0173] Step 505: Determine the first classification label from the multiple predicted labels based on the classification probabilities of the multiple predicted labels.

[0174] In this embodiment of the disclosure, a first classification label can be determined from multiple predicted labels based on their classification probabilities. For example, the predicted label with the highest classification probability can be used as the first classification label. The first classification label is used to indicate the dialogue intent of the target conversation.

[0175] Step 506: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database.

[0176] Among them, the candidate category labels in the candidate speech fragments are similar to the first category label, and the candidate input statements in the candidate speech fragments are similar to the target input statements.

[0177] Step 507: Based on the first category label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech fragment, determine the target response statement from each candidate response statement, and respond to the target input statement according to the target response statement.

[0178] The explanation of steps 506 to 507 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.

[0179] The dialogue processing method of this disclosure can classify each statement input by the user in the target dialogue based on deep learning technology to obtain a first classification label for indicating the dialogue intent of the target dialogue, thereby improving the accuracy of the classification results.

[0180] In one possible implementation of this disclosure, in a scenario where customer service personnel provide manual responses, the target response statement can be provided to the customer service personnel to assist them in quickly resolving user issues. The following is in conjunction with... Figure 6 The above process will be explained in detail.

[0181] Figure 6This is a schematic flowchart of the dialogue processing method provided in Embodiment Six of this disclosure.

[0182] like Figure 6 As shown, the dialogue processing method may include the following steps:

[0183] Step 601: Obtain the target input statement to be replied to.

[0184] Step 602: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0185] Step 603: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database.

[0186] Among them, the candidate category labels in the candidate speech fragments are similar to the first category label, and the candidate input statements in the candidate speech fragments are similar to the target input statements.

[0187] Step 604: Based on the first category label and the target input statement, and the matching degree between them and the candidate response statements in each candidate speech fragment, determine multiple target response statements from each candidate response statement.

[0188] For an explanation of steps 601 to 604, please refer to the relevant description in any embodiment of this disclosure.

[0189] In this embodiment of the disclosure, multiple target response statements can be determined from each candidate response statement based on the matching degree between the first classification label and the target input statement and the candidate response statements in each candidate speech fragment.

[0190] As an example, candidate response statements can be sorted from largest to smallest according to their matching degree, and a set number of candidate response statements at the top of the list can be used as the target response statements.

[0191] As another example, candidate response statements with a matching degree higher than a set similarity threshold can be used as target response statements.

[0192] Step 605: Sort the multiple target response statements in descending order of classification probability to obtain the first sorting sequence, and send the first sorting sequence to the target customer service so that the target customer service can reply to the target input statement according to the first sorting sequence.

[0193] Among them, the target customer service representative is the customer service staff who handle the user's input statements or questions in the target dialogue.

[0194] In this embodiment of the disclosure, multiple target reply statements can be sorted from largest to smallest according to their classification probability to obtain a first sorting sequence, and the first sorting sequence can be sent to the target customer service to assist the target customer service in replying to the target input statement according to the first sorting sequence.

[0195] Step 606: From each candidate script segment, determine the target script segment where each target reply statement is located; sort each target script segment according to the classification probability of each target reply statement to obtain a second sorting sequence; send the second sorting sequence to the target customer service so that the target customer service can reply to the target input statement according to the second sorting sequence.

[0196] In this embodiment of the disclosure, the target script segment containing each target response statement can be determined from each candidate script segment, and the target script segments can be sorted according to the classification probability of the target response statement in each target script segment to obtain a second sorting sequence. For example, multiple target script segments can be sorted from largest to smallest according to the classification probability of the target response statement in the multiple target script segments to obtain the second sorting sequence. Furthermore, the second sorting sequence can be sent to the target customer service representative to assist the target customer service representative in replying to the target input statement according to the second sorting sequence.

[0197] The dialogue processing method of this disclosure can not only push multiple reply statements to customer service personnel so that they can respond to users based on these multiple reply statements, but also push multiple script fragments to customer service personnel so that they can respond to users by combining the input statements and reply statements in the script fragments. This not only enables accurate recommendation of statements or script fragments, but also helps customer service personnel to quickly and accurately solve user problems, helps them understand business processes more quickly, and improves service efficiency.

[0198] In one possible implementation of this disclosure, in a scenario where customer service personnel provide manual responses, the script segment containing the target response statement can be provided to the customer service personnel to assist them in quickly resolving user issues. The following is in conjunction with... Figure 7 The above process will be explained in detail.

[0199] Figure 7 This is a schematic flowchart of the dialogue processing method provided in Embodiment 7 of this disclosure.

[0200] like Figure 7 As shown, the dialogue processing method may include the following steps:

[0201] Step 701: Obtain the target input statement to be replied to.

[0202] Step 702: Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0203] Step 703: Based on the first category label and the target input statement, obtain at least one candidate speech fragment from the speech database.

[0204] Among them, the candidate category labels in the candidate speech fragments are similar to the first category label, and the candidate input statements in the candidate speech fragments are similar to the target input statements.

[0205] Step 704: Based on the first category label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech fragment, determine the target response statement from each candidate response statement.

[0206] For an explanation of steps 701 to 704, please refer to the relevant description in any embodiment of this disclosure.

[0207] In any embodiment of this disclosure, the candidate dialogue segment may further include at least one associated response statement, wherein the associated response statement and the candidate response statement in the candidate dialogue segment are in the same round of dialogue, and the response time of the associated response statement is later than the response time of the candidate response statement. For example, the format of the candidate dialogue segment may be: candidate category label##candidate input statement##candidate response statement##associated response statement##associated response statement##…….

[0208] For example, for a certain candidate dialogue segment, the candidate response statement in the candidate dialogue segment is the i-th statement of the customer service reply in a certain dialogue, and the associated response statement can be the i+1-th statement of the customer service reply, the i+2-th statement of the customer service reply, and so on.

[0209] Step 705: Determine the target speech segment containing the target response statement from among the candidate speech segments.

[0210] In this embodiment of the disclosure, the target speech segment containing the target response statement can be determined from each candidate speech segment, that is, the target speech segment is a candidate speech segment containing the target response statement.

[0211] Step 706: Send the target script segment to the target customer service representative, or send the target response statement and at least one target related response statement from the target script segment to the target customer service representative.

[0212] In this embodiment of the disclosure, a target dialogue segment can be sent to the target customer service representative to assist the representative in responding to the target input data based on the target response statement in the target dialogue segment. Furthermore, it can also assist the target customer service representative in responding to subsequent user input statements or questions based on at least one associated response statement (referred to as the target associated response statement in this disclosure), that is, responding to the input statement following the target input statement in the target dialogue.

[0213] In this embodiment of the disclosure, the target response statement and at least one target associated response statement in the target dialogue segment can also be sent directly to the target customer service representative, so that the target customer service representative can respond according to the target response statement, respond to the target input statement, and respond to the input statement in the target dialogue that is located after the target input statement according to at least one target associated response statement, that is, respond to each input statement whose input time is later than the target input statement.

[0214] The dialogue processing method of this disclosure can not only push the reply statements corresponding to the current user's input statement or question to customer service personnel, but also push the reply statements corresponding to the user's subsequent possible input statements or questions to customer service personnel in advance. This can help customer service personnel quickly and accurately solve user problems, help customer service personnel understand business processes more quickly, and improve service efficiency.

[0215] In any embodiment of this disclosure, a large number of dialogue logs can be analyzed to identify user bottlenecks (i.e., input statements), and excellent scripts can be generated in batches to solve frequently occurring but relatively simple user problems. At the same time, when faced with complex user needs, candidate script fragments can be recommended in real time and accurately to assist customer service personnel in judging and quickly resolving user problems, helping customer service personnel to understand business processes more quickly, improve service efficiency, thereby enhancing users' service experience of the product, creating more business opportunities, and bringing potential commercial value to the enterprise.

[0216] The dialogue processing flow can be as follows Figure 8 As shown, the process is divided into an offline dialogue mining stage and an online dialogue recommendation stage. In the offline dialogue mining stage (i.e., the offline production stage), a large number of historical dialogue statements can be de-colloquialized. Then, based on the question-and-answer sentence pairs between customers and customer service representatives, the statements are stored in the Elasticsearch (a non-relational distributed full-text search framework that uses an inverted index storage method, suitable for complex search and full-text search scenarios) database for subsequent retrieval. After that, checkpoint identification (i.e., intent identification) and stage identification can be performed. Then, dialogue fragments corresponding to checkpoint questions (i.e., input statements) are extracted based on checkpoint and stage features and stored in the dialogue database. In the online dialogue recommendation stage, based on the input statements of online users, the user intent can be converted into text features and checkpoint features and input into the dialogue matching model to finally obtain high-scoring candidate dialogue fragments.

[0217] Key technical points include:

[0218] 1. De-colloquialize.

[0219] Marketing dialogue data is typically converted from speech recognition to text. This type of text is heavily influenced by colloquial speech, containing many meaningless interjections, repetitive phrases, and typos. Directly inputting it into a model for processing can cause semantic bias, indirectly affecting the quality of subsequent extracted dialogue. Therefore, this disclosure firstly involves de-colloquializing the raw dialogue data, including removing interjections, repeating words, correcting typos, and rewriting it in a colloquial style. The aim is to cleanse the data and produce clear and fluent dialogue content. Then, the dialogue data can be stored in a database in a question-and-answer format between customer and customer service representatives. This preserves historical dialogue data or statements for use in the subsequent dialogue extraction stage.

[0220] Since the dialogue data is quite large, an Elasticsearch database can be used to retrieve dialogue fragments more efficiently.

[0221] 2. Checkpoint recognition, i.e., dialogue intent recognition.

[0222] In marketing scenarios, conversations typically revolve around promoting a specific service or product. This process involves various user feedback regarding customer service responses or product / service offerings. This disclosure aims to extract scripts based on these feedback characteristics to proactively reassure and guide users, increasing the likelihood of a sale. Therefore, based on user conversation content, eight types of checkpoint tags have been summarized: no need, doubt about effectiveness, doubt about price, doubt about trust, doubt about business capabilities, unsubscribing, emotional agitation, and other. It should be noted that the above is only an example of eight checkpoint tags; in actual application, checkpoint tags can be flexibly adjusted according to the specific scenario.

[0223] Taking the "no need" tag as an example, when a user's dialogue statement (i.e., input statement) contains words such as "not needed," "not considering this product for the time being," or "there is no such need at present," then the dialogue statement belongs to the "no need" tag.

[0224] Based on a large amount of historical dialogue data, a small number of checkpoint labels can be manually annotated to pre-train a classification model for checkpoint recognition (referred to as the second classification model in this disclosure). The input of the classification model is a dialogue statement entered by the user. Through a series of semantic calculations, the output is the probability value of the dialogue statement belonging to each type of checkpoint label. The checkpoint label with the highest probability value is taken as the checkpoint label of the dialogue statement.

[0225] 3. Process identification.

[0226] A complete marketing conversation typically includes Figure 9The entire process illustrated shows that user feedback at each stage of the dialogue has different characteristics. Guiding users quickly to the closing stage requires tailoring appropriate scripts based on the characteristics of each stage. Therefore, dialogue stage identification is a relatively important part of script mining. Specifically, in real-world scenarios, depending on the customer service representative's responses, the dialogue may end at any stage. A successful sales pitch has a relatively long dialogue duration and can flow completely to the lead-up and contact stage. In the lead-up and contact stage, users typically show a willingness to continue learning about or trying the product or service, intending to engage in deeper cooperation. Therefore, when a dialogue successfully flows to the lead-up and contact stage, all user input statements and their corresponding customer service responses are considered valid scripts.

[0227] To accurately identify each stage of a dialogue, a stage recognition model can be pre-trained. Stage recognition is divided into two stages: dialogue segmentation and stage classification. Dialogue segmentation can be transformed into a classification task, where the end of each sentence in a round of dialogue may serve as a segmentation point. The pre-trained classification model outputs the probability value of the end of each sentence in a round of dialogue being a segmentation point. When the probability value is greater than a certain threshold, it indicates that the sentence will be re-divided into a natural segment (referred to as a text fragment in this disclosure).

[0228] Based on the divided natural segments (i.e. text fragments), each natural segment is input into a pre-trained segment classification model (referred to as the first classification model in this disclosure). Based on the text features of the natural segment, the probability value of the natural segment belonging to various dialogue segment labels is output through semantic calculation. The dialogue segment label with the highest probability is taken as the dialogue segment label of the natural segment.

[0229] 4. Script extraction.

[0230] like Figure 10 As shown, after completing the checkpoint identification and dialogue phase identification, if the current round of dialogue is detected to have successfully reached the call-in and contact-keeping stage, then any checkpoint question answered by the customer service representative in this round of dialogue is considered a valid script. The selected scripts are then stored in the script database in the form of checkpoint tag (referred to as the third category tag in this disclosure) ##customer checkpoint location text (referred to as historical input statement in this disclosure) ##customer service checkpoint reply (referred to as historical reply statement in this disclosure). This generates a candidate script fragment for checkpoint questions without specific needs, for use in the subsequent script matching stage. Other checkpoint questions are also selected using a similar targeted script selection method. Furthermore, the selection process can limit the depth of the selection range, and can select multiple rounds of responses after the checkpoint location as valid scripts, which can be flexibly configured according to specific needs.

[0231] 5. Matching of sales scripts.

[0232] The script library is built based on a large amount of historical dialogue data. When used online, the classification model used for checkpoint identification monitors in real time whether the current user's dialogue hits a certain checkpoint label. When it is identified that the probability of the user's input statement in the dialogue belonging to a certain checkpoint label is relatively high, the script matching model is used to obtain the best script for the current input statement.

[0233] The input to the dialogue matching model is the checkpoint label and the user's input statement, and the output is a set of excellent dialogues including confidence scores.

[0234] The script matching process can be as follows: Figure 11 As shown, firstly, based on the checkpoint tags and the checkpoint question (the user's input statement), similar dialogue fragments corresponding to checkpoint questions can be retrieved from the dialogue database. This process initially narrows down the dialogue range through literal features. Then, a dialogue matching model calculates the most suitable dialogue for the checkpoint question. This model is a pre-trained binary classification model. By concatenating the checkpoint tags, the checkpoint question, and the customer service checkpoint responses from candidate dialogue fragments into the text featureization process, the text is transformed into a high-dimensional feature vector. After a series of complex neural network calculations, the probability of a match is output. When the probability value is greater than a certain threshold, it indicates that the customer service checkpoint response in the current candidate dialogue fragment has successfully matched the checkpoint question. Therefore, a set of excellent candidate dialogues can be output according to the probability values, providing them to customer service personnel to assist in determining solutions to the current user's problem.

[0235] In summary, this solution possesses the capability to generate and recommend excellent sales scripts. Compared to existing technologies, this disclosure focuses on marketing scenario dialogues, finely segmenting multi-dimensional features of customer conversations. It selectively identifies effective scripts for different bottlenecks and stage characteristics, and uses semantic computing to match candidate script fragments for corresponding user bottlenecks, effectively resolving user issues. Furthermore, in complex scenarios, it assists inexperienced customer service personnel in making business judgments and selecting scripts, significantly improving their work efficiency and increasing the chances of closing deals. Moreover, the current script selection process is highly scalable and applicable to most service marketing scenarios, such as finance, advertising, and transportation / logistics, demonstrating strong versatility.

[0236] With the above Figures 1 to 7 Corresponding to the dialogue processing method provided in the embodiments, this disclosure also provides a dialogue processing apparatus. Since the dialogue processing apparatus provided in the embodiments of this disclosure is similar to the one described above… Figures 1 to 7 The dialogue processing method provided in the embodiments corresponds to the dialogue processing method provided in the embodiments of this disclosure, and therefore the implementation of the dialogue processing method is also applicable to the dialogue processing apparatus provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.

[0237] Figure 12This is a schematic diagram of the structure of the dialogue processing device provided in Embodiment 8 of this disclosure.

[0238] like Figure 12 As shown, the dialogue processing device 1200 may include: a first acquisition module 1201, a second acquisition module 1202, a third acquisition module 1203, and a first determination module 1204.

[0239] The first acquisition module 1201 is used to acquire the target input statement to be replied to.

[0240] The second acquisition module 1202 is used to acquire the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue.

[0241] The third acquisition module 1203 is used to acquire at least one candidate speech fragment from the speech database based on the first category label and the target input statement; wherein the candidate category label in the candidate speech fragment is similar to the first category label, and the candidate input statement in the candidate speech fragment is similar to the target input statement.

[0242] The first determining module 1204 is used to determine the target response statement from each candidate response statement based on the first classification label and the target input statement and the matching degree between the target response statement and the candidate response statements in each candidate speech fragment, so as to respond to the target input statement according to the target response statement.

[0243] In one possible implementation of this disclosure, the first determining module 1204 is configured to: for any candidate response statement in a candidate speech fragment, concatenate the first classification label, the target input statement, and the candidate response statement to obtain concatenated text; extract features from the concatenated text to obtain text features; classify the text features to obtain the classification probability of the candidate response statement, wherein the classification probability is used to indicate the matching degree between the candidate response statement and the target input statement; and determine the target response statement from each candidate response statement based on the classification probability of each candidate response statement.

[0244] In one possible implementation of this disclosure, the first determining module 1204 is configured to: determine the candidate response statement with the highest classification probability as the target response statement; or, determine the candidate response statement with a classification probability higher than a set threshold as the target response statement; or, sort the candidate response statements in descending order of their classification probabilities and determine the set number of candidate response statements at the top of the sort as the target response statements.

[0245] In one possible implementation of this disclosure, the script library is built through the following modules:

[0246] The fourth acquisition module is used to acquire historical dialogue statements from at least one round of dialogue.

[0247] The segmentation module is used to segment the historical dialogue statements in any round of dialogue to obtain at least one text fragment.

[0248] The classification module is used to classify each text fragment to obtain a second classification label for each text fragment. The second classification label is used to indicate the dialogue segment to which the text fragment belongs.

[0249] The generation module is used to generate at least one sample dialogue fragment based on the historical dialogue statements in the dialogue, provided that a category label exists in each of the second category labels.

[0250] A module is created to build a script library based on each sample script segment.

[0251] In one possible implementation of this disclosure, the classification module is configured to: input any text fragment into a first classification model for classification, so as to obtain the classification probabilities of multiple classification labels output by the first classification model; and determine a second classification label from the multiple classification labels based on the classification probabilities of the multiple classification labels; wherein the first classification model is trained based on sample text fragments, wherein the sample text fragments are marked with a first annotation label, which is used to indicate the dialogue segment to which the sample text fragments belong.

[0252] In one possible implementation of this disclosure, the generation module is configured to: group each historical dialogue statement in the dialogue to obtain at least one dialogue pair, wherein the dialogue pair includes a historical input statement and a historical response statement corresponding to the historical input statement; for any dialogue pair, obtain a third category label of the dialogue in which the historical input statement is located, wherein the third category label is used to indicate the dialogue intent of the dialogue in which the historical input statement is located; and generate sample speech fragments based on the third category label, the historical input statement and the historical response statement in the dialogue pair.

[0253] In one possible implementation of this disclosure, the fourth acquisition module is configured to: acquire a dialogue log; parse the dialogue log to obtain multiple initial dialogue statements in at least one round of dialogue; preprocess the multiple initial dialogue statements in each round of dialogue to obtain each historical dialogue statement in each round of dialogue; wherein the preprocessing includes at least one of the following: modal particle removal processing, repeated word removal processing, misspelling correction processing, and colloquial rewriting processing.

[0254] In one possible implementation of this disclosure, the second acquisition module 1202 is configured to: determine the target dialogue in which the target input statement is located; acquire at least one first input statement from the target dialogue; classify the target input statement and the at least one first input statement using a second classification model to obtain the classification probabilities of multiple predicted labels output by the second classification model; and determine a first classification label from the multiple predicted labels based on the classification probabilities of the multiple predicted labels. The second classification model is trained based on sample statements, wherein the sample statements are labeled with second annotation labels, which are used to indicate the dialogue intent of the dialogue in which the sample statements are located.

[0255] In one possible implementation of this disclosure, the number of target response statements is multiple, and the dialogue processing device 1200 may further include:

[0256] The first sorting module is used to sort multiple target response statements in descending order of classification probability to obtain the first sorting sequence.

[0257] The first sending module is used to send a first sorting sequence to the target customer service representative, so that the target customer service representative can reply to the target input statement according to the first sorting sequence.

[0258] or,

[0259] The second determination module is used to determine the target speech segment containing each target response statement from each candidate speech segment.

[0260] The second sorting module is used to sort the target speech fragments according to the classification probability of each target response statement to obtain the second sorting sequence.

[0261] The second sending module is used to send a second sorting sequence to the target customer service representative so that the target customer service representative can reply to the target input statement according to the second sorting sequence.

[0262] In one possible implementation of this disclosure, the candidate dialogue segment further includes: at least one associated response statement, wherein the associated response statement and the candidate response statement in the candidate dialogue segment are in the same round of dialogue, and the response time of the associated response statement is later than the response time of the candidate response statement; the dialogue processing device 1200 may further include:

[0263] The third determination module is used to determine the target speech segment containing the target response statement from each candidate speech segment.

[0264] The third sending module is used to send the target script segment to the target customer service representative, or to send the target response statement and at least one target related response statement from the target script segment to the target customer service representative.

[0265] Among them, the target response statement is used by the target customer service representative to respond to the target input statement; and at least one target associated response statement is used by the target customer service representative to respond to the input statement located after the target input statement in the target dialogue.

[0266] The dialogue processing apparatus of this disclosure acquires a first category label of the target dialogue containing the target input statement; wherein the first category label indicates the dialogue intent of the target dialogue; based on the first category label and the target input statement, it acquires at least one candidate dialogue segment from a dialogue library; wherein the candidate category labels in the candidate dialogue segments are similar to the first category label, and the candidate input statements in the candidate dialogue segments are similar to the target input statement; based on the matching degree between the first category label and the target input statement and the candidate response statements in each candidate dialogue segment, it determines a target response statement from each candidate response statement, and responds to the target input statement according to the target response statement. Thus, it can simultaneously combine the user-input statement and the dialogue intent of the dialogue containing the statement to determine the corresponding response statement, thereby improving the accuracy of response statement determination. Furthermore, in scenarios where customer service personnel provide manual responses, the determined response statement can be provided to customer service personnel to assist them in quickly resolving user issues, which not only improves the service efficiency of customer service personnel but also enhances the user experience.

[0267] To implement the above embodiments, this disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the dialogue processing method proposed in any of the above embodiments of this disclosure.

[0268] To implement the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the dialogue processing method proposed in any of the above embodiments of this disclosure.

[0269] To implement the above embodiments, this disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the dialogue processing method proposed in any of the above embodiments of this disclosure.

[0270] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0271] Figure 13A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device may include the server and client described in the above embodiments. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0272] like Figure 13 As shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in ROM (Read-Only Memory) 1302 or loaded from storage unit 1308 into RAM (Random Access Memory) 1303. The RAM 1303 may also store various programs and data required for the operation of the electronic device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. An I / O (Input / Output) interface 1305 is also connected to bus 1304.

[0273] Multiple components in electronic device 1300 are connected to I / O interface 1305, including: input unit 1306, such as keyboard, mouse, etc.; output unit 1307, such as various types of monitors, speakers, etc.; storage unit 1308, such as disk, optical disk, etc.; and communication unit 1309, such as network card, modem, wireless transceiver, etc. Communication unit 1309 allows electronic device 1300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0274] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods and processes described above, such as the dialogue processing methods described above. For example, in some embodiments, the dialogue processing methods described above can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the dialogue processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to perform the above-described dialogue processing method by any other suitable means (e.g., by means of firmware).

[0275] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0276] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0277] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0278] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0279] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0280] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers integrated with blockchain technology.

[0281] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0282] Deep learning is a new research direction in the field of machine learning. It learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process is very helpful in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to have analytical and learning capabilities like humans, and to recognize data such as text, images, and sound.

[0283] Cloud computing refers to a technology system that provides access to a shared pool of physical or virtual resources via a network. These resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for applications such as artificial intelligence and blockchain, as well as for model training.

[0284] According to the technical solution of this disclosure, a first category label of the target dialogue containing the target input statement is obtained; wherein the first category label is used to indicate the dialogue intent of the target dialogue; based on the first category label and the target input statement, at least one candidate dialogue segment is obtained from the dialogue library; wherein the candidate category label in the candidate dialogue segment is similar to the first category label, and the candidate input statement in the candidate dialogue segment is similar to the target input statement; based on the matching degree between the first category label and the target input statement and the candidate response statements in each candidate dialogue segment, the target response statement is determined from each candidate response statement, so as to reply to the target input statement according to the target response statement. Thus, it is possible to simultaneously combine the user-input statement and the dialogue intent of the dialogue containing the statement to determine the corresponding response statement, which can improve the accuracy of response statement determination. Furthermore, in scenarios where customer service personnel provide manual responses, the determined response statement can be provided to customer service personnel to assist them in quickly resolving user problems, which can not only improve the service efficiency of customer service personnel but also enhance the user experience.

[0285] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution proposed in this disclosure can be achieved, and this is not limited herein.

[0286] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A dialogue processing method, comprising: Retrieve the target input statement to be replied to; Obtain the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue; Based on the first classification label and the target input statement, at least one candidate speech fragment is obtained from the speech database; wherein, the candidate classification label in the candidate speech fragment is similar to the first classification label, and the candidate input statement in the candidate speech fragment is similar to the target input statement; Based on the first classification label and the target input statement, and the matching degree between the target input statement and the candidate response statements in each candidate speech segment, a target response statement is determined from each candidate response statement, so as to respond to the target input statement according to the target response statement; The script library is established through the following steps: Retrieve all historical dialogue statements from at least one round of dialogue; Divide the historical dialogue statements in any round of the dialogue to obtain at least one text fragment; Each of the text fragments is classified to obtain a second classification label for each text fragment, wherein the second classification label is used to indicate the dialogue segment to which the text fragment belongs; If a classification label is set in each of the second classification labels, at least one sample dialogue fragment is generated based on each historical dialogue statement in the dialogue. The speech database is established based on each of the sample speech segments; The step of dividing the historical dialogue statements in any round of the dialogue to obtain at least one text segment includes: Classify each historical dialogue statement in any round of the dialogue to obtain the classification probability of each historical dialogue statement; the classification probability of the historical dialogue statement is used to indicate the probability value that the end of the historical dialogue statement is a segment. Based on the classification probability of each historical dialogue statement, the historical dialogue statements in the dialogue are divided to obtain at least one text segment.

2. The method according to claim 1, wherein, The step of determining the target response statement from each of the candidate response statements based on the matching degree between the first classification label and the target input statement and the candidate response statements in each of the candidate speech segments includes: For any candidate response statement in any of the candidate speech fragments, the first category label, the target input statement, and the candidate response statement are concatenated to obtain concatenated text; Feature extraction is performed on the concatenated text to obtain text features; The text features are classified to obtain the classification probability of the candidate response statement, wherein the classification probability is used to indicate the matching degree between the candidate response statement and the target input statement; The target response statement is determined from the candidate response statements based on the classification probability of each candidate response statement.

3. The method according to claim 2, wherein, The step of determining the target response statement from the candidate response statements based on their classification probabilities includes: The candidate response statement with the highest classification probability is determined as the target response statement; or, Candidate response statements whose classification probability is higher than a set threshold are identified as the target response statements; or, The candidate response statements are sorted from largest to smallest according to the value of the classification probability, and the first set number of candidate response statements are determined as the target response statements.

4. The method according to claim 1, wherein, The process of classifying each of the text fragments to obtain a second classification label includes: For any of the aforementioned text fragments, the text fragments are input into a first classification model for classification, so as to obtain the classification probabilities of multiple classification labels output by the first classification model; Based on the classification probabilities of the multiple classification labels, a second classification label is determined from the multiple classification labels; The first classification model is trained based on sample text fragments, wherein the sample text fragments are labeled with a first annotation label, which is used to indicate the dialogue segment to which the sample text fragments belong.

5. The method according to claim 1, wherein, The step of generating at least one sample dialogue fragment based on each historical dialogue statement in the dialogue includes: The historical dialogue statements in the dialogue are grouped to obtain at least one dialogue pair, wherein the dialogue pair includes a historical input statement and a historical response statement corresponding to the historical input statement; For any of the dialogue pairs, obtain the third category label of the dialogue in which the historical input statement is located in the dialogue pair, wherein the third category label is used to indicate the dialogue intent of the dialogue in which the historical input statement is located; The sample dialogue fragment is generated based on the third category label, the historical input statements in the dialogue pair, and the historical response statements.

6. The method according to claim 1, wherein, The step of obtaining historical dialogue statements from at least one round of dialogue includes: Get the conversation log; The dialogue log is parsed to obtain multiple initial dialogue statements from at least one round of dialogue; Multiple initial dialogue statements in each round of the dialogue are preprocessed to obtain each historical dialogue statement in each round of the dialogue. The preprocessing includes at least one of the following: removal of modal particles, removal of repeated words, correction of misspellings, and colloquial rewriting.

7. The method according to any one of claims 1-6, wherein, The step of obtaining the first category label of the target dialogue where the target input statement is located includes: Determine the target dialog containing the target input statement; Obtain at least one first input statement from the target dialogue; A second classification model is used to classify the target input statement and the at least one first input statement to obtain the classification probabilities of multiple predicted labels output by the second classification model. Based on the classification probabilities of the multiple predicted labels, a first classification label is determined from the multiple predicted labels; The second classification model is trained based on sample statements, wherein the sample statements are labeled with a second label, which is used to indicate the dialogue intent of the dialogue in which the sample statement is located.

8. The method according to any one of claims 1-6, wherein, The number of target response statements is multiple, and the method further includes: The multiple target response statements are sorted from largest to smallest according to their classification probabilities to obtain a first sorting sequence; Send the first sorting sequence to the target customer service representative so that the target customer service representative can reply to the target input statement according to the first sorting sequence; or, From each of the candidate speech fragments, determine the target speech fragment containing each target response statement; Based on the classification probability of each target response statement, the target speech fragments are sorted to obtain a second sorting sequence; The second sorting sequence is sent to the target customer service representative so that the target customer service representative can reply to the target input statement according to the second sorting sequence.

9. The method according to any one of claims 1-6, wherein, The candidate dialogue segment further includes at least one associated response statement, wherein the associated response statement and the candidate response statement in the candidate dialogue segment are in the same round of dialogue, and the response time of the associated response statement is later than the response time of the candidate response statement; The method further includes: From each of the candidate speech fragments, determine the target speech fragment containing the target response statement; Send the target script segment to the target customer service representative, or send the target response statement and at least one target related response statement from the target script segment to the target customer service representative; The target response statement is used by the target customer service representative to respond to the target input statement; Wherein, the at least one target-related response statement is used by the target customer service representative to respond to the input statement located after the target input statement in the target dialogue.

10. A dialogue processing apparatus, comprising: The first acquisition module is used to acquire the target input statement to be replied to; The second acquisition module is used to acquire the first category label of the target dialogue where the target input statement is located; wherein, the first category label is used to indicate the dialogue intent of the target dialogue; The third acquisition module is used to acquire at least one candidate speech fragment from the speech database based on the first classification label and the target input statement; wherein the candidate classification label in the candidate speech fragment is similar to the first classification label, and the candidate input statement in the candidate speech fragment is similar to the target input statement; The first determining module is configured to determine a target response statement from each candidate response statement based on the matching degree between the first classification label and the target input statement and the candidate response statements in each candidate speech fragment, so as to respond to the target input statement according to the target response statement; The script library is established through the following modules: The fourth acquisition module is used to acquire historical dialogue statements from at least one round of dialogue; The segmentation module is used to segment each historical dialogue statement in any round of the dialogue to obtain at least one text fragment; A classification module is used to classify each of the text fragments to obtain a second classification label for each of the text fragments, wherein the second classification label is used to indicate the dialogue segment to which the text fragment belongs; The generation module is used to generate at least one sample dialogue fragment based on each historical dialogue statement in the dialogue, provided that a set category label exists in each of the second category labels. A module is established to build the speech library based on each of the sample speech segments; The partitioning module is used for: Classify each historical dialogue statement in any round of the dialogue to obtain the classification probability of each historical dialogue statement; the classification probability of the historical dialogue statement is used to indicate the probability value that the end of the historical dialogue statement is a segment. Based on the classification probability of each historical dialogue statement, the historical dialogue statements in the dialogue are divided to obtain at least one text segment.

11. The apparatus according to claim 10, wherein, The first determining module is used for: For any candidate response statement in any of the candidate speech fragments, the first category label, the target input statement, and the candidate response statement are concatenated to obtain concatenated text; Feature extraction is performed on the concatenated text to obtain text features; The text features are classified to obtain the classification probability of the candidate response statement, wherein the classification probability is used to indicate the matching degree between the candidate response statement and the target input statement; The target response statement is determined from the candidate response statements based on the classification probability of each candidate response statement.

12. The apparatus according to claim 11, wherein, The first determining module is used for: The candidate response statement with the highest classification probability is determined as the target response statement; or, Candidate response statements whose classification probability is higher than a set threshold are identified as the target response statements; or, The candidate response statements are sorted from largest to smallest according to the value of the classification probability, and the first set number of candidate response statements are determined as the target response statements.

13. The apparatus according to claim 10, wherein, The classification module is used for: For any of the aforementioned text fragments, the text fragments are input into a first classification model for classification, so as to obtain the classification probabilities of multiple classification labels output by the first classification model; Based on the classification probabilities of the multiple classification labels, a second classification label is determined from the multiple classification labels; The first classification model is trained based on sample text fragments, wherein the sample text fragments are labeled with a first annotation label, which is used to indicate the dialogue segment to which the sample text fragments belong.

14. The apparatus according to claim 10, wherein, The generation module is used for: The historical dialogue statements in the dialogue are grouped to obtain at least one dialogue pair, wherein the dialogue pair includes a historical input statement and a historical response statement corresponding to the historical input statement; For any of the dialogue pairs, obtain the third category label of the dialogue in which the historical input statement is located in the dialogue pair, wherein the third category label is used to indicate the dialogue intent of the dialogue in which the historical input statement is located; The sample dialogue fragment is generated based on the third category label, the historical input statements in the dialogue pair, and the historical response statements.

15. The apparatus according to claim 10, wherein, The fourth acquisition module is used for: Get the conversation log; The dialogue log is parsed to obtain multiple initial dialogue statements from at least one round of dialogue; Multiple initial dialogue statements in each round of the dialogue are preprocessed to obtain each historical dialogue statement in each round of the dialogue. The preprocessing includes at least one of the following: removal of modal particles, removal of repeated words, correction of misspellings, and colloquial rewriting.

16. The apparatus according to any one of claims 10-15, wherein, The second acquisition module is used for: Determine the target dialog containing the target input statement; Obtain at least one first input statement from the target dialogue; A second classification model is used to classify the target input statement and the at least one first input statement to obtain the classification probabilities of multiple predicted labels output by the second classification model. Based on the classification probabilities of the multiple predicted labels, a first classification label is determined from the multiple predicted labels; The second classification model is trained based on sample statements, wherein the sample statements are labeled with a second label, which is used to indicate the dialogue intent of the dialogue in which the sample statement is located.

17. The apparatus according to any one of claims 10-15, wherein, The number of target response statements is multiple, and the device further includes: The first sorting module is used to sort the multiple target response statements according to the classification probability from largest to smallest to obtain a first sorting sequence; The first sending module is used to send the first sorting sequence to the target customer service representative so that the target customer service representative can reply to the target input statement according to the first sorting sequence; or, The second determining module is used to determine the target speech segment where each target response statement is located from each of the candidate speech segments; The second sorting module is used to sort the target speech fragments according to the classification probability of each target response statement to obtain a second sorting sequence; The second sending module is used to send the second sorting sequence to the target customer service representative so that the target customer service representative can reply to the target input statement according to the second sorting sequence.

18. The apparatus according to any one of claims 10-15, wherein, The candidate dialogue segment further includes at least one associated response statement, wherein the associated response statement and the candidate response statement in the candidate dialogue segment are in the same round of dialogue, and the response time of the associated response statement is later than the response time of the candidate response statement; The device further includes: The third determining module is used to determine the target speech segment containing the target response statement from each of the candidate speech segments; The third sending module is used to send the target script segment to the target customer service representative, or to send the target reply statement and at least one target related reply statement in the target script segment to the target customer service representative; The target response statement is used by the target customer service representative to respond to the target input statement; Wherein, the at least one target-related response statement is used by the target customer service representative to respond to the input statement located after the target input statement in the target dialogue.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the dialogue processing method according to any one of claims 1-9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the dialogue processing method according to any one of claims 1-9.

21. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the dialogue processing method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Information pushing method and device

    CN112380331A

  • Session management method and device and electronic equipment

    CN115271024A