Method, device and storage medium for determining optimal speech sequence
By combining user characteristics and interaction process information, considering the hangup information and determining the optimal speech sequence, the problem of low prediction accuracy caused by the failure to consider the mid-way hangup in the prior art is solved, and more accurate speech prediction is achieved.
Patent Information
- Application Number
- CN202010833956.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-08-18
AI Technical Summary
In the prior art, the prediction method of the art can reach the final node by default, and the phenomenon of hang up in the middle is not considered, resulting in low prediction accuracy.
By determining the user characteristic information and dialing information of the target user, using a pre-trained prediction model, combining the interactive process information, including the current speech sequence, the speech sub-sequence and the subsequent speech sequence, considering the hanging up information to determine the optimal speech sequence.
When the user may hang up in the middle, the optimal speech sequence can still be obtained to improve the accuracy of speech prediction.
Smart Images

Figure CN114077656B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of conversation prediction, and particularly to a method, apparatus, and storage medium for determining an optimal conversation sequence. Background Art
[0002] Due to the development of artificial intelligence technology, there are now many voice communication systems in which robots communicate with people, especially in voice customer service, intelligent telemarketing, intelligent debt collection, intelligent speakers and other voice interaction scenarios, which are widely used.
[0003] After the robot can communicate with people by voice, how to improve the communication effect has become a difficult point (for example, in intelligent reminders, how to improve the efficiency of robot reminders). The main existing conversation optimization solutions are mainly to predict the conversion rate of each conversation. However, in the process of predicting the conversion rate, it is defaulted that all calls can reach the final node, so what is predicted is the optimal situation that can complete all processes. However, if the call is interrupted, for example, the user hangs up after saying the first sentence. In this case, since the hang-up information is not considered during the prediction, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction.
[0004] Regarding the technical problem in the above-mentioned existing technology that the existing conversation prediction method defaults that all calls can reach the final node, and since the possible phenomenon of hanging up midway is not considered, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present disclosure provide a method, apparatus, and storage medium for determining an optimal conversation sequence to at least solve the technical problem in the existing technology that the existing conversation prediction method defaults that all calls can reach the final node, and since the possible phenomenon of hanging up midway is not considered, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction.
[0006] According to one aspect of the embodiments of the present disclosure, a method for determining an optimal conversation sequence is provided. The optimal conversation sequence includes optimal conversations corresponding to respective conversation nodes for interacting with a target user, and the method includes: determining user characteristic information and call information of the target user; using a pre-trained prediction model to respectively determine the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in a plurality of conversation sequences according to the user characteristic information and the call information, wherein the conversations corresponding to the respective conversation nodes form the plurality of conversation sequences; and determining an optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and interaction process information for interacting with the target user, wherein the interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; the current conversation sequence is a sequence formed by the conversations corresponding to the current conversation node for interacting with the target user; the conversation subsequence is a sequence formed by the conversations corresponding to the previous conversation nodes for interacting with the target user; and the subsequent conversation sequence is a sequence formed by the conversations corresponding to the subsequent conversation nodes after the current conversation node.
[0007] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided. The storage medium includes a stored program, wherein the method described in any one of the above is executed by a processor when the program runs.
[0008] According to another aspect of the embodiments of the present disclosure, a device for determining an optimal conversation sequence is further provided. The optimal conversation sequence includes optimal conversations corresponding to respective conversation nodes for interacting with a target user, and the device includes: a first determination module for determining user characteristic information and call information of the target user; a second determination module for using a pre-trained prediction model to respectively determine the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in a plurality of conversation sequences according to the user characteristic information and the call information, wherein the conversations corresponding to the respective conversation nodes form the plurality of conversation sequences; and a third determination module for determining an optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and interaction process information for interacting with the target user, wherein the interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; the current conversation sequence is a sequence formed by the conversations corresponding to the current conversation node for interacting with the target user; the conversation subsequence is a sequence formed by the conversations corresponding to the previous conversation nodes for interacting with the target user; and the subsequent conversation sequence is a sequence formed by the conversations corresponding to the subsequent conversation nodes after the current conversation node.
[0009] According to another aspect of the embodiments of the present disclosure, there is also provided an apparatus for determining an optimal conversation sequence. The optimal conversation sequence includes optimal conversations corresponding to respective conversation nodes for interacting with a target user, and includes: a processor; and a memory connected to the processor for providing instructions for the processor to perform the following processing steps: determining user characteristic information and call information of the target user; using a pre-trained prediction model to respectively determine the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in a plurality of conversation sequences according to the user characteristic information and the call information, where the conversations corresponding to respective conversation nodes form the plurality of conversation sequences; and determining an optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and the predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and interaction process information for interacting with the target user, where the interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; the current conversation sequence is a sequence formed by conversations corresponding to the current conversation node for interacting with the target user; the conversation subsequence is a sequence formed by conversations corresponding to previous conversation nodes for interacting with the target user; and the subsequent conversation sequence is a sequence formed by conversations corresponding to subsequent conversation nodes after the current conversation node.
[0010] In the embodiments of the present disclosure, in the process of determining an optimal conversation sequence for interacting with a target user, not only the predicted hang-up rate and the predicted hang-up conversion rate of each conversation are considered, but also the interaction process information for interacting with the target user is combined, and the interaction process information includes the current conversation sequence, the conversation subsequence, and the subsequent conversation sequence, where the subsequent conversation sequence takes into account the hang-up information of hanging up halfway. Thus, in the present application, the optimal conversation prediction for each node takes into account the hang-up information. When the user cannot listen to the entire conversation, the optimal conversation sequence in the current situation can also be obtained. Furthermore, it solves the technical problem in the prior art that the existing conversation prediction method defaults that all calls can reach the final node. Since the phenomenon of hanging up halfway may not be considered, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings:
[0012] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure;
[0013] Figure 2It is a schematic flowchart of the method for determining the optimal conversation strategy sequence according to the first aspect of Embodiment 1 of the present disclosure;
[0014] Figure 3 It is a schematic diagram of each conversation strategy during the interaction according to Embodiment 1 of the present disclosure;
[0015] Figure 4 It is a schematic diagram for calculating the expected hang-up conversion rate according to Embodiment 1 of the present disclosure;
[0016] Figure 5 It is a schematic overall flowchart of the method for determining the optimal conversation strategy sequence according to Embodiment 1 of the present disclosure;
[0017] Figure 6 It is a schematic diagram of the device for determining the optimal conversation strategy sequence according to Embodiment 2 of the present disclosure; and
[0018] Figure 7 It is a schematic diagram of the device for determining the optimal conversation strategy sequence according to Embodiment 3 of the present disclosure. Detailed implementation manners
[0019] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0021] Embodiment 1
[0022] According to this embodiment, a method embodiment for determining an optimal speech sequence is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0023] The method embodiment provided by this embodiment can be executed in a server or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing the method of determining an optimal speech sequence is shown. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition to this, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0024] It should be noted that the above one or more processors and / or other data processing circuits can generally be referred to as "data processing circuits" in this article. The data processing circuit can be embodied in whole or in part as software, hardware, firmware, or any arbitrary combination thereof. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0025] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for determining the optimal conversation sequence in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for determining the optimal conversation sequence of the above application program. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories may be connected to the computing device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0026] The transmission device is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0027] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computing device.
[0028] It should be noted here that in some alternative embodiments, the above Figure 1 shown computing device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance, and is intended to show the types of components that may exist in the above computing device.
[0029] Under the above operating environment, according to the first aspect of this embodiment, a method for determining the optimal conversation sequence is provided. Figure 2 The flowchart of the method is shown. Refer to Figure 2 shown, the method includes:
[0030] S202: Determine the user characteristic information and call information of the target user;
[0031] S204: Using a pre-trained prediction model, respectively determine the predicted hang-up rate and the predicted hang-up conversion rate of each conversation script in multiple conversation script sequences according to the user feature information and the call information, where the conversation scripts corresponding to each conversation script node form multiple conversation script sequences; and
[0032] S206: According to the predicted hang-up rate and the predicted hang-up conversion rate of each conversation script in multiple conversation script sequences and the interaction process information of interacting with the target user, determine the optimal conversation script sequence for interacting with the target user.
[0033] Among them, the interaction process information includes: the current conversation script sequence of interacting with the target user, the conversation script subsequence of interacting with the target user, and the subsequent conversation script sequence of interacting with the target user; the current conversation script sequence is the sequence composed of the conversation scripts corresponding to the current conversation script node of interacting with the target user; the conversation script subsequence is the sequence composed of the conversation scripts corresponding to the previous conversation script nodes of interacting with the target user; and the subsequent conversation script sequence is the sequence composed of the conversation scripts corresponding to the subsequent conversation script nodes after the current conversation script node.
[0034] As described in the background art, when the robot can communicate with people by voice, how to improve the communication effect becomes a difficult point (for example, in intelligent reminders, how to improve the reminder efficiency of the robot). The main existing conversation script optimization solutions mainly predict the conversion rate of each conversation script. However, in the process of predicting the conversion rate, it is assumed that all calls can reach the final node, so what is predicted is the optimal situation that can complete all processes. However, if the call is interrupted, for example, the user hangs up after saying the first sentence. In this case, since the hang-up information is not considered during the prediction, the predicted conversation script sequence is not the optimal result, resulting in low accuracy of conversation script prediction.
[0035] In view of this, in the case where the intelligent voice robot needs to interact with the target user through a series of conversation scripts, the optimal conversation script sequence that can enable the target user to be converted can be pre-selected, so that the user can be converted with the highest probability. In this embodiment, first determine the user feature information and the call information of the target user. The user feature information is all feature information associated with the target user, including but not limited to user attributes, user behaviors, user historical call information, etc. The call information is information that can be obtained before the call, including but not limited to the phone number, call task, call activity, call time, etc. that can be obtained before the call.
[0036] Furthermore, using a pre-trained prediction model, respectively determine the predicted hang-up rate and the predicted hang-up conversion rate of each conversation script in multiple conversation script sequences according to the user feature information and the call information. Among them, refer to Figure 3As shown, there are three conversation nodes: conversation node 1, conversation node 2, and conversation node 3. Each conversation node includes at least one conversation. For example, conversation node 1 includes three conversations: S11, S12, and S13; conversation node 2 includes two conversations: S21 and S22; conversation node 3 includes two conversations: S31 and S32. At this time, the conversations in multiple conversation sequences are S11, S12, S13, S21, S22, S31, and S32. Using a pre-trained prediction model, according to the user characteristic information and interaction process information, the predicted hang-up rates and predicted hang-up conversion rates of S11, S12, S13, S21, S22, S31, and S32 are determined respectively.
[0037] Furthermore, according to the predicted hang-up rates and predicted hang-up conversion rates of S11, S12, S13, S21, S22, S31, and S32 and the interaction process information of interacting with the target user, the optimal conversation sequence for interacting with the target user is determined. The interaction process information of interacting with the target user includes: the current conversation sequence of interacting with the target user, the conversation subsequence of interacting with the target user, and the subsequent conversation sequence of interacting with the target user. The current conversation sequence is the sequence composed of the conversations corresponding to the current conversation node of interacting with the target user; the conversation subsequence is the sequence composed of the conversations corresponding to the previous conversation node of interacting with the target user; and the subsequent conversation sequence is the sequence composed of the conversations corresponding to the subsequent conversation nodes after the current conversation node.
[0038] Refer to Figure 3 As shown, assume that the current conversation for interacting with the target user is S21, and the conversation corresponding to the previous conversation node of interacting with the target user is S11. Then the current conversation sequence is S11->S21, and the current conversation subsequence is S11. Among them, when the target user hangs up after listening to the conversation S21 and does not have time to listen to the subsequent conversation, the subsequent conversation for interacting with the target user is 0 (represented by None). When the target user does not hang up after listening to the conversation S21 and hears the subsequent conversation, the subsequent conversations for interacting with the target user are S31 and S32. Therefore, the subsequent conversation sequences are S11->S21->None, S11->S21->S31, and S11->S21->S32.
[0039] In this embodiment, in the process of determining the optimal conversation sequence for interacting with the target user, not only the predicted hang-up rate and predicted hang-up conversion rate of each conversation are considered, but also the interaction process information of interacting with the target user is combined, and the interaction process information includes the current conversation sequence, conversation subsequence, and subsequent conversation sequence, where the subsequent conversation sequence takes into account the hang-up information of the mid-call hang-up. Thus, the optimal conversation prediction for each node in this application takes into account the hang-up information. When the user cannot listen to the entire call, the optimal conversation sequence for the current situation can also be obtained. Furthermore, it solves the technical problem in the prior art that the existing conversation prediction method defaults that all calls can reach the final node. Since the possible mid-call hang-up phenomenon is not considered, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction.
[0040] Optionally, the operation of determining the optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences and the interaction process information of interacting with the target user includes: determining the expected hang-up conversion rate of each conversation in multiple conversation sequences respectively according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences and the interaction process information of interacting with the target user; and respectively determining the conversation with the highest expected hang-up conversion rate corresponding to each conversation node as the optimal conversation, and generating the optimal conversation sequence.
[0041] Specifically, as shown in Figure 4 The optimal conversation sequence is determined through the expected hang-up conversion rate, that is, the expected hang-up conversion rate of each conversation in multiple conversation sequences is determined respectively, and then the conversation with the highest expected hang-up conversion rate corresponding to each conversation node is determined as the optimal conversation, and the optimal conversation sequence is generated. For example: the conversation with the highest expected hang-up conversion rate corresponding to conversation node 1 is S11, the conversation with the highest expected hang-up conversion rate corresponding to conversation node 2 is S21, and the conversation with the highest expected hang-up conversion rate corresponding to conversation node 3 is S32, then the optimal conversation sequence is S11->S21->S32. Even if the user hangs up after listening to conversation S21, the determined optimal conversation sequence has taken this result into account, so S11->S21 is also the optimal conversation sequence at this time. In this way, in this embodiment, not only the hang-up information is combined, but also the expected hang-up conversion rate is used to obtain the prediction result of each node, and the optimal conversation sequence is determined by using the hang-up information and the expected hang-up conversion rate, further solving the problem that the recommended conversation sequence after the user hangs up may not be the optimal conversation sequence.
[0042] Optionally, the operation of determining the expected hang-up conversion rate of each conversation statement in multiple conversation statement sequences according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the multiple conversation statement sequences and the interaction process information of interacting with the target user includes: determining the expected hang-up conversion rate of each conversation statement in the current conversation statement sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the current conversation statement sequence and the predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the subsequent conversation statement sequences; and determining the expected hang-up conversion rate of each conversation statement in the multiple conversation statement sequences respectively according to the determined expected hang-up conversion rate of each conversation statement in the current conversation statement sequence. Specifically, when the current conversation statements are S21 and S22, it is necessary to determine the expected hang-up conversion rate of S21 and the expected hang-up conversion rate of S22. When the current conversation statements are S31 and S32, it is necessary to determine the expected hang-up conversion rate of S31 and the expected hang-up conversion rate of S32. And so on until the expected hang-up conversion rate of all conversation statements is determined.
[0043] Optionally, the operation of determining the expected predicted hang-up conversion rate of each conversation statement in the current conversation statement sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the current conversation statement sequence and the predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the subsequent conversation statement sequences includes: determining the first hang-up conversion rate of each conversation statement in the current conversation statement sequence according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the current conversation statement sequence; determining the second hang-up conversion rate of each conversation statement in the subsequent conversation statement sequence according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of each conversation statement in the subsequent conversation statement sequence; and determining the expected hang-up conversion rate of each conversation statement in the current conversation statement sequence based on the first hang-up conversion rate of each conversation statement in the current conversation statement sequence and the second hang-up conversion rate of each conversation statement in the subsequent conversation statement sequence.
[0044] Specifically, referring to Figure 4 As shown, taking the current conversation statement as S21 as an example, in the case where the subsequent conversation statement is None, that is, when the user hangs up after listening to the conversation statement S21, the first hang-up conversion rate of S21 is determined according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of S21. Among them, the first hang-up conversion rate of S21 corresponds to the hang-up conversion rate of the subsequent conversation statement being None in Figure 4. In the case where the user does not hang up after listening to the conversation statement S21, first, the second hang-up conversion rate of S31 is determined according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of S31, where the second hang-up conversion rate of S31 corresponds to Figure 4 the hang-up conversion rate of the subsequent conversation statement being S31 in Figure Figure 4The subsequent conversation statement in it is the hang-up conversion rate of S32. Finally, determine the expected hang-up conversion rate of S21 based on the first hang-up conversion rate of S21, the second hang-up conversion rate of S31, and the second hang-up conversion rate of S32.
[0045] Optionally, the operation of determining the expected hang-up conversion rate of each conversation statement in the current conversation statement sequence based on the first hang-up conversion rate of each conversation statement in the current conversation statement sequence and the second hang-up conversion rate of each conversation statement in the subsequent conversation statement sequence includes: selecting the conversation statement with the highest second hang-up conversion rate from the subsequent conversation statement sequence; and determining the expected hang-up conversion rate of each conversation statement in the current conversation statement sequence based on the first hang-up conversion rate of each conversation statement in the current conversation statement sequence and the second hang-up conversion rate corresponding to the conversation statement with the highest second hang-up conversion rate in the subsequent conversation statement sequence.
[0046] Specifically, according to the above, compare the hang-up conversion rate of the subsequent conversation statement being S31 with the hang-up conversion rate of the subsequent conversation statement being S32, and select the one with the higher hang-up conversion rate. For example, the hang-up conversion rate of the subsequent conversation statement being S32. At this time, the expected hang-up conversion rate of S21 is the hang-up conversion rate of the subsequent conversation statement being None plus the hang-up conversion rate of the subsequent conversation statement being S32. That is, the expected hang-up conversion rate of the current conversation statement is equal to the sum of the hang-up conversion rate of its subsequent conversation statement being None and the higher hang-up conversion rate of the subsequent conversation statement not being None. In this way, the expected hang-up conversion rate of each conversation statement in the current conversation statement sequence can be accurately determined.
[0047] Optionally, determine the previous non-hang-up rate of each conversation statement in the subsequent conversation statement sequence through the following operations: determine the previous non-hang-up rate of each conversation statement in the subsequent conversation statement sequence according to the predicted hang-up rate of each conversation statement in the current conversation statement sequence.
[0048] Specifically, referring to Figure 4 As shown, taking the current conversation statement as S21 as an example, when the subsequent conversation statement is None, it means that the user hangs up after listening to the conversation statement S21. At this time, the previous non-hang-up rate of S21 is 1. When the subsequent conversation statement is S31, it means that the user does not hang up after listening to the conversation statement S21. At this time, the previous non-hang-up rate of S31 is 1 minus the predicted hang-up rate of S21. When the subsequent conversation statement is S32, it also means that the user does not hang up after listening to the conversation statement S21. At this time, the previous non-hang-up rate of S32 is also 1 minus the predicted hang-up rate of S21. In this way, the previous non-hang-up rate of each conversation statement in the subsequent conversation statement sequence can be accurately determined.
[0049] In addition, referring to Figure 1 As shown, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein when the program runs, the method of any one of the above is executed by a processor.
[0050] In addition, referring to Figure 5 as shown in
[0051] 1) Training process:
[0052] a) Obtain the call records, actual scripts, and corresponding user profile information of each predicted script.
[0053] b) Process the obtained data to form training data available for the model.
[0054] c) Use real-time user profile features and user script features for training and publish a prediction service.
[0055] 2) Prediction process:
[0056] a) Obtain the user to be called and their user profile features.
[0057] b) For each predicted script node, construct its predicted script and all its sub-paths.
[0058] c) Process the obtained data to form training data available for the model.
[0059] d) Use the trained model to predict the hang-up rates and hang-up conversion rates of all sub-paths.
[0060] e) For each predicted script node, calculate its expected hang-up conversion rate, and select the combination of the optimal expected hang-up conversion rates of all predicted nodes for recommended broadcast.
[0061] Exemplarily, assume there is a sample as shown in Table 1 below for model training, and use deep learning multi-objective joint training to simultaneously learn two objectives: predicting the hang-up probability and the hang-up conversion rate.
[0062] Table 1
[0063]
[0064] During prediction, taking the calculation of the expected hang-up conversion rate of script node 2 as an example, assume that the optimal expected hang-up conversion rate script of script node 1 is S11, as Figure 4 Calculate the expected hang-up conversion rate of S2X: According to Figure 4 Calculate, select the one with a larger expected hang-up probability between S21 and S22 as the optimal script corresponding to the predicted node. Assume that S22 is larger. In this way, the script combination of S11 and S22 is obtained. At this time, calculate the expected hang-up probability of S3X according to this method. Assume that S32 is the best, then the final recommended combination is S11, S22, S32. Even if the user hangs up at S22, since this result has been considered during prediction, S11 and S22 are also the optimal script combinations at this time.
[0065] In this embodiment, hang-up information is taken into account for prediction of each node, thereby increasing the model learning capability. When predicting, even when the user cannot hear the entire conversation, the optimal combination of words for the current situation can be obtained.
[0066] In summary, the verification method of the intermediate indicator of speech comparison proposed in the present invention can produce the following effects:
[0067] 1) Using the hang-up information and the expected hang-up conversion rate, the optimal combination of words (corresponding to the optimal sequence of words mentioned above) is calculated to solve the problem that the combination recommendation after the user hangs up may not be optimal;
[0068] 2) Use deep learning joint training to predict the hang-up rate and the hang-up conversion rate, so that the two tasks share the underlying information and improve the accuracy of the two tasks. If two separate models are used to predict the hang-up rate and the conversion rate respectively, it also belongs to an implementation method of this embodiment;
[0069] 3) Regardless of whether any model such as logistic regression, tree type model, deep learning, etc. is used, or various changes are made to the features (such as feature intersection, feature mapping), and feature information is added, as long as the sample construction method and the expected hang-up conversion rate calculation method are as shown in this embodiment, they all belong to the implementation method of this embodiment.
[0070] 4) User portrait features refer to all features that can be associated with a user, including but not limited to user attributes, user behavior, user dialing history information, etc.
[0071] 5) The technical solution proposed in this application is applicable to all intelligent voice speech prediction scenarios, including but not limited to intelligent outbound calls, inbound calls, voice collection, follow-up visits, telemarketing, etc., and is also applicable to text multi-round dialogue speech prediction scenarios.
[0072] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0074] Embodiment 2
[0075] Figure 6 Fig. shows a device 600 for determining an optimal conversation sequence according to the present embodiment. The optimal conversation sequence includes the optimal conversations corresponding to each conversation node for interacting with the target user. The device 600 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 6 As shown, the device 600 includes: a first determination module 610, configured to determine the user characteristic information and call information of the target user; a second determination module 620, configured to use a pre-trained prediction model to respectively determine the predicted hang-up rate and predicted hang-up conversion rate of each conversation in a plurality of conversation sequences according to the user characteristic information and call information, where the conversations corresponding to each conversation node form a plurality of conversation sequences; and a third determination module 630, configured to determine the optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and the interaction process information for interacting with the target user, where the interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; the current conversation sequence is the sequence formed by the conversations corresponding to the current conversation node for interacting with the target user; the conversation subsequence is the sequence formed by the conversations corresponding to the previous conversation nodes for interacting with the target user; and the subsequent conversation sequence is the sequence formed by the conversations corresponding to the subsequent conversation nodes after the current conversation node.
[0076] Optionally, the third determination module 630 includes: a determination sub-module, configured to respectively determine the expected hang-up conversion rate of each conversation in the plurality of conversation sequences according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and the interaction process information for interacting with the target user; and a generation sub-module, configured to determine the conversation with the highest expected hang-up conversion rate corresponding to each conversation node as the optimal conversation and generate the optimal conversation sequence.
[0077] Optionally, the determination sub-module includes: a first determination unit, configured to determine the expected hang-up conversion rate of each conversation in the current conversation sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the current conversation sequence, the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the subsequent conversation sequence, and the interaction process information of interacting with the target user; and a second determination unit, configured to determine the expected hang-up conversion rate of each conversation in multiple conversation sequences respectively according to the determined expected hang-up conversion rate of each conversation in the current conversation sequence.
[0078] Optionally, the first determination unit includes: a first determination subunit, configured to determine the first hang-up conversion rate of each conversation in the current conversation sequence according to the previous non-hang-up rate, predicted hang-up rate, and predicted hang-up conversion rate of each conversation in the current conversation sequence; a second determination subunit, configured to determine the second hang-up conversion rate of each conversation in the subsequent conversation sequence according to the previous non-hang-up rate, predicted hang-up rate, and predicted hang-up conversion rate of each conversation in the subsequent conversation sequence; and a third determination subunit, configured to determine the expected hang-up conversion rate of each conversation in the current conversation sequence according to the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate of each conversation in the subsequent conversation sequence.
[0079] Optionally, the third determination subunit includes: a selection component, configured to select the conversation with the highest second hang-up conversion rate from the subsequent conversation sequence; and a determination component, configured to determine the expected hang-up conversion rate of each conversation in the current conversation sequence according to the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate corresponding to the conversation with the highest second hang-up conversion rate in the subsequent conversation sequence.
[0080] Optionally, the apparatus 600 further includes a fourth determination module, configured to determine the previous non-hang-up rate of each conversation in the subsequent conversation sequence by: determining the previous non-hang-up rate of each conversation in the subsequent conversation sequence according to the predicted hang-up rate of each conversation in the current conversation sequence.
[0081] Thus, according to this embodiment, in the process of determining the optimal conversation sequence for interacting with the target user, not only the predicted hang-up rate and predicted hang-up conversion rate of each conversation are considered, but also the interaction process information of interacting with the target user is combined. And this interaction process information includes the current conversation sequence, conversation subsequence, and subsequent conversation sequence for interacting with the target user, where the subsequent conversation sequence takes into account the hang-up information of the mid-call hang-up. Thus, the optimal conversation prediction for each node in this application takes into account the hang-up information. When the user cannot listen to the entire call, the optimal conversation sequence for the current situation can also be obtained. Furthermore, it solves the technical problem in the prior art that the existing conversation prediction method defaults that all calls can reach the final node. Since the possible mid-call hang-up phenomenon is not considered, the predicted conversation sequence is not the optimal result, resulting in low accuracy of conversation prediction.
[0082] Embodiment 3
[0083] Figure 7 Fig. 7 shows a device 700 for determining an optimal conversation sequence according to this embodiment. The optimal conversation sequence includes the optimal conversations corresponding to each conversation node for interacting with the target user. The device 700 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 7 As shown, the device 700 includes: a processor 710; and a memory 720, connected to the processor 710, for providing instructions for the processor 710 to perform the following processing steps: determining the user characteristic information and call information of the target user; using a pre-trained prediction model to respectively determine the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences and the interaction process information of interacting with the target user according to the user characteristic information and interaction process information, where the conversations corresponding to each conversation node constitute multiple conversation sequences; and determining the optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences, where the interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; the current conversation sequence is the sequence composed of the conversations corresponding to the current conversation node for interacting with the target user; the conversation subsequence is the sequence composed of the conversations corresponding to the previous conversation nodes for interacting with the target user; and the subsequent conversation sequence is the sequence composed of the conversations corresponding to the subsequent conversation nodes after the current conversation node.
[0084] Optionally, the operation of determining an optimal conversation sequence for interacting with a target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences and the interaction process information of interacting with the target user includes: determining the expected hang-up conversion rate of each conversation in multiple conversation sequences respectively according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences; and determining the conversation with the highest expected hang-up conversion rate corresponding to each conversation node as the optimal conversation, and generating an optimal conversation sequence.
[0085] Optionally, the operation of determining the expected hang-up conversion rate of each conversation in multiple conversation sequences respectively according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in multiple conversation sequences and the interaction process information of interacting with the target user includes: determining the expected hang-up conversion rate of each conversation in the current conversation sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the current conversation sequence and the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the subsequent conversation sequences; and determining the expected hang-up conversion rate of each conversation in multiple conversation sequences respectively according to the determined expected hang-up conversion rate of each conversation in the current conversation sequence.
[0086] Optionally, the operation of determining the expected hang-up conversion rate of each conversation in the current conversation sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the current conversation sequence and the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the subsequent conversation sequences includes: determining the first hang-up conversion rate of each conversation in the current conversation sequence according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of each conversation in the current conversation sequence; determining the second hang-up conversion rate of each conversation in the subsequent conversation sequences according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of each conversation in the subsequent conversation sequences; and determining the expected hang-up conversion rate of each conversation in the current conversation sequence according to the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate of each conversation in the subsequent conversation sequences.
[0087] Optionally, the operation of determining the expected hang-up conversion rate of each conversation in the current conversation sequence according to the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate of each conversation in the subsequent conversation sequences includes: selecting the conversation with the highest second hang-up conversion rate from the subsequent conversation sequences; and determining the expected hang-up conversion rate of each conversation in the current conversation sequence according to the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate corresponding to the conversation with the highest second hang-up conversion rate in the subsequent conversation sequences.
[0088] Optionally, the previous non-hang-up rate of each conversation statement in the subsequent conversation statement sequence is determined by the following operations: according to the predicted hang-up rate of each conversation statement in the current conversation statement sequence, the previous non-hang-up rate of each conversation statement in the subsequent conversation statement sequence is determined.
[0089] Therefore, according to this embodiment, in the process of determining the optimal conversation statement sequence for interacting with the target user, not only the predicted hang-up rate and the predicted hang-up conversion rate of each conversation statement are considered, but also the interaction process information of interacting with the target user is combined, and the interaction process information includes the current conversation statement sequence, the conversation statement subsequence, and the subsequent conversation statement sequence, where the subsequent conversation statement sequence takes into account the hang-up information of hanging up midway. Therefore, each node's optimal conversation statement prediction in this application considers the hang-up information. Even when the user cannot listen to the entire conversation, the optimal conversation statement sequence in the current situation can still be obtained. Furthermore, it solves the technical problem in the prior art that the existing conversation statement prediction method defaults that all calls can reach the final node. Since the phenomenon of hanging up midway may not be considered, the predicted conversation statement sequence is not the optimal result, resulting in low accuracy of conversation statement prediction.
[0090] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0091] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0092] In the several embodiments provided by this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0093] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, in each embodiment of the present invention, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0095] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0096] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for determining an optimal conversation strategy sequence, the optimal conversation strategy sequence including optimal conversation strategies corresponding to respective conversation strategy nodes for interacting with a target user, characterized in that, Including: Determine the user characteristic information and call information of the target user; Using a pre-trained prediction model, respectively determine the predicted hang-up rate and predicted hang-up conversion rate of each call in multiple call sequences according to the user characteristic information and the call information, where the calls corresponding to the respective call nodes constitute the multiple call sequences; And According to the predicted hang-up rate and predicted hang-up conversion rate of each call in the multiple call sequences and the interaction process information of interacting with the target user, determine the optimal call sequence for interacting with the target user, where The interaction process information includes: the current call sequence for interacting with the target user, the call subsequence for interacting with the target user, and the subsequent call sequence for interacting with the target user; The current call sequence is the sequence composed of the calls corresponding to the current call node for interacting with the target user; The call subsequence is the sequence composed of the calls corresponding to the previous call nodes for interacting with the target user; and The subsequent call sequence is the sequence composed of the calls corresponding to the subsequent call nodes after the current call node.
2. The method according to claim 1, wherein The operation of determining the optimal call sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each call in the multiple call sequences includes: Respectively determine the expected hang-up conversion rate of each call in the multiple call sequences according to the predicted hang-up rate and predicted hang-up conversion rate of each call in the multiple call sequences and the interaction process information of interacting with the target user; and Respectively determine the call with the highest expected hang-up conversion rate corresponding to each call node as the optimal call, and generate the optimal call sequence.
3. The method according to claim 1, wherein The operation of respectively determining the expected hang-up conversion rate of each call in the multiple call sequences according to the predicted hang-up rate and predicted hang-up conversion rate of each call in the multiple call sequences and the interaction process information of interacting with the target user includes: Determine the expected hang-up conversion rate of each call in the current call sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each call in the current call sequence and the predicted hang-up rate and predicted hang-up conversion rate of each call in the subsequent call sequence; and Respectively determine the expected hang-up conversion rate of each call in the multiple call sequences according to the determined expected hang-up conversion rate of each call in the current call sequence.
4. The method according to claim 3, wherein The operation of determining the expected hang-up conversion rate of each call in the current call sequence according to the predicted hang-up rate and predicted hang-up conversion rate of each call in the current call sequence and the predicted hang-up rate and predicted hang-up conversion rate of each call in the subsequent call sequence includes: Determine the first hang-up conversion rate of each call in the current call sequence according to the previous non-hang-up rate, predicted hang-up rate and predicted hang-up conversion rate of each call in the current call sequence; Determine the second hang-up conversion rate of each conversation in the subsequent conversation sequence according to the previous non-hang-up rate, predicted hang-up rate, and predicted hang-up conversion rate of each conversation in the subsequent conversation sequence; Determine the expected hang-up conversion rate of each conversation in the current conversation sequence based on the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate of each conversation in the subsequent conversation sequence.
5. The method according to claim 4, wherein The operation of determining the expected hang-up conversion rate of each conversation in the current conversation sequence based on the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate of each conversation in the subsequent conversation sequence includes: Select the conversation with the highest second hang-up conversion rate from the subsequent conversation sequence; and Determine the expected hang-up conversion rate of each conversation in the current conversation sequence based on the first hang-up conversion rate of each conversation in the current conversation sequence and the second hang-up conversion rate corresponding to the conversation with the highest second hang-up conversion rate in the subsequent conversation sequence.
6. The method according to claim 4, characterized in that, Determine the previous non-hang-up rate of each conversation in the subsequent conversation sequence through the following operations: Determine the previous non-hang-up rate of each conversation in the subsequent conversation sequence according to the predicted hang-up rate of each conversation in the current conversation sequence.
7. A storage medium, characterized in that, The storage medium includes a stored program, wherein the method according to any one of claims 1 to 6 is executed by a processor when the program runs.
8. An apparatus for determining an optimal speech sequence, the optimal speech sequence including optimal speeches corresponding to respective speech nodes for interacting with a target user, characterized in that Including: A first determination module for determining the user characteristic information and call information of the target user; A second determination module for using a pre-trained prediction model to respectively determine the predicted hang-up rate and predicted hang-up conversion rate of each conversation in a plurality of conversation sequences according to the user characteristic information and the call information, wherein the conversations corresponding to each conversation node constitute the plurality of conversation sequences; And A third determination module for determining the optimal conversation sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and the interaction process information of interacting with the target user, wherein The interaction process information includes: the current conversation sequence for interacting with the target user, the conversation subsequence for interacting with the target user, and the subsequent conversation sequence for interacting with the target user; The current conversation sequence is a sequence composed of the conversations corresponding to the current conversation node for interacting with the target user; The conversation subsequence is a sequence composed of the conversations corresponding to the previous conversation nodes for interacting with the target user; and The subsequent conversation sequence is a sequence composed of the conversations corresponding to the subsequent conversation nodes after the current conversation node.
9. The device according to claim 8, characterized in that, The third determination module includes: A determination sub-module for respectively determining the expected hang-up conversion rate of each conversation in the plurality of conversation sequences according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation in the plurality of conversation sequences and the interaction process information of interacting with the target user; and A generation sub-module for respectively determining the conversation with the highest expected hang-up conversion rate corresponding to each conversation node as the optimal conversation, and generating the optimal conversation sequence.
10. A device for determining an optimal conversation sequence, where the optimal conversation sequence includes optimal conversations corresponding to each conversation node for interacting with a target user, characterized in that Comprising: A processor; And A memory, connected to the processor, for providing instructions for the processor to process the following processing steps: Determine the user characteristic information and call information of the target user; Utilize a pre-trained prediction model to respectively determine the predicted hang-up rate and predicted hang-up conversion rate of each conversation technique in a plurality of conversation technique sequences according to the user characteristic information and the call information, wherein the conversation techniques corresponding to the respective conversation technique nodes constitute the plurality of conversation technique sequences; And Determine an optimal conversation technique sequence for interacting with the target user according to the predicted hang-up rate and predicted hang-up conversion rate of each conversation technique in the plurality of conversation technique sequences and the interaction process information of interacting with the target user, wherein The interaction process information includes: the current conversation technique sequence of interacting with the target user, the conversation technique subsequence of interacting with the target user, and the subsequent conversation technique sequence of interacting with the target user; The current conversation technique sequence is a sequence constituted by the conversation techniques corresponding to the current conversation technique node of interacting with the target user; The conversation technique subsequence is a sequence constituted by the conversation techniques corresponding to the previous conversation technique nodes of interacting with the target user; and The subsequent conversation technique sequence is a sequence constituted by the conversation techniques corresponding to the subsequent conversation technique nodes after the current conversation technique node.
Citation Information
Patent Citations
Intelligent outbound processing method and device, computer equipment and storage medium
CN110351443A
Enhanced call return in a wireless telephone network
US20090042550A1