Method, apparatus, and storage medium for determining conversation tactics

By determining the user's current interaction node user portrait and interaction process information during the interaction process, and using pre-trained models to calculate the target speech, the problem of insufficient speech accuracy in the interaction process in the prior art is solved, and an interactive experience that is more accurate and close to user needs is achieved.

CN114005439BActive Publication Date: 2025-06-13BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010732572.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-06-13
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

The prior art lacks reference to all the interactive content generated from the beginning of the interaction to the present when determining the speech during the interaction process, resulting in insufficient accuracy of the speech.

Method used

By determining the user's current interaction node user portrait during the interaction process, and obtaining the interaction process information from the beginning of the interaction to the current interaction node, using a pre-trained model to calculate the user portrait and interaction process information, the target speech at the current interaction node is determined.

Benefits of technology

It improves the accuracy of the speech, makes it closer to the needs of users, provides high-quality user experience, and solves the problem of insufficient accuracy of speech in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005439B_ABST
    Figure CN114005439B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, and storage medium for determining a conversation strategy. The method includes: determining a user profile of a user at a current interaction node during an interaction with the user; obtaining interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe node conversation strategies adopted at at least one interaction node during the interaction, response information of the user to the node conversation strategies, and the sequential relationship between the node conversation strategies and the response information; and calculating the user profile and the interaction process information using a pre-trained model to determine a target conversation strategy for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation strategy conversion results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction technologies, and in particular, to a method, an apparatus, and a storage medium for determining a conversation strategy. Background Art

[0002] Due to the development of artificial intelligence technology, there are now many voice communication systems in which robots communicate with people, especially in voice customer service, intelligent telemarketing, intelligent collection, intelligent speakers and other voice interaction scenarios. After the robot can communicate with people by voice, it is necessary to continuously optimize the robot's conversation strategy to improve the communication effect. The existing conversation strategy selection methods mainly select a conversation strategy for each answer of the user, without considering all the question-and-answer information in the entire interaction process (that is, all the interaction content from the start of the interaction to the current time). Therefore, the accuracy of the finally determined conversation strategy remains to be improved.

[0003] In view of the above technical problem that the method for determining a conversation strategy in the interaction process in the prior art lacks reference to all the interaction content generated from the start of the interaction to the current time, thus affecting the accuracy of the conversation strategy, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, an apparatus, and a storage medium for determining a conversation strategy, so as to at least solve the technical problem that the method for determining a conversation strategy in the interaction process in the prior art lacks reference to all the interaction content generated from the start of the interaction to the current time, thus affecting the accuracy of the conversation strategy.

[0005] According to one aspect of the embodiments of the present disclosure, a method for determining a conversation strategy is provided, including: during the interaction with a user, determining a user profile of the user at the current interaction node; obtaining interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe at least one node conversation strategy adopted at an interaction node in the interaction process, the answer information of the user to the node conversation strategy, and the sequential relationship between the node conversation strategy and the answer information; and using a pre-trained model to calculate the user profile and the interaction process information to determine a target conversation strategy for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation strategy conversion result.

[0006] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided, where the storage medium includes a stored program, and when the program runs, the method described in any one of the above is executed by a processor.

[0007] According to another aspect of the embodiments of the present disclosure, there is also provided an apparatus for determining a conversation statement, including: a portrait determination module, configured to determine a user portrait of a user at a current interaction node during an interaction with the user; a data acquisition module, configured to acquire interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe node conversation statements adopted by at least one interaction node during the interaction, response information of the user to the node conversation statements, and the sequential relationship between the node conversation statements and the response information; and a conversation statement determination module, configured to calculate the user portrait and the interaction process information by using a pre-trained model, and determine a target conversation statement for interacting with the user at the current interaction node, where the model is trained based on the user portrait, the interaction process information, and the corresponding conversation statement conversion results.

[0008] According to another aspect of the embodiments of the present disclosure, there is also provided an apparatus for determining a conversation statement, including: a processor; and a memory connected to the processor, configured to provide instructions for the processor to perform the following processing steps: determining a user portrait of a user at a current interaction node during an interaction with the user; acquiring interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe node conversation statements adopted by at least one interaction node during the interaction, response information of the user to the node conversation statements, and the sequential relationship between the node conversation statements and the response information; and calculating the user portrait and the interaction process information by using a pre-trained model, and determining a target conversation statement for interacting with the user at the current interaction node, where the model is trained based on the user portrait, the interaction process information, and the corresponding conversation statement conversion results.

[0009] In the embodiments of the present disclosure, first, a user portrait of the current interaction node and interaction process information are determined, and then a model is used to predict the conversation statement adopted by the current interaction node according to the user portrait and the interaction process information. Thus, compared with the prior art, the present solution can combine real-time user portrait information and interaction process information (question and answer information) during the process of determining the conversation statement. Therefore, the finally determined target conversation statement can be closer to the user's needs, more accurate, and can provide a high-quality experience for the user. Furthermore, it solves the technical problem in the prior art that the method for determining the conversation statement during the interaction lacks reference to all interaction contents generated from the start of the interaction to the current, thus affecting the accuracy of the conversation statement. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present disclosure, and constitute a part of this application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure, and do not constitute an improper limitation to the present disclosure. In the drawings:

[0011] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure;

[0012] Figure 2 It is a schematic flowchart of the method for determining the conversation words according to the first aspect of Embodiment 1 of the present disclosure;

[0013] Figure 3 It is a schematic diagram of the interaction node according to Embodiment 1 of the present disclosure;

[0014] Figure 4 It is a schematic diagram of model training and prediction according to Embodiment 1 of the present disclosure;

[0015] Figure 5 It is a schematic diagram of the device for determining the conversation words according to Embodiment 2 of the present disclosure; and

[0016] Figure 6 It is a schematic diagram of the device for determining the conversation words according to Embodiment 3 of the present disclosure. Detailed implementation manners

[0017] In order to enable those skilled in the art of the present technology to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present disclosure.

[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0019] Embodiment 1

[0020] According to this embodiment, an embodiment of a method for determining the conversation words is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.

[0021] The method embodiment provided in this embodiment may be executed on a server or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computing device for implementing a method for determining a conversation strategy. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computing device may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.

[0022] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0023] The memory may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for determining a conversation strategy in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for determining a conversation strategy of the above-mentioned application program. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories may be connected to the computing device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0024] The transmission device is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computing device. In one example, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0025] The display can be, for example, a touch-screen Liquid Crystal Display (LCD), which enables users to interact with the user interface of the computing device.

[0026] It should be noted here that in some alternative embodiments, the above Figure 1 shown computing device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific concrete instance and is intended to show the types of components that may exist in the above computing device.

[0027] Under the above operating environment, according to the first aspect of this embodiment, a method for determining conversation words is provided. This method can be applied, for example, to the server of an intelligent robot interaction system. The robot of this interaction system can handle interaction processes such as telephone return visits, marketing, complaints, and consultations. Figure 2 The flow diagram of this method is shown. Refer to Figure 2 shown, this method includes:

[0028] S202: During the interaction with the user, determine the user profile at the current interaction node;

[0029] S204: Obtain the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation words adopted by at least one interaction node during the interaction process, the user's answer information to the node conversation words, and the sequential relationship between the node conversation words and the answer information; and

[0030] S206: Use a pre-trained model to calculate the user profile and the interaction process information, and determine the target conversation words for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation word conversion results.

[0031] As described in the background art, due to the development of artificial intelligence technology, there have emerged many voice communication systems between robots and humans. In particular, voice interaction scenarios such as voice customer service, intelligent telemarketing, intelligent debt collection, and intelligent speakers are widely used. After the robot can communicate with people by voice, it is necessary to continuously optimize the robot's conversation skills to improve the communication effect. The existing method of selecting conversation skills mainly selects conversation skills for each answer of the user, without considering all the question-and-answer information in the entire interaction process (that is, all the interaction content from the start of the interaction to the current time). Therefore, the accuracy of the finally determined conversation skills remains to be improved.

[0032] To address the technical problems in the background art, in step S202 of the technical solution of this embodiment, during the interaction between the robot and the user, the system server first determines the user profile at the current interaction node. Among them, the interaction process may include multiple interaction nodes. In a specific example, as Figure 3 shown, it includes an opening interaction node, a guiding method interaction node, a credit line conversation skill interaction node, a marketing selling point interaction node, etc. Each interaction node includes multiple node conversation skills. For example, the opening node includes name verification (verifying name and identity) conversation skills and self-introduction conversation skills. Among them, the current interaction node is the interaction node at the current time in the interaction process. For example, during the interaction with the user, the interaction processes of the opening and the guiding method have been completed, and the interaction of the credit line conversation skill node is about to start. Then the current interaction node corresponds to the credit line conversation skill node. The server needs to determine the user profile of the current interaction node, that is, the user's behavior or characteristics before the current interaction node. In a specific example, taking telephone debt collection as an example, the user profile may include, for example but not limited to, the user's gender, age, statistical characteristics of historical purchase records, login and access behavior records, loan records, etc.

[0033] Further, in step S204, the server obtains the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation phrases adopted by at least one interaction node during the interaction, the response information of the user to the node conversation phrases, and the sequential relationship between the node conversation phrases and the response information. That is, the server obtains the call record between the robot and the user from the start of the interaction to the current interaction node. In a specific example, the node conversation phrase at the opening interaction node is "self-introduction", and then the user can make corresponding response information to this "self-introduction" phrase. The response information can be intention information indicating the user's intention (for example: affirmative intention when the user is interested, or other intentions when the user is not interested). In addition, the user's response information can also be specific text information (for example: specific response words, and then converted into corresponding text). Then the server can select an appropriate node conversation phrase at the guiding interaction node for guidance according to the user's response information. For example, if the phrase adopted at the guiding node is "return visit guidance", then "self-introduction", the response information, and "return visit guidance" constitute the interaction process information, and the corresponding order is: self-introduction - response information - return visit guidance. In addition, during the process of selecting the interaction node, it can be selected according to the user's response information, and finally sorted in order according to the robot's node conversation phrases and the user's response information to form the interaction process information.

[0034] Finally, in step S206, the server uses the pre-trained model to calculate the user portrait and the interaction process information to determine the target conversation phrase for interacting with the user at the current interaction node, that is, to determine the target conversation phrase adopted by the quota conversation phrase interaction node (the current interaction node). The model is trained based on the user portrait, the interaction process information, and the corresponding conversation phrase conversion results. Specifically, during the model training process, reference is made to Figure 4 As shown, first, the call record (historical call record) of each predicted conversation phrase is obtained, including information such as robot questions, user answers, and conditional judgments. Further, the user portrait information (gender, age, historical purchase record statistics characteristics, login access behavior records, loan records) at the time point corresponding to each call record (the current interaction node) is obtained. Then, based on the user portrait information, the call record, and the conversion result (whether the conversion is successful) of this predicted conversation phrase, the model is trained. In addition, the modeling method of the model is, for example but not limited to, modeling methods such as logistic regression, tree type models, and deep learning. After the model training is completed, this model can be used to predict the conversation phrase according to the user portrait and the interaction process information to obtain the target conversation phrase.

[0035] In this way, the server first determines the user profile and interaction process information of the current interaction node, and then uses the model to predict the conversation strategy adopted for the current interaction node based on the user profile and interaction process information. Therefore, compared with the prior art, in the process of determining the conversation strategy, this solution can combine real-time user profile information and interaction process information (question-and-answer information). As a result, the final determined target conversation strategy can be closer to the user's needs, more accurate, and can provide a better experience for the user. Furthermore, it solves the technical problem in the prior art that the method for determining the conversation strategy in the interaction process lacks reference to all the interaction content generated from the beginning to the current interaction, thus affecting the accuracy of the conversation strategy.

[0036] Optionally, using a pre-trained model to calculate the user profile and interaction process information to determine the target conversation strategy for interacting with the user at the current interaction node includes: determining first feature information corresponding to the user profile, determining second feature information corresponding to the interaction process information; and using the pre-trained model to calculate the first feature information and the second feature information to determine the target conversation strategy for interacting with the user at the current interaction node.

[0037] Specifically, in the operation of using a pre-trained model to calculate the user profile and interaction process information to determine the target conversation strategy for interacting with the user at the current interaction node, the server first determines the first feature information corresponding to the user profile and the second feature information corresponding to the interaction process information, that is, performs feature representation on the user profile and interaction process information. The representation method can be, for example, the feature representation method in the prior art. Among them, when the user profile is discrete, for example, Embedding encoding (Onehot can also be used) can be used to determine the first feature information, and when the user profile is a continuous feature, it is directly used. For the interaction process information, for example, Embedding encoding (Onehot can also be used) can also be used for processing, and then the Embedding encodings corresponding to each interaction record (node conversation strategy and the user's answer information) are combined into a sequence array (corresponding to the second feature information). Further, the server uses the pre-trained model to calculate the first feature information and the second feature information to determine the target conversation strategy for interacting with the user at the current interaction node. Among them, in the process of training this model, after determining the call record and the user profile, the obtained data (user profile and call record) is also subjected to feature processing to form training data available for the model, and then the model is trained using the feature data.

[0038] Optionally, the response information corresponds to the user's intention information. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: representing the node conversation and intention information in the interaction process information using a preset identifier, determining the interaction information sequence corresponding to the interaction process information, and determining the second feature information corresponding to the interaction process information, including: determining the second feature information corresponding to the interaction information sequence.

[0039] Specifically, the user's response information corresponds to the user's intention information. The intention information, for example, includes a positive intention and other intentions. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, the server also represents the node conversation and intention information in the interaction process information using a preset identifier, and determines the interaction information sequence corresponding to the interaction process information. Refer to Figure 3As shown, in a specific example, for instance, the identifier adopted for node conversation scripts can be represented by SXY, where X represents the serial number of the interaction node. For example, the sequence of the opening remarks, the guiding method, and the quota conversation script is 1, 2, and 3 in turn. Y represents the serial number of the node conversation script in each interaction node. For example, in the opening remarks interaction node, the sequence of the "surname identity verification" conversation script and the "self-introduction" conversation script is 1 and 2 in turn. Then, the identifier corresponding to the "surname identity verification" node conversation script in the opening remarks interaction node is S11, and the identifier corresponding to the "self-introduction" conversation script is S12. Similarly, in the guiding method interaction node, the identifier corresponding to the "polite guidance" conversation script is S21, the identifier corresponding to the "direct marketing" conversation script is S22, and the identifier corresponding to the "return visit guidance" conversation script is S23, etc. The intention information can be represented by IXY for example. The identifier corresponding to the positive intention is I11, and the identifier corresponding to other intentions is I12. In addition, it can also include the condition judgment identifier JXY. The identifier corresponding to meeting the condition (within 10 days of the repayment time) can be J11 for example, and the identifier corresponding to not meeting the condition (the repayment time is other times) is J21. By using the preset identifiers to represent the node conversation scripts and intention information in the interaction process information, the corresponding interaction information sequence can be obtained. For example, the form of the interaction information sequence can be: S11->S21 (indicating that the interaction process is to first conduct surname identity verification and then polite guidance), S12->S22->J12->S41 (self-introduction - direct marketing - the user does not meet the condition - positive benefit notification). And, in the operation of determining the second feature information corresponding to the interaction process information, the server determines the second feature information corresponding to the interaction information sequence. That is: converting the interaction information sequence into the corresponding second feature information. In addition, during the model training process, training is carried out according to this second feature information. Thus, the user's intention can be introduced in the process of determining the features, so that the result can be closer to the user. In addition, the text information of the interaction process can be first converted into the corresponding identifier, and then the corresponding second feature information can be generated according to the identifier, so that the generation of feature information can be facilitated. In addition, using the identifier to represent the node unifies the information expression methods of the robot questions, user answers, condition judgments, etc., which is convenient for subsequent processing.

[0040] Optionally, the answer information corresponds to the user's answer text. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: using the preset identifier to represent the node conversation script in the interaction process information to generate a sentence vector corresponding to the answer text; and determining the interaction information sequence corresponding to the interaction process information according to the sentence vector and the node conversation script represented by the identifier, and determining the second feature information corresponding to the interaction process information, including: determining the second feature information corresponding to the interaction information sequence.

[0041] Specifically, the response information can also correspond to the user's response text, that is, the user's specific interaction content. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, the server can also use a preset identifier to represent the node conversation in the interaction process information, that is, the above-mentioned use of the SXY identifier to represent the node conversation, and generate a sentence vector corresponding to the response text. The method of generating the sentence vector can be, for example, the method in the prior art, such as generating by a BERT model, or other models can also be used, which is not specifically limited here. After representing the node conversation with an identifier and converting the response text into a corresponding sentence vector, an interaction information sequence corresponding to the interaction process information is determined according to the sentence vector and the node conversation represented by the identifier. For example: connecting the identifier and the sentence vector in sequence. Then, in the operation of determining the second feature information corresponding to the interaction process information, the server determines the second feature information corresponding to the interaction information sequence. In addition, during the model training process, training is performed according to this second feature information. Thus, the sentence vector of the user's response information can be directly introduced during the generation of the feature information, so it can be closer to the user's needs and facilitate the generation of the feature information. In addition, using an identifier to represent the node unifies the information expression methods of robot questions, user answers, conditional judgments, etc., which is convenient for subsequent processing.

[0042] Optionally, use a pre-trained model to calculate the first feature information and the second feature information to determine the target conversation for interacting with the user at the current interaction node, including: using a pre-trained model to calculate the first feature information and the second feature information to determine the probability values corresponding to multiple conversations at the current interaction node; and according to the probability values, determine the target conversation for interacting with the user at the current interaction node from multiple conversations.

[0043] Specifically, in the operation of using a pre-trained model to calculate the first feature information and the second feature information to determine the target conversation for interacting with the user at the current interaction node, the server first uses a pre-trained model to calculate the first feature information and the second feature information to determine the probability values corresponding to multiple conversations at the current interaction node. For example: the current interaction node is a credit limit conversation interaction node, and the credit limit conversation interaction node includes two node conversations, "Credit limit inquiry 1" and "Credit limit inquiry 2". After model calculation, the probability values corresponding to the two conversations can be obtained, and then according to the probability values, the target conversation for interacting with the user at the current interaction node is determined from multiple conversations. For example: select the one with the larger probability value as the target conversation. The target conversation can be clearly and quickly determined through the probability values.

[0044] Optionally, the method further includes training a model according to the following operations: obtaining the call records of the user and the user profile of the user; determining the call record feature information corresponding to the call records and the user feature information corresponding to the user profile; and training the model according to the call record feature information, the user feature information, and the conversion result of the call records.

[0045] Specifically, as shown in Figure 4 During the process of training the model, first obtain the call records (historical call records) of each predicted conversation script and the user profile information (gender, age, statistical characteristics of historical purchase records, login access behavior records, loan records). Then, perform feature representation on the user profile information and the call records. Further, train the model according to the user profile information features, the call record features, and the conversion result (whether the conversion is successful) of the predicted conversation script. After the model training is completed, the model can be used to predict the conversation script according to the user profile and the interaction process information to obtain the target conversation script.

[0046] In addition, it should be supplemented that, as shown in the schematic diagram of Figure 4 model training and model prediction shown in

[0047] Model training includes the following steps:

[0048] 1. Obtain the call records of each predicted conversation script. In addition to the call attribute information (call time, call duration, what task, what activity, what robot), the path information of the dialogue access (corresponding to the call record) also needs to be obtained. The call record includes information such as robot questions, user answers, and conditional judgments.

[0049] 2. Obtain the user profile information (gender, age, statistical characteristics of historical purchase records, login access behavior records, loan records) at the time point corresponding to each call record.

[0050] 3. Perform feature processing on the obtained data to form training data available for the model. Among them, identify the call records to form a sequence of dialogue access node paths (corresponding to the above-mentioned interaction information sequence).

[0051] 4. Use the model with the real-time features of the user profile and the features of the user dialogue access path information for training and publish the prediction service.

[0052] Model prediction includes the following steps:

[0053] 1. Obtain the above-mentioned information of the current conversation (corresponding to the interaction process information), including information such as robot questions, user answers, and conditional judgments, to form a sequence of dialogue access node paths (represented by identification to generate an interaction information sequence).

[0054] 2. Obtain the real-time portrait information of the current user and save the features for easy use in the next training.

[0055] 3. Perform feature processing on the acquired data to form training data that can be used for the model.

[0056] 4. Use the real-time features of the user portrait and the context of the user conversation + each predicted speech to be predicted for prediction (model prediction), and select the predicted speech with the highest conversion rate for recommendation.

[0057] In a specific example, reference Figure 3 As shown, the interaction nodes include opening remarks, guidance methods, quota words, and marketing selling points. Each interaction node includes multiple node words (subsequent nodes are omitted). In addition, it also includes judgment nodes such as conditional judgment and intention judgment that will lead to different subsequent words. Among them, the node words are represented by the SXY logo, the intention is represented by the IXY logo, and the judgment condition is represented by the JXY logo. Table 1 shows the interaction information sequence after the logo representation, the user portrait, and the call record table for model training whether it is converted. Table 1 is as follows:

[0058] Table 1

[0059] Call record User characteristics (user profile) Interaction information sequence Conversion flag User 1 …… S11 -> S21 1 User 2 …… S12 -> S22 -> J12 -> S41 0 User 3 …… S11 -> S23 -> S31 -> I12 -> S42 1

[0060] Referring to Table 1, the training process is as follows:

[0061] 1. Get all dial records, including dial records, conversion identifiers, and interaction information sequences.

[0062] 2. Obtain the user portrait information at the corresponding time point of each call record to form user feature xxx (user portrait).

[0063] 3. Discrete features in user portrait features use Embedding encoding (Onehot can also be used), and continuous features are used directly. For interactive information sequences, each node uses Embedding encoding (Onehot can also be used), and then all nodes form an Embedding encoded sequence array. (Feature representation of data)

[0064] 4. Train the conversion model for the above data and release the prediction service. The user features are modeled using conventional DNN (or other models such as DeepFM), and the node process is modeled using Attention (or RNN or other sequential models) (based on user feature prediction). Then, the various parts are combined (can be connected in parallel) for modeling. (The process of modeling based on user portrait features, call record features, and the final conversion results).

[0065] The model predicts the following:

[0066] 1. Obtain the previous information (interaction process information) of the current conversation, and add each node to be predicted to form node process information.

[0067] 2. Obtain the real-time portrait information of the user at the current interaction node, form user feature xxx information, and save the features for future training use.

[0068] 3. Process the obtained data to form training data available for the model. The processing method is the same as that in the training stage.

[0069] 4. Use the real-time features of the user portrait and the features of the user interaction process information to predict the conversion rate of the words of each node to be predicted, and use the words of the node with the highest conversion rate as the target words.

[0070] The present solution has the following advantages:

[0071] By using real-time features of calls, it solves the problem that only offline features were used before, and the features may have changed during prediction.

[0072] After introducing user dialogue node information, each recommended word can utilize the previous information of the conversation, combined with the previous words of the robot and the previous answers of the user in the current conversation for recommendation. Therefore, each word prediction is the optimal combination prediction at present, solving the problem that the previous solutions of [opening remarks] and [guiding methods] were independently predicted without correlation. For example, [opening remarks S11] and [guiding method S21] were the best when independently predicted, but the combination of [opening remarks S11] and [guiding method S22] was better than the combination of [opening remarks S11] and [guiding method S21], and such a situation could not be recommended, reducing the model effect.

[0073] The present solution considers the global information of the robot, enables a model to learn the difference in the effects of all word combinations, directly enables the model to learn the combination with the optimal effect, and achieves the optimal effect of the robot. The present solution generates the following benefits:

[0074] For offline training, using real-time user features can obtain more accurate user information to improve the model prediction ability, such as information that is likely to change, such as whether the user has logged in to the APP today in telemarketing.

[0075] Introducing user dialogue node information will increase more conversation information and can solve the problem of optimal recommendation of word combinations. For example, the best of [guiding method S2X] may have different results with or without referring to [opening remarks]. Using all branch nodes as dialogue state nodes unifies the information expression methods of robot questions, user answers, conditional judgments, etc., facilitating subsequent processing.

[0076] In addition, this solution uses real-time user session information to address possible changes during user information prediction. Whether using any model such as logistic regression, tree-type models, deep learning, or making various changes to features (such as feature crossing, feature mapping), adding feature information, as long as the sample construction method is as shown in this solution, it only counts as an algorithm application of the modeling method of this solution.

[0077] User portrait features refer to all features that can be associated with a user, including but not limited to user attributes, user behaviors, user historical call information, etc.

[0078] Whether multiple conversation scripts are calculated in parallel or using a matrix method during conversation script prediction is just an implementation method for accelerating prediction in this patent.

[0079] This solution is applicable to all intelligent voice conversation script prediction scenarios, including but not limited to intelligent outbound calls, inbound calls, voice collections, callbacks, telemarketing, etc., and is also applicable to text multi-turn conversation script prediction scenarios.

[0080] In addition, referring to Figure 1 As shown, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program runs, the method described in any one of the above is executed by a processor.

[0081] Thus, according to this embodiment, first, the user portrait and interaction process information of the current interaction node are determined, and then the model is used to predict the conversation script adopted for the current interaction node based on the user portrait and interaction process information. Thus, compared with the prior art, this solution can combine real-time user portrait information and interaction process information (question and answer information) during the process of determining the conversation script. Therefore, the finally determined target conversation script can be closer to the user's needs, more accurate, and can provide a high-quality experience for the user. Furthermore, it solves the technical problem in the prior art that the method for determining the conversation script during the interaction process lacks reference to all interaction content generated from the start of the interaction to the current, thus affecting the accuracy of the conversation script.

[0082] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0084] Embodiment 2

[0085] Figure 5 Fig. shows a device 500 for determining a conversation statement according to the present embodiment. The device 500 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 5 As shown, the device 500 includes: a portrait determination module 510, configured to determine a user portrait of the user at the current interaction node during the interaction with the user; a data acquisition module 520, configured to acquire interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation statements adopted by at least one interaction node during the interaction process, the response information of the user to the node conversation statements, and the sequential relationship between the node conversation statements and the response information; and a conversation statement determination module 530, configured to calculate the user portrait and the interaction process information by using a pre-trained model, and determine a target conversation statement for interacting with the user at the current interaction node, where the model is trained based on the user portrait, the interaction process information, and the corresponding conversation statement conversion result.

[0086] Optionally, the conversation statement determination module 530 includes: a feature determination sub-module, configured to determine first feature information corresponding to the user portrait and second feature information corresponding to the interaction process information; and a conversation statement determination sub-module, configured to calculate the first feature information and the second feature information by using a pre-trained model, and determine a target conversation statement for interacting with the user at the current interaction node.

[0087] Optionally, the response information corresponds to the intention information of the user. The device further includes: a first conversion module, configured to, after acquiring the interaction process information generated from the start of the interaction to the current interaction node, use a preset identifier to represent the node conversation statements and the intention information in the interaction process information, determine an interaction information sequence corresponding to the interaction process information, and the feature determination sub-module includes: a first feature determination unit, configured to determine second feature information corresponding to the interaction information sequence.

[0088] Optionally, the response information corresponds to the user's response text, and the device further includes: a second conversion module, configured to, after obtaining the interaction process information generated from the start of the interaction to the current interaction node, represent the node conversation in the interaction process information using a preset identifier, and generate a sentence vector corresponding to the response text; and a sequence determination module, configured to determine an interaction information sequence corresponding to the interaction process information according to the sentence vector and the node conversation represented by the identifier, and a feature determination sub-module, including: a second feature determination unit, configured to determine second feature information corresponding to the interaction information sequence.

[0089] Optionally, the conversation determination sub-module includes: a calculation unit, configured to calculate the first feature information and the second feature information using a pre-trained model to determine probability values corresponding to multiple conversations at the current interaction node; and a determination unit, configured to determine a target conversation for interacting with the user at the current interaction node from the multiple conversations according to the probability values.

[0090] Optionally, the device 500 further includes a model training module, configured to train the model according to the following operations: obtain the call records of the user and the user profile of the user; determine call record feature information corresponding to the call records and user feature information corresponding to the user profile; and train the model according to the call record feature information, the user feature information, and the conversion result of the call records.

[0091] Therefore, according to this embodiment, first, the user profile and the interaction process information of the current interaction node are determined, and then the model is used to predict the conversation adopted at the current interaction node according to the user profile and the interaction process information. Thus, compared with the prior art, in the process of determining the conversation in this solution, real-time user profile information and interaction process information (question and answer information) can be combined. Therefore, the finally determined target conversation can be closer to the user's needs, more accurate, and can provide a good experience for the user. Furthermore, it solves the technical problem in the prior art that the method for determining the conversation in the interaction process lacks reference to all the interaction content generated from the start of the interaction to the current, thus affecting the accuracy of the conversation.

[0092] Embodiment 3

[0093] Figure 6 Shows a device 600 for determining a conversation according to this embodiment, and the device 600 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 6As shown, the device 600 includes: a processor 610; and a memory 620, connected to the processor 610, for providing instructions for the processor 610 to process the following processing steps: during the interaction with the user, determining the user profile of the user at the current interaction node; obtaining the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation words adopted by at least one interaction node during the interaction, the response information of the user to the node conversation words, and the sequential relationship between the node conversation words and the response information; and using a pre-trained model to calculate the user profile and the interaction process information to determine the target conversation words for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation word conversion results.

[0094] Optionally, using a pre-trained model to calculate the user profile and the interaction process information to determine the target conversation words for interacting with the user at the current interaction node includes: determining the first feature information corresponding to the user profile, determining the second feature information corresponding to the interaction process information; and using the pre-trained model to calculate the first feature information and the second feature information to determine the target conversation words for interacting with the user at the current interaction node.

[0095] Optionally, the response information corresponds to the intention information of the user. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: using a preset identifier to represent the node conversation words and the intention information in the interaction process information, determining the interaction information sequence corresponding to the interaction process information, and determining the second feature information corresponding to the interaction process information, including: determining the second feature information corresponding to the interaction information sequence.

[0096] Optionally, the response information corresponds to the response text of the user. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: using a preset identifier to represent the node conversation words in the interaction process information, generating a sentence vector corresponding to the response text; and determining the interaction information sequence corresponding to the interaction process information according to the sentence vector and the node conversation words represented by the identifier, and determining the second feature information corresponding to the interaction process information, including: determining the second feature information corresponding to the interaction information sequence.

[0097] Optionally, using a pre-trained model to calculate the first feature information and the second feature information to determine the target conversation words for interacting with the user at the current interaction node includes: using the pre-trained model to calculate the first feature information and the second feature information to determine the probability values corresponding to multiple conversation words at the current interaction node; and determining the target conversation words for interacting with the user at the current interaction node from the multiple conversation words according to the probability values.

[0098] Optionally, the memory 620 is further configured to provide instructions for the processor 610 to process the following training model: obtain the call records of the user and the user profile of the user; determine the call record feature information corresponding to the call records and the user feature information corresponding to the user profile; and train the model according to the call record feature information, the user feature information, and the conversion result of the call records.

[0099] Thus, according to this embodiment, first, the user profile of the current interaction node and the interaction process information are determined, and then the model is used to predict the conversation strategy adopted for the current interaction node according to the user profile and the interaction process information. Therefore, compared with the prior art, in the process of determining the conversation strategy, this solution can combine real-time user profile information and interaction process information (question and answer information). As a result, the final determined target conversation strategy can be closer to the user's needs, more accurate, and can provide a better experience for the user. Furthermore, it solves the technical problem in the prior art that the method for determining the conversation strategy in the interaction process lacks reference to all the interaction content generated from the beginning to the current of the interaction, thus affecting the accuracy of the conversation strategy.

[0100] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0101] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0102] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0103] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0105] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0106] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for determining a conversation strategy, characterized in that, it includes: During the interaction with the user, determining the user profile of the user at the current interaction node; Obtaining the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation strategies adopted by at least one interaction node during the interaction, the response information of the user to the node conversation strategies, and the sequential relationship between the node conversation strategies and the response information; and Using a pre-trained model to calculate the user profile and the interaction process information to determine the target conversation strategy for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation strategy conversion results; Using a pre-trained model to calculate the user profile and the interaction process information to determine the target conversation strategy for interacting with the user at the current interaction node includes: Determining the first feature information corresponding to the user profile, and determining the second feature information corresponding to the interaction process information; and Using a pre-trained model to calculate the first feature information and the second feature information to determine the target conversation strategy for interacting with the user at the current interaction node; The response information corresponds to the intention information of the user. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: using a preset identifier to represent the node conversation strategies and the intention information in the interaction process information, determining the interaction information sequence corresponding to the interaction process information, and Determining the second feature information corresponding to the interaction process information includes: determining the second feature information corresponding to the interaction information sequence.

2. The method according to claim 1, characterized in that, The response information corresponds to the response text of the user. After obtaining the interaction process information generated from the start of the interaction to the current interaction node, it further includes: Using a preset identifier to represent the node conversation strategies in the interaction process information to generate a sentence vector corresponding to the response text; and Determining the interaction information sequence corresponding to the interaction process information according to the sentence vector and the node conversation strategies represented by the identifier, and Determining the second feature information corresponding to the interaction process information includes: determining the second feature information corresponding to the interaction information sequence.

3. The method according to claim 1, characterized in that, Using a pre-trained model to calculate the first feature information and the second feature information to determine the target conversation strategy for interacting with the user at the current interaction node includes: Using a pre-trained model to calculate the first feature information and the second feature information to determine the probability values corresponding to multiple conversation strategies at the current interaction node; and According to the probability values, determining the target conversation strategy for interacting with the user at the current interaction node from the multiple conversation strategies.

4. The method according to claim 1, characterized in that, It further includes training the model according to the following operations: Obtain the call records of the user and the user profile of the user; Determine the call record feature information corresponding to the call records and the user feature information corresponding to the user profile; And Train the model according to the call record feature information, the user feature information, and the conversion result of the call records.

5. A storage medium, characterized in that the storage medium includes a stored program, wherein, when the program runs, the method according to any one of claims 1 to 4 is executed by a processor.

6. A device for determining a conversation script, characterized in that it includes: A profile determination module, configured to determine the user profile of the user at the current interaction node during the interaction with the user; A data acquisition module, configured to acquire the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation scripts adopted by at least one interaction node during the interaction, the response information of the user to the node conversation scripts, and the sequence relationship between the node conversation scripts and the response information; And A conversation script determination module, configured to calculate the user profile and the interaction process information by using a pre-trained model, and determine the target conversation script for interacting with the user at the current interaction node, where the model is trained based on the user profile, the interaction process information, and the corresponding conversation script conversion result; Calculating the user profile and the interaction process information by using a pre-trained model to determine the target conversation script for interacting with the user at the current interaction node includes: A feature determination sub-module, configured to determine the first feature information corresponding to the user profile and determine the second feature information corresponding to the interaction process information; and A conversation script determination sub-module, configured to calculate the first feature information and the second feature information by using a pre-trained model, and determine the target conversation script for interacting with the user at the current interaction node; After the response information corresponds to the intention information of the user and the interaction process information generated from the start of the interaction to the current interaction node is acquired, it further includes: representing the node conversation scripts and the intention information in the interaction process information by using a preset identifier, determining the interaction information sequence corresponding to the interaction process information, and Determining the second feature information corresponding to the interaction process information includes: determining the second feature information corresponding to the interaction information sequence.

7. A device for determining a conversation script, characterized in that it includes: A processor; And A memory, connected to the processor, for providing instructions for the processor to perform the following processing steps: During the interaction with the user, determine the user profile of the user at the current interaction node; Acquire the interaction process information generated from the start of the interaction to the current interaction node, where the interaction process information is used to describe the node conversation scripts adopted by at least one interaction node during the interaction, the response information of the user to the node conversation scripts, and the sequence relationship between the node conversation scripts and the response information; And Calculating the user portrait and the interaction process information using a pre-trained model to determine a target conversation statement for interacting with the user at the current interaction node, where the model is trained based on the user portrait, the interaction process information, and the corresponding conversation statement conversion results; Calculating the user portrait and the interaction process information using a pre-trained model to determine a target conversation statement for interacting with the user at the current interaction node, including: Determining first feature information corresponding to the user portrait and determining second feature information corresponding to the interaction process information; and Calculating the first feature information and the second feature information using a pre-trained model to determine a target conversation statement for interacting with the user at the current interaction node; After obtaining the response information corresponding to the user's intent information and the interaction process information generated from the start of the interaction to the current interaction node, it further includes: representing the node conversation statement and the intent information in the interaction process information using a preset identifier, determining an interaction information sequence corresponding to the interaction process information, and Determining second feature information corresponding to the interaction process information includes: determining the second feature information corresponding to the interaction information sequence.

Citation Information

Patent Citations

  • Dialogue automatic reply system based on deep learning and reinforcement learning

    CN106448670A

  • Multi-round dialogue intelligent voice interaction system and device

    CN110209791A