Talk skill recommendation method and device and related equipment

By using the speech search model and user patience analysis in the customer service system, the most appropriate customer service speech is recommended, which solves the problem that the formulation of speech in the existing technology depends on manual experience and subjectivity, and achieves more efficient and personalized customer service speech recommendations.

CN120067309APending Publication Date: 2025-05-30CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510142589.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The formulation of existing customer service speeches relies on manual experience, making it difficult to fully cover business scenarios, and is highly subjective, resulting in inconsistent answers and affecting user satisfaction.

Method used

By obtaining the current session text and interaction information of the target user, input it into the speech search model, output a matching target speech set, and select the most suitable speech from it based on the user's patience to recommend it.

Benefits of technology

It improves the objectivity and comprehensiveness of speech recommendations, enhances flexibility and personalization, ensures that recommended speeches can meet users' needs, avoid user dissatisfaction, and significantly improves the quality and user experience of speech recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067309A_ABST
    Figure CN120067309A_ABST
Patent Text Reader

Abstract

The invention provides a verbal skill recommendation method and device and related equipment, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining a current session text input by a target user in a current round of dialogue and interaction information related to the current session text; the current session text is input into a verbal skill retrieval model, a target verbal skill set matched with the current session text is output, the target verbal skill set comprises multiple verbal skill of different contents, the verbal skill retrieval model is obtained by training an initial retrieval model based on a target verbal skill library, and verbal skill used for recommendation is stored in the target verbal skill library; determining the user tolerance of the target user according to the interaction information related to the current session text; and according to the determined user tolerance, determining a target verbal skill matched with the user tolerance from the target verbal skill set, and recommending the target verbal skill to the target object. According to the invention, objectivity and comprehensiveness of verbal skill recommendation are improved, and flexibility and personalization of verbal skill recommendation are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the continuous progress of Internet technology, the status of customer service in technology companies has become increasingly prominent. At present, customer service is mainly provided by human customer service and robot customer service based on artificial intelligence technology, and the conversation script plays a crucial role in customer service. However, at present, the formulation of customer service conversation scripts mainly relies on manual experience to summarize and refine a relatively fixed conversation script template.

[0003] However, the current customer service conversation scripts are generally formulated manually based on past experience. This method makes it difficult for the conversation scripts to cover a vast number of business scenarios, and their applicability is limited. At the same time, due to strong subjectivity, they can often only provide standardized answers and cannot be adjusted appropriately according to the user's patience and demand differences, which easily leads to user dissatisfaction and thus affects the accuracy of conversation script recommendations and user satisfaction.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present disclosure provides a conversation script recommendation method, apparatus and related equipment, which improve the objectivity and comprehensiveness of conversation script recommendations, and at the same time enhance the flexibility and personalization of conversation script recommendations.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0007] According to one aspect of the present disclosure, there is provided a conversation script recommendation method, the method comprising: obtaining the current conversation text input by a target user in the current round of conversation and interaction information related to the current conversation text; inputting the current conversation text into a conversation script retrieval model to output a target conversation script set that matches the current conversation text, wherein the target conversation script set includes multiple conversation scripts with different contents, and the conversation script retrieval model is obtained by training an initial retrieval model based on a target conversation script library, and the target conversation script library stores conversation scripts for recommendation; determining the user patience of the target user according to the interaction information related to the current conversation text; and determining a target conversation script that matches the user patience from the target conversation script set, and recommending the target conversation script to a target object.

[0008] In some embodiments, the interaction information related to the current session text includes the time information of the current session text, the previous session text of the current session text in the current conversation, and the time information of the previous session text; determining the user patience of the target user according to the interaction information related to the current session text includes: determining whether there is a previous interaction for the current session text in the current conversation; when there is no previous interaction, determining the user patience of the target user according to the time information of the current session text; when there is a previous interaction, determining the user patience of the target user according to the time information of the current session text, the previous session text of the current session text in the current conversation, and the time information of the previous session text.

[0009] In some embodiments, before determining the user patience of the target user according to the time information of the current session text, the method further includes: presetting the correspondence between time periods and user patience; determining the user patience of the target user according to the time information of the current session text includes: obtaining the preset correspondence between time periods and user patience according to the time information of the current session text; determining the user patience of the target user according to the preset correspondence between time periods and user patience.

[0010] In some embodiments, determining the user patience of the target user according to the time information of the current session text, the previous session text of the current session text in the current conversation, and the time information of the previous session text includes: determining the time influence value of the target user according to the time information of the current session text and the time information of the previous session text; determining the content influence value of the target user according to the current session text and the previous session text; determining the user patience of the target user according to the time influence value of the target user and the content influence value of the target user.

[0011] In some embodiments, when there are multiple previous session texts of the current session text in the current conversation, determining the time influence value of the target user according to the time information of the current session text and the time information of the previous session text includes: determining the total duration of the previous session and the average time difference of the target user's replies according to the time information of the current session text and the time information of the previous session text, where the average time difference of the target user's replies is calculated according to the time information of two adjacent session texts input by the target user in the current conversation; determining the average interval time between each session text according to the total duration of the previous session and the total number of previous session texts; determining the time influence value of the target user according to the average time difference of the target user's replies and the average interval time between each session text.

[0012] In some embodiments, determining the content influence value of the target user according to the current session text and the previous session text includes: clustering the current session text and the previous session text based on semantics to obtain a clustered session text cluster; determining a first sub-session text cluster in the clustered session text cluster where the number of session texts meets a preset index, and a second sub-session text cluster grouped with the current session text; and determining the content influence value of the target user according to the first sub-session text cluster and the second sub-session text cluster.

[0013] In some embodiments, determining the content influence value of the target user according to the first sub-session text cluster and the second sub-session text cluster includes: determining the text overlap relevance according to the session texts in the first sub-session text cluster and the session texts in the second sub-session text cluster; determining the session intersection relevance according to the number of session texts in the first sub-session text cluster and the number of texts in the second sub-session text cluster; and determining the content influence value of the target user according to the text overlap relevance and the session intersection relevance.

[0014] In some embodiments, the conversation strategy retrieval model is trained as follows: obtaining sample conversation texts with labels, where the labels annotate the standard conversation strategies of the sample conversation texts; inputting the sample conversation texts with labels into an initial retrieval model to retrieve sample conversation strategies matching the sample conversation texts from a target conversation strategy library;

[0015] evaluating the sample conversation strategies using the standard conversation strategies of the sample conversation texts to obtain an evaluation result; adjusting the initial retrieval model according to the evaluation result, and re-performing retrieval based on the sample conversation texts; until the evaluation result meets a preset condition, stopping the training of the initial retrieval model to obtain a trained conversation strategy retrieval model.

[0016] In some embodiments, when the target conversation strategy library is a document set containing sample conversation strategies, inputting the sample conversation texts with labels into the initial retrieval model to retrieve sample conversation strategies matching the sample conversation texts from the target conversation strategy library includes: determining a set of sample keywords based on the sample conversation texts; determining a target directory item based on the first relevance between the set of sample keywords and the target item corresponding to each document in the document set; determining a target abstract based on the second relevance between the set of sample keywords and the abstract corresponding to the target directory item; and determining a sample conversation strategy matching the sample conversation texts based on the third relevance between the set of sample keywords and the conversation semantics corresponding to the target abstract.

[0017] In some embodiments, evaluating the sample speech script using the standard speech script of the sample conversation text to obtain an evaluation result includes: determining a set of standard keywords based on the standard speech script; determining the standard degree of the sample speech script according to the sample keyword set, the standard keyword set, the sample speech script, and the standard speech script; determining the abnormality degree of the sample speech script according to the fourth correlation degree between the sample keyword set and the target summary; and determining the accuracy of the sample speech script according to the standard degree and the abnormality degree of the sample speech script.

[0018] In some embodiments, determining a target speech script that matches the user patience from the target speech script set according to the determined user patience includes: cutting the speech scripts in the target speech script set according to the determined user patience to obtain an answer speech script for the current conversation text; and determining a target speech script that matches the user patience according to the determined user patience and the answer speech script of the current conversation text.

[0019] According to another aspect of the present disclosure, there is also provided a speech script recommendation device, including: a first acquisition module, configured to acquire a current conversation text input by a target user in the current round of conversation and interaction information related to the current conversation text; a retrieval module, configured to input the current conversation text into a speech script retrieval model and output a target speech script set that matches the current conversation text, where the target speech script set includes multiple speech scripts with different contents, and the speech script retrieval model is obtained by training an initial retrieval model based on a target speech script library, and the target speech script library stores speech scripts for recommendation; a first determination module, configured to determine the user patience of the target user according to the interaction information related to the current conversation text; and a second determination module, configured to determine a target speech script that matches the user patience from the target speech script set according to the determined user patience, and recommend the target speech script to a target object.

[0020] According to another aspect of the present disclosure, there is also provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein the processor is configured to execute the speech script recommendation method according to any one of the above through executing the executable instructions.

[0021] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the speech script recommendation method according to any one of the above.

[0022] According to another aspect of the present disclosure, there is also provided a computer program product, including: a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the speech script recommendation method according to any one of the above.

[0023] A method, device, and related equipment for recommended conversation scripts provided in an embodiment of the present disclosure. The method includes: obtaining the current conversation text input by a target user in the current round of conversation and interaction information related to the current conversation text; inputting the current conversation text into a conversation script retrieval model to output a set of target conversation scripts that match the current conversation text, where the set of target conversation scripts includes multiple conversation scripts with different contents, and the conversation script retrieval model is obtained by training an initial retrieval model based on a target conversation script library, and the target conversation script library stores conversation scripts for recommendation; determining the user patience of the target user according to the interaction information related to the current conversation text; determining, according to the determined user patience, a target conversation script that matches the user patience from the set of target conversation scripts, and recommending the target conversation script to a target object. By constructing a conversation script retrieval model trained based on a target conversation script library, it is possible to quickly output a set of conversation scripts (target conversation script set) that contains multiple different contents and matches the current conversation text input by the target user in the current round of conversation. This effectively solves the problems in the prior art that the formulation of conversation scripts depends on manual experience, it is difficult to comprehensively cover business scenarios, and the subjectivity is strong, resulting in inconsistent answers.

[0024] Furthermore, the method also accurately determines the user patience of the target user by analyzing the interaction information related to the current conversation text, and accordingly selects the conversation script that most matches the user patience from the set of target conversation scripts for recommendation. This improves the objectivity and comprehensiveness of the conversation script recommendation, enhances the personalization and flexibility of the conversation script recommendation, ensures that the recommended conversation script can meet the user's needs for the effectiveness of the answer, avoids the situation where users are dissatisfied due to repeated recommendation of the same standard answer, and thus significantly improves the quality of the conversation script recommendation and the user experience.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0027] Figure 1 A schematic diagram of the system architecture of a method for recommended conversation scripts in an embodiment of the present disclosure is shown;

[0028] Figure 2 A flowchart of a method for recommended conversation scripts in an embodiment of the present disclosure is shown;

[0029] Figure 3Flowchart of a method for determining the user patience of a target user in an embodiment of the present disclosure;

[0030] Figure 4 Flowchart of a method for determining the user patience of a target user in an embodiment of the present disclosure;

[0031] Figure 5 Flowchart of a method for determining the user patience of a target user in an embodiment of the present disclosure;

[0032] Figure 6 Flowchart of a method for determining the content influence value of a target user in an embodiment of the present disclosure;

[0033] Figure 7 Flowchart of a method for training a conversation strategy retrieval model in an embodiment of the present disclosure;

[0034] Figure 8 Flowchart of a method for retrieving a sample conversation strategy that matches a sample conversation text in an embodiment of the present disclosure;

[0035] Figure 9 Flowchart of a method for obtaining an evaluation result in an embodiment of the present disclosure;

[0036] Figure 10 Flowchart of a method for determining a target conversation strategy that matches the user patience in an embodiment of the present disclosure;

[0037] Figure 11 Flowchart of a specific implementation of a conversation strategy recommendation method in an embodiment of the present disclosure;

[0038] Figure 12 Schematic diagram of a conversation strategy recommendation device in an embodiment of the present disclosure;

[0039] Figure 13 Block diagram of the structure of an electronic device in an embodiment of the present disclosure. Detailed implementation

[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0041] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0042] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:

[0043] The Retrieval-Augmented Generation (RAG) model is a hybrid model architecture that combines the advantages of information retrieval and text generation. It aims to overcome the limitations of a single method by combining a retrieval component and a generation component to provide more accurate and contextually relevant responses. First, RAG uses a retrieval model to find the most relevant information fragments from a large-scale document or knowledge source for the input query. This retrieval step helps the model focus on the most relevant context, thereby improving the relevance and accuracy of the generated content. Next, the retrieved relevant information is passed to a generation model, such as a sequence-to-sequence (seq2seq) model under the Transformer architecture. The generation model generates the final answer based on this information. Since the generation stage has targeted context support, it can generate more precise and coherent answers.

[0044] The following will describe in detail the specific implementation manners of the embodiments of the present disclosure with reference to the accompanying drawings.

[0045] Figure 1 The following shows a schematic diagram of an exemplary application system architecture to which the conversation strategy recommendation method in the embodiments of the present disclosure can be applied. As Figure 1 shown, the system architecture may include a terminal device 101, a network 102, and a server 103.

[0046] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103, which can be a wired network or a wireless network.

[0047] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hypertext Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0048] The terminal device 101 can be various electronic devices, including but not limited to smartphones, tablets, laptop portable computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.

[0049] Optionally, the clients of the application programs installed in different terminal devices 101 are the same, or the clients of the same type of application programs based on different operating systems. Depending on the different terminal platforms, the specific form of the client of the application program can also be different. For example, the client of the application program can be a mobile client, a PC client, etc.

[0050] The server 103 can be a server that provides various services, such as a background management server that supports the operations performed by the user using the terminal device 101. The background management server can analyze and process data such as received requests, and feedback the processing results to the terminal device.

[0051] Optionally, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0052] Those skilled in the art can know that Figure 1 the numbers of the terminal devices, networks, and servers in

[0053] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. The embodiments of the present disclosure do not limit this.

[0054] In some embodiments, the conversation strategy recommendation method provided in the embodiments of the present disclosure can be executed by the terminal device of the above system architecture; in other embodiments, the conversation strategy recommendation method provided in the embodiments of the present disclosure can be executed by the server in the above system architecture; in other embodiments, the conversation strategy recommendation method provided in the embodiments of the present disclosure can be implemented by the terminal device and the server in the above system architecture through interaction.

[0055] Figure 2 shows a flowchart of a conversation strategy recommendation method in the embodiments of the present disclosure. As Figure 2 shown, the conversation strategy recommendation method provided in the embodiments of the present disclosure includes the following steps:

[0056] S202, obtain the current conversation text input by the target user in the current round of conversation and interaction information related to the current conversation text.

[0057] In this embodiment, the target user refers to the user who is having a customer service conversation, that is, the object that needs to be provided with customer service. The current conversation text is the text content input by the target user in this round during the customer service conversation, which can be a question, a request, or any other form of information exchange. The interaction information is additional information related to the current conversation text, which may include but is not limited to previous historical conversation records, real-time behavior data, etc. Among them, the historical conversation record is to view the interaction records of the user with the customer service system before to understand their past needs, preferences, and any unresolved problems. The real-time behavior data is to monitor the behavior of the user during the customer service interaction, such as the waiting time, the speed of sending messages, etc., to evaluate the user's patience and other emotional states.

[0058] In some embodiments, in a modern customer service system, the content input by the target user in the current conversation not only involves text information, but also includes voice and video content. Therefore, if the question raised by the target user is in text form (such as through an online chat window or a social media platform), the system can directly obtain this text as the current conversation text. For the case of voice input (such as through a telephone customer service or a voice assistant), speech recognition technology (Automatic Speech Recognition, ASR) is required to convert the voice signal into text format as the current conversation text. When the target user raises a question in video form, first, the system needs to use image recognition technology to identify the video content and separate the audio from the video, and then convert the extracted voice into text using the same method as for voice input. In addition, if the video contains text information (such as subtitles, bullet screens, etc.), these texts can also be extracted as supplementary information. Finally, the system combines the recognized video content and the converted voice text to form the current conversation text.

[0059] S204, input the current conversation text into the conversation strategy retrieval model, and output a set of target conversation strategies that match the current conversation text, where the set of target conversation strategies includes multiple conversation strategies with different contents, and the conversation strategy retrieval model is obtained by training an initial retrieval model based on a target conversation strategy library, and the target conversation strategy library stores the conversation strategies used for recommendation.

[0060] In this embodiment, the conversation strategy retrieval model is a model trained based on machine learning or deep learning, and its function is to retrieve the set of conversation strategies that best match the current conversation text from the conversation strategy library. This model is trained by analyzing a large amount of historical conversation data so that it can accurately identify and match suitable answers. The set of target conversation strategies is a set of conversation strategies with different contents output by the conversation strategy retrieval model according to the current conversation text, and they are considered as the set of answers most likely to meet the user's needs. The target conversation strategy library stores the data set of conversation strategies used for recommendation to users. These conversation strategies are carefully selected or compiled to cover as many service scenarios as possible and improve the efficiency and quality of problem-solving.

[0061] Specifically, the conversation strategy retrieval model can be trained based on the RAG (Retrieval-augmented Generation) model and a document collection. Among them, the RAG model is an artificial intelligence technology that combines retrieval and generation capabilities. By introducing an external knowledge base retrieval mechanism, it overcomes problems such as the limited storage capacity of large language models, the difficulty in obtaining the latest information in real time, and the lack of knowledge in specific fields, thus enabling the generation of more accurate, detailed, and targeted answers. The document collection is a repository that is maintained in real time and includes various answer information, that is, the target conversation strategy library. This library can contain rich materials such as historical conversation records, frequently asked questions and answers, product manuals, service guides, etc. Since the quality of the document collection directly affects the effect of the retrieval model, it needs to be continuously updated and optimized to ensure that its content is up-to-date, comprehensive, and covers as many business scenarios as possible.

[0062] S206. Determine the user patience level of the target user according to the interaction information related to the current conversation text.

[0063] In this embodiment, the user patience level is used to measure the tolerance of the user during the process of waiting for the customer service response or other services, reflecting the objective patience level of the user when asking this question. If the user patience level is poor, the expected answer will be more accurate. The user patience level is a value between 0 and 1. The larger the value, the better the user patience level, that is, the more patient the user is. The smaller the value, the worse the user patience level. The user patience level can be inferred based on various factors, including but not limited to the user's historical behavior (such as the waiting time and communication frequency in previous conversations), the interaction pattern in the current conversation (for example, continuously sending multiple messages in a short period of time may indicate an increase in the user's urgency), and the language style used (such as language expressing impatience or eagerness), that is, the determination of the user patience level is related to whether there is a previous interaction of the user. By analyzing this interaction information, the system can dynamically evaluate the user patience level of each user, so as to provide services that better meet their expectations. Interaction information refers to all additional information related to the current conversation, which is not limited to the text content input by the user currently, but also includes the historical records and behavior data of the user during the entire customer service interaction process. For example, whether the user has tried to contact the customer service multiple times to solve the same problem before, the time interval between each message sent by the user, and whether the user uses an urgent or dissatisfied tone, etc.

[0064] For example, to more clearly understand step S206, the following uses a specific example to illustrate how to determine the user patience level (Patience Level, PL) of the user according to the user's interaction information, and how this process affects the subsequent conversation strategy recommendation.

[0065] Example scenario: Suppose a user asks a question in the customer service system: "Why can't I access the Internet?"

[0066] In S204, it obtains multiple standard answers, such as "device setting problem, fee problem, etc." In the related technology, all the answers will be generated into a script to recommend to the user, such as "You can try the following methods to solve the problem: 1. Confirm whether the Internet function is turned on. 2. Confirm whether the phone card is in arrears. ... etc." However, due to the different user patience PL, the same user will have different reactions to the above existing answers. If the user patience PL is large, it means that the user is very patient and will continue to ask patiently after getting the above answer. If the user patience PL is small, it means that the user is impatient and will complain after getting the above answer, believing that it is an invalid reply, resulting in a decrease in user experience.

[0067] Therefore, this embodiment will be based on the collected interaction information, specifically including historical conversation records: check the user's past interactions with the customer service system. For example, it is found that the user has asked similar questions three times in the past week, and each time has not received a satisfactory solution. Current conversation behavior: analyze the user's behavior pattern in this conversation. For example, after sending the first message, the user repeatedly sends the same or similar questions every 30 seconds, showing an attitude of eagerly seeking help. Language style and tone: Note that the language used by the user contains some words or sentences that express impatience, such as "Please solve it quickly" and "I have been waiting for a long time." Based on the above collected information, it is determined that the target user has low user patience (small PL). Frequent message sending, multiple attempts to contact customer service in a short period of time, and the use of eager language all indicate that the user hopes to get effective help quickly, rather than receiving a long list of possible solutions.

[0068] Specifically, low user patience (PL≤5) may include the following situations:

[0069] Example of interactive information: The user sends multiple similar messages in rapid succession and clearly states that the problem needs to be solved urgently; System judgment: Due to these behavioral characteristics, the system recognizes that the user's patience is low; Recommended adjustment of the wording: In this case, the system will not directly provide detailed troubleshooting steps (such as "You can try the following methods to solve the problem: 1. Confirm whether the Internet function is turned on. 2. Confirm whether the phone card is in arrears..."), but will choose to provide the simplest and quickest inspection suggestion: "Please pull down the phone interface to check whether the network switch is turned on." This concise and clear answer can quickly respond to user needs and reduce user dissatisfaction.

[0070] High user patience (PL>5) can include the following situations:

[0071] Example of interaction information: The user asks questions in a relatively calm manner, without showing a sense of urgency, and is even willing to provide more background information to help the customer service better understand the problem. System judgment: Based on this interaction pattern, the system believes that the user has a relatively high level of patience. Adjustment of recommended conversation tactics: At this time, the system can choose to provide a more comprehensive answer: "You can try the following methods to solve this problem: 1. Confirm whether the Internet access function is turned on. 2. Confirm whether the phone card is overdue. 3. Try restarting the device...". Such detailed guidance can help the user completely solve the problem and at the same time meet their need for detailed information.

[0072] In this embodiment, dynamically adjusting the recommended conversation tactics strategy according to the user's patience can not only improve the efficiency of problem-solving, but also effectively enhance the user experience and satisfaction. This method ensures that no matter what state the user is in, they can obtain the support and services most suitable for their current needs.

[0073] S208. According to the determined user patience, determine the target conversation tactics that match the user patience from the target conversation tactics set, and recommend the target conversation tactics to the target object.

[0074] In this embodiment, the target conversation tactics set is a set of potential answer sets retrieved by the conversation tactics retrieval model according to the current conversation text. Each conversation tactic is generated based on the document information in the target conversation tactics library maintained within the system, aiming to solve the user's specific problems or respond to their queries. The target conversation tactics that match the user patience are based on the user patience evaluation results. The system will select one or several of the most suitable conversation tactics for the current user state as the final reply. If the system detects that the user's patience is relatively low, it may choose a direct and concise answer to quickly solve the problem; on the contrary, if the user's patience is relatively high, it can choose a more detailed and considerate explanation, providing more background information and support. The target object refers to the entity that actually interacts with the user, such as a human customer service or a robot customer service, and can also be the target user himself / herself. The system recommends the best conversation tactics selected according to the user patience to this "target object" so that they can effectively respond to the user's needs. For the robot customer service, recommending the target conversation tactics to the target object means automatically selecting and sending the most suitable conversation tactics; for the human customer service, recommending the target conversation tactics to the target object means providing a suggested reply for their reference or direct use. In some embodiments, the recommended method can be directly displayed in the chat window or read out through voice synthesis technology, depending on the specific implementation of the customer service system and the user's preference settings.

[0075] In this embodiment, by constructing a conversation strategy retrieval model trained based on a target conversation strategy library, it is possible to quickly output a set of conversation strategies (target conversation strategy set) containing multiple different contents that match the current conversation text input by the target user in the current round of conversation. This effectively solves the problems in the prior art that conversation strategy formulation depends on manual experience, it is difficult to comprehensively cover business scenarios, and subjectivity leads to inconsistent answers. Further, this method also accurately determines the user patience of the target user by analyzing the interaction information related to the current conversation text, and accordingly selects the conversation strategy that best matches the user patience from the target conversation strategy set for recommendation. This improves the objectivity and comprehensiveness of conversation strategy recommendation, enhances the personalization and flexibility of conversation strategy recommendation, ensures that the recommended conversation strategies can meet the user's needs for answer effectiveness, avoids the situation of causing user dissatisfaction due to repeatedly recommending the same standard answer, and thus significantly improves the quality of conversation strategy recommendation and the user experience.

[0076] Figure 3 The flowchart of a method for determining the user patience of a target user in an embodiment of the present disclosure is shown, as Figure 3 As shown, the method for determining the user patience of a target user according to the interaction information related to the current conversation text provided in the embodiment of the present disclosure includes the following steps:

[0077] S302, determine whether there is a previous interaction in the current conversation text in the current round of conversation.

[0078] In this embodiment, the previous interaction refers to whether there has been any interaction within the same conversation cycle before the user presents the current conversation text. Such interaction can be an inquiry, a greeting (such as "Hello"), providing information, or verification, etc. Any interaction before the first input information within a question cycle is regarded as a previous interaction. The question cycle is a time period calculated from the establishment of the customer service channel. For example, when the user dials the customer service phone, initiates a request through the online chat window, or starts other forms of customer service communication methods, it marks the start of a new question cycle. Within this cycle, all interactions with the customer service system are regarded as related.

[0079] The interaction information related to the current conversation text includes the time information of the current conversation text, the previous conversation text of the current conversation text in the current round of conversation, and the time information of the previous conversation text. Among them, the time information of the current conversation text records, for example, the specific timestamp when the user sends each message, which helps to analyze the user's waiting time and response speed. The previous conversation text and the time information of the previous conversation text include all the interaction contents and their occurrence times of the user before this in the current question cycle.

[0080] In some embodiments, certain customer service representatives will conduct login guidance at the beginning of a call. For example, bank customer service representatives will first guide the user to enter their ID number for verification after the call is connected. After verification is successful, they will then connect to a human customer service representative or an AI customer service representative. At this time, the start of the authentication will be counted as a question cycle, and the verification information at this time will also be retained by the customer service system for subsequent processing of the answers to obtain the target conversation script.

[0081] It should be noted that although the user's identity information is retained here, this information will not be provided to other personnel such as human customer service representatives. If the existing method is used instead of the method of this proposal, the system will also obtain the user's identity information. Therefore, this embodiment can avoid adding the process of obtaining and storing identity information and does not involve the problem of user data leakage.

[0082] S304. When there is no previous interaction, determine the user patience level of the target user according to the time information of the current conversation text.

[0083] In this embodiment, if no previous interaction is found during the current conversation cycle, it means that this is the target user's first input in this cycle. In this case, the system mainly relies on the time information of the current conversation text to evaluate the user's patience level.

[0084] In this embodiment, the patience level of the user can be inferred by analyzing the specific time point when the target user sends a message (for example, a specific time period of the day: morning, afternoon, or evening), as well as other factors that may affect the user's patience.

[0085] Specifically, the influence of time information on user patience is as follows:

[0086] Time period of the day: In the morning (such as from 6 am to 12 noon), users may be in a state of starting a new day, usually with a relatively positive and peaceful mindset. If they receive a slow response from the customer service during this time period, they may have a certain degree of tolerance because their schedule is relatively loose; in the afternoon (such as from 12 noon to 6 pm), which is a time when many people are working or studying, users may be busier. If the problem is not resolved quickly, they may feel more impatient and have a lower user patience level; in the evening (such as from 6 pm to late at night), after work or school, people may have experienced a day of fatigue and hope to solve the problem quickly in order to rest or relax. Therefore, users in this time period tend to expect a faster response.

[0087] Differences between weekdays and weekends / holidays: On weekdays, especially during working hours, users may be dealing with work tasks or other important matters, so they have higher requirements for the response speed of customer service and lower patience. On weekends or holidays, users are usually more relaxed and may be more accepting even if they have to wait a bit longer, with relatively higher patience.

[0088] Special periods: Such as non-peak periods near midnight or early morning. Although the number of users contacting customer service is small during these periods, users who choose to contact customer service at these times may be in urgent situations, so their patience may be extremely low and they need to get a response as soon as possible.

[0089] S306. When there is a previous interaction, determine the user patience of the target user according to the time information of the current session text, the previous session text of the current session text in this round of conversation, and the time information of the previous session text.

[0090] In this embodiment, when there is a previous interaction, the system not only considers the time information of the current session text, but also combines the content and time information of the previous session text for comprehensive analysis. It should be noted that the previous session text can be a question or non-question.

[0091] For example, if the target user just completed identity verification a few minutes ago and then immediately asked a question, this may indicate that the target user hopes to get help as soon as possible and has low patience. On the other hand, if the previous interaction shows that the target user just had a long friendly conversation or provided detailed background information before asking a specific question, then this may indicate that the user is more relaxed and willing to spend more time solving the problem, with higher patience.

[0092] Figure 4 The flowchart showing a method for determining the user patience of the target user in an embodiment of the present disclosure is as Figure 4 shown. The method for determining the user patience of the target user according to the time information of the current session text provided in the embodiment of the present disclosure includes the following steps:

[0093] S402. Preset the correspondence between time periods and user patience.

[0094] In this embodiment, in the relationship between time periods and user patience, the closer the time period is to the user's active time, the greater the user patience.

[0095] For example, the relationship between the preset time periods and user patience is as shown in Table 1:

[0096] Table 1

[0097] Time period User patience 0:00 - 5:00 0.1 5:00 - 9:00 0.5 9:00 - 17:00 0.9 17:00 - 23:00 0.4 23:00 - 24:00 0.1

[0098] In this embodiment, the time period and user patience shown in Table 1 are only examples. During actual application, the time period and user patience can be set flexibly.

[0099] S404. Obtain the correspondence between the preset time period and user patience according to the time information of the current conversation text.

[0100] S406. Determine the user patience of the target user according to the correspondence between the preset time period and user patience.

[0101] In this embodiment, obtain the time AT of the current conversation text 0 , if AT 0 is 20:10:59, then it belongs to the period from 17:00 to 23:00, and the user patience PL is obtained as 0.4.

[0102] In this embodiment, it can not only better understand the specific needs and emotional states of users, but also dynamically adjust the response strategy, thereby improving the overall service quality and enhancing the user experience. This method is particularly suitable for customer service scenarios that require high personalization and fast response.

[0103] Figure 5 The flowchart of a method for determining the user patience of a target user in an embodiment of the present disclosure is shown. As Figure 5 shown, the method for determining the user patience of a target user provided in the embodiment of the present disclosure according to the time information of the current conversation text, the previous conversation text of the current conversation text in this round of conversation, and the time information of the previous conversation text includes the following steps:

[0104] S502. Determine the time influence value of the target user according to the time information of the current conversation text and the time information of the previous conversation text.

[0105] In this embodiment, the time influence value is a measurement index used to evaluate the influence of time factors on user patience. It is based on the time information of the current conversation text (for example, the specific time when the message is sent) and the time information of the previous conversation text (that is, the time when the previous interaction occurred). By analyzing these time data, the system can judge the user's waiting time and their interaction frequency, so as to infer the user's patience level within a specific time period.

[0106] The specific calculation method may include: waiting time: the time interval from when the user first sends a message to when a reply is received. A shorter waiting time usually corresponds to a higher user patience. Interaction frequency: the frequency at which the user repeatedly sends messages in a short period of time. Frequent message sending may indicate a lower user patience of the user. Time period of the day: As mentioned above, the psychological state and expected response speed of users may be different within different time periods.

[0107] In some embodiments, when there are multiple pieces of previous session texts of the current session text in the current round of conversation, determining the time influence value of the target user according to the time information of the current session text and the time information of the previous session texts includes: determining the total duration of the previous session and the average time difference of the target user's responses according to the time information of the current session text and the time information of the previous session texts, where the average time difference of the target user's responses is calculated according to the time information of two adjacent session texts input by the target user in the current round of conversation; determining the average interval time between each session text according to the total duration of the previous session and the total number of previous session texts; and determining the time influence value of the target user according to the average time difference of the target user's responses and the average interval time between each session text.

[0108] Specifically, the time influence value

[0109] where represents the average time difference of the target user's responses, and n 1 is the total number of contents of the previous interaction, T 1 is the total duration of the contents of the previous interaction, is the average interval time between each session text.

[0110] is calculated according to the actual interaction interval time (i.e., time difference) of the target user. It can be understood that in a round of conversation, there is both a single reply session in which the customer service responds to the target user's session, multiple consecutive reply sessions in which the customer service responds to the target user's session, a single reply session in which the target user responds to the customer service's session, and multiple consecutive reply sessions in which the target user responds to the customer service's session. The time difference characterizes the response duration of the user, whether it is a response to the customer service's session or a response to its own session. It should be noted that if an interaction session is a response to the customer service's session, then the previous session is the customer service's session. If an interaction session is also the session input by the user himself / herself in the previous session, then the previous session is the session input by himself / herself. The time influence value TI characterizes the ratio between the average interaction interval time of the target user and the average interaction interval time of all session participants in the current round of conversation. This value is between 0 and 1. The smaller this value is, the shorter the actual response time of the user is, the more anxious the user is, the greater the impact on the user's patience, and the lower the user's patience.

[0111] For a certain interaction session of the target user, take as the standard interval time of the user's interaction, and Δt as the actual interaction interval time of the user, Characterizes the ratio between the actual interaction interval time of the user and the standard. This value is between 0 and 1. The smaller this value, the shorter the actual response time of the user, the more anxious the user is, and the greater the impact on the user's patience, resulting in a lower user patience level.

[0112] S504. Determine the content influence value of the target user based on the current session text and the previous session text.

[0113] In this embodiment, the content influence value is an indicator used to evaluate the impact of the conversation content itself on the user's patience. It takes into account the content characteristics of the current session text (the information newly input by the user) and the previous session text (the previous communication records).

[0114] Specific content analysis may include: Language style and tone: For example, using urgent, dissatisfied, or complaining language may indicate a lower user patience level; while friendly and polite language may imply a higher user patience level. Problem complexity: For simple problems compared to complex problems, users may have a higher user patience level waiting for an answer. Historical interaction background: If users have tried to solve the same problem multiple times in the past but have not succeeded, their user patience level may decrease.

[0115] S506. Determine the user patience level of the target user based on the time influence value and the content influence value of the target user.

[0116] In this embodiment, the user patience level of the target user determined by integrating the time influence value and the content influence value reflects the tolerance level of the user towards waiting and service quality during the entire customer service interaction process. The following is a formula for determining the user patience level of the target user provided by this embodiment:

[0117] User patience level

[0118] where TI represents the time influence value and CI represents the content influence value.

[0119] In some embodiments, there are also multiple possible implementation methods for comprehensive evaluation: Different weights can be set for the time influence value and the content influence value according to the actual situation, and then the weighted average value is calculated as the final user patience level score. As the conversation progresses, new interaction information will be taken into consideration, thereby updating the user's user patience level score in real time to ensure that the recommended conversation techniques always meet the user's immediate needs. Based on the calculated user patience level score, the system can select the most suitable conversation technique for the current situation to improve the efficiency of problem-solving and user satisfaction.

[0120] In this embodiment, not only the sense of urgency in terms of time is considered, but also the emotions and intentions at the content level are deeply analyzed, so as to provide a more accurate and personalized service experience. This method helps to improve the overall customer service quality and enhance the user experience.

[0121] Figure 6 The flowchart of a method for determining the content influence value of a target user in an embodiment of the present disclosure is shown, as Figure 6 shown, the steps for determining the content influence value of a target user according to the current session text and the previous session text provided in the embodiment of the present disclosure are as follows:

[0122] S602, cluster the current session text and the previous session text based on semantics to obtain the clustered session text clusters.

[0123] In this embodiment, semantic clustering is to group similar session texts together by analyzing the content and meaning (i.e., semantics) of the texts. Using natural language processing (NLP) techniques, such as word vector models (e.g., Word2Vec, BERT, etc.) or topic models (such as LDA), the similarity between texts can be calculated and clustering can be performed accordingly. The session text clusters are multiple groups formed after semantic clustering, and each group contains session texts that are semantically similar. These clusters can help identify the main topics or problem types that the user is concerned about.

[0124] Specifically, the semantics of each interaction content and the current session text can be identified through existing semantic recognition methods. Through existing clustering methods, all interaction contents and the contents in the current session text are clustered according to semantics, so that the interaction contents with similar semantics are grouped into one category.

[0125] S604, determine the first sub-session text cluster whose number of session texts in the clustered session text clusters meets a preset index, and the second sub-session text cluster that is clustered into the same category as the current session text.

[0126] In this embodiment, the first sub-session text cluster refers to those clusters that contain a large number of session texts, usually indicating that the problems or topics in these clusters are the key points that the user repeatedly mentions or pays special attention to. The "preset index" here can be a specific quantity threshold for screening out the most representative clusters. The second sub-session text cluster refers to a group of texts that are semantically closest to the current session text. Specifically, it is a cluster determined by comparing the similarity between the current session text and other historical session texts, indicating that the current session text may be closely related to other problems in this cluster.

[0127] S606, determine the content influence value of the target user according to the first sub-session text cluster and the second sub-session text cluster.

[0128] In this embodiment, if the number of interaction contents in a category is larger, it indicates that the semantics of this category is the main semantics of the user's previous interactions. And the larger the number of interactions grouped with the current conversation text into one category, it indicates that the user has been involved in relevant contents multiple times.

[0129] In some embodiments, determining the content influence value of the target user according to the first sub-conversation text cluster and the second sub-conversation text cluster includes: determining the text overlap relevance according to the conversation texts in the first sub-conversation text cluster and the conversation texts in the second sub-conversation text cluster; determining the conversation cross relevance according to the number of conversation texts in the first sub-conversation text cluster and the number of texts in the second sub-conversation text cluster; and determining the content influence value of the target user according to the text overlap relevance and the conversation cross relevance.

[0130] In this embodiment, the text overlap relevance is an index for measuring the similarity between the conversation texts in the first sub-conversation text cluster and the conversation texts in the second sub-conversation text cluster. It reflects the overlapping degree or common points of these texts in semantics. The conversation cross relevance is an index determined based on the number of conversation texts in the first sub-conversation text cluster and the number of texts in the second sub-conversation text cluster, and is used to measure the interaction degree or the correlation strength between these two clusters.

[0131] Specifically, the text overlap relevance can be calculated by comparing the word frequency distributions of all documents in the two clusters, using TF-IDF (Term Frequency-Inverse Document Frequency) or other methods based on the vector space model to calculate the similarity score between texts. It is also possible to use a pre-trained language model (such as BERT) to calculate the similarity between sentences. A high text overlap relevance indicates that the current conversation text is highly consistent with the previously discussed topic, meaning that the user has been trying to solve the same or similar problems, which may reduce the user's patience; while a low relevance may indicate that this is a new and independent problem.

[0132] The conversation cross relevance can consider the size ratio of the two clusters and the number of conversations directly shared between them. For example, if the first sub-conversation text cluster is very large and contains many repeated questions, while the second sub-conversation text cluster is relatively small but has a significant intersection with the former, it may indicate that the user has a very high attention to a specific topic. A higher conversation cross relevance may indicate that the user repeatedly encounters the same type of problems, and in this case, the user's patience is usually lower; on the contrary, if the cross degree is low, it may represent that the user raises a new question, and the user's patience is relatively high.

[0133] Finally, by combining the text overlap relevance and the conversation cross - relevance, the system can obtain a comprehensive score, namely the content impact value. This value reflects the overall impact of the conversation content on the user's patience. Different weights can be set for the text overlap relevance and the conversation cross - relevance according to the actual application scenario, and then the weighted sum is used to obtain the content impact value. In addition, other factors can be introduced for adjustment, such as the language style used by the user, historical interaction records, etc.

[0134] Specifically, determine the content impact value

[0135] sim is the semantic similarity between the class with the most interaction content and the class where the current conversation text is located, n 2 is the number of interactions grouped into the same class as the current conversation text, n 3 is the number of interactions in the class with the most interaction content, characterizes the degree to which the current conversation text is mentioned in the previous interaction. The larger this value, the more it is mentioned in the previous interaction.

[0136] Among them, sim is obtained through existing similarity calculation methods. The larger this value, the more the content of the current conversation text is the main content of the previous interaction. Characterizing the situation of the previous interaction involving the current conversation text from the two dimensions of content and the number of mentions, the larger this value, the more it is involved. And if the user still asks questions, it means that the user's patience is decreasing.

[0137] In this embodiment, according to the calculated content impact value, the customer service system can select the most appropriate response strategy. For a higher content impact value (i.e., the user shows lower patience), the system may give priority to recommending concise and clear answers; for a lower content impact value, it can choose to provide more detailed information and support.

[0138] Figure 7 Show a method flow chart for training a conversation strategy retrieval model in an embodiment of the present disclosure, as Figure 7 shown, the conversation strategy retrieval model provided in the embodiment of the present disclosure is trained as follows:

[0139] S702, obtain sample conversation texts with labels, where the labels annotate the standard conversation strategies of the sample conversation texts.

[0140] In this embodiment, the labeled sample conversation text refers to the customer service conversation record that has been manually annotated or processed by automated tools. Each sample conversation text is attached with a label, which indicates the standard words (i.e., ideal responses) recommended for this text. These samples are used to train and evaluate retrieval models. Standard words refer to answers or solutions that are considered to be best practices in a specific situation. They are usually carefully selected based on historical data, expert opinions, or user feedback, aiming to improve problem solving efficiency and user satisfaction.

[0141] S704, input the sample conversation text with labels into the initial retrieval model, and retrieve the sample speech matching the sample conversation text from the target speech library.

[0142] In this embodiment, the initial retrieval model is a preliminary constructed model used to find the most appropriate answer from the target speech library based on the input conversation text. It may be based on machine learning algorithms or other information retrieval technologies. The target speech library is a data set that stores a large number of predefined speeches, which cover various common customer service scenarios and service requests. The target speech library is the main data source for the retrieval model. Sample speech When labeled sample conversation text is input into the initial retrieval model, the model searches the target speech library and returns a set of candidate speeches that are considered to be most relevant to the input text, called sample speech. For example, sample speech is selected from a document collection based on sample conversation text through the RAG model.

[0143] S706, using the standard speech of the sample conversation text to evaluate the sample speech to obtain an evaluation result.

[0144] In this embodiment, the evaluation result is a score obtained by comparing the similarity or accuracy between the sample words output by the retrieval model and the standard words attached to the sample conversation text. The evaluation process can use a variety of indicators, such as precision, recall, F1 score, etc., to measure the performance of the model. The purpose is to determine how well the retrieval model performs in selecting appropriate words and provide a basis for subsequent adjustments.

[0145] S708, adjusting the initial retrieval model according to the evaluation result, and re-searching based on the sample conversation text.

[0146] In this embodiment, adjusting the initial retrieval model means identifying the deficiencies of the model based on the evaluation results and optimizing it accordingly. This may involve adjusting model parameters, improving feature extraction methods, adding more training data, etc. The adjusted model is tested again using the same sample conversation text to check whether its performance has improved.

[0147] S710. Keep training the initial retrieval model until the evaluation result meets the preset conditions, and then stop the training to obtain a trained conversation retrieval model.

[0148] In this embodiment, the preset conditions refer to a set of pre-set performance indicators or thresholds. Only when the evaluation result of the model reaches or exceeds these criteria is the model considered to be trained. For example, the precision rate reaches more than 90%, or the evaluation result does not improve significantly after several consecutive iterations. Once the model meets the preset conditions, it can be considered that it has been fully trained and can effectively retrieve high-quality conversation texts from the target conversation library for actual customer service scenarios.

[0149] In this embodiment, by using the trained retrieval model to obtain answers from the maintained document collection including various answer information, the objectivity and comprehensiveness of the answers are ensured, and the problems in the prior art, such as different answers to the same question due to different subjectivities of the processors and the reduction of answer accuracy, are avoided.

[0150] In some embodiments, in order to improve the retrieval efficiency of the document collection, text semantic chunking is performed on each document in the document collection, the abstract of the document is obtained based on the semantics of each chunk using the abstract extraction algorithm, and then the table of contents of the entire document collection is extracted based on the abstracts of each document, thereby forming a four-layer structure of table of contents - abstract - semantics - document. Figure 8 The flowchart of a method for retrieving a sample conversation text matching a sample conversation is shown in the embodiments of the present disclosure, as Figure 8 shown. When the target conversation library is a document collection containing sample conversations, the steps of inputting the labeled sample conversation text into the initial retrieval model and retrieving a sample conversation matching the sample conversation text provided in the embodiments of the present disclosure are as follows:

[0151] S802. Determine a sample keyword set based on the sample conversation text.

[0152] In this embodiment, the sample keyword set is a set of keywords or phrases extracted from the sample conversation text, and these keywords can represent the core theme or main content of the conversation. Natural language processing (NLP) techniques, such as word frequency statistics, TF-IDF (term frequency - inverse document frequency), named entity recognition (NER), or more advanced pre-trained models (such as BERT), are usually used to automatically extract keywords. The keyword set is used for matching and association calculations in subsequent steps to help the system better understand the content of the conversation text and find the most relevant documents, table of contents items, abstracts, and ultimately conversation texts.

[0153] In some embodiments, determining the sample keyword set may include the following process: Extract keywords from the sample conversation text, label them as original keywords, and at the same time label the key degree as 1; Expand the original keywords with synonyms, label them as derivative keywords, and at the same time label the key degree as the synonym degree of the word and the keyword, which is calculated when determining whether it is a synonym; Integrate the original keywords and derivative keywords into a keyword set.

[0154] S804. Determine the target directory entry based on the first association degree between the sample keyword set and the target item corresponding to each document in the document set.

[0155] In this embodiment, the first association degree is used to evaluate the similarity or relevance between the sample keyword set and the target item corresponding to each document in the document set. The target item can be a document title, a chapter title, or other forms of classification labels. This association degree can be quantified by calculating methods such as the cosine similarity and Jaccard similarity coefficient between the keyword set and the target item. The target directory entry can be, based on the first association degree score, select one or more target items with the highest scores as the "target directory entry". These directory entries represent the document categories or topics most relevant to the content of the sample conversation text. The relevance RE of the sample keyword set to each directory entry i The calculation formula is as follows:

[0156]

[0157] where i is the identifier of the directory entry, j is the element identifier of the sample keyword set, is the association between element j and directory entry i, w j is the key degree of element j, and n 1 is the total number of elements in the sample keyword set.

[0158] In this embodiment, by narrowing the search scope and focusing on specific documents or parts that are most likely to contain useful information, the retrieval efficiency is improved.

[0159] S806. Determine the target abstract based on the second association degree between the sample keyword set and the abstract corresponding to the target directory entry.

[0160] In this embodiment, the second correlation degree is used to evaluate the similarity or relevance between the sample keyword set and the abstract corresponding to the selected target directory item. An abstract is a highly condensed summary of the core content of a document, aiming to quickly convey the main information. Similarly, methods such as cosine similarity and Jaccard similarity coefficient can be used to calculate the correlation degree between the keyword set and the abstract. The target abstract can be selected as the "target abstract" based on the second correlation degree score, which is the highest score. This abstract is extracted from the document most relevant to the sample conversation text and can provide a concise overview of the core content of the document.

[0161] In this embodiment, by further refining the search results, it is ensured that the selected abstract directly reflects the issues or needs that the user cares about, providing a basis for the final script recommendation.

[0162] S808. Based on the third correlation degree between the sample keyword set and the script semantics corresponding to the target abstract, the sample script that matches the sample conversation text is determined.

[0163] In this embodiment, the third correlation degree is used to evaluate the similarity or relevance between the sample keyword set and the script semantics corresponding to the target abstract. Script semantics refer to the meaning and context background of a specific answer or solution. A pre-trained language model (such as BERT) can be used to calculate the semantic similarity between sentences, or other deep learning methods can be used to capture subtle semantic differences. The sample script can be selected as the "sample script" based on the third correlation degree score, which is the highest score. These scripts are refined from the target abstract and are the best responses specifically designed for the issues in the sample conversation text. In this embodiment, it is ensured that the answer finally recommended to the user not only matches the user's query superficially but also is highly consistent semantically at a deep level, thereby improving the user experience and satisfaction.

[0164] In summary, in this embodiment, based on the four-layer structure of directory-abstract-semantics-document, the number of retrievals can be reduced and the retrieval efficiency can be improved.

[0165] Figure 9 The flowchart of a method for obtaining an evaluation result in an embodiment of the present disclosure is shown. As Figure 9 shown, the steps for evaluating the sample script using the standard script of the sample conversation text to obtain an evaluation result provided in the embodiment of the present disclosure include the following:

[0166] S902. Determine the standard keyword set based on the standard script.

[0167] In this embodiment, the standard keyword set is a set of keywords or phrases extracted from the standard speech, and these keywords can represent the core theme or main content of the standard speech. Similar to S802, natural language processing (NLP) techniques are usually used to automatically extract keywords, such as word frequency statistics, TF-IDF, named entity recognition (NER), or more advanced pre-trained models (such as BERT). The standard keyword set is used for matching and association calculations in subsequent steps, helping the system better understand the content of the standard speech and evaluate the quality of the sample speech.

[0168] S904. Determine the standard degree of the sample speech according to the sample keyword set, the standard keyword set, the sample speech, and the standard speech.

[0169] In this embodiment, the standard degree of the sample speech is used to evaluate the similarity or consistency between the sample speech and the standard speech. It considers the overlapping degree between the sample keyword set and the standard keyword set, as well as the proximity of the sample speech and the standard speech in content. Specifically, through keyword matching: comparing the overlapping situation between the sample keyword set and the standard keyword set, methods such as Jaccard similarity coefficient and cosine similarity can be used. Content matching: By analyzing the actual text content of the sample speech and the standard speech, evaluate their semantic similarity, and the semantic similarity between sentences can be calculated using pre-trained language models (such as BERT). The following is a formula for determining the standard degree of the sample speech provided in this embodiment:

[0170]

[0171] Among them, sim is the similarity between the sample keyword set and the standard keyword set, S is the identifier of the standard speech, j is the element identifier of the sample keyword set, RE s is the relevance between each keyword in the sample keyword set and the standard speech, is the relevance between the sample keyword set and the standard speech, is the relevance between element j and the standard speech S, w j is the key degree of element j, n 1 is the total number of elements in the sample keyword set.

[0172] Through this embodiment, it can be ensured that the sample speech is as close as possible to the standard speech in content, thus ensuring that the answers recommended to users are of high quality and meet expectations.

[0173] S906. Determine the abnormality degree of the sample speech according to the fourth relevance between the sample keyword set and the target summary.

[0174] In this embodiment, the fourth correlation degree is used to evaluate the similarity or relevance between the sample keyword set and the target abstract. The target abstract is a highly condensed summary of the core content of the document, aiming to quickly convey the main information. This correlation degree can be quantified by calculating methods such as the cosine similarity and Jaccard similarity coefficient between the keyword set and the target abstract. The abnormality degree of the sample speech can be determined based on the fourth correlation degree score, which reflects the deviation degree of the sample speech from the target abstract. If there are significant differences between the keywords of the sample speech and those of the target abstract, it indicates that the sample speech may be abnormal and may not fully meet the specific needs or problem background of the user.

[0175] Specifically, the abnormality degree of the sample speech can be determined through the following steps: Determine the relevance RE between each document corresponding to the target abstract and the standard speech according to the standard keyword set k , where k is the target abstract identifier. Determine the RE k for each target abstract k and the relationship between RE k . RE k represents the standard relevance score (or standard threshold) expected to be achieved for the k-th target abstract. If RE k ∈ [RE k + ε], then determine that the outlier corresponding to this target abstract is 0; otherwise, the outlier is 1. Here, ε is a pre-set deviation threshold.

[0176] Determine the abnormality degree of the sample speech

[0177] where n k is the outlier of the k-th target abstract, and K is the total number of target abstracts.

[0178] Through this embodiment, it can be detected whether the sample speech deviates from the main content described in the target abstract, so as to avoid providing irrelevant or incorrect answers.

[0179] S908. Determine the accuracy of the sample speech according to the standard degree and abnormality degree of the sample speech.

[0180] In this embodiment, the accuracy of the sample speech combines the standard degree and abnormality degree of the sample speech, and is used to measure the overall quality of the sample speech. Different weights can be set for the standard degree and abnormality degree according to the actual application scenario, and then the weighted average value is calculated as the final accuracy score. Specific thresholds can also be set. Only when the standard degree is higher than a certain threshold and the abnormality degree is lower than another threshold, the sample speech is considered to have high accuracy.

[0181] Specifically, the accuracy of the sample speech

[0182] Among them, S is the standard degree of the sample conversation; E is the abnormality degree of the sample conversation;

[0183] Through this embodiment, it can be ensured that the answer finally recommended to the user is not only close to the standard conversation in form, but also highly relevant to the user's specific question in content, thereby improving the overall service quality and user experience.

[0184] Through the above steps, the system can comprehensively evaluate the quality of the sample conversation, ensure that the customer service response provided meets the standards and fits the actual needs of the user, thereby improving the service efficiency and user satisfaction. This method emphasizes a multi-dimensional evaluation mechanism, which helps to achieve precise and personalized conversation recommendation.

[0185] Figure 10 The flowchart of a method for determining a target conversation matching the user's patience degree in an embodiment of the present disclosure is shown, as Figure 10 As shown, the steps for determining a target conversation matching the user's patience degree from a set of target conversations according to the determined user's patience degree in the embodiment of the present disclosure may include the following steps:

[0186] S1002, according to the determined user's patience degree, clip the conversations in the set of target conversations to obtain the answer conversation for the current conversation text.

[0187] In this embodiment, the user's patience degree is used to measure the tolerance degree of the user during the waiting for the customer service response or other service processes, and reflects the objective patience degree of the user when asking this question. It is determined by analyzing information such as the user's interaction behavior, language style, and historical conversation records. The preset threshold is a reference value set by the system to determine whether the user's patience degree is high enough to receive a detailed answer. If the user's patience degree is greater than this threshold, it is considered that the user has a high patience; otherwise, it is considered that the user's patience degree is low.

[0188] If the user's patience degree PL is greater than the preset threshold, it means that the user has a greater patience degree when asking questions, and the answer obtained in S202 can be used as the final response answer. If the user's patience degree PL is not greater than the threshold, it means that the user has a lower patience degree when asking questions. At this time, obvious answers that can be excluded can be deleted from the answer obtained in S202 for the user.

[0189] Suppose the user reports "unable to access the Internet" and provides some identification information, such as a phone number or an ID card number. The system searches for relevant information based on this identification information and decides which answer branches need to be retained or clipped. Identification information: For example, the phone number or ID card number provided by the user. These identifications can help the system quickly locate the specific situation of the user, such as the account status, arrears situation, etc.

[0190] In some embodiments, the cutting process can be performed based on a preset response module, such as presetting the confirmation process of various answer branches, and each process is stored as a module separately to form a module library. The module library is a database that stores various answer branches and their confirmation processes. Each answer branch corresponds to an independent module, which includes a series of condition checks and confirmation steps. When cutting, according to the answer branch obtained in S202, the modules corresponding to each answer branch are determined in turn, and the modules are executed to determine whether the answer branch is cut. If the current condition does not meet the execution requirements of the module corresponding to a certain answer branch, then no confirmation is performed, and the answer branch is determined to be a reserved answer branch and is not cut. If it is confirmed that it is invalid according to the execution result of the module corresponding to a certain answer branch, the answer branch is cut. If it is confirmed that it is valid according to the execution result of the module corresponding to a certain answer branch, the answer branch is determined to be a reserved answer branch and is not cut. Finally, all the reserved answer branches are integrated into the answer words of the current conversation text.

[0191] For example, if the user is in arrears, the system can directly confirm that "phone card arrears" is one of the reasons for the problem. At this time, "1. Confirm whether the Internet access function is turned on." This type of possible answer branch can be cut off because it is obviously invalid in this case. Module library execution: For each answer branch, the system will call the corresponding module in turn for confirmation. If the current conditions do not meet the module execution requirements corresponding to a certain answer branch, the answer branch is considered to be a reserved answer branch and will not be cut. If it is confirmed to be valid based on the module execution result corresponding to a certain answer branch, the answer branch is retained and will not be cut. If the confirmation result shows that a certain answer branch is invalid (for example, the user has confirmed that the Internet access function is turned on), the answer branch will be cut off. Integration of reserved answer branches: After the above cutting process, all reserved answer branches will be integrated into the final response answer. For example, if the system confirms that the user's phone card is in arrears and other possible reasons have been ruled out, the final response answer will be "phone card arrears".

[0192] A specific example is as follows: Assume that the user reports that he cannot access the Internet and provides a phone number as an identifier. The system finds out through query that the user's phone card is in arrears, and the user has completed the login guide (entered the ID number) at the beginning of the question cycle. Based on this information, the system will perform the following steps: Initial answer set: It may be "You can try the following methods to solve the problem: 1. Confirm whether the Internet function is turned on. 2. Confirm whether the phone card is in arrears. 3. Restart the device to try..."; Since the system confirms that the user's phone card is in arrears, the answer branch "Confirm whether the phone card is in arrears" is valid, while the two answer branches "Confirm whether the Internet function is turned on" and "Restart the device to try" may be invalid or irrelevant, so they will be cut off; after cutting, the system only retains the answer branch "Phone card is in arrears" and returns it to the user as the final answer.

[0193] S1004, determining a target speech that matches the user's patience based on the determined user's patience and the answer speech of the current conversation text.

[0194] In this embodiment, an artificial intelligence method can be used to implement it, such as generating a target speech that matches the user's patience through a large language model. Specifically, the user's patience PL is used by the large language model to determine the tone of the target speech that matches the user's patience. For example, the smaller the user's patience PL value, the more friendly the tone of the response speech generated by the large language model is, in order to soothe the user's anxious psychology. For example, "Hello, your phone card is currently in arrears. Please pay the fee and confirm again whether you can access the Internet." The larger the user's patience PL value, the more official the tone of the response speech generated by the large language model, such as "Your phone card is currently in arrears. Please reply to the Internet after paying the fee."

[0195] In this embodiment, the user's patience for the current question is evaluated and the answer is tailored so that the response ultimately recommended to the user meets the user's demand for answer validity, avoiding recommending the same standard answer to the user every time and causing user dissatisfaction.

[0196] Figure 11 A flowchart of a method for recommending a speech technique according to an embodiment of the present disclosure is shown. Figure 11 As shown, the speech recommendation method provided by the embodiment of the present disclosure may include:

[0197] S111, obtaining the user's question.

[0198] S112, inputting the obtained user question into a pre-trained speech retrieval model to obtain an answer to the question.

[0199] S113, determining user patience.

[0200] S114. Cut the answer to the question according to the user's patience to obtain the response answer.

[0201] S115. Generate a response script based on the user's patience and the response answer.

[0202] S116. Recommend the response script to the user.

[0203] It should be noted that in the technical solution of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations. In the embodiments of the present disclosure, various types of data such as personal identity data, operation data, and behavior data related to individuals, customers, and groups have been authorized.

[0204] Based on the same inventive concept, an embodiment of the present disclosure also provides a script recommendation device as described in the following embodiments. Since the principle of solving problems in this device embodiment is similar to that of the above method embodiment, the implementation of this device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be elaborated.

[0205] Figure 12 Show a schematic diagram of a script recommendation device in an embodiment of the present disclosure, as Figure 12 shown, the device includes: a first acquisition module 121, a retrieval module 122, a first determination module 123, and a second determination module 124;

[0206] The first acquisition module 121 is configured to acquire the current conversation text input by the target user in the current round of conversation and the interaction information related to the current conversation text; the retrieval module 122 is configured to input the current conversation text into a script retrieval model and output a set of target scripts that match the current conversation text, where the set of target scripts includes multiple scripts with different contents, and the script retrieval model is obtained by training an initial retrieval model based on a target script library, and the target script library stores scripts for recommendation; the first determination module 123 is configured to determine the user patience of the target user according to the interaction information related to the current conversation text; the second determination module 124 is configured to determine a target script that matches the user patience from the set of target scripts according to the determined user patience, and recommend the target script to the target object.

[0207] In some embodiments, the interaction information related to the current session text includes the time information of the current session text, the previous session text of the current session text in the current round of conversation, and the time information of the previous session text; the first determination module 123 is configured to: determine whether there is a previous interaction for the current session text in the current round of conversation; when there is no previous interaction, determine the user patience of the target user according to the time information of the current session text; when there is a previous interaction, determine the user patience of the target user according to the time information of the current session text, the previous session text of the current session text in the current round of conversation, and the time information of the previous session text.

[0208] In some embodiments, the first determination module 123 is further configured to: preset the correspondence between time periods and user patience; the first determination module 123 is configured to: obtain the preset correspondence between time periods and user patience according to the time information of the current session text; determine the user patience of the target user according to the preset correspondence between time periods and user patience.

[0209] In some embodiments, the first determination module 123 is configured to: determine the time influence value of the target user according to the time information of the current session text and the time information of the previous session text; determine the content influence value of the target user according to the current session text and the previous session text; determine the user patience of the target user according to the time influence value of the target user and the content influence value of the target user.

[0210] In some embodiments, when there are multiple previous session texts of the current session text in the current round of conversation, the first determination module 123 is configured to: determine the total duration of the previous session and the average time difference of the target user's replies according to the time information of the current session text and the time information of the previous session text, where the average time difference of the target user's replies is calculated according to the time information of two adjacent session texts input by the target user in the current round of conversation; determine the average interval time between each session text according to the total duration of the previous session and the total number of previous session texts; determine the time influence value of the target user according to the average time difference of the target user's replies and the average interval time between each session text.

[0211] In some embodiments, the first determination module 123 is configured to: cluster the current session text and the previous session text based on semantics to obtain a clustered session text cluster; determine a first sub-session text cluster in the clustered session text cluster whose number of session texts meets a preset index, and a second sub-session text cluster clustered with the current session text; determine the content influence value of the target user according to the first sub-session text cluster and the second sub-session text cluster.

[0212] In some embodiments, the first determination module 123 is configured to: determine the text overlap relevance based on the conversation texts in the first sub-conversation text cluster and the conversation texts in the second sub-conversation text cluster; determine the conversation intersection relevance based on the number of conversation texts in the first sub-conversation text cluster and the number of texts in the second sub-conversation text cluster; and determine the content influence value of the target user based on the text overlap relevance and the conversation intersection relevance.

[0213] In some embodiments, the retrieval module 122 is configured to: obtain sample conversation texts with labels, where the labels annotate the standard phrases of the sample conversation texts; input the sample conversation texts with labels into an initial retrieval model, and retrieve sample phrases that match the sample conversation texts from a target phrase library; evaluate the sample phrases using the standard phrases of the sample conversation texts to obtain an evaluation result; adjust the initial retrieval model according to the evaluation result, and re-perform retrieval based on the sample conversation texts; until the evaluation result meets a preset condition, stop training the initial retrieval model to obtain a trained phrase retrieval model.

[0214] In some embodiments, when the target phrase library is a document set containing sample phrases, the retrieval module 122 is configured to: determine a set of sample keywords based on the sample conversation texts; determine a target directory item based on a first relevance between the set of sample keywords and a target item corresponding to each document in the document set; determine a target abstract based on a second relevance between the set of sample keywords and an abstract corresponding to the target directory item; and determine a sample phrase that matches the sample conversation texts based on a third relevance between the set of sample keywords and the phrase semantics corresponding to the target abstract.

[0215] In some embodiments, the retrieval module 122 is configured to: determine a set of standard keywords based on the standard phrases; determine the standard degree of the sample phrases according to the set of sample keywords, the set of standard keywords, the sample phrases, and the standard phrases; determine the abnormality degree of the sample phrases according to a fourth relevance between the set of sample keywords and the target abstract; and determine the accuracy of the sample phrases according to the standard degree and the abnormality degree of the sample phrases.

[0216] In some embodiments, the second determination module 124 is configured to: clip the phrases in the target phrase set according to the determined user patience to obtain the answer phrases for the current conversation text; and determine the target phrases that match the user patience according to the determined user patience and the answer phrases for the current conversation text.

[0217] It should be noted here that the examples and application scenarios implemented by each module in the above device embodiments are the same as the corresponding steps in the method embodiments, but are not limited to the content disclosed in the above method embodiments. It should be noted that the above modules, as part of the device, can be executed in a computer system such as a set of computer-executable instructions.

[0218] Those skilled in the art can understand that various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module" or "system" here.

[0219] Based on the same inventive concept, embodiments of the present disclosure also provide an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the conversation recommendation method of any one of the above via executing the executable instructions. Since the principle of solving problems in this electronic device embodiment is similar to that of the above method embodiment, the implementation of this electronic device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be elaborated.

[0220] Next, refer to Figure 13 to describe the electronic device 1300 according to this embodiment of the present disclosure. Figure 13 The shown electronic device 1300 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0221] As Figure 13 shown, the electronic device 1300 is presented in the form of a general-purpose computing device. The components of the electronic device 1300 may include but are not limited to: at least one of the above processing units 1310, at least one of the above storage units 1320, and a bus 1330 connecting different system components (including the storage unit 1320 and the processing unit 1310).

[0222] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1310, so that the processing unit 1310 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1310 may execute the following steps of the above method embodiment: obtaining the conversation recommendation requirement information of the business system; generating the business process of the business system according to the conversation recommendation requirement information; and obtaining at least one functional component from the low-code environment platform according to the business process and the conversation recommendation requirement information to obtain the conversation recommendation, so that the low-code environment platform executes the conversation recommendation.

[0223] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 13201 and / or a cache storage unit 13202, and may further include a read-only storage unit (ROM) 13203.

[0224] The storage unit 1320 may also include a program / utilities 13204 having a set (at least one) of program modules 13205. Such program modules 13205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0225] The bus 1330 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0226] The electronic device 1300 may also communicate with one or more external devices 1340 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1300, and / or may communicate with any device that enables the electronic device 1300 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 1350. Moreover, the electronic device 1300 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1360. As shown in the figure, the network adapter 1360 communicates with other modules of the electronic device 1300 through the bus 1330. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0227] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0228] Based on the same inventive concept, embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned speech recommendation method of any one of the above is implemented. Since the principle of solving problems in the embodiments of this computer-readable storage medium is similar to that of the above method embodiments, the implementation of the embodiments of this computer-readable storage medium can refer to the implementation of the above method embodiments, and the repeated parts will not be elaborated.

[0229] More specific examples of the computer-readable storage medium in the present disclosure may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0230] In the present disclosure, the computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, on which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, and this readable medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0231] Optionally, the program code included on the computer-readable storage medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0232] In specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0233] Based on the same inventive concept, embodiments of the present disclosure also provide a computer program product, including: a computer program or instruction, which when executed by a processor implements the speech recommendation method in any one of the above method embodiments. Since the principle of solving problems in this computer program product embodiment is similar to that of the above method embodiments, the implementation of this computer program product embodiment can refer to the implementation of the above method embodiments, and the repeated parts will not be elaborated.

[0234] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0235] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0236] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0237] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A method for recommending speech skills, characterized in that: The method comprises: Obtain the current conversation text entered by the target user in this round of conversation and the interaction information related to the current conversation text; Input the current conversation text into a speech retrieval model, and output a target speech set matching the current conversation text, wherein the target speech set includes multiple speech with different contents, and the speech retrieval model is obtained by training an initial retrieval model based on a target speech library, and the target speech library stores speech for recommendation; Determine the user patience of the target user based on the interaction information related to the current conversation text; According to the determined user patience, a target speech matching the user patience is determined from the target speech set, and the target speech is recommended to the target object.

2. The method for recommending words according to claim 1, characterized in that: The interactive information related to the current conversation text includes time information of the current conversation text, a previous conversation text of the current conversation text in the current round of dialogue, and time information of the previous conversation text; determining the user patience of the target user according to the interactive information related to the current conversation text includes: In this round of dialogue, determine whether there is a previous interaction in the current conversation text; When there is no previous interaction, determining the user patience of the target user according to the time information of the current conversation text; When there is a previous interaction, the user patience of the target user is determined according to the time information of the current conversation text, the previous conversation text of the current conversation text in the current round of dialogue, and the time information of the previous conversation text.

3. The method for recommending words according to claim 2, characterized in that: Before determining the user patience of the target user according to the time information of the current conversation text, the method further includes: Pre-set the relationship between time periods and user patience; The determining the user patience of the target user according to the time information of the current conversation text includes: Acquire a correspondence between a preset time period and user patience according to the time information of the current conversation text; The user patience of the target user is determined according to the correspondence between the preset time period and the user patience.

4. The method for recommending words according to claim 2, characterized in that: The determining of the user patience of the target user according to the time information of the current conversation text, the previous conversation text of the current conversation text in the current round of dialogue, and the time information of the previous conversation text includes: Determine the time impact value of the target user according to the time information of the current conversation text and the time information of the previous conversation text; Determine the content impact value of the target user according to the current conversation text and the previous conversation text; The user patience of the target user is determined according to the time influence value of the target user and the content influence value of the target user.

5. The method for recommending words according to claim 4, characterized in that: When there are multiple preceding conversation texts of the current conversation text in the current round of conversation, determining the time influence value of the target user according to the time information of the current conversation text and the time information of the preceding conversation text includes: Determine the total duration of the previous conversation and the average time difference between the target user's reply according to the time information of the current conversation text and the time information of the previous conversation text, wherein the average time difference between the target user's reply is calculated according to the time information of two adjacent conversation texts input by the target user in the current round of conversation; Determine the average interval between each conversation text according to the total duration of the preceding conversation and the total number of preceding conversation texts; The time impact value of the target user is determined according to the average time difference of the target user's reply and the average interval time between each conversation text.

6. The method for recommending speech techniques according to claim 4, characterized in that: The determining the content impact value of the target user according to the current conversation text and the previous conversation text includes: Cluster the current conversation text and the previous conversation text based on semantics to obtain a clustered conversation text cluster; Determining, in the clustered conversation text cluster, a first sub-conversation text cluster whose number of conversation texts meets a preset index, and a second sub-conversation text cluster that is clustered into the same category as the current conversation text; The content influence value of the target user is determined according to the first sub-conversation text cluster and the second sub-conversation text cluster.

7. The method for recommending speech techniques according to claim 6, characterized in that: The determining the content influence value of the target user according to the first sub-session text cluster and the second sub-session text cluster comprises: Determining text overlap relevance based on the conversation texts in the first sub-conversation text cluster and the conversation texts in the second sub-conversation text cluster; Determining the conversation cross-correlation according to the number of conversation texts in the first sub-conversation text cluster and the number of texts in the second sub-conversation text cluster; The content influence value of the target user is determined according to the text overlap relevance and the session cross relevance.

8. The method for recommending speech techniques according to claim 1, characterized in that: The speech retrieval model is obtained by the following training: Acquire a sample conversation text with a label, wherein the label marks a standard speech of the sample conversation text; Input the sample conversation text with the label into the initial retrieval model, and retrieve the sample speech matching the sample conversation text from the target speech library; Using the standard speech of the sample conversation text to evaluate the sample speech to obtain an evaluation result; Adjust the initial retrieval model based on the evaluation results and perform retrieval based on the sample conversation text again; Until the evaluation results meet the preset conditions, stop training the initial retrieval model and obtain a trained speech retrieval model.

9. The method for recommending speech techniques according to claim 8, characterized in that: When the target speech library is a document collection containing sample speech, the sample conversation text with the label is input into the initial retrieval model, and sample speech matching the sample conversation text is retrieved from the target speech library, including: Determining a sample keyword set based on the sample conversation text; Determining a target catalog item based on a first degree of association between the sample keyword set and a target item corresponding to each document in the document set; Determining a target summary based on a second degree of association between the sample keyword set and the summary corresponding to the target directory item; Based on the third association between the sample keyword set and the meaning of the words corresponding to the target summary, a sample word matching the sample conversation text is determined.

10. The method for recommending speech techniques according to claim 9, characterized in that: The step of evaluating the sample speech by using the standard speech of the sample conversation text to obtain an evaluation result includes: Determining a standard keyword set based on the standard speech; Determining the standardization of the sample speech according to the sample keyword set, the standard keyword set, the sample speech, and the standard speech; Determining the abnormality of the sample speech according to a fourth correlation between the sample keyword set and the target summary; The accuracy of the sample speech is determined based on the standardization of the sample speech and the abnormality of the sample speech.

11. The method for recommending speech techniques according to claim 1, characterized in that: According to the determined user patience, a target speech matching the user patience is determined from the target speech set, including: According to the determined user patience, cutting the words in the target word set to obtain the answer words of the current conversation text; According to the determined user patience and the answer words of the current conversation text, a target word matching the user patience is determined.

12. A speech recommendation device, characterized in that: The device comprises: The first acquisition module is used to acquire the current conversation text input by the target user in the current conversation and the interaction information related to the current conversation text; A retrieval module, used for inputting the current conversation text into a speech retrieval model, and outputting a target speech set matching the current conversation text, wherein the target speech set includes a plurality of speech with different contents, and the speech retrieval model is obtained by training an initial retrieval model based on a target speech library, and the target speech library stores speech for recommendation; A first determination module, configured to determine the user patience of the target user according to the interaction information related to the current conversation text; The second determination module is used to determine a target speech that matches the user's patience from the target speech set according to the determined user's patience, and recommend the target speech to the target object.

13. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the speech recommendation method described in any one of claims 1 to 11 by executing the executable instructions.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for recommending speech techniques described in any one of claims 1 to 11 is implemented.

15. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the speech recommendation method described in any one of claims 1 to 11.

Citation Information

Cited By

  • Session processing method and device

    CN121166852A

  • A conversation processing method and apparatus

    CN121166852B