Intelligent response method, device and equipment based on user session attention information
By acquiring historical communication records and identifying conversation focus information, knowledge-based decision-making information is generated, solving the problem of inaccurate response information in intelligent dialogue scenarios, achieving accurate and efficient intelligent responses, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CITIC-PRUDENTIAL LIFE INSURANCE CO LTD
- Filing Date
- 2025-07-01
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, the response information generated in intelligent dialogue scenarios is not accurate enough, resulting in low response efficiency, high communication bandwidth consumption, and poor user experience.
By acquiring historical communication records, identifying conversation focus information, generating knowledge decision information and decision categories, storing them in an intelligent knowledge base, and responding accurately to queries, supporting both text and voice responses.
It improves the accuracy and efficiency of intelligent response, enhances user experience, and reduces communication bandwidth usage.
Smart Images

Figure CN120780811B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to an intelligent response method and apparatus based on user session attention information. Background Technology
[0002] Currently, with the continuous development of artificial intelligence, intelligent dialogue scenarios are being applied more and more widely. Effectively and accurately identifying the dialogue motivations in intelligent dialogue can effectively reveal the habits and purposes of the dialogue partners, providing an information foundation for subsequent intelligent dialogue or data processing. The typical approach is as follows: First, extract the business query text. Then, directly utilize the response model to generate the corresponding response text. Finally, generate the response text and send it to the corresponding user.
[0003] However, when using the above method, the following technical problems often arise:
[0004] The generated responses to query text are often inaccurate, leading to disorganized responses and inefficient responses. This not only consumes significant communication bandwidth but also results in a poor user experience.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion that follows. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure propose intelligent response methods, apparatuses, and devices based on user session attention information to solve one or more of the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide an intelligent response method based on user session attention information, comprising: acquiring historical communication records, wherein the historical communication records include: communication records after communication in an intelligent response application and communication records after communication with human customer service, wherein the intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios; determining session attention information corresponding to the historical communication records; generating knowledge decision information and decision categories for the session attention information; storing the knowledge decision information and decision categories in an intelligent knowledge base; in response to receiving a response query information sent by the intelligent response application, querying the knowledge decision information and decision categories corresponding to the response query information from the intelligent knowledge base as response information; in response to determining that the response should be in text form, sending the response information to the intelligent response application and displaying it as text to the corresponding user; in response to determining that the response should be in voice form, sending the response voice corresponding to the response information to the intelligent response application and playing it as voice to the corresponding user.
[0009] Secondly, some embodiments of this disclosure provide an intelligent response device based on user session attention information, comprising: an acquisition unit configured to acquire historical communication records, wherein the historical communication records include: communication records after communication in an intelligent response application and communication records after communication with human customer service, wherein the intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios; a determination unit configured to determine session attention information corresponding to the historical communication records; a generation unit configured to generate knowledge decision information and decision categories for the session attention information; a storage unit configured to store the knowledge decision information and decision categories in an intelligent knowledge base; a query unit configured to, in response to receiving a response query information sent by the intelligent response application, query the knowledge decision information and decision categories corresponding to the response query information from the intelligent knowledge base as response information; a first sending unit configured to, in response to determining that a response is to be given in text form, send the response information to the intelligent response application and display it to the corresponding user as text; and a second sending unit configured to, in response to determining that a response is to be given in voice form, send the response voice corresponding to the response information to the intelligent response application and play it to the corresponding user as voice.
[0010] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0011] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0012] The above embodiments of this disclosure have the following beneficial effects: By utilizing an intelligent knowledge base, the intelligent response method based on user session focus information in some embodiments of this disclosure can accurately and efficiently provide intelligent responses to inquiring users, improving response efficiency. Specifically, the reason for the inaccuracy and low efficiency of related intelligent responses is that the generated response information is often inaccurate for the query text, leading to chaotic response information and low response efficiency. In this case, not only is a large amount of communication bandwidth consumed, but the user experience is also poor. Based on this, the intelligent response method based on user session focus information in some embodiments of this disclosure first obtains historical communication records, including communication records after communication in the intelligent response application and communication records after communication with human customer service. The intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios. Here, by obtaining historical communication records within a historical time period, it is possible to subsequently understand the main insights of the communicating user (i.e., the main focus during the communication process). Here, intelligent responses can be achieved through the intelligent response application, greatly improving the work efficiency of customer service and related tasks. Here, the intelligent response application supports deployment in multiple application scenarios, providing great convenience. Then, the conversation focus information corresponding to the aforementioned historical communication records can be accurately determined. Here, generating conversation focus information facilitates the subsequent generation of knowledge decision information and decision categories, helping the intelligent knowledge base provide more decision knowledge about each decision category and achieve intelligent and accurate responses. Next, knowledge decision information and decision categories for the aforementioned conversation focus information are generated for subsequent storage in the intelligent knowledge base, providing accurate decision knowledge during subsequent response queries. Then, the aforementioned knowledge decision information and decision categories are stored in the intelligent knowledge base. Furthermore, in response to receiving a response query from the aforementioned intelligent response application, the corresponding knowledge decision information and decision category are retrieved from the aforementioned intelligent knowledge base as response information. Here, the obtained response information helps the intelligent response application achieve accurate replies. Further, in response to determining that the response is in text form, the aforementioned response information is sent to the aforementioned intelligent response application and displayed as text to the corresponding user. Finally, in response to determining that the response is in voice form, the corresponding response voice is sent to the aforementioned intelligent response application and played as voice to the corresponding user. Here, through voice and text responses, diverse responses are supported, improving the user experience. In summary, by leveraging historical communication records, key insights from each user over a historical period can be accurately collected and summarized into knowledge-based decision-making information. This helps to achieve accurate and intelligent responses and improves reply efficiency. Attached Figure Description
[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0014] Figure 1 This is a flowchart of some embodiments of the intelligent response method based on user session attention information according to the present disclosure;
[0015] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the intelligent response device based on user session attention information according to this disclosure;
[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] refer to Figure 1The diagram illustrates a flow 100 of some embodiments of the intelligent response method based on user session attention information according to this disclosure. This intelligent response method based on user session attention information includes the following steps:
[0024] Step 101: Obtain historical communication records after communication in the smart reply application.
[0025] In some embodiments, the executing entity (e.g., an electronic smart device) can obtain historical communication records via wired or wireless means. These historical communication records include communication records following communication within an intelligent response application and communication records following communication with human customer service. The intelligent response application can be an application (APP) that supports intelligent voice response. Intelligent voice response can be an automatic reply to questions or inquiries raised by users. Historical communication records can be records of communication conducted within the intelligent response application over a historical period. In practice, historical communication records can be communication information in text or voice format. In practice, historical communication records can be communication records generated by target users within a target historical period. For example, a historical communication record could be "What is the interest rate of the target product today?". The aforementioned intelligent response application supports the following scenarios: intelligent customer service assistant scenario and robot follow-up scenario. The intelligent customer service assistant scenario can be a scenario where automatic replies are achieved through an intelligent customer service assistant. The robot follow-up scenario can be a scenario where, based on the existing human telephone follow-up evaluation channels, an AI-powered outbound call robot automatically obtains the request ticket information requiring follow-up, effectively solving the problems of high workload and repetitiveness in telephone follow-ups. The intelligent response application may support, but is not limited to, at least one of the following: outbound call task management, automatic dialing, simple intelligent communication, simple intelligent intent analysis, and intelligent telemarketing assistance. The aforementioned intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios. Online intelligent text response customer service can be an intelligent customer service scenario that responds in text form. Intelligent voice response scenarios can be intelligent customer service scenarios that respond in audio form. The user asking the question can be a user inquiring about the content.
[0026] Step 102: Determine the conversation attention information corresponding to the above historical communication records.
[0027] In some embodiments, the aforementioned executing entity may determine the session focus information corresponding to the aforementioned historical communication records. The session focus information may be the communication content in the historical communication records that indicates the main communication direction between the user and the intelligent response application (i.e., the communication content that the user is primarily concerned with). For example, the session focus information may be complaint information from a user's complaint against a target in the historical communication records.
[0028] As an example, firstly, the aforementioned implementing entity can utilize an intent recognition model to determine the communication intent corresponding to the aforementioned historical communication records. Then, the communication intent is identified as conversational focus information. The intent recognition model can be a neural network model based on an LSTM model used for intent recognition.
[0029] In some optional implementations of certain embodiments, the aforementioned historical communication records may include a sequence of communication information in a multimodal format. The communication information may be a message sent by a user communicating with a human customer service representative, a message sent by a user making an inquiry in an intelligent reply application, a reply from a human customer service representative, or a reply from an intelligent reply application. Specifically, a communication message may be a single message sent. Each communication message corresponds to a sender. The sender may be a respondent (e.g., a human customer service representative or an intelligent reply application) or an inquirer (the user making the communication). The sequence of communication information is ordered according to the time of its sending. The communication information may be in audio, text, or image format. For audio communication information, it is generally input by the user making the communication. Image-based communication information may be in the form of emojis, related chat logs, or screenshots of product interfaces, etc.
[0030] The aforementioned implementing entity can determine the session attention information corresponding to the aforementioned historical communication records by including the following steps:
[0031] The first step is to determine the source of the aforementioned historical communication records through data tracing. The source of record generation can be any source that generated the historical communication record. For example, if the historical communication record was generated by the intelligent customer service within an e-commerce application, the corresponding source is the e-commerce intelligent customer service source. If the historical communication record was generated by a chat application, the corresponding source is the chat application source. If the historical communication record was generated by a human customer service representative within an e-commerce application, the corresponding source is the e-commerce human customer service source.
[0032] The second step involves splitting the aforementioned historical communication records based on their source to obtain a first historical communication sub-record of intelligent customer service responses and a second historical communication sub-record of human customer service responses. The first historical communication sub-record can be a historical record of communication conducted via intelligent responses. The second historical communication sub-record can be a historical record of communication conducted via human responses.
[0033] As an example, the aforementioned executing entity can, based on the source of the record generation, divide the aforementioned historical communication records into a classification method to obtain a first historical communication sub-record of intelligent customer service response and a second historical communication sub-record of human customer service response.
[0034] Third, for the first historical communication sub-record, perform the following first information binding steps:
[0035] Sub-step 1 involves determining the set of intelligent communication sub-records output by the corresponding intelligent response object in the first historical communication sub-record. The intelligent response object can be an intelligent customer service representative or an intelligent application. The set of intelligent communication sub-records is often in sequential form.
[0036] Sub-step 2 involves querying the cached data corresponding to each intelligent communication sub-record from the response cache corresponding to the intelligent response object. The cached data includes: at least one information identifier corresponding to at least one communication message, overall communication semantic information in vector form corresponding to at least one communication message, and the intelligent communication sub-record. The information identifier can represent the record location and identity information corresponding to the intelligent communication sub-record. Each communication message has a corresponding information identifier. The information identifier can be the location of the communication message recorded when responding to it. At least one communication message can be the response object of the intelligent communication sub-record. That is, the intelligent response object responds to the overall communication semantic information of at least one communication message to generate the intelligent communication sub-record. The overall communication semantic information can represent the overall semantic content corresponding to at least one communication message. At least one communication message can be a sequence of communication messages continuously sent by the sending object.
[0037] In practice, firstly, each piece of communication information in at least one communication message can be transcribed into text to generate communication text information, resulting in at least one piece of communication text information. Then, the at least one piece of communication text information is sequentially combined to obtain a communication text information sequence. Next, the communication text information sequence is input into the overall communication semantic information generation model corresponding to the intelligent response object to generate the overall communication semantic information generation model corresponding to the aforementioned at least one communication message. The overall communication semantic information generation model can be an RNN-based encoding and decoding model. During the training process of the RNN-based encoding and decoding model, the training labels can be the overall communication text in text form corresponding to at least one communication message. The overall communication text can be the communication text generated after summarizing the main semantics of at least one piece of communication information.
[0038] Sub-step 3: For each cached data, perform the following binding steps:
[0039] The first sub-step involves inputting the cached data corresponding to the intelligent communication sub-record into the overall communication semantic information generation model to generate overall intelligent communication semantic information in vector form.
[0040] The second sub-step involves sequentially concatenating the aforementioned overall intelligent communication semantic information and the corresponding overall communication semantic information according to their temporal relationships, to obtain the concatenated overall communication semantic information.
[0041] Sub-step 4: Determine the sequence of overall communication semantic information corresponding to the intelligent communication sub-record set.
[0042] Sub-step 5 involves combining each adjacent number of spliced overall communication semantic information in the above-mentioned spliced overall communication semantic information sequence to obtain an overall communication semantic information group sequence.
[0043] Sub-step 6: Using the communication semantic direction classification model, determine the communication semantic direction and direction probability corresponding to each overall communication semantic information group in the above overall communication semantic information group sequence, obtaining the communication semantic direction sequence and direction probability sequence. The communication semantic direction can represent the communication direction of the semantic content corresponding to the overall communication semantic information group. For example, in the e-commerce field, the corresponding communication semantic direction can be one of the following: price communication direction, claim direction, function consultation direction, or logistics direction. The communication semantic direction classification model can be a classification model based on multiple concatenated convolutional layers. The communication semantic direction classification model can be trained based on a training sample set labeled with each communication semantic direction. The training method can be a conventional deep learning training method. Each communication semantic direction has a corresponding direction probability. The direction probability can represent the accuracy of the communication semantic direction. The direction probability can be a value between 0 and 1. The higher the direction probability, the more accurate the corresponding communication semantic direction.
[0044] Sub-step 7: Select communication semantic directions with a probability less than the target value from the communication semantic direction sequence and use them as target communication semantic directions to obtain at least one target communication semantic direction.
[0045] Sub-step 8: Based on at least one target communication semantic direction and the remaining communication semantic direction sub-sequence set, generate at least one communication semantic direction sequence with different corresponding communication directions.
[0046] Sub-step 9 generates at least one candidate session attention information corresponding to at least one communication semantic direction sequence. Each communication semantic direction has corresponding candidate session attention information.
[0047] As an example, the aforementioned execution entity can input the overall communication semantic information group corresponding to each communication semantic direction into the decoding model to obtain candidate session attention information. The decoding model can be a network model used to output session attention semantic content. For example, the decoding model can be a decoding model based on the existence-time output of deconvolutional layers.
[0048] Fourth step: In response to determining that the second historical communication sub-record is empty, at least one candidate session attention information is identified as session attention information.
[0049] Optionally, based on at least one target communication semantic direction and a set of remaining communication semantic direction subsequences, at least one communication semantic direction sequence with different corresponding communication directions is generated, including:
[0050] The first step, for each target communication semantic direction, is to perform the following breakdown steps:
[0051] Sub-step 1: Determine the overall communication semantic information group corresponding to the target communication semantic direction, and use it as the target overall communication semantic information group.
[0052] Sub-step 2: Determine the remaining communication semantic directions corresponding to the existing directional proximity relationships of the above target communication semantic direction, and obtain the first remaining communication semantic direction in the left proximity relationship and the second remaining communication semantic direction in the right proximity relationship.
[0053] Sub-step 3: Determine the overall communication semantic information group corresponding to the first residual communication semantic direction and the overall communication semantic information group corresponding to the second residual communication semantic direction, and use them as the first overall communication semantic information group and the second overall communication semantic information group, respectively.
[0054] Sub-step 4 involves sequentially concatenating the various overall communication semantic information in the first overall communication semantic information group to obtain the first concatenated information, and sequentially concatenating the various overall communication semantic information in the second overall communication semantic information group to obtain the second concatenated information.
[0055] Sub-step 5 involves selecting a first overall communication semantic information group from the target overall communication semantic information group whose vector similarity to the first concatenated information is higher than the target similarity, and selecting a second overall communication semantic information group from the target overall communication semantic information group whose vector similarity to the second concatenated information is higher than the target similarity. The target similarity can be set based on historical experience. For example, the target similarity can be 0.5.
[0056] Sub-step 6: Add the first overall communication semantic information group to the first overall communication semantic information group corresponding to the first splicing information to obtain the added first overall communication semantic information group; and add the second overall communication semantic information group to the second overall communication semantic information group corresponding to the second splicing information to obtain the added second overall communication semantic information group.
[0057] Sub-step 7: The communication semantic direction corresponding to the first overall communication semantic information group after the addition and the communication semantic direction corresponding to the first overall communication semantic information group after the addition are respectively determined as the remaining communication semantic directions, so as to replace the first and second remaining communication semantic directions in the remaining communication semantic direction sub-sequence set.
[0058] The second step is to determine the obtained set of remaining communication semantic direction subsequences after replacement as at least one communication semantic direction sequence.
[0059] The aforementioned steps, as one of the inventive points of this disclosure, solve another technical problem: "In historical communication records, the information is massive and the semantic content of each piece of information is semantically intertwined, leading to limitations in model understanding and output lag and resource waste when directly inputting and outputting based on conventional neural network models." Based on this, this disclosure first divides historical communication records into a sequence of communication information, using this fine-grained approach to accurately generate conversational focus information. Then, by determining the source of the record generation, it can effectively and accurately distinguish between communication records with intelligent customer service responses and those with human customer service responses (i.e., the first and second historical communication sub-records, respectively). Here, during the intelligent customer service's response process, caching of cached data enables effective storage of cached data corresponding to each response, facilitating subsequent high-efficiency semantic analysis. Next, based on the semantic relationships stored in the cached data, by applying the overall communication semantic information generation model, the intelligent communication sub-record splicing, the communication semantic direction classification model, and the processing of communication semantic directions under different communication directions, it is possible to accurately generate at least one candidate conversation focus information in the intelligent response scenario.
[0060] Optionally, the above method further includes:
[0061] The first step, in response to determining that the second historical communication sub-record is not empty, is to determine the sequence of human-operated communication sub-records corresponding to the second historical communication sub-record. These human-operated communication sub-records can be records of communication conducted through human customer service. The individual human-operated communication sub-records in the sequence are ordered based on the chronological order in which they were sent.
[0062] The second step is to filter out the human customer service replies from the above human communication sub-record sequence as the target human communication sub-record, thus obtaining the target human communication sub-record sequence.
[0063] Third, for each target human communication sub-record in the above target human communication sub-record sequence, perform the following combined steps:
[0064] Sub-step 1: Determine a group of human communication records centered on the position corresponding to the target human communication sub-record, with the number of context records being the target number. The position corresponding to the target human communication sub-record can be its sequence position within the aforementioned human communication sub-record sequence. The target number can be a pre-set number. The human communication record group includes: the target human communication sub-record, a first human communication sub-sequence preceding the target human communication sub-record in the target number of human communication sub-record sequences, and a second human communication sub-sequence preceding the target human communication sub-record in the target number of human communication sub-record sequences. For example, the human communication sub-record sequence is {Record 1, Record 2, Record 3, Record 4, Record 5, Record 6…Record 100}. For the target human communication sub-record being Record 5, corresponding to a target number of 3, the first human communication sub-record sequence can include {Record 2, Record 3, Record 4}. The second human communication sub-record sequence can include {Record 6, Record 7, Record 8}.
[0065] Sub-step 2 involves inputting each manual communication record from the aforementioned manual communication record group into the overall communication semantic information generation model to generate vector-based manual communication semantic information, thus obtaining the manual communication semantic information group. The manual communication semantic information can characterize the semantic information of the record content corresponding to the manual communication record.
[0066] Sub-step 3 involves inputting the aforementioned group of semantic information from human communication into the attention information generation model to obtain the first attention information. This first attention information includes, but is not limited to, at least one of the following: communication topic and communication topic content. The attention information generation model can be trained based on a text set and its corresponding attention tag set. The attention information generation model can be a model incorporating an attention mechanism. Through the attention mechanism, the attention content corresponding to each human communication sub-record can be quickly identified. For example, the attention information generation model can be a Transformer model.
[0067] The fourth step involves combining a predetermined number of artificial communication semantic information groups from the obtained sequence to form a combined sequence of artificial communication semantic information groups. The predetermined number can be determined based on the number of sequences corresponding to the original sequence of artificial communication semantic information groups. For example, the predetermined number could be one-fifth of the total number of sequences.
[0068] The fifth step is to generate the second point of interest information corresponding to each combination of artificial communication semantic information in the above sequence of artificial communication semantic information combinations, thus obtaining the second point of interest information sequence. The generation of the second point of interest information can be found in the generation of the first point of interest information.
[0069] The sixth step is to determine the obtained first concern information sequence, second concern information sequence, and at least one candidate session concern information as session concern information.
[0070] One of the inventive points of this disclosure is to solve the problem of "chaotic responses between human customer service representatives and users, leading to an inability to effectively organize semantics, and the difficulty and time-consuming nature of semantic organization when simultaneously organizing the semantics of all human communication sub-record sequences." Based on this, this disclosure first uses the human communication sub-records of human customer service replies as a foundation to determine groups of human communication records within a certain response range (i.e., the number of contexts). Within each group of human communication records, there are relevant key conversational content from the human customer service replies. Therefore, by determining the key conversational content locally based on these groups, it is possible to quickly grasp the key conversational content without understanding excessive redundant semantic content. Furthermore, by combining a predetermined number of human communication semantic information groups, it is possible to avoid the leakage of relevant attention information between different human communication record groups, thus preventing issues with the accuracy of leaked attentional information.
[0071] In some optional implementations of certain embodiments, the aforementioned executing entity may determine the session attention information corresponding to the aforementioned historical communication records, including the following steps:
[0072] The first step is to generate the corresponding audio text of the aforementioned historical communication records, since the records are in audio format, using speech-to-text technology.
[0073] The second step is to extract the set of scenario keywords from the aforementioned voice text. These scenario keywords can be keywords from various business scenarios. The scenario keyword set can be empty or not. Various business scenarios can refer to the different business scenarios supported by the intelligent voice response application. For example, these scenarios could include: credit investigation scenarios and insurance scenarios.
[0074] As an example, the aforementioned execution entity can utilize TF-IDF (Term Frequency-Inverse Document Frequency) technology to extract the set of scene keywords from the aforementioned speech text.
[0075] The third step, in response to the determination that the keyword set for the aforementioned scenario is empty, is to extract the text semantic feature information and text intent information corresponding to the aforementioned speech text. The text semantic feature information can be in vector form, representing the semantic content of the speech text. The text intent information can represent the main inquiry intent of the speech text regarding the target business-related content.
[0076] As an example, firstly, the aforementioned speech text is segmented to generate a word set. Then, the word set is input into a pre-trained text feature extraction model to generate semantic text features. Finally, the semantic text features are input into a pre-trained text intent generation model to generate text intent information. In practice, the text feature extraction model can be a Transformer encoding model. The text intent generation model can be a text classification task based on LSTM (Long-Short Term Memory). The text feature extraction model and the text intent generation model can be trained simultaneously using a pre-acquired text training set based on gradient descent.
[0077] The fourth step is to identify the aforementioned text semantic features and text intent information as conversational attention information.
[0078] In some optional implementations of certain embodiments, the execution entity may extract the text semantic feature information and text intent information corresponding to the aforementioned speech text, including the following steps:
[0079] The first step involves inputting the aforementioned speech character vector sequence into at least one business information output model targeting the target business to generate at least one business intent output information. This business information output model includes a text semantic feature extraction layer and at least one business intent output layer. The business intent output information includes a business output intent and an intent probability. The target business can be an intelligent voice response business. At least one business information output model corresponds to a sub-business within at least one sub-business included in the target business. For example, the intelligent voice response business includes at least one sub-business within at least one voice response scenario. That is, different sub-businesses correspond to different voice response scenarios. For example, at least one voice response scenario can include: a hospital response scenario, a gate control response scenario, and a logistics response scenario. The business information output model can be a neural network model that outputs intent information for the corresponding response scenario. That is, different response scenarios correspond to different neural network models that output intent information. The model structure of the business information output model can differ for different response scenarios. The business intent output information can be the intent output result in a corresponding response scenario. The text semantic feature extraction layer can be a network layer used to extract text semantic features. Text semantic features can represent the vector semantic content of the combined vectors corresponding to the speech character vector sequence. Text semantic features can be information in vector form. There is a one-to-one correspondence between at least one business intent output layer and at least one business intent output information. The business intent output layer can be a network layer that outputs the business intent in a corresponding response scenario. For example, at least one business intent output layer can be a different output layer corresponding to a fully connected layer. The business output intent can be a predicted intent in a business scenario. The intent probability can characterize the accuracy of the business output intent. The intent probability can be a value between 0 and 1; the higher the value, the higher the accuracy of the business output intent.
[0080] The second step is to determine the output information corresponding to the above text semantic feature information extraction layer as the above text semantic feature information.
[0081] The third step is to select the business intent output information with the highest probability of corresponding intent from at least one of the above business intent output information, and use it as the target business intent output information.
[0082] The fourth step is to determine the business output intent corresponding to the above target business intent output information as the above text intent information.
[0083] In some optional implementations of certain embodiments, the aforementioned execution entity may extract a set of scene keywords from the aforementioned speech text, including the following steps:
[0084] The first step is to segment the above speech text into characters to obtain a speech character sequence.
[0085] As an example, the aforementioned execution entity can divide the speech text into segments according to each character, obtaining a speech character sequence. The order of the speech characters in the speech character sequence corresponds to the order of the characters in the speech text.
[0086] The second step is to vectorize the above speech character sequence to obtain a speech character vector sequence. The speech character vectors can represent the semantic features of the corresponding speech character content.
[0087] As an example, the aforementioned execution entity can perform word embedding processing on each speech character to generate a speech character vector, resulting in a speech character vector sequence.
[0088] The third step involves using a cosine similarity-based character vector traversal algorithm to match each of the obtained full-scene key character vector sets with the aforementioned speech character vector sequence, generating matching results and matching positions. Each full-scene key character vector set corresponds to a specific scene keyword. Each full-outbound call scene key character vector set corresponds to a specific outbound call scene keyword. Each outbound call scene keyword includes at least one key character. There is a one-to-one correspondence between the key character in the at least one key character and the full-outbound call scene key character vector in the full-scene outbound call scene key character vector set. There is also a one-to-one correspondence between the full-scene outbound call scene key character vector sets in the full-scene outbound call scene key character vector set and the outbound call scene keywords in the full-scene outbound call scene keyword set. The full-scene outbound call scene keyword set can be all keywords in the outbound call scene. The full-scene outbound call scene key character vector set represents the outbound call scene key character vector sets corresponding to all keywords in the outbound call scene. The character vector traversal algorithm based on cosine similarity can be implemented by calculating the cosine similarity between the key character vector group and the speech character vector sequence in the full outbound call scenario, and then sequentially traversing each speech character vector in the speech character vector sequence. The matching result can be the sum of the cosine similarities. The matching position can be the vector traversal position in the speech character vector sequence. For example, the speech character vector sequence can be {first speech character vector, second speech character vector, third speech character vector, fourth speech character vector, fifth speech character vector}. For the key character vector group in the full outbound call scenario, which corresponds to 2 vectors, firstly, the first vector combination corresponding to {first speech character vector, second speech character vector} is combined with the first and second vectors in the key character vector group of the full outbound call scenario, and the sum of the first cosine similarities is obtained, which is used as the matching result corresponding to the first vector combination. The matching position corresponding to the first vector combination is determined as {sequence position 1, sequence position 2}. Then, the cosine similarity of the second vector combination corresponding to {second speech character vector, third speech character vector} with the first and second vectors in the key character vector group of the full outbound call scenario is calculated separately, and the sum of the second cosine similarities is used as the matching result for the second vector combination. The matching positions corresponding to the second vector combination are determined as {sequence position 2, sequence position 3}. Next, the cosine similarity of the third vector combination corresponding to {third speech character vector, fourth speech character vector} with the first and second vectors in the key character vector group of the full outbound call scenario is calculated separately, and the sum of the third cosine similarities is used as the matching result for the third vector combination. The matching positions corresponding to the third vector combination are determined as {sequence position 3, sequence position 4}.Next, the fourth vector combination corresponding to {fourth speech character vector, fifth speech character vector} is summed with the first and second vectors in the key character vector group of the full outbound call scenario using cosine similarity. The sum of the fourth cosine similarities is used as the matching result corresponding to the fourth vector combination. The matching position corresponding to the fourth vector combination is determined as {sequence position 4, sequence position 5}.
[0089] The fourth step is to remove empty matching results from the obtained matching result set to obtain the removed matching result set.
[0090] The fifth step is to determine the matching position corresponding to each matching result in the set of matching results after the above removal, and obtain the matching position set.
[0091] Step 6: For each matching position in the above set of matching positions, perform the following second generation step:
[0092] Sub-step 1: Select the matching word corresponding to the above matching position from the above speech character sequence as the target matching word.
[0093] Sub-step 2: Filter out sub-voices from the above target business inquiry voices that are in the same time period as the above target matching words.
[0094] As an example, the aforementioned entity can use target speech processing software to perform speech segmentation on the target business inquiry speech to generate sub-speech that is in the same time period as the target matching word.
[0095] Sub-step 3 involves retrieving the speech feature information corresponding to the target scene keywords at the matching positions from the speech feature information database, using this as the second speech feature information. The speech feature information database stores speech feature information corresponding to keywords for various outbound call scenarios. Each speech feature information in the database is a standardized vector representing the semantic content of the speech features for the outbound call scenario keywords.
[0096] Sub-step 4 involves combining the first speech feature information corresponding to the aforementioned sub-speech and the speech character vector group corresponding to the aforementioned target matching word in a target order to generate a first combined vector. The target order can be the order in which the first speech feature information is on the left and the speech character vector group is on the right.
[0097] As an example, the aforementioned execution entity can utilize a speech feature extraction model based on Mel Frequency Cepstrum Coefficient (MFCC) to extract the speech feature information corresponding to the aforementioned sub-speech, which serves as the first speech feature information.
[0098] Sub-step 5 involves combining the aforementioned second speech feature information and the corresponding full-scale scene key character vector group in target order to generate a second combined vector.
[0099] Sub-step 6 generates the vector similarity between the first combined vector and the second combined vector. The vector similarity characterizes the content similarity between the semantic content of the vectors corresponding to the first combined vector and the semantic content of the vectors corresponding to the second combined vector. A higher vector similarity indicates a higher content similarity between the corresponding vector semantic contents.
[0100] As an example, the aforementioned execution entity can input the first combined vector and the second combined vector into multiple convolutional layers that are sequentially downsampled to obtain vector similarity.
[0101] Sub-step 7: In response to determining that the above vector similarity is higher than the target similarity, the above target matching word is identified as a scene keyword. The target similarity can be a pre-set similarity. For example, the target similarity can be 0.7.
[0102] Step 103: Generate knowledge decision information and decision categories for the above-mentioned conversation focus information.
[0103] In some embodiments, the aforementioned executing entity can generate knowledge decision information and decision categories for the aforementioned session-related information. The knowledge decision information can be decision information used to assist intelligent response applications in providing knowledge responses. The decision category can be a category of knowledge response. The knowledge decision information can also be decision information used for data analysis or internal company analysis of the session-related information. For example, the knowledge decision information could be: "Complaint categories are divided into service quality complaints, product quality complaints, service attitude complaints, fee dispute complaints, and compliance complaints. The characteristic description for service quality complaints is 'service delay / error / incompleteness,' and the corresponding case is '10 instances of express delivery being delayed by 3 days within six months.' The characteristic description for product quality complaints is 'product defects / functional incompatibility / safety hazards,' and the corresponding case is '5 instances of abnormal battery overheating in mobile phones within six months.' The characteristic description for service attitude complaints is 'employee indifference / inappropriate language / insufficient professionalism,' and the corresponding case is '8 instances of customer service interrupting customer statements during calls within six months.' The characteristic description for fee dispute complaints..." The description of "billing errors / hidden charges / price fraud" corresponds to a case where "automatic deductions for membership activation occurred 4 times within six months without prior notification." The characteristic description for compliance complaints is "violation of laws and regulations / contract terms," and the case for compliance complaints is "unauthorized modifications to the user agreement occurred 5 times within six months." The corresponding decision category can be an inquiry category for the complaint. In practice, the decision category can be, but is not limited to, at least one of the following: inquiry category, complaint category. Each decision category can be customized according to the specific business scenario. For inquiry scenarios, the corresponding decision categories can be various decision categories under each inquiry type. For complaint scenarios, the corresponding decision categories can be various decision categories under each complaint category.
[0104] In some optional implementations of certain embodiments, the aforementioned executing entity may generate knowledge decision information and decision categories for the aforementioned session concern information, including the following steps:
[0105] The first step is to generate the decision categories mentioned above based on the textual intent information.
[0106] As an example, the aforementioned executing entity can input textual intent information into a decision category classification model to generate the aforementioned decision categories. The decision category classification model can be a neural network model that determines the decision category corresponding to the intent. In practice, the decision category classification model can consist of multiple concatenated convolutional layers or multiple concatenated residual layers.
[0107] As another example, the aforementioned implementing entity can query the intent-to-decision-category association table to determine the decision category corresponding to the text intent information. The intent-to-decision-category association table represents the correspondence between intents and decision categories.
[0108] The second step is to generate prompts based on the aforementioned decision categories and textual semantic features to generate knowledge decision information. These prompts can be prompt instructions. The templates for the prompts can be pre-set by relevant technical personnel based on historical experience.
[0109] The third step is to input the above prompts into a pre-trained large language model to obtain knowledge decision information. For example, the large language model could be a GPT model.
[0110] Step 104: Store the above knowledge decision information and the above decision categories into the intelligent knowledge base.
[0111] In some embodiments, the aforementioned executing entity may store the aforementioned knowledge decision information and decision categories in an intelligent knowledge base. The intelligent knowledge base may store various responses to different queries. This intelligent knowledge base, combined with technologies such as full-text search, vector models, and AIGC large-scale models, provides AI capabilities and enables intelligent question answering.
[0112] Step 105: In response to receiving the reply query information sent by the aforementioned intelligent reply application, query the knowledge decision information and decision category corresponding to the aforementioned reply query information from the aforementioned intelligent knowledge base, and use them as the reply information.
[0113] In some embodiments, in response to receiving a response query information sent by the aforementioned intelligent response application, the executing entity can query knowledge decision information and decision categories corresponding to the response query information from the aforementioned intelligent knowledge base as response information. The response query information can be a query question raised by the user in the intelligent response application. For example, the response query information could be "How's the weather today?".
[0114] In some optional implementations of certain embodiments, the executing entity can query knowledge decision information and decision categories corresponding to the above-mentioned response query information from the intelligent knowledge base, including the following steps:
[0115] The first step is to determine that the above-mentioned response query information is in voice form, and then use speech-to-text technology to generate the corresponding query voice text.
[0116] The second step is to generate the query decision category corresponding to the above query voice text.
[0117] As an example, firstly, the executing entity can determine the query text intent corresponding to the query voice text. Then, it determines the decision category corresponding to the query text intent, which is then used as the query decision category.
[0118] The third step is to perform category vectorization on the above query decision categories to generate query decision category vectors.
[0119] As an example, the aforementioned execution entity can input the query decision category into the word embedding model to achieve category vectorization and obtain the query decision category vector.
[0120] The fourth step is to determine the set of decision category vectors stored in the intelligent knowledge base. Each decision category vector uniquely corresponds to a decision category. The intelligent knowledge base stores a set of decision information. Each piece of decision information includes: a decision category, a set of knowledge decision information, a decision category vector, and a set of knowledge decision vectors. That is, a decision category has a set of knowledge decision information. Under each decision category, there is stored knowledge decision information for each category.
[0121] The fifth step is to select at least one decision category vector from the aforementioned set of decision category vectors whose corresponding vector similarity is higher than the target similarity. The target similarity can be a pre-set similarity value. For example, the target similarity can be 0.5.
[0122] Step 6: For each decision category vector, perform the following first generation step:
[0123] Sub-step 1: Determine the knowledge decision vector set corresponding to the above decision category vectors. Sub-step 2: Determine the sum similarity between the above knowledge decision vector set and the query text feature information corresponding to the above query speech text. The query text feature information represents the semantic content of the speech text in the query speech text and is in vector form.
[0124] As an example, firstly, the aforementioned executing entity can determine the similarity between each knowledge decision vector in the knowledge decision vector set and the feature information of the query text, thus obtaining a similarity set. Then, the similarity sets are summed to obtain the total similarity.
[0125] Step 7: Select the decision category vector whose total similarity satisfies the target condition from at least one of the above decision category vectors, and use it as the target decision category vector. The target condition can be the decision category vector with the highest total similarity.
[0126] Step 8: Select at least one knowledge decision vector from the knowledge decision vector set corresponding to the target decision category vector above, whose similarity to the query text feature information above is higher than the target similarity.
[0127] Step 9: Input at least one of the above-mentioned knowledge decision vectors and the above-mentioned query text feature information into the attention-based knowledge decision generation model to generate the target knowledge decision vector. The attention-based knowledge decision generation model can be the knowledge decision vector extracted based on the most similar feature information to the query text feature information, obtained by fusing the similarity feature information corresponding to each knowledge decision vector and the full feature information corresponding to the query text feature information. In practice, the attention-based knowledge decision generation model can include: a multi-head attention mechanism layer and an encoding and decoding layer. The encoding and decoding layer can include: an encoding layer and a decoding layer. The encoding layer is used to fuse and compress the input feature information. The decoding layer is used to generate the target knowledge decision vector based on the compressed feature information. The encoding layer can include: a vector scale unification layer, a concatenation layer, and multiple concatenated downsampling convolutional layers. The vector scale unification layer can be a network layer that unifies the vector scale of each input feature information. The concatenation layer can be a network layer that concatenates and fuses the feature information after vector scale unification. The decoding layer can include: multiple concatenated upsampling convolutional layers and an output layer. The output layer can be a fully connected layer.
[0128] As an example, at least one knowledge decision vector and the query text feature information are input into a multi-head attention mechanism layer to generate at least one decision semantic weight corresponding to at least one knowledge decision vector. The decision semantic weight characterizes the importance of the decision semantic feature information in the knowledge decision vector. The decision semantic weight can be a value between 0 and 1; the higher the value, the more useful the decision semantic feature information in the knowledge decision vector is for the subsequent knowledge decision corresponding to the query text feature information. Then, at least one decision semantic weight, at least one knowledge decision vector, and the query text feature information are input into an encoding and decoding layer to obtain the target knowledge decision vector.
[0129] Step 10: Generate the knowledge decision information corresponding to the target decision category vector and the knowledge decision information corresponding to the target knowledge decision vector.
[0130] Step 106: In response to the determination to reply in text form, the above response information is sent to the above smart reply application and displayed to the corresponding user as text.
[0131] In some embodiments, in response to determining that a reply should be given in text form, the aforementioned executing entity may send the reply information to the aforementioned smart reply application via wired or wireless means, and display it to the corresponding user as text.
[0132] Step 107: In response to determining that a reply will be given in the form of voice, the corresponding voice reply information is sent to the smart reply application and played to the corresponding user.
[0133] In some embodiments, in response to determining that a response will be given in the form of voice, the executing entity may send the corresponding voice response to the intelligent response application for playback to the corresponding user. In practice, text-to-speech technology can be used to generate the voice response corresponding to the aforementioned response information. This text-to-speech technology can be a technique for generating corresponding audio based on text. In practice, this text-to-speech technology can be based on the AudioLM audio generation model.
[0134] In some optional implementations of certain embodiments, after step 107, the steps further include:
[0135] The first step is to determine the vocabulary type corresponding to each scenario keyword in the above scenario keyword set based on the keyword type classification table, in response to the determination that the above scenario keyword set is not empty.
[0136] The vocabulary type can be defined by the category to which the vocabulary belongs. Each control instruction has a corresponding vocabulary type. That is, each control instruction has at least one closely related scenario keyword. Through at least one scenario keyword, it can be clearly known which control instruction will be executed. The control instruction can be an instruction that instructs the intelligent response application to perform a corresponding control operation. In practice, the control instruction can be, but is not limited to, at least one of the following: a control instruction to execute timeout control, a control instruction to execute interruption control, a control instruction to execute hang-up control, and a control instruction to execute manual control. The vocabulary type corresponding to the control instruction to execute timeout control is the first vocabulary type. The first vocabulary type can represent the entire vocabulary group corresponding to at least one scenario keyword related to the execution of timeout control operation. The vocabulary type corresponding to the control instruction to execute interruption control is the second vocabulary type. The second vocabulary type can represent the entire vocabulary group corresponding to at least one scenario keyword related to the execution of timeout control operation. The vocabulary type corresponding to the control instruction to execute hang-up control is the third vocabulary type. The third vocabulary type can represent the entire vocabulary group corresponding to at least one scenario keyword related to the execution of timeout control operation. The keyword type classification table can be a table that represents the relationship between vocabulary types and scenario keywords, pre-set by relevant technical personnel based on various control operations.
[0137] As an example, the aforementioned executing entity can determine the vocabulary type corresponding to each scenario keyword in the aforementioned scenario keyword set by querying the keyword type classification table.
[0138] The second step is to determine the subset of scene keywords corresponding to each word type, thus obtaining the scene keyword subset group.
[0139] The third step is to determine the target scenario keyword subset as the subset of scenarios with the highest number of corresponding keywords in the above scenario keyword subset groups.
[0140] Fourth, in response to determining that the aforementioned subset of target scenario keywords is the first vocabulary type, a control instruction representing the execution of timeout control is sent to the aforementioned intelligent reply application. Here, timeout control can be a control operation that limits the call duration.
[0141] Fifth, in response to determining that the subset of target scenario keywords in the aforementioned scenario keyword set is the second vocabulary type, a control command representing the execution of interruption control is sent to the aforementioned intelligent reply application. Here, interruption control can be a control operation to interrupt a call.
[0142] Step six: In response to determining that the target scenario keyword subset in the aforementioned scenario keyword set is a third vocabulary type, a control command representing the execution of hang-up control is sent to the aforementioned intelligent reply application. Hang-up control can be a control operation for handling call hang-up.
[0143] Step 7: In response to determining that the target scenario keyword subset in the aforementioned scenario keyword set is the fourth vocabulary type, a control instruction representing the execution of a transfer to human intervention is sent to the aforementioned intelligent response application. Transfer to human intervention can be a control operation that transfers a call to a human operator.
[0144] In some optional implementations of certain embodiments, before determining the vocabulary type corresponding to each scenario keyword in the scenario keyword set according to the keyword type classification table in response to determining that the scenario keyword set is not empty, the method further includes:
[0145] The first step, for each scene keyword, is to perform the following third generation step:
[0146] Sub-step 1 involves determining the user's identity address and location address sent by the intelligent reply application to the user seeking assistance. The user's identity address can be the user's registered address, i.e., the address on the user's ID card. Here, the user is authenticated in the intelligent reply application. Alternatively, the user's identity address can be the outbound call address used by the user. For example, the outbound call address can be the location of the mobile phone number used for the outbound call. The user's location address can be the IP address of the current user's location.
[0147] Sub-step 2 involves determining, from the target address knowledge graph, the set of addresses whose edge distances between the aforementioned user identity addresses are less than the target distance, as the first neighboring address set; and determining, from the target address knowledge graph, the set of addresses whose edge distances between the aforementioned user location addresses are less than the target distance, as the second neighboring address set. The target address knowledge graph can be a knowledge graph storing the relationships between various addresses. These relationships can include: address proximity relationships and address membership relationships. The target distance can be a pre-set distance value. For example, the target distance can be determined by the number of edges in the knowledge graph. If there are 2 edges between two nodes in the corresponding knowledge graph, the distance between the two nodes is 2. If there are 3 edges between two nodes in the corresponding knowledge graph, the distance between the two nodes is 3. The first neighboring address set can be the set of addresses centered on the user's outbound call address, with the number of edge intervals less than the target distance (i.e., the target edge interval number). The second neighboring address set can be the set of addresses centered on the user's location address, with the number of edge intervals less than the target distance (i.e., the target edge interval number).
[0148] Sub-step 3 involves sorting the first nearest-neighbor address set and the user identity address according to edge distance to generate a first address sequence, and sorting the second nearest-neighbor address set and the user location address to generate a second address sequence. The edge distance corresponding to the user identity address is "0". The edge distance corresponding to the user location address is also "0".
[0149] As an example, the aforementioned execution entity can sort the first neighboring address set and the user's outbound call address in ascending order of edge distance to obtain the first address sequence.
[0150] As an example, the aforementioned execution entity can sort the second nearest address set and the user location address in ascending order of edge distance to obtain the second address sequence.
[0151] Sub-step 4 involves retrieving the first timbre data sequence corresponding to the first address sequence and the second timbre data sequence corresponding to the second address sequence from the target timbre database. The target timbre database can be a database storing dialect timbre data for various regions. In practice, the target timbre database can be obtained by relevant technical personnel in various regions after preprocessing the collected timbre audio. Each region corresponds to a specific address. For example, each region can be divided by prefecture-level city. Both the first and second addresses can also be divided by prefecture-level city. There is a one-to-one correspondence between the first timbre data in the first timbre data sequence and the first address in the first address sequence. Similarly, there is a one-to-one correspondence between the second timbre data in the second timbre data sequence and the second address in the second address sequence.
[0152] Sub-step 5: Determine the business inquiry sub-voice corresponding to the above scenario keywords. The business inquiry sub-voice can be a voice segment whose corresponding voice content includes outbound call scenario keywords.
[0153] As an example, the aforementioned implementing entity can use audio processing software to slice the target business inquiry speech to generate various speech segments. Speech segments whose corresponding content includes outbound call scenario keywords are identified as business inquiry sub-speech.
[0154] Sub-step 6 involves determining the first similarity set between the aforementioned first timbre data sequence and the timbre data corresponding to the aforementioned business inquiry sub-speech. The first similarity represents the similarity between the semantic content of the timbre features corresponding to the timbre data. The first similarity can be a value between 0 and 1. The larger the value, the more similar the semantic content of the timbre features between the first timbre data and the timbre data corresponding to the business inquiry sub-speech. The timbre data corresponding to the business inquiry sub-speech can be determined by a relevant timbre feature semantic information extraction model. The timbre feature semantic information extraction model can be a neural network model that extracts the semantic content of timbre features in speech. In practice, the timbre feature semantic information extraction model can be an extraction model based on Mel-frequency cepstral coefficients.
[0155] Sub-step 7 involves determining the second similarity set between the aforementioned second timbre data sequence and the corresponding timbre data of the aforementioned business query sub-voice. Similarly, the explanation of the second similarity can be found in the explanation of the first similarity. Further details will not be elaborated upon here.
[0156] Sub-step 8 involves selecting similarities from the first and second similarity sets that have a similarity value within the top target number, thus obtaining a similarity sequence. The target number can be a predetermined number. The similarity sequence can include either the first or second similarities. The number of each similarity in the similarity sequence is the target number.
[0157] Sub-step 9 involves determining the similarity lexicon corresponding to each address in the address sequence corresponding to the above similarity sequence, thus obtaining a similarity lexicon sequence. Each address has a uniquely corresponding and maintained similarity lexicon. This similarity lexicon stores semantically similar word groups. That is, the semantic content of each word in each word group is similar.
[0158] Sub-step 10: Using the above-mentioned similar word library sequence, determine the similar words corresponding to the above-mentioned scenario keywords to obtain similar word groups, wherein the number of words corresponding to the above-mentioned similar word groups is a predetermined number.
[0159] As an example, firstly, the aforementioned execution entity can obtain a set of similar words by filtering similar words that are closely related to the keywords in the outbound calling scenario from each similar word library in the similar word library sequence. That is, each set of similar words has a corresponding similar word library. Finally, each similar word in the obtained set of similar words is identified as a similar word group.
[0160] The second step is to determine the obtained set of similar phrases and set of scene keywords as the scene keyword set.
[0161] The above embodiments of this disclosure have the following beneficial effects: By utilizing an intelligent knowledge base, the intelligent response method based on user session focus information in some embodiments of this disclosure can accurately and efficiently provide intelligent responses to inquiring users, improving response efficiency. Specifically, the reason for the inaccuracy and low efficiency of related intelligent responses is that the generated response information is often inaccurate for the query text, leading to chaotic response information and low response efficiency. In this case, not only is a large amount of communication bandwidth consumed, but the user experience is also poor. Based on this, the intelligent response method based on user session focus information in some embodiments of this disclosure first obtains historical communication records, including communication records after communication in the intelligent response application and communication records after communication with human customer service. The intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios. Here, by obtaining historical communication records within a historical time period, it is possible to subsequently understand the main insights of the communicating user (i.e., the main focus during the communication process). Here, intelligent responses can be achieved through the intelligent response application, greatly improving the work efficiency of customer service and related tasks. Here, the intelligent response application supports deployment in multiple application scenarios, providing great convenience. Then, the conversation focus information corresponding to the aforementioned historical communication records can be accurately determined. Here, generating conversation focus information facilitates the subsequent generation of knowledge decision information and decision categories, helping the intelligent knowledge base provide more decision knowledge about each decision category and achieve intelligent and accurate responses. Next, knowledge decision information and decision categories for the aforementioned conversation focus information are generated for subsequent storage in the intelligent knowledge base, providing accurate decision knowledge during subsequent response queries. Then, the aforementioned knowledge decision information and decision categories are stored in the intelligent knowledge base. Furthermore, in response to receiving a response query from the aforementioned intelligent response application, the corresponding knowledge decision information and decision category are retrieved from the aforementioned intelligent knowledge base as response information. Here, the obtained response information helps the intelligent response application achieve accurate replies. Further, in response to determining that the response is in text form, the aforementioned response information is sent to the aforementioned intelligent response application and displayed as text to the corresponding user. Finally, in response to determining that the response is in voice form, the corresponding response voice is sent to the aforementioned intelligent response application and played as voice to the corresponding user. Here, through voice and text responses, diverse responses are supported, improving the user experience. In summary, by leveraging historical communication records, key insights from each user over a historical period can be accurately collected and summarized into knowledge-based decision-making information. This helps to achieve accurate and intelligent responses and improves reply efficiency.
[0162] Further reference Figure 2As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an intelligent response device based on user session attention information. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this smart response device based on user session attention information can be specifically applied to various electronic devices.
[0163] like Figure 2 As shown, an intelligent response device 200 based on user conversation attention information includes: an acquisition unit 201, a determination unit 202, a first generation unit 203, a storage unit 204, a query unit 205, a second generation unit 206, and a sending unit 207. The acquisition unit 201 is configured to acquire historical communication records, including communication records after communication via an intelligent response application and communication records after communication with human customer service. The intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios. The determination unit 202 is configured to determine the conversation attention information corresponding to the historical communication records. The generation unit 203 is configured to generate knowledge decision information and decision categories for the aforementioned conversation attention information. The storage unit 204 is configured to store the knowledge decision information and decision categories in the intelligent response device 206. The system includes a knowledge base and a query unit 205 configured to, in response to receiving a response query information sent by the intelligent response application, query the intelligent knowledge base for knowledge decision information and decision category corresponding to the response query information, and use this as response information; a first sending unit 206 configured to, in response to determining that a response will be given in text form, send the response information to the intelligent response application and display it as text to the corresponding user; and a second sending unit 207 configured to, in response to determining that a response will be given in voice form, send the response voice corresponding to the response information to the intelligent response application and play it as voice to the corresponding user.
[0164] It is understandable that the units described in the device 200 based on user session attention information are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein based on user session attention information, and will not be repeated here.
[0165] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0166] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0167] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0168] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0169] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0170] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0171] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire historical communication records, wherein the historical communication records include: communication records after communication via an intelligent response application and communication records after communication with human customer service, wherein the aforementioned intelligent response application supports the following scenarios: online intelligent text response customer service and intelligent voice response scenarios; determine the session attention information corresponding to the aforementioned historical communication records; generate knowledge decision information and decision categories for the aforementioned session attention information; store the aforementioned knowledge decision information and decision categories in an intelligent knowledge base; in response to receiving a response query information sent by the aforementioned intelligent response application, query the knowledge decision information and decision categories corresponding to the aforementioned response query information from the intelligent knowledge base as response information; in response to determining that a response should be given in text form, send the aforementioned response information to the aforementioned intelligent response application and display it to the corresponding user as text; in response to determining that a response should be given in voice form, send the response voice corresponding to the aforementioned response information to the aforementioned intelligent response application and play it to the corresponding user as voice.
[0172] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0174] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a determination unit, a generation unit, a storage unit, a query unit, a first sending unit, and a second sending unit. The names of these units do not necessarily limit the specific unit; for example, an acquisition unit may also be described as a "unit for acquiring historical communication records."
[0175] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0176] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A smart response method based on user session attention information, comprising: Access historical communication records; Determine the conversation attention information corresponding to the historical communication records; Generate knowledge-based decision information and decision categories for the session's focus information; The knowledge decision information and the decision category are stored in an intelligent knowledge base; In response to receiving a response query information sent by the intelligent response application, the system queries the intelligent knowledge base for knowledge decision information and decision category corresponding to the response query information, and uses this as response information. The step of determining the conversation attention information corresponding to the historical communication record includes: The source of the historical communication records was determined by tracing the data. Based on the source of the record generation, the historical communication record is split to obtain a first historical communication sub-record of intelligent customer service response and a second historical communication sub-record of human customer service response; For the first historical communication sub-record, perform the following first information binding steps: Determine the set of intelligent communication sub-records output by the corresponding intelligent reply object in the first historical communication sub-record; Retrieve the cached data corresponding to each smart communication sub-record from the reply cache corresponding to the smart reply object; For each cached data, perform the following binding steps: input the smart communication sub-record corresponding to the cached data into the overall communication semantic information generation model to generate overall smart communication semantic information in vector form; according to the time relationship, concatenate the overall smart communication semantic information and the corresponding overall communication semantic information to obtain the concatenated overall communication semantic information; Determine the sequence of overall communication semantic information corresponding to the intelligent communication sub-record set; The overall communication semantic information group sequence is obtained by combining each adjacent number of spliced overall communication semantic information sequences in the spliced overall communication semantic information sequence. Using a communication semantic direction classification model, the communication semantic direction and direction probability corresponding to each overall communication semantic information group in the overall communication semantic information group sequence are determined, thus obtaining the communication semantic direction sequence and the direction probability sequence. The communication semantic direction sequence is filtered to select the communication semantic direction whose corresponding direction probability is less than the target value, and then used as the target communication semantic direction to obtain at least one target communication semantic direction; Based on at least one target communication semantic direction and a set of remaining communication semantic direction subsequences, generate at least one communication semantic direction sequence with different corresponding communication directions; Generate at least one candidate session attention information corresponding to at least one communication semantic direction sequence; In response to determining that the second historical communication sub-record is empty, at least one candidate session attention information is identified as session attention information.
2. The method according to claim 1, wherein, Determining the session attention information corresponding to the historical communication records includes: In response to the fact that the historical communication record is in voice format, speech-to-text technology is used to generate the corresponding voice text of the historical communication record; Extract the set of scene keywords from the spoken text; In response to determining that the scene keyword set is empty, extract the text semantic feature information and text intent information corresponding to the speech text; The semantic features of the text and the intent information of the text are identified as conversational attention information.
3. The method according to claim 2, wherein, The generation of knowledge decision information and decision categories for the session attention information includes: The decision category is generated based on the text intent information; Generate prompts that generate knowledge decision information based on the decision category and the text semantic feature information; The prompt information is input into a pre-trained large language model to obtain knowledge decision information.
4. The method according to claim 2, wherein, The step of retrieving knowledge decision information and decision categories from the intelligent knowledge base that correspond to the response query information includes: In response to determining that the reply query information is in voice form, speech-to-text technology is used to generate the query voice text corresponding to the reply query information; Generate the query decision category corresponding to the query voice text; The query decision categories are vectorized to generate query decision category vectors; Determine the set of decision category vectors stored in the intelligent knowledge base; Select at least one decision category vector from the set of decision category vectors whose corresponding vector similarity is higher than the target similarity; For each decision category vector, perform the following first generation step: Determine the knowledge decision vector set corresponding to the decision category vector; Determine the summation similarity between the knowledge decision vector set and the feature information of the query text corresponding to the query speech text; From the at least one decision category vector, select the decision category vector whose total similarity satisfies the target condition, and use it as the target decision category vector; Filter at least one knowledge decision vector from the knowledge decision vector set corresponding to the target decision category vector, whose similarity to the query text feature information is higher than the target similarity. The at least one knowledge decision vector and the query text feature information are input into an attention-based knowledge decision generation model to generate a target knowledge decision vector. Generate knowledge decision information corresponding to the target decision category vector and knowledge decision information corresponding to the target knowledge decision vector.
5. The method according to claim 2, wherein, The extraction of the scene keyword set from the speech text includes: The speech text is divided into characters to obtain a speech character sequence; The speech character sequence is vectorized to obtain a speech character vector sequence; For each full scene key character vector group in the acquired full scene key character vector group set, the character vector traversal algorithm based on cosine similarity is used to perform character matching between the full scene key character vector group and the speech character vector sequence to generate matching results and matching positions. Each full scene key character vector group has corresponding scene keywords. Remove empty matches from the obtained match result set to get the match result set after removal; Determine the matching position corresponding to each matching result in the removed matching result set to obtain the matching position set; For each matching position in the set of matching positions, perform the following second generation step: The matching word corresponding to the matching position is selected from the speech character sequence and used as the target matching word; Filter out sub-voices from the target business inquiry voice that are in the same time period as the target matching word; The speech feature information corresponding to the target scene keyword at the matching position is retrieved from the speech feature information database and used as the second speech feature information. The first speech feature information corresponding to the sub-speech and the speech character vector group corresponding to the target matching word are combined in target order to generate a first combined vector; The second speech feature information and the corresponding full-scene key character vector group are combined in target order to generate a second combined vector; Generate the vector similarity between the first combined vector and the second combined vector; In response to determining that the vector similarity is higher than the target similarity, the target matching word is identified as a scene keyword.
6. The method according to claim 5, wherein, The extraction of text semantic feature information and text intent information corresponding to the speech text includes: The speech character vector sequence is input into at least one business information output model for the target business to generate at least one business intent output information, wherein the business information output model includes: a text semantic feature information extraction layer and at least one business intent output layer, and the business intent output information includes: business output intent and intent probability; The output information corresponding to the text semantic feature information extraction layer is determined as the text semantic feature information; From the at least one business intent output information, select the business intent output information with the highest corresponding intent probability as the target business intent output information; The business output intent corresponding to the target business intent output information is determined as the text intent information.
7. The method according to claim 2, wherein, The method further includes: In response to determining that the scene keyword set is not empty, the vocabulary type corresponding to each scene keyword in the scene keyword set is determined according to the keyword type classification table; Determine the subset of scene keywords corresponding to each word type to obtain the scene keyword subset group; The subset of scene keywords with the highest number of corresponding keywords in the aforementioned subset of scene keywords is determined as the target scene keyword subset; In response to determining that the subset of keywords in the target scenario is a first vocabulary type, a control instruction representing the execution of timeout control is sent to the intelligent reply application; In response to determining that the target scenario keyword subset in the scenario keyword set is a second vocabulary type, a control command representing the execution of interruption control is sent to the intelligent reply application; In response to determining that the target scenario keyword subset in the scenario keyword set is a third vocabulary type, a control command representing the execution of hang-up control is sent to the smart reply application; In response to determining that the target scenario keyword subset in the scenario keyword set is a fourth vocabulary type, a control instruction representing the transfer of execution to manual control is sent to the intelligent response application.
8. A smart response device based on user session attention information, comprising: The acquisition unit is configured to acquire historical communication records; The determining unit is configured to determine the session attention information corresponding to the historical communication record; The generation unit is configured to generate knowledge decision information and decision categories for the session's focus information; The storage unit is configured to store the knowledge decision information and the decision category into an intelligent knowledge base; The query unit is configured to, in response to receiving a response query information sent by the intelligent response application, query the intelligent knowledge base for knowledge decision information and decision category corresponding to the response query information, as response information; The determining unit is further configured to: The source of the historical communication records was determined by tracing the data. Based on the source of the record generation, the historical communication record is split to obtain a first historical communication sub-record of intelligent customer service response and a second historical communication sub-record of human customer service response; For the first historical communication sub-record, perform the following first information binding steps: Determine the set of intelligent communication sub-records output by the corresponding intelligent reply object in the first historical communication sub-record; Retrieve the cached data corresponding to each smart communication sub-record from the reply cache corresponding to the smart reply object; For each cached data, perform the following binding steps: input the smart communication sub-record corresponding to the cached data into the overall communication semantic information generation model to generate overall smart communication semantic information in vector form; according to the time relationship, concatenate the overall smart communication semantic information and the corresponding overall communication semantic information to obtain the concatenated overall communication semantic information; Determine the sequence of overall communication semantic information corresponding to the intelligent communication sub-record set; The overall communication semantic information group sequence is obtained by combining each adjacent number of spliced overall communication semantic information sequences in the spliced overall communication semantic information sequence. Using a communication semantic direction classification model, the communication semantic direction and direction probability corresponding to each overall communication semantic information group in the overall communication semantic information group sequence are determined, thus obtaining the communication semantic direction sequence and the direction probability sequence. The communication semantic direction sequence is filtered to select the communication semantic direction whose corresponding direction probability is less than the target value, and then used as the target communication semantic direction to obtain at least one target communication semantic direction; Based on at least one target communication semantic direction and a set of remaining communication semantic direction subsequences, generate at least one communication semantic direction sequence with different corresponding communication directions; Generate at least one candidate session attention information corresponding to at least one communication semantic direction sequence; In response to determining that the second historical communication sub-record is empty, at least one candidate session attention information is identified as session attention information.
9. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Intelligent question answering method and system based on traditional model and domain knowledge base enhancement
CN118657217A
Business question answering method and system based on large language model and scene label
CN119537553A
Intelligent question and answer method and device, electronic equipment and storage medium
CN119760094A