Context-aware speech-to-text conversion
The system improves IVR speech-to-text accuracy and reduces resource consumption by employing a general-purpose language model to generate and augment text strings, addressing inefficiencies in existing IVR systems with multiple conversation stage-specific models.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2026-03-19
AI Technical Summary
Existing speech-to-text systems in interactive voice response (IVR) systems face inaccuracies and increased computing resource requirements due to the need for multiple conversation stage-specific predictive models, leading to inefficiencies in resource usage and administrative tasks.
Implement a system that utilizes a data repository with predictive acoustic and language models, including a general-purpose language model, to generate, inspect, and augment candidate text strings based on user voice input, reducing the need for multiple conversation stage-specific models by leveraging context and confidence scoring to improve transcription accuracy.
Enhances speech-to-text accuracy in IVR systems while minimizing computing resource requirements and administrative tasks by using a general-purpose language model, thus optimizing system performance and reducing complexity.
Smart Images

Figure 0007833248000005 
Figure 0007833248000006 
Figure 0007833248000007
Abstract
Description
Technical Field
[0002] ,
[0004] , , , , , , ,
[0003]
[0001] The embodiments described in this specification generally relate to speech to text conversion, and more specifically, to speech to text conversion according to context.
Background Art
[0002] To improve the operation of computer systems, data structures have been adopted. A data structure refers to the organization of data in a computer environment for improving computer system operation. Types of data structures include containers, lists, stacks, queues, tables, graphs, etc. Data structures have been adopted to improve computer system operation, for example, from the perspectives of algorithm efficiency, memory usage efficiency, maintainability, and reliability.
[0003] Artificial intelligence (AI) refers to the intelligence shown by machines. The research of artificial intelligence (AI) includes search and mathematical optimization, neural networks, probability, etc. Artificial intelligence (AI) solutions include functions derived from research in a wide variety of scientific and technological fields spanning computer science, mathematics, psychology, linguistics, statistics, and neuroscience. Machine learning is described as a research field that gives computers the ability to learn without explicitly programming them.
Summary of the Invention
[0004] In one embodiment, a method is provided that overcomes the shortcomings of the prior art and provides further advantages. The method may include, for example, determining prompt data to present to a user during the execution of an interactive voice response (IVR) session, storing text-based data defining the prompt data in a data repository, presenting the prompt data to the user, receiving voice string data returned by the user in response to the prompt data, generating a plurality of candidate text strings related to the returned voice string of the user, inspecting the text-based data defining the prompt data, augmenting the plurality of candidate text strings according to the results of the inspection and providing a plurality of augmented candidate text strings related to the returned voice string data, evaluating each of the plurality of augmented candidate text strings related to the returned voice string data, and selecting one of the augmented candidate text strings as the returned transcript related to the returned voice string data.
[0005] In another embodiment, a computer program product can be provided. The computer program product may include a computer-readable storage medium that is readable by one or more processing circuits and stores instructions that are executed by one or more processors to perform a method. The method may include, for example, determining prompt data to present to a user in the execution of an interactive voice response (IVR) session, storing text-based data defining the prompt data in a data repository, presenting the prompt data to the user, receiving voice string data returned by the user in response to the prompt data, generating a plurality of candidate text strings related to the returned voice string of the user, examining the text-based data defining the prompt data, augmenting the plurality of candidate text strings in accordance with the results of the examination and providing a plurality of augmented candidate text strings related to the returned voice string data, evaluating each of the plurality of augmented candidate text strings related to the returned voice string data, and selecting one of the augmented candidate text strings as a returned transcript related to the returned voice string data.
[0006] In a further embodiment, a system can be provided. The system may, for example, include memory. Furthermore, the system may include one or more processors communicating with the memory. Furthermore, the system may include program instructions executable by one or more processors via the memory to perform a method. The method may, for example, include, in the execution of an interactive voice response (IVR) session, determining prompt data to present to a user and storing text-based data defining the prompt data in a data repository; presenting the prompt data to the user; receiving voice string data returned by the user in response to the prompt data; generating a plurality of candidate text strings related to the returned voice string of the user; examining the text-based data defining the prompt data; augmenting the plurality of candidate text strings in accordance with the results of the examination and providing a plurality of augmented candidate text strings related to the returned voice string data; evaluating each of the plurality of augmented candidate text strings related to the returned voice string data; and selecting one of the augmented candidate text strings as a returned transcript related to the returned voice string data.
[0007] Further features are realized by the technologies described herein. Other embodiments and aspects, including but not limited to methods, computer program products, and systems, are described in detail herein and are considered to be part of the claimed invention.
[0008] One or more aspects of the present invention are specifically cited and expressedly claimed in the claims at the end of this specification. The foregoing, as well as other objects, features, and advantages of the present invention, will become apparent from the following detailed description made in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows a system comprising an enterprise system that runs an interactive voice response (IVR) application and a plurality of UE devices, according to one embodiment. [Figure 2] This figure shows a predictive model according to one embodiment. [Figure 3] This figure shows a predictive model according to one embodiment. [Figure 4] A flowchart illustrating a method performed by an enterprise system interoperating with a UE device according to one embodiment. [Figure 5] This figure shows a user interface according to one embodiment. [Figure 6] This figure shows a dialogue decision tree for guiding an IVR session according to one embodiment. [Figure 7] A flowchart illustrating a method performed by an enterprise system interoperating with a UE device according to one embodiment. [Figure 8] This figure shows a computing node according to one embodiment. [Figure 9] This figure shows a cloud computing environment according to one embodiment. [Figure 10] This figure shows an abstraction model layer according to one embodiment. [Modes for carrying out the invention]
[0010] Figure 1 shows a system 100 for converting user voice data. System 100 may include an enterprise system 110 having an associated data repository and user equipment (UE) devices 120A to 120Z. System 100 may include a number of devices. These devices may be computing node-based devices connected by a network 190. Network 190 may be a physical network, a virtual network, or both. A physical network may be, for example, a physical telecommunications network connecting a number of computing nodes or systems such as computer servers and computer clients. A virtual network may be, for example, a logical virtual network formed by combining a number of physical networks or a portion thereof. In another example, a number of virtual networks may be defined on a single physical network.
[0011] According to one embodiment, the enterprise system 110 can be located outside of the UE devices 120A to 120Z. According to one embodiment, the enterprise system 110 can be co-located with one or more of the UE devices 120A to 120Z.
[0012] Each of the different UE devices 120A to 120Z can be associated with a different user. With respect to UE devices 120A to 120Z, in one embodiment, one or more of the computer devices of UE devices 120A to 120Z can be computing node devices provided by a client computer. For example, a mobile device (e.g., a smartphone or tablet), laptop, smartwatch, or PC that runs one or more programs (e.g., including a web browser for opening and displaying web pages).
[0013] The embodiments described herein acknowledge that challenges remain in accurately recognizing speech being transcribed. The embodiments also acknowledge that one approach in situations where inaccuracies are observed in speech-to-text is to additionally deploy a specifically trained predictive model, trained on training data specific to that particular situation. For example, the embodiments acknowledge that an interactive voice response (IVR) system can be characterized by N conversation stages, each corresponding to a specific node in a dialog tree. One approach to improving the accuracy of an IVR system is to provide N specific conversation stage predictive models, each corresponding to a specific conversation stage in the IVR system, and train each of the N conversation stage predictive models on its respective historical training data. The embodiments acknowledge that while such an approach can improve the accuracy of speech-to-text and remain useful in augmenting the embodiments described herein, it can significantly increase the system's computing resource requirements as well as its administrative tasks. Additional processes include separately recording training data for N different conversation stages, separately training and maintaining N different conversation stage-specific predictive models, and separately querying each of the N different conversation stage-specific predictive models at runtime.
[0014] The data repository 108 can store various types of data. The data repository 108 can store predictive models queried by the enterprise system 110 in the model area 2121. The predictive models stored in the model area 2121 may include predictive acoustic models 9002 that represent one or more language models. The predictive acoustic model 9002 can respond to query data provided by the speech data to return candidate text strings associated with the input speech data. As shown in Figure 2, the predictive acoustic model 9002 can be trained using a training dataset consisting of speech clips mapped to a phoneme set.
[0015] A predictive model in model domain 2121 may also include a predictive language model 9004 that represents one or more language models. One or more language models may include one or more text-based language models. As shown in Figure 3, the predictive language model 9004 can be trained with training data consisting of text strings that define the overall language. In some use cases, different language models can be provided for different topic domains. A predictive language model 9004 trained with training data consisting of text strings can learn patterns present in the language, such as terms that appear commonly in relation to each other.
[0016] The predictive language model 9004 can be provided by a general language model or a conversation stage-specific language model. In one embodiment, the data repository in the model region 2121 can store a control registry that specifies the state and attributes of the predictive language model associated with each conversation stage in the interactive voice response (IVR) application 111.
[0017] The data repository 108 in the decision data structure area 2122 can store, for example, a dialog decision tree and a decision table for returning action decisions made by the enterprise system 110. In one use case, the enterprise system 110 can run an interactive voice response (IVR) application 111 to run an IVR session that can present voice prompts to the user. In one use case, a speech-to-text process 113 can be incorporated into the interactive voice response application (IVR). The prompt data for the IVR application provided by the virtual agent (VA) can be guided by the decision data structure provided by the dialog tree. The dialog decision tree can provide guidance for the conversational flow between the user and the VA.
[0018] In the logging area 2123, the data repository 108 can store conversation logging data for past IVR sessions. The conversation logging data may include text-based prompt data from the virtual agent (VA) and text-based response data of user input to the prompt data converted using text-to-speech conversion. For logged IVR sessions, the conversation data may include tags that specify the conversation stage and corresponding dialog tree node for each segmented conversation segment. The conversation logging data may also include user ID tags, start tags, end tags, etc.
[0019] In one embodiment, enterprise system 110 can execute an IVR application 111. The enterprise system 110 can execute various processes. The enterprise system 110 executing the IVR application 111 can include the enterprise system 110 executing a prompt process 112 and a speech-to-text process 113. The speech-to-text process 113 can include various processes such as a generation process 114, an inspection process 115, a reinforcement process 116, and an action determination process 117.
[0020] The enterprise system 110 executing the prompt process 112 can include the enterprise system 110 presenting prompt data to the user on the UE device. The enterprise system 110 executing the prompt process 112 can include the enterprise system 110 using a dialog decision tree to guide the conversation between the virtual agent and the user. Different prompts can be presented according to the stage of customer support. Various inputs can be provided to determine a specific VA prompt. Examples of inputs that can be used to determine the prompt data presented by the VA include parameter values indicating the stage of the IVR session, the user's past response data, the user's current sentiment, and the like. When prompt data is presented from the VA by executing the prompt process 112 of the IVR application 111, the user can return response data. The response data can be provided by the user's voice data and can be obtained by the voice input device of the user's UE device.
[0021] When the enterprise system 110 receives voice data from a user, it can process and receive the voice data by executing a voice-to-text process 113. The enterprise system 110 executing the voice-to-text process 113 can include the enterprise system 110 executing a generation process 114, an inspection process 115, a reinforcement process 116, and an action determination process 117.
[0022] The enterprise system 110 executing the generation process 114 can include the enterprise system 110 generating a candidate text string from the voice data input. The enterprise system 110 executing the generation process 114 can include the enterprise system 110 querying a predictive acoustic model 9002 to return the candidate text string. The predictive acoustic model 9002 can be trained using training data of the voice data and can be optimized to respond to the voice data. In one embodiment, the predictive acoustic model 9002 can execute various sub-processes such as voice clip segmentation, phoneme classification, and speaker identification.
[0023] The enterprise system 110 executing the inspection process 115 may include the enterprise system 110 inspecting the context of the voice data received from the user. The context of the received voice data may include prompt data presented to the user prior to the voice data received from the user. Each embodiment in this specification recognizes that inspecting prompt data preceding the received voice data is useful in improving the speech-to-text conversion of the received voice data. The enterprise system 110 executing the inspection process 115 may include the enterprise system 110 inspecting prompt data presented to the user prior to the voice data received from the user. The enterprise system 110 executing the inspection process 115 may include the enterprise system 110 inspecting the prompt data to identify the sentence attribute of the preceding prompt data.
[0024] The enterprise system 110 executing the augmentation process 116 may include the enterprise system 110 augmenting the candidate text strings provided by the generation process 114 using one or more results of the inspection process 115. The enterprise system 110 executing the augmentation process 116 may include the enterprise system 110 adding text data to the candidate text strings generated by the generation process 114 depending on one or more results of the inspection process. The enterprise system 110 executing the augmentation process 116 may include the enterprise system 110 providing the first to nth candidate text strings by adding text data to the first to nth candidate text strings generated by the generation process 114 depending on the attributes of the prompt data determined by the execution of the inspection process 115.
[0025] The execution of the action decision process 117 by the enterprise system 110 may include the enterprise system 110 making an action decision in which it selects from among the first to the Nth candidate text strings provided using the augmented text strings returned by the execution of the augmentation process 116. The execution of the action decision process 117 by the enterprise system 110 may include the enterprise system 110 querying a predictive language model 9004, which can be configured as a language model, using each augmented candidate text string. The predictive language model 9004 can be configured to return a confidence score associated with each augmented candidate text string in response to query data defined by the augmented candidate text strings, where the confidence score indicates the determined likelihood that the augmented candidate text string accurately represents the user's utterance intent. The execution of the action decision process 117 by the enterprise system 110 may include selecting the augmented candidate text string with the highest confidence score as a transcribed utterance of the user's speech.
[0026] The operation method of the enterprise system 110 in cooperation with the UE device 120A will be described with reference to the flowchart in Figure 4. In block 1201, the UE device 120A can send user-defined registration data to the enterprise system 110 as the recipient. The registration data can be user-defined registration data defined using a user interface, such as the user interface 300 shown in Figure 5. The user interface 300 can be a user interface displayed on the display of the UE device 120A and may include an area 302 for text-based data input by the user and an area 304 for presenting text-based data, graphical data, or both to the user. The registration data may include, for example, the user's contact data and user permission data that allows the enterprise system 110 to use various user data, including voice data as defined herein. Upon receiving the registration data, the enterprise system 110 can proceed to block 1101. In block 1101, the enterprise system 110 can send the registration data received from the user and have it stored in the data repository 108. The registered data can be stored in block 1081 by the data repository 108.
[0027] In block 1102, the enterprise system 110 can send an installation package for reception and installation by the UE device 120A. In block 1202, the UE device 120A can receive and install the installation package. The installation package data may include an installation package to be installed on the UE device 120A. The installation package may include, for example, a library of executable code that can enhance the performance of the UE device 120A to run within the system 100. In some embodiments, the provisioning data that the UE device 120A receives from the enterprise system 110 may be minimal, and the UE device 120A can operate as a thin client. In other embodiments, the method shown in the flowchart of Figure 4 may not have block 1102, and the UE device 120A can use web page data sent to the UE device 120A at the start of a communication session to prepare for optimal operation within the system 100. For example, in some embodiments, the IVR functionality described herein can be performed during a web browsing session. During a web browsing session, preparation data that enables the UE device 120A to operate optimally within system 100 is received along with the web pages returned from enterprise system 110 during the browsing session. In other embodiments, the method shown in the flowchart of Figure 4 does not necessarily have block 1102, and the UE device 120A can perform minimal preparations to send voice data to enterprise system 110. The registration process shown in blocks 1201, 1101, and 1081 can be informal, for example, when enterprise system 110 registers a user as a guest user.
[0028] Once the user of UE device 120A is registered with system 100 and permission is presented to enterprise system 110, UE device 120A can proceed to block 1203. In block 1203, UE device 120A can send chat initiating data to enterprise system 110. In block 1203, chat initiating data can be sent by UE device 120A when the user of UE device 120A initiates appropriate control (for example, by starting a voice call, clicking the voice chat button on the displayed user interface 300, or both).
[0029] Upon receiving chat start data, the enterprise system 110 can proceed to block 1103. In block 1103, the enterprise system 110 can execute prompt process 112 to present prompt data to the user. The prompt data presented to the user may be in the form of audio prompt data output on the audio output device of the UE device 120A, or it may include text prompt data displayed in area 302 (Figure 5) on the user interface 300 of the UE device 120A, or both. In the initial pass of block 1103, the prompt data may include, for example, predetermined baseline greeting data. Upon determining the prompt data in block 1103, the enterprise system 110 can proceed to block 1104.
[0030] In one embodiment, the enterprise system 110 can use a dialogue decision tree 3002, as defined in Figure 6, to determine the VA prompt data to present to the user, which is determined in block 1103. In block 1103, the enterprise system 110 can return an artificial intelligence (AI) response decision. That is, it can determine an intelligently generated response based on the most recently received voice data from the user as a response to present to the user by the VA defined by the enterprise system 110. Alternatively, the response can be controlled by the content of the root node if the IVR session has just started. In one embodiment, the enterprise system 110 can return an AI response decision by referring to a dialogue decision tree 3002, as shown in Figure 6.
[0031] The segment of the dialogue decision tree to be activated after the initial greeting is indicated by node 3011 of the dialogue decision tree 3002. The dialogue decision tree 3002 in Figure 6 can control the conversational flow between the VA and the user in a customer service scenario, and each node can define each conversational stage of the IVR session. Referring to the dialogue decision tree in Figure 6, the nodes can encode content to be examined in order to determine the question asked by the VA. The edges between nodes can define the intent of the IVR session, which can be determined by examining a transcript of the user's response to the VA's question. Given a given set of user response data, the enterprise system 110 can use semantic analysis to classify the response data into an intent selected from a set of candidate intents. Each embodiment of this specification recognizes that determining the appropriate decision state may depend on whether an accurate transcript of the user's speech is available. The enterprise system 110 can be configured to refer to the dialogue decision tree 3002 to return an action decision regarding the next VA question to be presented to the user participating in the current IVR session. The VA voice response can be predetermined, as shown in node 3011, which specifies a predetermined VA response, "What problem are you having?", for an intent indicated by an edge titled "Product". In other scenarios, such as those indicated by intents referenced by edges titled "Software" and "Hardware", the VA response can be selected from a menu of candidate question sets, for example, question set A on node 3021 or question set B on node 3022.The enterprise system 110 can be configured to determine IVR prompt data to present to the user based on various parameter values (for example, parameter values indicating the stage of the IVR session indicated by the currently active node in the dialogue decision tree 3002), the user's past response data, and the user's current mood. In some scenarios, the user's past response data, current mood, or both can be used to select from encoded question candidates for conversation stages defined by the dialogue tree nodes. In some scenarios, the user's past response data, current mood, or both can be used to modify reference prompt data associated with a particular conversation stage node. In some embodiments, the enterprise system 110 may invoke a wide variety of dialogue decision trees in response to specific audio data received from the user. In the middle of an IVR session, the enterprise system 110 can disable a first dialogue decision tree and enable a second dialogue decision tree. The dialogue decision tree 3002 may include nodes 3011, 3021-3022, 3031-3035, and 3041-3050 that define each conversation stage of the IVR session.
[0032] Enterprise system 110 running IVR application 111 can perform natural language processing (NLP) to extract NLP output parameter values from received user voice data as well as prompt data from VA. Enterprise system 110 can perform one or more of the following processes: a topic classification process to determine the topic of a message and output one or more topic NLP output parameter values; a sentiment analysis process to determine the sentiment parameter values of a message (e.g., polar sentiment NLP output parameters "negative," "positive," or non-polar sentiment NLP output parameters (e.g., "anger," "disgust," "fear," "joy," and / or "sadness"), or both); and one or more other classification processes to output one or more other NLP output parameter values (e.g., one or more "social tendency" NLP output parameters, one or more "writing style" NLP output parameters, or one or more "part of speech" NLP output parameter values, or a combination thereof). Part-of-speech tagging methods can include, for example, constraint grammar, brill tagger, Baum-Welch algorithm (forward-backward algorithm), and the Viterbi algorithm, which can employ hidden Markov models. Hidden Markov models can be implemented using the Viterbi algorithm. Brill taggers can learn a set of rule patterns and apply these patterns rather than optimizing statistical quantities. Applying natural language processing can also include performing sentence segmentation. Sentence segmentation can include identifying the end of a sentence, for example, by searching for periods while considering periods that indicate contractions.
[0033] The Enterprise System 110 performing natural language processing may include (a) performing topic classification on an incoming message and outputting one or more topic NLP output parameters, (b) performing sentiment classification on an incoming message and outputting one or more sentiment NLP output parameter values, or (c) performing other NLP classifications on an incoming message and outputting one or more other NLP output parameters. Topic analysis for topic classification and outputting NLP output parameter values may include topic segmentation to identify multiple topics within a message. Topic analysis may apply one or more of various techniques (e.g., Hidden Markov Models (HMMs), artificial chains, passage similarities using word co-occurrence, topic modeling, clustering). Sentiment analysis for sentiment classification and outputting one or more sentiment NLP parameters may determine the speaker's or writer's attitude towards a particular topic or the overall contextual polarity of a document. Attitude may be the author's judgment or evaluation, emotional state (the author's emotional state at the time of writing), or intended emotional communication (the emotional effect the author wants to have on the reader). In one embodiment, sentiment analysis can classify the polarity of a given text in terms of whether the expressed opinion is positive, negative, or neutral. Advanced sentiment classification can classify beyond the polarity of a given text. Advanced sentiment classification can classify emotional states as sentiment classifications. Sentiment classifications include the classifications of "anger," "disgust," "fear," "joy," and "sadness."
[0034] In block 1104, the enterprise system 110 can send prompt data to the UE device 120A for presentation to the user via the UE device 120A. The prompt data sent in block 1104 can be the prompt data determined in block 1103. In one embodiment, the prompt data determined in block 1103 can be text-based prompt data, and the data sent in block 1104 can be synthesized speech-based data synthesized from the text-based prompt data. The enterprise system 110 can generate synthesized speech-based data from the text-based data using a text-to-speech conversion process.
[0035] Upon receiving prompt data, the user of UE device 120A can send user-defined reply voice data using the voice input device of UE device 120A in block 1204. Upon receiving voice data, the enterprise system 110 can process the received voice data by executing blocks 1105, 1106, 1107, and 1108.
[0036] In generation block 1105, the enterprise system 110 can execute generation process 114 to generate one or more candidate text strings related to the speech string data transmitted in block 1204 and received by the enterprise system 110. The enterprise system 110 executing generation block 1105 may include the enterprise system 110 executing generation process 114 to query the predictive acoustic model 9002 in model region 2121. The predictive acoustic model 9002 can be configured to output candidate text strings corresponding to the received speech data defined by the speech string.
[0037] The predictive acoustic model 9002 can perform various processes, including, for example, phoneme segmentation, phoneme classification, or speaker identification, or a combination thereof. In one embodiment, the predictive acoustic model 9002 can employ a Hidden Markov Model. The Hidden Markov Model can be employed such that each phoneme has a different output distribution. A Hidden Markov Model for a sequence of phonemes can be created by concatenating each Hidden Markov Model, which has been individually trained for separate phonemes. In addition to or instead of this, the predictive acoustic model 9002 can employ dynamic time warping (DTW), a neural network, or both. The predictive acoustic model 9002 can be trained using training data consisting of past speech clips that define a set of phonemes in a given language. The past speech clips may be from non-users of system 100 or from users of system 100. The predictive acoustic model 9002 can return candidate text strings corresponding to the input speech string without identifying the user. However, in some embodiments, the user associated with the input speech string can be determined in order to optimize the performance of the predictive acoustic model 9002. When optimizing the performance of the predictive acoustic model 9002 with respect to the detected current user, the predictive acoustic model 9002 can be trained using training data specific to that current user.
[0038] Upon completion of block 1105, the enterprise system 110 has generated multiple candidate text strings related to the voice string data sent by the user in block 1204. In response to the completion of generation block 1105, the enterprise system 110 can proceed to inspection block 1106. In inspection block 1106, the enterprise system 110 can execute inspection process 115 to inspect one or more candidate text strings generated in block 1105 and preceding prompt data related to the voice data string sent in block 1204.
[0039] Examining the preceding prompt data may involve performing text parsing on the preceding prompt data using natural language processing. This may include grammatical parsing, such as sentence segmentation, topic segmentation, and part-of-speech tagging. Sentence segmentation may include identifying the end of a sentence, for example, by searching for periods while considering periods that indicate abbreviations. Topical segmentation may involve assigning topics to consecutive sets of words in a sentence. A sentence can be segmented by identifying topics associated with different groups of words in an identified sentence. Part-of-speech tagging may involve tagging words in a sentence as belonging to a specific part of speech, for example, tagging nouns, verbs, adjectives, adverbs, and pronouns in a sentence. Enterprise system 110 can use text segmentation parameter values (e.g., specifying sentence segmentation, topic details, and part-of-speech tagging) to generate augmented candidate text strings in block 1107. Text parsing can be performed using various natural language processing tools. Examples of natural language processing tools include WATSON DISCOVERY™, WATSON NATURAL LANGUAGE UNDERSTANDING™, and WATSON ASSISTANT™ from International Business Machines Corporation.
[0040] In one embodiment, in block 1106, the enterprise system 110 can discard specific data. The enterprise system 110 can apply rule-based criteria to discard, for example, all sentences except the last sentence defined by prompt data from prompt data that has been determined to contain multiple sentences. The enterprise system 110 can apply rule-based criteria to discard, for example, all words except the set of words defining the last topic from the last identified sentence that has been determined to contain multiple topics. Once block 1106 is complete, the enterprise system 110 can proceed to reinforcement block 1107.
[0041] In reinforcement block 1107, the enterprise system 110 can execute reinforcement process 116 to reinforce the candidate text string generated in block 1105. To reinforce the candidate text string, the enterprise system 110 can apply various rules and transform the prompt data into transformed data to be added as prepended data to the candidate text string associated with the user response data received in block 1204.
[0042] The transformation rules in reinforcement block 1107 include, for example, the following: (a) Replacing third-person pronouns in prompt data with first-person pronouns (e.g., "your" → "my", "you" → "I"); (b) Changing the text of prompt data that defines a question into a statement (e.g., "What state are you traveling to" → "I am traveling to the state"); (c) Rephrasing user instructions in prompt data into first-person declarations (e.g., "Please state your destination" → "My destination is"); (d) If a transformation cannot be performed, the segment of the context is passed to the next stage as is or not used at all. The enterprise system 110 can perform transformations (b) and (c) by identifying text strings in the prompt data text that match template text strings and using a mapping decision data structure that maps the transformed text to the template text string. The mapping decision data structure for performing transformation (b) may include a mapping data structure stored in the decision data structure area 2122 of the data repository 108, as shown in Table 1. [Table 1]
[0043] For text strings containing third-person pronouns, if a template match is identified using Table 1, conversion rule (b) can be prioritized over conversion rule (a).
[0044] The mapping decision data structure for performing transformation (c) may include the mapping data structure stored in the decision data structure area 2122 of the data repository 108, as shown in Table 2. [Table 2]
[0045] For strings containing third-person pronouns, if a template match is identified using Table 2, the transformation rule (c) that rephrases the instructional statement in the prompt data into a first-person declaration can be preferred over the transformation rule (a) that replaces the third-person pronoun in the prompt data with a first-person pronoun. This significantly reduces the number of transformation text strings and further simplifies the mapping decision data structure, making prediction and maintenance easier. Referring to Tables 1 and 2, the enterprise system 110 performing transformations (b) and (c) may include (i) providing the prompt data text string for natural language processing part-of-speech tagging and applying part-of-speech tags to the words in the text string, and (ii) identifying a match between the text string of the prompt data text string and a template text string stored in the data repository 108 in which one or more words in the string are represented as part-of-speech in wildcard form.
[0046] As described in relation to conversion processes (a) to (d), once the text defining the prompt data is converted in block 1107, the enterprise system 110 may further use the converted text in block 1107 to augment the candidate text string generated in block 1105. The enterprise system 110 augmenting the candidate text string using the converted text may include the enterprise system 110 prepending the converted text obtained from conversion processes (a) to (d) to the beginning of the candidate text string generated in block 1105. The enterprise system 110 augmenting the candidate text string in block 1107 may include lengthening the candidate text string, i.e., adding text to the beginning of the candidate text string. The enterprise system 110 augmenting the candidate text string in block 1107 may include providing an augmented candidate text string, for example, the candidate text string having the prefix text obtained by converting the prompt data through conversion processes (a) to (d).
[0047] Once the augmentation block 1107 is complete, the enterprise system 110 can proceed to the decision block 1108. In the decision block 1108, the enterprise system 110 can make a selection from the augmented candidate text strings provided using the augmentation process in block 1107. To execute block 1108, the enterprise system 110 can execute the action decision process 117 to select a specific augmented candidate text string from the set of candidate text strings as a transcript obtained from the received audio data. The enterprise system 110 executing block 1108 may include querying the predictive language model 9004 using the augmented candidate text strings provided in block 1107.
[0048] The predictive language model 9004 can be configured to return confidence parameter values associated with each candidate text string, including both the unreinforced and reinforced text strings. The confidence parameter values may have one or more classifications. The predictive language model 9004 can be a language model configured to return one or more confidence parameter values in response to query data containing candidate text strings. The confidence parameter values can indicate the likelihood that the candidate text string represents the intended content of the voice data received from the user.
[0049] The predictive language model 9004, configured as a language model, can provide a probability distribution for a word sequence. Given a word sequence of length m, the predictive language model 9004 can assign probabilities to the entire sequence. The predictive language model 9004 can employ a neural network to represent words distributively as a nonlinear combination of weights in a neural network to approximate a language function. The neural network architecture employed can be, for example, a feedforward neural network or a recurrent neural network. The neural network defining the predictive language model 9004 can be trained to predict the probability distribution across a lexicon using a neural network training algorithm, for example, stochastic radiant descent with backpropagation. The training dataset for training the predictive language model 9004 can include text strings that define a language.
[0050] The predictive language model 9004 described herein can be trained as a general-purpose language model or a conversation stage-specific language model. The general-purpose language model can be trained using training data from a common topic domain (e.g., a common topic domain of the current IVR application, the company of the current IVR application, or the industry of the current IVR application and / or a combination thereof). The general-purpose language model described herein can also be provided by a commercially available (COTS: commercial off-the-shelf) language model, i.e., an out-of-box language model commonly trained in a major language (e.g., English). The training data for training a general-purpose language model provided by a COTS general-purpose language model can consist, for example, of thousands to millions of common text strings in a given language.
[0051] The predictive language model 9004, trained as a general-purpose language model, is available as a pre-trained model, which can reduce the tasks associated with custom training of predictive language models.
[0052] Providing the predictive language model 9004 as a general-purpose language model may include training the predictive language model 9004 using training data generally relevant to the IVR application 111. Such training data may include, for example, general conversation logs of IVR sessions that do not require parsing and tracking conversation data for association with specific conversation stages and dialogue tree nodes. Providing the predictive language model 9004 as a general-purpose language model may also include training the predictive language model 9004 using training data generally relevant to the company associated with the IVR application 111. Such training data may consist of text strings from product specifications of products offered by the company, including, for example, service products, training documents, and procedures. Providing the predictive language model 9004 as a general-purpose language model may also include training the predictive language model 9004 using training data generally relevant to the industry associated with the IVR application 111. Such training data may include, for example, text strings from information technology textbooks if the topic domain is information technology, or text strings from medical textbooks if the topic domain is medicine. Providing the predictive language model 9004 as a general-purpose language model may also include using a pre-trained COTS general-purpose language model. Providing the predictive language model 9004 as a general-purpose language model by applying training data that is generally relevant to the IVR application 111, its associated company, or associated industry, or a combination thereof, may include (a) applying additional training data to a pre-trained COTS general-purpose language model, or (b) applying training data to an untrained general-purpose language model. In use case (a), a COTS general-purpose language model can be used as a starting point, and the COTS general-purpose language model can be further trained with text training data specific to the IVR application, company, or industry, or a combination thereof, thereby adapting the general-purpose language model to be relevant to the IVR application, company, or industry, or a combination thereof.
[0053] In one embodiment, the predictive language model 9004 can be configured as a conversation stage-specific language model. The predictive language model 9004 can be provided as a conversation stage-specific topic domain language model dedicated to a topic domain of a particular conversation stage in an IVR session, for example, a conversation related to a certain dialogue stage in an IVR session. This dialogue stage can be associated with a node in the dialogue decision tree 3002, as shown in Figure 6. Training data for training a conversation stage-specific language model for a particular conversation stage can include usage data history related to that particular conversation stage. The conversation stage-specific language model can be selectively trained using training data for a particular conversation stage in an IVR session (for example, defined by a specific dialogue tree node in the dialogue tree that controls the operation of the IVR session). While the embodiments of this specification may benefit from deploying conversation stage-specific language models, their use should be acknowledged as involving additional programmatic complexity and the consumption of additional computing resources, for example, in terms of training data collection and application.
[0054] In some use cases described herein, the use of conversation stage-specific language models can be avoided. In some use cases described herein, the use of conversation stage-specific language models can be managed considering the complexity and computing resource costs associated with such use.
[0055] The embodiments of this specification recognize the complexity and computing resource consumption associated with the use of conversation stage-specific language models. Because the training data for training conversation stage-specific language models is inherently limited, deploying conversation stage-specific language models may require not only deploying, storing, and querying numerous models, but also iteratively collecting and maintaining usage data history for application as training data, and applying such training data for the optimization of the models (multiple models if multiple conversation stage-specific language models are deployed). The embodiments of this specification recognize that benefits in terms of complexity and computing resource savings can be obtained by using the general-purpose language models described herein. By using general-purpose language models, the complexity and computing resource consumption associated with the deployment, storage, querying, or training of conversation stage-specific language models (for example, conversation stage-specific language models related to each conversation stage in an IVR session, corresponding to each node in the dialogue decision tree 3002 in Figure 6) can be reduced.
[0056] In block 1108, the enterprise system 110 can query the predictive language model 9004 using multiple candidate text strings related to the input voice string data entered by the user in block 1204. Embodiments of this specification can facilitate the use of a predictive language model 9004 configured as a general-purpose language model. The use of a general-purpose language model can reduce the computing resource requirements of system 100. Embodiments of this specification recognize that augmenting text strings in block 1107 can facilitate the use of a general-purpose language model in block 1108, thereby reducing reliance on conversation stage-specific language models in block 1108. Embodiments of this specification recognize that applying augmented text strings with additional words increases the likelihood that the general-purpose language model will return reliable results. For example, longer input text strings with more words are more likely to match past text strings used as training data to train the language model than shorter text strings. Using a general-purpose language model instead of conversation stage-specific language models can reduce the utilization of computing resources associated with deploying, training, and updating multiple conversation stage-specific language models. To determine which predictive language models to query in blocks 1107 and 1108, enterprise system 110 can examine the control registry in model region 2121. An example of a control registry is shown in Table 4. In one embodiment, the enterprise system 110 can be configured to selectively execute the augmentation block 1107 when the general-purpose language prediction model is active, or when performance monitoring for the current conversation stage is being performed, or both.
[0057] The multiple candidate text strings input to the predictive language model 9004 in block 1108 are each candidate text string generated by the predictive acoustic model 9002 in block 1105 according to the input speech string data, and block 110 7Each augmented candidate text string returned by the enterprise system 110 may include the following: For each input candidate text string input to the predictive language model 9004, the predictive language model 9004 may return one or more confidence parameter values. One or more confidence parameter values may include, for example, a context confidence parameter value, a transcription confidence parameter value, and a domain confidence parameter value. The predictive language model 9004 may return a context confidence parameter value higher than the threshold if the input text string, consisting of multiple consecutive words, strongly matches a past text string used as training data to train the predictive model. The predictive language model 9004 may return a transcription confidence parameter value higher than the threshold if the individual words defining the input text string strongly match individual words in a past text string used as training data to train the predictive language model 9004. The predictive language model 9004 can return a topic domain confidence parameter value higher than the threshold if one or more words defining the input text string strongly match one or more individual words characterizing the current topic domain associated with the current IVR session (e.g., industry topic domain, corporate topic domain, or conversation topic domain, or a combination thereof).
[0058] In block 1108, the enterprise system 110 receives the returned confidence parameter values and can aggregate the confidence parameter values for each candidate input text string. Then, in block 1108, it can return an action decision to select the candidate input text string with the highest aggregated confidence score as the transcript to be returned for the input voice data string sent in block 1204. Aggregating confidence parameter values can include, for example, providing the mean of the values, providing the weighted mean of the values, or providing the geometric mean of the values.
[0059] Enterprise System 110 is Block 1109 The returned text-based transcript can be sent and stored in the data repository 108 in block 1083. Then, in block 1083, the data repository 108 can store the returned transcript in its logging area 2123. Block 110 9 When sending the text returned, the enterprise system 110 can tag the returned transcript with an identifier for the current conversation stage as metadata. The conversation at this conversation stage can be mapped to a node identifier in the current IVR session's dialog decision tree, as described with reference to the dialog decision tree 3002 shown in Figure 6. The enterprise system 110 running the IVR application 111 can then use this returned transcript to derive the user's intent, for example, by semantic analysis. Once the intent is derived, the enterprise system 110 running the IVR application 111 can use, for example, the dialog decision tree 3002 shown in Figure 6 to advance the IVR session to the appropriate next conversation stage. As an example of deriving an intent, the enterprise system 110 running the IVR application 111 can provide a correspondence score for the returned transcript in relation to several candidate intents, such as candidate intents associated with edges in the dialog decision tree 3002 in Figure 6. Embodiments of this specification are described in block 110 9 If an incorrect transcription is returned, the enterprise system 110 recognizes that this may lead to the deriving of an incorrect intent regarding the user and potentially advancing the current IVR session to an inappropriate next stage.
[0060] Once block 1109 is complete, the enterprise system 110 can proceed to block 1110. In block 1110, the enterprise system 110 can perform training of the predictive language model 9004. Training in block 1110 may include updating the training of the predictive language model 9004 (Figure 3) using the transcript sent in block 1109 and stored in the data repository 108. The data repository 108 can respond to requests for stored logging data in the receive / response block 1084. Training of the predictive model stored in the model region 2121 can be performed continuously in the background, concurrently with the execution of other processes, such as the processes defined by the loop of blocks 1103-1111. Although training can be performed in block 1110, the embodiments herein can provide reliability with little to no updating of the training of the predictive language model 9004. In some use cases where the predictive language model 9004 is configured as a common general-purpose language model for multiple conversation stages, the execution of the IVR application 111 by the enterprise system 110 may include the enterprise system 110 executing a lightweight training procedure for training the predictive language model 9004. The lightweight training procedure may include applying the session logging data from the logging area 2123 obtained from the completed session as training data to the common general-purpose language model at the end of the IVR session (block 1111). To determine whether or not to perform training in block 1110, the enterprise system 110 may examine the control registry of the model area 2121. An example of a control registry is shown in Table 4. In one embodiment, the enterprise system 110 may be configured to selectively perform training in block 1110 if a conversation stage-specific language model is active for the current conversation stage.In one embodiment, the enterprise system 110 can be configured to skip training in block 1110 if only the general-purpose language model is active for the current conversation stage. In one embodiment, the enterprise system 110 can be configured to use the lightweight training procedure described above if a common general-purpose language model is active for one or more conversation stages. In this case, at the end of the IVR session (block 1111), the session logging data from the logging area 2123 obtained from the completed session is applied as training data for the common general-purpose language model.
[0061] As defined herein, features such as augmentation of candidate text strings by performing augmentation process 116 allow the predictive language model 9004 to be configured as a general-purpose language model that can be reliably used to return confidence parameter values associated with input candidate text strings without iteratively updating the training of the predictive language model 9004. On the other hand, if the predictive language model 9004 is configured as a conversation stage-specific language model, there may be limited specific conversation stage training data for training the predictive language model, and the reliable use of the model may depend on iteratively updating the training of the predictive model. Therefore, in some embodiments, if the predictive language model 9004 used in block 1108 is configured as a conversation stage-specific language model, the training operation in block 1110 can be performed, and the predictive language model 9004 used in block 1108 can be configured as a conversation stage-specific language model. General purpose If configured as a language model, training in block 1110 can be avoided. Also, in some embodiments, even if the predictive language model 9004 used in block 1108 is configured as a general-purpose language model, the training operation in block 1110 can still be performed.
[0062] Once block 1110 is complete, the enterprise system 110 can proceed to block 1111. In block 1111, the enterprise system 110 can determine whether the current IVR session has ended, for example, due to a user selection or a timeout. As long as the current IVR session has not ended, the enterprise system 110 can iteratively execute the loop of blocks 1103-1111. In the subsequent execution of prompt block 1103, i.e., after the initial greeting, the enterprise system 110 can use the returned transcript sent in block 1109 to derive, for example, the intent and the appropriate next step, and further adapt the prompt data determined in block 1103 and presented in block 1104. The enterprise system 110 can use the previously returned transcript sent in block 1109 for natural language processing to extract topic parameter values and sentiment parameter values, and then use the derived topic parameter values or sentiment parameter values or both to select from stored candidate text strings that define the prompt data. Additionally, the derived topic parameter values and sentiment parameter values can be used to append the stored text to a text string for use as prompt data (for example, prompt data tailored to the specific user sentiment detected).
[0063] If enterprise system 110 determines that the session has ended in block 1111, enterprise system 110 can proceed to block 1112. In block 1112, enterprise system 110 can return to the stage before block 1103 and wait for the next chat start data.
[0064] Further embodiments of the embodiments described herein will be described with reference to Example 1.
[0065] (Example 1) In block 1103, the enterprise system 110 running the IVR application 111 determines the prompt data "What state are you traveling to?" and sets this prompt data to block 108 2The data is stored in the data repository 108. The enterprise system 110 running the IVR application processes the text-based prompt data for text-to-speech conversion and presents the synthesized speech-based prompt data to the user in block 1104. The user sends speech string data in block 1204. The enterprise system 110 running the IVR application 111 supplies this speech string data to the predictive acoustic model 9002 as query data. In order for the enterprise system 110 to generate candidate text strings, the predictive acoustic model 9002 outputs candidate text strings (a) "I'll ask her" and (b) "Alaska" in block 1105. In block 1106, the enterprise system 110 running the IVR application 111 examines the stored prompt data stored in block 1082 by text analysis and extracts data that characterizes the prompt data. In block 1107, the enterprise system 110 running the IVR application 111 uses data characterizing the prompt data and the content of the prompt data to augment the candidate text strings. In block 1107, the enterprise system 110 determines that the prefix text to add to the candidate text string is "I am traveling to state of". In block 1107, the enterprise system 110 generates the augmented candidate text strings by adding the prefix text to the previously determined candidate text strings. In block 1107, the enterprise system 110 can generate the augmented text strings (c) "I am traveling to state of I'll ask her" and (d) "I am traveling to state of Alaska".In block 1108, the enterprise system 110 running the IVR application 111 queries the predictive language model 9004, configured as a general-purpose language model, using multiple candidate text strings. The candidate text strings include the aforementioned candidate text strings: (a) "I'll ask her", (b) "Alaska", (c) "I'm traveling to state of I'll ask her", and (d) "I am traveling to state of Alaska". The predictive language model 9004 can be configured to return confidence parameter values in response to query data. The predictive language model 9004 can return context confidence parameter values, transcription confidence parameter values, and domain confidence parameter values. The predictive language model 9004 can return confidence parameter values as shown in Table 3. [Table 3]
[0066] Referring to Table 3, which shows the example data, it can be seen that candidate string (d) can be selected because the context confidence parameter value is strong. Furthermore, it can be seen that without candidate text string reinforcement, which adds text to the beginning of candidate text strings, the enterprise system 110 running the IVR application 111 is more likely to select candidate text string (a) "I'll ask her" than candidate text string (b) "Alaska".
[0067] <Conclusion of Example 1> In one embodiment, as mentioned in Example 1, the enterprise system 110 can be configured to perform the augmentation described in block 1107 for each user input voice string, i.e., each candidate text string output by the predictive acoustic model 9002. In another embodiment, the enterprise system 110 can selectively apply augmentation in block 1107 only to voice strings that meet certain conditions, for example. Referring to Table 3, it can be seen that the predictive language model 9004, configured as a general-purpose language model, can output confidence parameter values for the unreinforced candidate text strings (a) and (b) output by the predictive acoustic model 9002 for the user's input speech string. In one embodiment, the enterprise system 110 can be configured not to perform reinforcement in block 1107 for the received speech string based on the condition that the confidence parameter values returned from the predictive language model 9004 for one or more candidate text strings satisfy a threshold. In another embodiment, the enterprise system 110 can be configured not to perform reinforcement in block 1107 for the received speech string based on the condition that the aggregated confidence parameter values returned from the predictive language model 9004 (which can be configured as a common general-purpose language model) for one or more candidate text strings output by the predictive acoustic model 9002 satisfy a threshold of 0.55. Referring to the data exemplified in Table 3, since the aggregated confidence parameter values for either candidate text string (a) or (b) do not satisfy the threshold, reinforcement in block 1107 is actually performed. Referring to the example described based on the data illustrated in Table 3, if any of the unreinforced candidate text strings have a confidence parameter value of 0.55 after aggregation, the enterprise system 110 avoids performing reinforcement in block 1107 and instead uses the unreinforced candidate text column with the highest score.
[0068] Alternative or additional conditions can be used to trigger the execution of augmentation in block 1107. Embodiments of this specification recognize that the speech-to-text conversion may be less reliable if the number of words, phonemes, or both in the speech string is small. According to one embodiment, the enterprise system 110 can be configured not to perform augmentation in block 1107 for an incoming speech string based on the condition that the number of words in one or more candidate text strings returned from the predictive acoustic model 9002 meets a threshold. According to one embodiment, the enterprise system 110 can be configured not to perform augmentation in block 1107 for an incoming speech string based on the condition that the number of phonemes in one or more candidate text strings returned from the predictive acoustic model 9002 meets a threshold. According to one embodiment, the enterprise system 110 can predict for each returned text string language Model 900 4 The enterprise system 110 can be configured to perform augmentation in block 1107 on a received speech string based on the condition that the confidence parameter value returned from is below a threshold. In one embodiment, the enterprise system 110 can be configured to perform augmentation in block 1107 on a received speech string based on the condition that the word count of each returned text string returned from the predictive acoustic model 9002 is below a threshold. In one embodiment, the enterprise system 110 can be configured to perform augmentation in block 1107 on a received speech string based on the condition that the phoneme count of each returned text string returned from the predictive acoustic model 9002 is below a threshold. In one embodiment, the enterprise system 110 can be configured to perform augmentation in block 1107 on a received speech string based on a condition that depends on one or more of the following: (a) the confidence parameter value (output by the predictive language model 9004), (b) the word count of one or more candidate word strings returned from the predictive acoustic model 9002, or (c) the phoneme count of one or more candidate word strings returned from the predictive acoustic model 9002.
[0069] Embodiments of this specification recognize that by configuring the IVR application 111 so that augmentation is performed only selectively in block 1107 based on observed conditions, the operating speed can be improved and computing resources can be saved (which may become more important as the number of concurrently running instances of the IVR application 111 increases).
[0070] In some embodiments, the enterprise system 110 can store and update a control registry in the model area 2121 of the data repository 108, which specifies attributes of one or more predictive language models associated with each conversation stage of the IVR application 111. An example of control registry data is shown in Table 4. [Table 4]
[0071] The features defined herein facilitate the use of a common general-purpose language prediction model across multiple nodes in an IVR session. Using a common general-purpose language prediction model makes it easier to minimize or eliminate the use of conversation-stage specific language prediction models to return predictions relevant to the conversation stage.
[0072] In some embodiments, the IVR application 111 can be configured to first deploy a predictive language model 9004, configured as a general-purpose language model, for each possible conversation stage of an IVR session (for example, mapped to each node of the dialogue decision tree 3002 in Figure 6), and the same common general-purpose language model can be deployed for each stage and node. However, while the system 100 is deploying, the enterprise system 110 can monitor the performance of the general-purpose language model deployed commonly for each node, for example, by inspecting the user voice string input in the next conversation stage. The enterprise system 110 can perform natural language processing to monitor for keywords indicating that the transcription in the previous stage was incorrect (e.g., "I did not ask that question") or for the presence of negative user sentiment below a low threshold. In addition to or instead of this, the enterprise system 110 can monitor performance by monitoring the confidence parameter values output by the predictive language model 9004 for candidate text strings, including augmented candidate text strings as shown in Table 3. The enterprise system 110 can iteratively score each conversation stage (mapped to a node) over a time window of one or more IVR sessions for the same or different users. Then, if the confidence score of a particular conversation stage and dialogue decision tree node falls below a low threshold over a time window of one or more sessions, the enterprise system 110 can deploy a conversation stage-specific language model for that conversation stage and dialogue decision tree node.
[0073] When deploying a new conversation stage-specific language model for a particular conversation stage, the enterprise system 110 can train the new conversation stage-specific language model for that stage using past transcripts returned for that stage, which are stored in the data repository 108 and tagged with metadata indicating the conversation stage and the dialogue decision tree node for that stage. In some use cases, depending on performance monitoring, when a conversation stage-specific model is deployed for a particular conversation stage and dialogue tree node, the enterprise system 110 can mine conversation data from past sessions particularly relevant to that particular conversation stage and dialogue tree node, or the conversation logging data history in the logging area 2123, and selectively use the selectively obtained conversation data to train the newly deployed conversation stage-specific prediction model.
[0074] In another example, the execution of the IVR application 111 by the enterprise system 110 may include (a) using a conversation stage-specific language model to return speech-to-text conversion for a particular conversation stage mapped to a dialog tree node, (b) monitoring the performance of a general-purpose language model for that particular stage across one or more IVR sessions, and (c) decommissioning the conversation stage-specific language model based on the condition that a general-purpose language model, which can be commonly applied to multiple dialog tree nodes, is producing confidence results that exceed a threshold. This decommissioning may involve deleting the model to save computing resources. The same process can then be performed for multiple conversation stages mapped to different IVR dialog tree nodes.
[0075] In some embodiments, the enterprise system 110 can store multiple models, such as a first general-purpose language model and a second conversation stage-specific model, for each conversation stage mapped to a dialogue decision tree node. The enterprise system 110 running the IVR application 111 can query both models using the ensemble model technique in block 1108.
[0076] The enterprise system 110 can be configured to selectively instantiate conversation stage-specific predictive models, for example, only when necessary, during the execution of the IVR application 111. Invoking conversation stage-specific predictive models as needed can be done in response to performance monitoring. For example, audio strings indicating transcription defects or user sentiment may be monitored. Alternatively, or in addition to this, performance monitoring may include monitoring confidence levels returned by the predictive language model 9004.
[0077] In another embodiment, the enterprise system 110 may be configured such that, when a conversation stage-specific predictive model is instantiated and deployed as needed, limited to a specific conversation stage mapped to a particular dialogue tree node, performance monitoring is performed, and if it is determined that the common general-purpose language predictive model can provide satisfactory performance, the conversation stage service provision for that particular conversation stage can be reverted back to the service provision by the common general-purpose language predictive model.
[0078] In one use case, the enterprise system 110 can be configured so that, in the initial session of the IVR application 111, all conversation stages of the IVR session are serviced by a common general-purpose language prediction model for each node. When a user voice string is received for a particular conversation stage, the performance of the common general-purpose language prediction model can be monitored, that is, for the current session or over a time window of multiple sessions. Monitoring the performance of the prediction language model 9004 may include monitoring the confidence parameter values output by the prediction language model 9004, configured as a common general-purpose language model, for candidate text strings, including augmented candidate text strings as shown in Table 3. If the monitoring results indicate that the performance for a particular conversation stage is unsatisfactory according to predetermined criteria, a conversation stage-specific prediction model can be instantiated for that particular conversation stage.
[0079] After instantiating a conversation-stage specific predictive model for a particular conversation stage, the enterprise system 110 can continue to monitor the output of a common general-purpose language model for that particular conversation stage, for example, over a time window of one or more sessions, although the common general-purpose language model may be deactivated for the purpose of returning transcripts (and may be activated in an ensemble model configuration). Monitoring the performance of the predictive language model 9004 may include monitoring the confidence parameter values output by the predictive language model 9004, configured as a common general-purpose language model, for candidate text strings, including augmented candidate text strings as shown in Table 3. The enterprise system 110 can be configured to return service for a particular conversation stage if it determines that the performance of the general-purpose language model (e.g., by monitoring the confidence parameter values) is satisfactory.
[0080] The enterprise system 110 can be configured to deactivate a conversation-stage specific prediction model for a particular conversation stage and return the service provision for that conversation stage to the general-purpose common language prediction model if performance monitoring of the common general language prediction model indicates that the common general language prediction model is satisfactory for that conversation stage. This reduces the computing resources spent on collecting, organizing, and applying training data for specific conversation-stage specific prediction models.
[0081] In some scenarios, the performance level of a common general-purpose language prediction model may improve over time after the initial execution of the IVR application 111. If the common general-purpose language prediction model is trained using IVR application-specific training data (e.g., conversation log data) obtained from multiple IVR sessions of a particular IVR application, the performance of the prediction language model 9004 configured as a common general-purpose language prediction model may improve across multiple instances of the IVR application 111 executed to run multiple IVR sessions (particularly by the augmentation process 116 that enhances the performance of such a common general-purpose language prediction model).
[0082] In some use cases, the conversation stage-specific predictive model may outperform the common general-purpose language predictive model for a certain period after the initial execution of the IVR application for specific conversation stages. However, after multiple session histories in the IVR application 111, the common general-purpose language predictive model may become smarter with more training and thus become a better-performing option than the conversation stage predictive model, in addition to achieving the complexity reduction and computing resource advantages described herein. The embodiments herein facilitate switching between the common general-purpose language predictive model and the conversation stage-specific predictive model with respect to service delivery for specific conversation stages through performance monitoring.
[0083] Referring to Table 4, various statuses are possible. For conversation stage A001, the common general-purpose language prediction model is active, and the conversation stage-specific language prediction model has never been instantiated for this conversation stage. Referring to conversation stage A003, the conversation stage-specific prediction model has been instantiated and is active, and the common general-purpose language prediction model is inactive for the purpose of returning a transcript and performing IVR decisions. However, because there is a possibility of switching to the common general-purpose language prediction model to provide service for this particular conversation stage A003, it continues to generate output for the purpose of monitoring its performance. Referring to conversation stage A006, both the common general language prediction model and the conversation stage-specific prediction model are active.
[0084] Further embodiments of the embodiments described herein will be described with reference to the flowchart in Figure 7. Referring to block 7102, the predictive acoustic model 9002 can be trained, for example, using voice data of a user of system 100 or another user. Referring to block 7104, the enterprise system 110 running the IVR application 111 can determine prompt data to send to the user, for example, depending on a derived intent, dialogue, or entity. In block 7106A, the enterprise system 110 can put text-based prompt data into text-to-speech conversion to send synthesized voice prompt data to the user, and in block 7106C, the user can send voice string data to the enterprise system 110 in response to the synthesized voice prompt data. The enterprise system 110 running the IVR application 111 can query the predictive acoustic model 9002 using the received speech string data and return candidate text strings associated with the speech string data. In block 7106B, the enterprise system 110 running the IVR application 111 can store the prompt data as context in the data repository 108. In block 7106D, the enterprise system 110 running the IVR application 111 can subject the IVR prompt data to text analysis and extract data that characterizes this text-based prompt data. In block 7110, the enterprise system 110 running the IVR application 111 can determine a prefix text to add to the candidate text string output by the predictive acoustic model 9002 in block 7102. In block 7114, the enterprise system 110 running the IVR application 111 can add the prefix text determined in block 7110 to the candidate text string output by the predictive acoustic model 9002 in block 7102. In block 7114, the enterprise system 110 running the IVR application 111 can query the predictive language model 9004 using multiple candidate text strings. These multiple text strings may include candidate text strings output by the predictive acoustic model 9002, and enhanced versions of these candidate text strings modified by adding prefix text determined in block 7110. In block 7114, the predictive language model 9004 can return one or more sets of confidence parameter values in response to the query. In block 7116, the enterprise system 110 running the IVR application 111 can examine the returned confidence values and select the candidate text string with the highest score as the returned transcript.Then, in the next iteration of block 7104, the enterprise system 110 running the IVR application 111 can use this returned transcript to determine the next prompt data to present to the user.
[0085] Referring to block 7102, the predictive acoustic model 9002 can be trained. (a) The developer user may provide the predictive acoustic model 9002 by training a predictive acoustic model, or may provide the predictive acoustic model 9002 using a commercially available acoustic model (COTS). Referring to block 7104, the voice model can be used as the conversational context. A commercially available voice chat solution (e.g., Watson Assistant™ for voice interaction) can be deployed to facilitate the IVR session. Referring to block 7106A, the system can prompt the user and initiate an STT recognition request. Referring to block 7106B, the last prompt can be saved as the context. Referring to block 7106C, an existing API (e.g., a WebSocket connection) can be used for recognition. Referring to block 7106D, the context is passed along with the user voice. Referring to block 7108, the system splits the context string into the minimum useful components. The context string can be sent through typical grammatical analysis (sentence, part of speech). If the context is a single short sentence, splitting may not be performed. If the context contains multiple sentences, the last sentence can be used. If the context contains multiple clauses within a single sentence, the last clause can be used. Referring to block 7110, grammar tree manipulation can be performed on the context string to make it suitable for prefixing. Various possible transformations include: (a) replacing pronouns: e.g., "your" → "my", "you" → "I", (b) changing a question into a statement: "What state are you traveling to" → "I am traveling to the state", (c) truncating a statement using a template, for example: "Please state your destination" → "My destination is", (d) if the operation cannot be performed, the context segment can be passed as is or not used at all. Referring to block 7112, the predictive acoustic model 9002, configured as an STT-based model, can transcribe speech into hypotheses of candidate text strings.Referring to block 7114, the hypothesis text is (a) first evaluated as is, and (b) then evaluated together with the prefixed contextual text. Each evaluated hypothesis can be scored by "context-match" confidence and transcription confidence. For example, (i) "I'm traveling to the state of Alaska" has high context-matching confidence and moderate transcription confidence, (ii) "I'm traveling to the state I'll ask her" has low context-matching confidence and high transcription confidence (the phrase "to the state I'll ask her" has low context-matching confidence (grammatically irregular)), and (iii) "I'll ask her" has moderate context-matching confidence and high transcription confidence. Context confidence can be derived from the appropriateness of the grammatical structure (part of speech, grammar). Domain confidence can also be scored (whether the transcription contains domain-specific words or phrases). The transcription score can be a weighted average or geometric mean of context confidence, domain confidence, and transcription confidence. The transcription candidate with the highest transcription score can be returned.
[0086] Embodiments of this specification recognize that speech-to-text services can be trained for specific tasks and perform very well when specifically trained. Embodiments of this specification recognize that multiple models can be used and that an orchestrator can be provided to select the appropriate model for the appropriate task. Embodiments of this specification recognize that a multi-model approach can impose a significant development burden on developers and may incur computing resource costs, for example, in training multiple models.
[0087] For example, a chatbot (e.g., VA) could present the following prompt: "What state are you traveling to?" and the user responds: "Alaska." A general-purpose language model that has not undergone custom training might transcribe "I'll ask her." A conversational stage-specific language model using a specific grammar or language model might transcribe "Alaska." Embodiments of this specification recognize that using multiple models can impose development and computing resource burdens, such as the burden of storing and applying training data for multiple models. Embodiments of this specification can facilitate the use of general-purpose language models in chatbots (e.g., IVR applications).
[0088] Embodiments of this specification can facilitate the use of a common, single, general-purpose model for multiple conversation stages in an IVR application, and in some use cases, for all conversation stages. Such an architecture facilitates simplified training. Training can be achieved by using text data collected from conversation logs with a single general-purpose language model. Embodiments of this specification recognize that AI services (e.g., chat, transcription) have been developed as independent microservices rather than to serve a consistent common goal.
[0089] Embodiments of this specification can facilitate the use of a common, single, general-purpose model for each of multiple conversational stages in an IVR application, and in some use cases, for each of all conversational stages. Embodiments of this specification can supplement a “speech recognition” application program interface (API) with a context message containing prompting text, which is used to supplement recognized speech string data. The speech recognition system can use the above-described context message during the hypothesis text evaluation stage to ensure that the transcribed response makes sense in the context of the initial request. The context message can be converted into a prefix text string combined with an utterance hypothesis. In the example described, the context message sent to the user may be “What state are you traveling to?”. During evaluation, this message can be converted to “I'm traveling to the state of 'The hypothesis'.” In this way, “I'm traveling to the state of Alaska” can be selected and “I'm traveling to the state I'll ask her” can be rejected.
[0090] Embodiments of this specification recognize that converting chatbot prompt data into prefixed text and adding the prefixed text to candidate text strings associated with input speech string data facilitates the use of a single common general-purpose language model for multiple conversational stages of an IVR stage (e.g., all conversational stages). In one embodiment, longer text strings with added words are more likely to match the text string history of past training data used to train the general-purpose language model. The common general-purpose language model can be used without training or can be trained by a simplified training procedure. The simplified training procedure may include simply extracting relevant data from a conversational log associated with the entire IVR session and applying the entire conversational log to the common model, rather than individually storing and applying special conversational data associated with each specific conversational stage mapped to different IVR dialogue decision tree nodes.
[0091] Certain embodiments of this specification can provide a variety of technical computing advantages and practical applications, including computing advantages for addressing problems arising from the domain of computer systems. For example, embodiments of this specification can provide a machine learning model training procedure for use in an IVR application having multiple communication stages, each of which is mappable to a node in a dialogue decision tree. Embodiments of this specification can feature a single common predictive model provided by a general-purpose language model for use in the first to nth conversation stages of an IVR session. By using a common predictive model for the first to nth conversation stages, the need to organize and maintain separate training data for multiple conversation stage-specific language models corresponding to each conversation stage of an IVR application is reduced. The machine learning training procedure of this specification can reduce the complexity of the design for developers and reduce the utilization of computing resources, for example, by reducing the work of maintaining and applying training data. Machine learning procedures for use in IVR applications as defined herein may include simply extracting general conversation log data from conversation logs of completed IVR sessions and applying the conversation log data as training data without performing computationally intensive tasks such as separately organizing and storing separate training data for different conversation stages or separately training multiple different conversation form-specific language models. Embodiments herein may include inspecting text-based data that defines prompt data presented to a user by a chatbot such as a VA. Such inspection may include subjecting the text data to text parsing (e.g., grammatical parsing to perform part-of-speech tagging, which tags the text-based data that defines the prompt data). Based on the inspection results, for example, the IVR application may use the part-of-speech tags to transform the text-based data that defines the prompt data and provide the transformed text.Candidate text strings associated with user voice strings transmitted in response to prompt data can be augmented using the converted text to generate augmented candidate text strings. The augmented candidate text strings can be evaluated using a predictive model configured as a language model. In one embodiment, the language model can be a general-purpose language model deployed in common across first to n conversational stages, either without the need for associated training procedures or through the lightweight training procedures described herein, which reduce development complexity and computing resource utilization. Various decision data structures can be used to drive artificial intelligence (AI) decision-making. The decision data structures defined herein can be updated by machine learning, improving accuracy and reliability over time and iteratively without resource-intensive rule-intensive processing. Machine learning processes can be implemented to improve accuracy and reduce reliance on rule-based criteria, thus reducing computational overhead. To improve computational accuracy, each embodiment may feature a computing platform that exists solely in the realm of computer networks, such as an artificial intelligence platform or a machine learning platform. Embodiments of this specification may utilize data structuring processes, such as processes for converting unstructured data into a format optimized for computer processing. Embodiments of this specification may include an artificial intelligence processing platform featuring an improved process for converting unstructured data into a structured format that enables computer-based analysis and decision-making. Embodiments of this specification may include both specific configurations for collecting rich data into a data repository and additional specific configurations for updating such data and using such data to drive artificial intelligence decision-making. Specific embodiments can be implemented using various types of cloud platforms / data centers, including Software-as-a-Service (SaaS), Platform-as-a-Service (PaaS), Database-as-a-Service (DBaaS), and combinations thereof, based on the type of subscription.
[0092] Figures 8-10 illustrate various forms of computing, including computer systems and cloud computing, relating to one or more embodiments as defined herein.
[0093] While this disclosure includes a detailed description of cloud computing, it should be understood that the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in combination with any other type of computing environment that is currently known or may be developed in the future.
[0094] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and deployed with minimal administrative effort or interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0095] The characteristics are as follows:
[0096] On-demand self-service: Cloud consumers can unilaterally prepare computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.
[0097] Broad network access: Computing power is available over the network and accessible through standard mechanisms. This facilitates utilization by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).
[0098] Resource pooling: A provider's computing resources are pooled and delivered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated as needed. Generally, consumers have a sense of location independence because they do not manage or know the exact location of the resources provided. However, consumers may be able to identify the location at a higher level of abstraction (e.g., country, state, data center).
[0099] Rapid Elasticity: Computing power can be prepared quickly and flexibly, allowing it to scale out automatically and immediately, and to be quickly released and scale in immediately. To consumers, the computing power available for preparation often appears unlimited and can be purchased in any quantity at any time.
[0100] Service Measurement: Cloud systems leverage metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0101] The service model is as follows:
[0102] Software as a Service (SaaS): The functionality offered to consumers is the ability to use the provider's applications running on a cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for configuring a limited number of user-specific applications.
[0103] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired to cloud infrastructure using programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, and storage, but they can control the deployed applications and, in some cases, the configuration of their hosting environment.
[0104] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provision of processors, storage, networking, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they can control the operating system, storage, and deployed applications, and in some cases, partially control certain network components (e.g., host firewalls).
[0105] The deployment model is as follows:
[0106] Private Cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by that organization or a third party and can reside on-premises or off-premises.
[0107] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can reside on-premises or off-premises.
[0108] Public Cloud: This cloud infrastructure is provided to a large number of people or large industry groups and is owned by organizations that sell cloud services.
[0109] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While maintaining the unique entities of each model, they are bound together by standards or individual technologies to achieve data and application portability (e.g., cloud bursting for load balancing across clouds).
[0110] Cloud computing environments are service-oriented environments that emphasize statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure, which includes a network of interconnected nodes.
[0111] Figure 8 shows a schematic diagram of an example of a computing node. Note that computing node 10 is merely an example of a computing node suitable for use as a cloud computing node, and is not intended to imply any limitation on the scope or functionality of the embodiments of the present invention described herein. In any case, computing node 10 can implement, perform, or both of the functions described herein. Computing node 10 can be implemented as a cloud computing node in a cloud computing environment, or as a computing node in a computing environment other than a cloud computing environment.
[0112] A computer system 12 resides within the computing node 10. The computer system / server 12 can operate with many other general-purpose or dedicated computing system environments or configurations. Examples of well-known computing systems, environments, or configurations or combinations suitable for use with the computer system 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of these systems or devices.
[0113] Computer system 12 can be described in general terms with computer system executable instructions, such as program processes executed by the computer system. Generally, a program process can include routines, programs, objects, components, logic, data structures, etc., that perform a specific task or implement a specific data type. Computer system 12 can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program processes can be stored in both local and remote computer system storage media, including memory storage devices.
[0114] As shown in Figure 8, the computer system 12 within the computing node 10 is shown as a computing device. The components of the computer system 12 are not particularly limited, but may include one or more processors 16, system memory 28, and a bus 18 connecting various system components, including the system memory 28, to the processors 16. In one embodiment, the computing node 10 is a computing node in a non-cloud computing environment. In one embodiment, the computing node 10 is a computing node in a cloud computing environment as defined herein in relation to Figures 9-10.
[0115] Bus 18 represents one or more of several types of bus structures, including memory buses or memory controllers using any of the various bus architectures, peripheral buses, accelerated graphics ports (AGP), and processor or local buses. As a non-exclusive example, such architectures include the Industry Standard Architecture (ISA) bus, Microchannel Architecture (MCA) bus, Expansion ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0116] The computer system 12 generally includes various computer system-readable media. Such media can be any available media accessible to the computer system 12 and include both volatile and non-volatile media, and both removable and non-removable media.
[0117] The system memory 28 may include computer system-readable media as volatile memory (such as RAM 30 or cache memory 32 or both). The computer system 12 may further include other removable / non-removable volatile / non-volatile computer system-readable media. As an example only, the storage system 34 may be provided for reading and writing to a non-removable non-volatile magnetic medium (not shown; commonly referred to as a “hard drive”). Also, although not shown, a magnetic disk drive for reading and writing to removable non-volatile magnetic disks (e.g., “floppy disks”) and an optical disk drive for reading and writing to removable non-volatile optical disks (such as CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these examples, each may be connected to the bus 18 by one or more data medium interfaces. As further illustrated and described below, the memory 28 may include at least one program product having a set (e.g., at least one) of program processes configured to perform the functions of embodiments of the present invention.
[0118] As a non-limiting example, one or more programs 40 having a set (at least one) of program processes 42 can be stored in memory 28, as well as the operating system, one or more application programs, other program processes, and program data. One or more programs 40 including program processes 42 can generally perform the functions defined herein. In one embodiment, the enterprise system 110 can include one or more computing nodes 10 and one or more programs 40 for performing the functions described with reference to the enterprise system 110 and the IVR application 111, as shown in the flowchart of Figure 4. In one embodiment, the enterprise system 110 can include one or more computing nodes 10 and one or more programs 40 for performing the functions described with reference to the enterprise system 110 and the IVR application 111, as shown in the flowchart of Figure 7. In one embodiment, one or more UE devices among a plurality of UE devices can include one or more computing nodes 10 and one or more programs 40 for performing the functions described with reference to the UE devices, as shown in the flowchart of Figure 4. In one embodiment, one or more of the multiple UE devices may include one or more computing nodes 10 and one or more programs 40 for performing the functions described with reference to the UE devices, as shown in the flowchart of Figure 7. In one embodiment, the computing node-based system and devices shown in Figure 1 may include one or more programs for performing the functions described with reference to such computing node-based systems and devices.
[0119] Furthermore, the computer system 12 can communicate with one or more external devices 14 such as a keyboard, pointing device, and display 24, one or more devices that enable interaction between the user and the computer system 12, or any device that enables communication between the computer system 12 and one or more other computing devices (e.g., a network card or modem), or a combination thereof. Such communication can be performed via the input / output (I / O) interface 22. In addition, the computer system 12 can communicate with one or more networks (such as a local area network (LAN), a general-purpose wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof) via the network adapter 20. As shown in the figure, the network adapter 20 communicates with other components of the computer system 12 via the bus 18. Although not shown in the figure, other hardware components, software components, or both can be used in conjunction with the computer system 12. Examples of such components, but not limited to, include microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems. In addition to having an external device 14 and a display 24 that can be configured to provide user interface functions, or alternatively, the computing node 10 may, in one embodiment, include a display 25 connected to the bus 18, or other output devices such as one or more audio output devices connected to the bus 18, or both. In one embodiment, the display 25 may be configured as a touchscreen display and configured to provide user interface functions. For example, the display 25 may facilitate virtual keyboard functions and total data input. Also, in one embodiment, the computer system 12 may include one or more sensor devices 27 connected to the bus 18.Alternatively, one or more sensor devices 27 may be connected via an I / O interface 22. In one embodiment, one or more sensor devices 27 may include a Global Positioning Sensor (GPS) device and be configured to provide the location of the computing node 10. Alternatively, or in addition to this, one or more sensor devices 27 may in one embodiment include, for example, one or more of a camera, a gyroscope, a temperature sensor, a humidity sensor, a pulse sensor, a blood pressure (bp) sensor, or an audio input device. The computer system 12 may include one or more network adapters 20. In Figure 9, the computing node 10 is shown as being implemented within a cloud computing environment and is therefore referred to as a cloud computing node in relation to Figure 9.
[0120] Here, Figure 9 shows an exemplary cloud computing environment 50. As shown in the figure, the cloud computing environment 50 includes one or more cloud computing nodes 10. Local computer devices used by cloud consumers (e.g., PDAs or mobile phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof) can communicate with these nodes. The nodes 10 can communicate with each other. The nodes 10 can be grouped physically or virtually (not shown) in one or more networks, such as the private, community, public, or hybrid clouds or a combination thereof. This allows the cloud computing environment 50 to provide infrastructure, platforms, or software as a service, or a combination thereof, without requiring cloud consumers to maintain resources on their local computer devices. Note that the types of computer devices 54A-N shown in Figure 9 are merely examples, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.
[0121] Here, Figure 10 shows a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 9). It should be understood that the components, layers, and functions shown in Figure 10 are merely illustrative, and the embodiments of the present invention are not limited to these. As illustrated, the following layers and corresponding functions are provided.
[0122] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, a reduced instruction set computer (RISC) architecture-based server 62, server 63, blade server 64, storage 65, and a network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0123] The virtualization layer 70 provides an abstraction layer. From this layer, virtual entities such as virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75 can be provided.
[0124] As an example, the management layer 80 can provide the following functions: Resource preparation 81 enables the dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 82 enables cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only protection of data and other resources but also identification and verification of cloud consumers and tasks. The user portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 enables the allocation and management of cloud computing resources to ensure that requested service levels are met. Service Level Agreement (SLA) planning and execution 85 enables the pre-arrangement and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.
[0125] Workload layer 90 provides examples of functions available in a cloud computing environment. Examples of workloads and functions available from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and a processing component 96 for returning transcribed text related to speech string data, as defined herein. The processing component 96 can be implemented using one or more programs 40 as described in Figure 8.
[0126] The present invention may be a system, method, or computer program product or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0127] A computer-readable storage medium can be a tangible device capable of holding and storing instructions used by an instruction execution device. Examples of computer-readable storage media include electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or appropriate combinations thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, erasable programmable ROM (EPROM or flash memory), static random access memory (SRAM), CD-ROMs, DVDs, memory sticks, floppy disks, mechanically encoded devices with instructions recorded on punch cards or grooved raised structures, and appropriate combinations thereof. The computer-readable storage media used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical pulses passing through optical fiber cables), or electrical signals transmitted through wires.
[0128] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing device / processing device. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, LAN, WAN, or wireless network, or a combination thereof). The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives computer-readable program instructions from the network and transfers them for storage in a computer-readable storage medium within each computing device / processing device.
[0129] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed as a standalone software package, either entirely on the user's computer or partially on the user's computer. Alternatively, they can be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including LANs and WANs, or it may be connected to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), and programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of computer-readable program instructions in order to customize the electronic circuits for the purpose of performing aspects of the present invention.
[0130] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. Each block in a flowchart or block diagram, or both, and combinations of blocks in a flowchart or block diagram, or both, are executable by computer-readable program instructions.
[0131] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine. This creates a means for these instructions, executed via such a computer or other programmable data processing device processor, to perform functions / operations identified in one or more blocks in a flowchart or block diagram, or both. These computer-readable program instructions can further be stored in a computer-readable storage medium that can be instructed to function in a particular manner to a computer, programmable data processing device, or other device, or a combination thereof. Thus, the computer-readable storage medium containing the instructions constitutes a product containing instructions for performing functions / operations identified in one or more blocks in a flowchart or block diagram, or both.
[0132] Alternatively, a computer execution process may be generated by loading computer-readable program instructions into a computer, another programmable device, or other device, and having a series of operational steps executed on that computer, other programmable device, or other device. This ensures that the instructions executed on the computer, other programmable device, or other device perform functions / operations identified by one or more blocks in a flowchart, block diagram, or both.
[0133] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions for performing a specific logical function. In some other implementations, the functions shown within a block may be executed in an order different from the order shown in each diagram. For example, depending on the functions involved, two consecutively shown blocks may actually be achieved as a single process, executed simultaneously or nearly simultaneously, executed in a manner that partially or entirely overlaps in time, or the blocks may be executed in reverse order. Each block in a block diagram or flowchart or both, and combinations of multiple blocks in a block diagram or flowchart or both, are executable by a dedicated hardware-based system that performs a specific function or operation, or executes a combination of dedicated hardware and computer instructions.
[0134] The terms used herein are intended solely to describe specific embodiments and are not intended to limit them. In this specification, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context makes it clear otherwise. Furthermore, the terms “comprise” (and any form of “comprise,” such as “comprises,” “comprising,” etc.), “have” (and any form of “have,” such as “has,” “having,” etc.), “include” (and any form of “include,” such as “includes,” “including,” etc.), and “contain” (and any form of “contains,” “containing,” etc.) are understood to be open-ended linking verbs. As a result, a method or device that “comprises,” “has,” “includes,” or “contains” one or more steps or elements has, but is not limited to having only, those one or more steps or elements. Similarly, an element of a step or device of a method that “comprises,” “has,” “includes,” or “contains” one or more features has, but is not limited to having only, those one or more features. The term “based on” as used herein encompasses relationships where elements are partially based and relationships where elements are fully based. A method, product, and system described as having a certain number of elements may be implemented using fewer or more elements than that specific number. Furthermore, a device or structure configured in a particular manner may be configured in at least that particular manner, but may also be configured in other manners not described.
[0135] Numerical and other values described herein, whether explicitly stated or derived intrinsically from the descriptions herein, are intended to be qualified by the term “about.” The term “about” as used herein is, but is not limited to, defining numerical boundaries of the qualified value to include the numerical values and ranges below that of the numerical value qualified by the term. In other words, numerical values may include actual values explicitly stated, and other values that are decimal, fractional, or other multiples of actual values as suggested, stated, or both implied in this disclosure.
[0136] The corresponding structures, materials, actions, and equivalents of all means-plus-function elements or step-plus-function elements in the following claims are intended to include any structures, materials, or actions for performing a function in combination with other specifically claimed elements, where applicable. The descriptions herein are presented for illustrative and explanatory purposes only and are not intended to be exhaustive or limit to the disclosed forms. Many changes and modifications will be apparent to those skilled in the art without departing from the scope of this disclosure. These embodiments have been selected and described to best illustrate the principles and practical applications of one or more embodiments described herein and to enable others skilled in the art to understand one or more embodiments described herein with various modifications suitable for their particular intended use.
Claims
1. In the execution of an interactive voice response (IVR) session, the prompt data to be presented to the user is determined, and text-based data defining that prompt data is stored in a data repository. Presenting the aforementioned prompt data to the user, In response to the prompt data, the system receives the voice string data returned by the user, To generate a plurality of candidate text strings related to the returned voice string of the user, Examining the text-based data that defines the prompt data, which includes performing text parsing using natural language processing on the text-based data; The present invention provides a plurality of augmented candidate text strings related to the returned speech string data, wherein the augmentation includes transforming the text-based data that defines the prompt data and adding the transformed text-based data to the plurality of candidate text strings. Evaluating each of the multiple augmented candidate text strings associated with the returned speech string data, Selecting one of the aforementioned augmented candidate text strings as the returned transcript related to the returned audio string data, Computer implementation methods, including those mentioned above.
2. The computer implementation method according to claim 1, wherein the evaluation includes querying a prediction model provided by a general-purpose language model using each of the plurality of augmented candidate text strings associated with the returned speech string data, and examining the returned confidence parameter value obtained as a result of said query.
3. The computer implementation method according to claim 1, comprising: applying the plurality of candidate text strings as query data to a predictive language model used to perform the evaluation, prior to the inspection and the augmentation; verifying the performance of the predictive language model in response to the application of the plurality of candidate text strings as query data to the predictive language model; determining, based on the verification, that the predictive language model does not perform satisfactorily with respect to the user's returned speech string data; and selectively performing the inspection and the augmentation in response to the determination.
4. The computer implementation method according to claim 1, comprising: verifying the performance of a predictive language model used to perform the evaluation using the plurality of candidate text strings before the inspection and the augmentation; determining, based on the verification, that the predictive language model does not work satisfactorily with respect to the user's returned speech string data; and selectively performing the inspection and the augmentation in accordance with the determination.
5. The computer implementation method according to claim 1, wherein the augmentation includes identifying a specific text string in the text-based data that defines the prompt data, which is referenced in a mapping data structure stored in a data repository, the mapping data structure maps the text string to a converted text string, and the augmentation includes using a specific converted text string associated with the specific text string in the mapping data structure.
6. The computer implementation method according to claim 1, wherein the inspection includes subjecting the text-based data defining the prompt data to natural language processing for assigning part-of-speech tags to the text-based data, and the augmentation includes identifying a specific text string in the text-based data defining the prompt data that matches a template text string in a mapping data structure stored in a data repository, the template text string stored in the data repository includes one or more terms expressed in wildcard form as parts of speech.
7. The computer implementation method according to claim 1, wherein the presentation includes presenting the prompt data to the user in synthesized speech using text-to-speech conversion.
8. The computer implementation method according to claim 1, wherein the inspection includes subjecting the text-based data defining the prompt data to natural language processing for assigning part-of-speech tags to each term in the text-based data, and transforming the text-based data defining the prompt data using the part-of-speech tags.
9. The computer implementation method according to claim 1, wherein the augmentation includes converting the text-based data defining the prompt data to provide converted prompt data, and prepending the converted prompt data to the beginning of each of the candidate text strings of the plurality of candidate text strings.
10. The evaluation described above includes querying a predictive model provided by a specific general-purpose language model using each of the multiple augmented candidate text strings associated with the returned speech string data, and examining the returned confidence parameter values obtained as a result of said query. The prompt data presented to the user includes determining the prompt data for the first conversation stage of the IVR session. The method includes determining second prompt data for a second conversation stage of the IVR session in response to the returned transcript, and storing second text-based data defining the second prompt data in the data repository. The computer implementation method according to claim 1, further comprising: presenting the second prompt data to the user; receiving a second returned speech string data from the user in response to the second prompt data; generating a second plurality of candidate text strings related to the second returned speech string of the user; performing an inspection of text-based data defining the second prompt data; augmenting the second plurality of candidate text strings according to the results of the inspection to provide a second plurality of augmented candidate text strings related to the second returned speech string data; querying a specific general-purpose language model using each of the second plurality of augmented candidate text strings according to the results of the inspection to return confidence data; and selecting one of the second plurality of augmented candidate text strings as a second returned transcript related to the second user speech string data.
11. The method includes maintaining a control registry that defines status information for a predictive language model associated with each conversation stage of an IVR session, including a specific conversation stage associated with the prompt data, The computer implementation method according to claim 1, comprising: analyzing status data of the control registry associated with the particular conversation stage to identify a particular predictive language model referenced in the control registry; and performing the evaluation using the particular predictive language model referenced in the control registry.
12. The method includes maintaining a control registry that defines status information for a predictive language model associated with each conversation stage of an IVR session, including a specific conversation stage associated with the prompt data, The computer implementation method according to claim 1, further comprising: monitoring the performance of a common general-purpose language model used to return a transcript related to a particular conversation stage in a subsequent session of an IVR application executing an IVR session; and, if such monitoring of the performance of the common general-purpose language model indicates that the common general-purpose language model does not produce a satisfactory transcript, instantiating a conversation stage-specific language model to return a transcript related to the particular conversation stage.
13. The method includes maintaining a control registry that defines status information for a predictive language model associated with each conversation stage of an IVR session, including a specific conversation stage associated with the prompt data, The method includes monitoring the performance of a common general-purpose language model used to return a transcript related to a particular conversation stage in a subsequent session of an IVR application that runs an IVR session, and, if such monitoring of the performance of the common general-purpose language model indicates that the common general-purpose language model does not produce a satisfactory transcript, instantiating a conversation stage-specific language model to return a transcript related to that particular conversation stage. The computer implementation method according to claim 1, further comprising: performing continuous monitoring of the performance of the common general-purpose language model with respect to a particular conversation stage after instantiation; and, if the continuous monitoring determines that the common general-purpose language model outputs a satisfactory transcript for the particular conversation stage, returning the service provision for the particular conversation stage to the common general-purpose language model.
14. The evaluation described above includes querying a predictive model provided by a specific general-purpose language model using each of the multiple augmented candidate text strings associated with the returned speech string data, and examining the returned confidence parameter values obtained as a result of said query. The prompt data presented to the user includes determining the prompt data for the first conversation stage of the IVR session. The method includes determining second prompt data for a second conversation stage of the IVR session in response to the returned transcript, and storing second text-based data defining the second prompt data in the data repository. The method further includes presenting the second prompt data to the user; receiving a second returned speech string data from the user in response to the second prompt data; generating a second set of candidate text strings related to the second returned speech string of the user; performing an inspection of text-based data defining the second prompt data; augmenting the second set of candidate text strings according to the results of the inspection to provide a second set of augmented candidate text strings related to the second returned speech string data; querying a specific general-purpose language model using each of the second set of augmented candidate text strings according to the results of the inspection to return confidence data; and selecting one of the second set of augmented candidate text strings as a second returned transcript related to the second user speech string data. The computer implementation method according to claim 1, wherein the IVR application executing the IVR session queries the specific general-purpose language model in common to return the returned transcript relating to the first conversation stage of the IVR application and to return the second returned transcript relating to the second conversation stage of the IVR application.
15. The computer implementation method according to claim 1, wherein generating the plurality of candidate text strings related to the returned speech string of the user includes querying a predictive acoustic model.
16. The inspection includes providing the text-based data defining the prompt data for part-of-speech tagging and providing part-of-speech tags associated with the text-based data defining the prompt data, The computer implementation method according to claim 1, wherein the reinforcement includes converting the text-based data defining the prompt data using the part-of-speech tag to provide a converted text string, and adding the converted text string to the beginning of a text string among a plurality of candidate text strings.
17. The presenting includes presenting the prompt data to the user in synthesized speech using text-to-speech conversion. The computer implementation method according to claim 1, wherein the evaluation includes querying a prediction model provided by a general-purpose language model using each of the plurality of augmented candidate text strings associated with the returned speech string data.
18. The presenting includes presenting the prompt data to the user in synthesized speech using text-to-speech conversion. Generating the plurality of candidate text strings related to the returned speech string of the user includes querying a predictive acoustic model, The computer implementation method according to claim 1, wherein the evaluation includes querying a prediction model provided by a general-purpose language model using each of the plurality of augmented candidate text strings associated with the returned speech string data, and examining the returned confidence parameter value obtained as a result of said query.
19. The computer implementation method according to claim 1, wherein the augmentation includes identifying a specific text string in the text-based data that defines the prompt data.
20. The computer implementation method according to claim 1, wherein the enhancement includes identifying a specific text string in the text-based data that defines the prompt data, and converting the specific text string into a converted string using data stored in a data repository.
21. The computer implementation method according to claim 1, wherein the augmentation includes prepending the converted prompt data to the beginning of each of the candidate text strings of the plurality of candidate text strings.
22. The computer implementation method according to claim 1, wherein the inspection includes subjecting the text-based data defining the prompt data to natural language processing for assigning part-of-speech tags to the text-based data.
23. The computer implementation method according to claim 1, wherein the inspection includes providing the text-based data defining the prompt data for part-of-speech tagging to provide part-of-speech tags associated with the text-based data defining the prompt data.
24. A computer program for performing a method, wherein the method is: In the execution of an interactive voice response (IVR) session, the prompt data to be presented to the user is determined, and text-based data defining that prompt data is stored in a data repository. Presenting the aforementioned prompt data to the user, In response to the prompt data, the system receives the voice string data returned by the user, To generate a plurality of candidate text strings related to the returned voice string of the user, Examining the text-based data that defines the prompt data, which includes performing text parsing using natural language processing on the text-based data; The present invention provides a plurality of augmented candidate text strings related to the returned speech string data, wherein the augmentation includes transforming the text-based data that defines the prompt data and adding the transformed text-based data to the plurality of candidate text strings. Evaluating each of the multiple augmented candidate text strings associated with the returned speech string data, Selecting one of the aforementioned augmented candidate text strings as the returned transcript related to the returned audio string data, A computer program that includes [this].
25. Memory and At least one processor that communicates with the memory, A system comprising: a program instruction executable by one or more processors via memory for performing a method, wherein the method is In the execution of an interactive voice response (IVR) session, the prompt data to be presented to the user is determined, and text-based data defining that prompt data is stored in a data repository. Presenting the aforementioned prompt data to the user, In response to the prompt data, the system receives the voice string data returned by the user, To generate a plurality of candidate text strings related to the returned voice string of the user, Examining the text-based data that defines the prompt data, which includes performing text parsing using natural language processing on the text-based data; The present invention provides a plurality of augmented candidate text strings related to the returned speech string data, wherein the augmentation includes transforming the text-based data that defines the prompt data and adding the transformed text-based data to the plurality of candidate text strings. Evaluating each of the multiple augmented candidate text strings associated with the returned speech string data, Selecting one of the aforementioned augmented candidate text strings as the returned transcript related to the returned audio string data, A system that includes this.
Citation Information
Patent Citations
Multi-slot dialogue system and method
JP2008506156A
Bidirectional probabilistic natural language rewriting and selection
JP2019070799A
Voice recognition device, voice recognition method and voice recognition program
WO2008004666A1