Systems and methods for constructing a state-specific multi-turn contextual language understanding system
By inferring or configuring semantic patterns and rules that depend on the dialogue state through the dialogue construction platform, the complexity of building multi-turn dialogue systems is solved, and the development of multi-turn dialogue systems is simplified, reducing the need for resources and expert knowledge.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2017-06-12
- Publication Date
- 2026-07-14
AI Technical Summary
Building a multi-turn dialogue system requires a great deal of expert knowledge, time and resources. Existing technologies lack a simple and scalable way to create multi-turn dialogue systems, making dialogue system development complex and prone to failure.
It provides a dialogue building platform that, by inferring or configuring semantic patterns and rules that depend on the dialogue state, forms a state-specific, multi-turn contextual language understanding system, reducing the need for expert knowledge and resources from the builders.
It simplifies the construction process of multi-turn dialogue systems, reduces development complexity and resource requirements, and makes it easier for third-party developers to build multi-turn contextual language understanding systems specific to dialogue states.
Smart Images

Figure CN116663569B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with international application number PCT / US2017 / 036926, international application date June 12, 2017, entered the Chinese national phase on December 17, 2018, Chinese national application number 201780037738.7, and invention title "System and method for constructing a state-specific multi-turn contextual language understanding system". Background Technology
[0002] Language understanding systems, personal digital assistants, agents, and artificial intelligence (AI) are transforming how users interact with computers. Developers of computers, web services, and / or applications are constantly striving to improve human-computer interaction. However, building such systems requires significant expert knowledge, time, money, and other resources.
[0003] The aspects disclosed herein have been implemented in consideration of the foregoing and other general considerations. Furthermore, although relatively specific problems may have been discussed, it should be understood that these aspects should not be limited to solving the specific problems identified in the background section or elsewhere in this disclosure. Summary of the Invention
[0004] In its general description, this disclosure relates to systems and methods for constructing state-specific multi-turn contextual language understanding systems. More specifically, the systems and methods disclosed herein infer or are configured to infer state-specific decoding configurations, such as dialogue-state-dependent semantic patterns and / or dialogue-state-dependent rules, from single-slot language understanding systems. Thus, the systems and methods disclosed herein require only the information necessary for forming single-slot language understanding models or single-slot rules from the builder or the author of the dialogue-state-specific multi-turn contextual language understanding system. Accordingly, compared to systems and methods that do not infer or provide the ability to infer state-specific patterns and / or state-specific rules, the systems and methods disclosed herein for constructing state-specific multi-turn contextual language understanding systems can reduce the expert knowledge, time, and resources required to construct application-specific state-specific multi-turn contextual language understanding systems.
[0005] One aspect of this disclosure relates to a system with a platform for building application-specific, dialogue-state-based, multi-turn contextual language understanding systems. The system includes at least one processor and a memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0006] Receive information from the builder to construct a single-slot language understanding model;
[0007] Configure a single-slot language understanding model based on information;
[0008] Based on semantic patterns that depend on dialogue state, the decoding of the single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for each determined dialogue state; and
[0009] Implement a constrained single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0010] Another aspect of this disclosure relates to a system with a platform for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and a memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0011] Receive information from the builder to create single-slot rules;
[0012] Single-slot rules are formed based on information;
[0013] Based on single-slot rules, infer semantic patterns that depend on the dialogue state for different dialogue states;
[0014] Based on single-slot rules and dialogue-state-dependent semantic patterns, dialogue-state-dependent rules are derived for different dialogue states; and
[0015] Implement dialogue-state-dependent rules to form a dialogue-state-specific multi-turn contextual language understanding system.
[0016] Another aspect of this disclosure relates to a system with a platform for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and a memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0017] Receive information from the builder to create single-slot rules;
[0018] Single-slot rules are formed based on information;
[0019] Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, forming a first provisioning capability based on single-slot rules and user input from the dialogue with the user during decoding;
[0020] Provides the ability to derive dialogue-state-dependent rules to form a second provisioning capability based on dialogue-state-dependent semantic patterns, single-slot rules, and user input from the dialogue with the user during decoding; and
[0021] Implement single-slot rules, first provisioning capabilities, and second provisioning capabilities to form a multi-turn contextual language understanding system specific to the dialogue state.
[0022] Another aspect of this disclosure relates to a system with a platform for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and a memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0023] Receive information from the builder to create a composite single-slot language understanding system based on a combination of single-slot rules and single-slot models learned through machine learning;
[0024] Information is used to form a combined single-slot language understanding system, which includes single-slot rules and single-slot language understanding models learned by machine learning;
[0025] For decoding that depends on dialogue state, the combined single-slot language understanding system is adjusted to form an adjusted combined single-slot language understanding model; and
[0026] Implement a modified combined single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0027] Another aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0028] Receive information from the builder to construct a single-slot language understanding model;
[0029] Configure a single-slot language understanding model based on information;
[0030] Based on semantic patterns that depend on dialogue state, the decoding of the single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for each determined dialogue state; and
[0031] Implement a constrained single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0032] Another aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0033] Receive information from the builder to create single-slot rules;
[0034] Single-slot rules are formed based on information;
[0035] Based on single-slot rules, infer dialogue-state-dependent semantic patterns for different dialogue states;
[0036] Based on single-slot rules and dialogue-state-dependent semantic patterns, dialogue-state-dependent rules are derived for different dialogue states; and
[0037] Implement dialogue-state-dependent rules to form a dialogue-state-specific multi-turn contextual language understanding system.
[0038] A further aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0039] Receive information from the builder to create single-slot rules;
[0040] Single-slot rules are formed based on information;
[0041] Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, forming a first provisioning capability based on single-slot rules and user input from the dialogue with the user during decoding;
[0042] Provides the ability to derive dialogue-state-dependent rules to form a second provisioning capability based on dialogue-state-dependent semantic patterns, single-slot rules, and user input from the dialogue with the user during decoding; and
[0043] Implement single-slot rules, first provisioning capabilities, and second provisioning capabilities to form a multi-turn contextual language understanding system specific to the dialogue state.
[0044] Additional aspects of this disclosure include a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. This method includes:
[0045] Receive information from the builder to create a composite single-slot language understanding system based on a combination of single-slot rules and single-slot models learned through machine learning;
[0046] Information is used to form a combined single-slot language understanding system, which includes single-slot rules and single-slot language understanding models learned by machine learning;
[0047] For decoding that depends on dialogue state, the combined single-slot language understanding system is adjusted to form an adjusted combined single-slot language understanding model; and
[0048] Implement a modified combined single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0049] This invention provides a simplified summary of a series of concepts, which are further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0050] Non-limiting and non-exhaustive embodiments are described with reference to the following figures.
[0051] Figure 1A This is a schematic diagram illustrating a dialog building platform utilized by a builder via a client computing device, according to various aspects of this disclosure.
[0052] Figure 1B This is a schematic diagram illustrating a dialog building platform used by a builder via a client computing device, according to various aspects of this disclosure.
[0053] Figure 2A This is a simplified block diagram illustrating a state-specific contextual language understanding system built via a dialogue building platform according to various aspects of this disclosure.
[0054] Figure 2B This is a simplified block diagram illustrating a rule-based, state-specific contextual language understanding system constructed via a dialogue building platform according to various aspects of this disclosure.
[0055] Figure 2C This is a simplified block diagram illustrating a combined, state-specific contextual language understanding system constructed via a dialogue building platform according to various aspects of this disclosure.
[0056] Figure 3 This is a flowchart illustrating a method for constructing a state-specific contextual language understanding system according to various aspects of this disclosure.
[0057] Figure 4A This is a schematic diagram illustrating a user interface provided by a dialogue building platform according to various aspects of this disclosure.
[0058] Figure 4B This is a schematic diagram illustrating a user interface provided by a dialogue building platform according to various aspects of this disclosure.
[0059] Figure 5 This is a block diagram illustrating an example physical component of a computing device from which various aspects of this disclosure can be practiced.
[0060] Figure 6A This is a simplified block diagram of a mobile computing device that can be implemented using various aspects of this disclosure.
[0061] Figure 6B This disclosure is available in various aspects that can be utilized in practice. Figure 6A A simplified block diagram of a mobile computing device is shown in the figure.
[0062] Figure 7 This is a simplified block diagram of a distributed computing system in which various aspects of this disclosure can be implemented.
[0063] Figure 8 The illustrations depict various aspects of this disclosure that can be utilized in the creation of tablet computing devices. Detailed Implementation
[0064] The following detailed description is illustrated with reference to the accompanying drawings, which form part of the detailed description and are shown therein by way of illustrating specific aspects or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the spirit or scope of this disclosure. Therefore, the following detailed description should not be considered limiting, and the scope of this disclosure is defined by the claims and their equivalents.
[0065] Advances in machine learning, language understanding, and artificial intelligence are transforming how users interact with computers. Digital assistant applications, such as Siri, Google Now, and Microsoft Cortana, are examples of this shift in human-computer interaction.
[0066] The next wave of computing is widely expected to revolve around Conversation as a Platform (CaaP). As an example of CaaP, bot platforms for building conversational agents for applications and services have emerged. Bots primarily target messaging workflows such as Messenger, WebChat, and Skype, with their main form of interaction being text (typing), although touch and voice can be used as additional modalities, such as speech recognition. Users can interact with bots through multi-turn conversations. However, CaaP will only become a reality when third-party applications and services can easily build bots to serve customers. This will scale to thousands of applications and services.
[0067] However, due to the complexity of multi-turn dialogues, building multi-turn dialogue systems has become a challenging task. Recently, various platforms have been deployed to enable third-party creation of natural language experiences for single-turn experiences (e.g., leveraging artificial intelligence, language understanding intelligent services, and api.ia). However, these previously deployed platforms are limited to single-turn experiences because building multi-turn dialogue platforms requires significant time and resources to consider an exponentially different number of scenarios for any given task.
[0068] For example, consider a simple agent or bot that provides ferry timetables between Seattle and various destinations (e.g., Bremerton, Bainbridge Islands, etc.). In this example, some sample dialogues expected to be handled by the agent / bot are listed below:
[0069] Scene #1
[0070] [User Wheel 1]: I would like to know the ferry timetable.
[0071] [Robot Wheel 1]: Okay, I can provide you with timetables between Seattle and various destinations. Where would you like to go?
[0072] [User Round 2]: I want to go from Seattle to Bremerton.
[0073] [Robot Wheel 2]: Which date do you want?
[0074] [User Round 3]: Tomorrow.
[0075] [Robot Ferry 3]: This is the ferry service that starts at 8:00 am...
[0076] Scene #2
[0077] [User Wheel 1]: I would like to know the ferry timetable;
[0078] [Robot Wheel 1]: Okay, I can provide you with timetables between Seattle and various destinations. Where would you like to go?
[0079] [User Wheel 2]: Seattle.
[0080] [Robot Wheel 2]: Where would you like to start?
[0081] [User Round 3]: Bremerton.
[0082] There are more scenarios that can be added. The challenge here is how to build language understanding models (or rules) that will label "going to location" and "from location". In scenario #1, the user utterances (user turn 1, user turn 2, and user turn 3) provide enough information to the multi-turn dialogue system for the model to extract "going to location" (Bremerton) and "from location" (Seattle). However, in scenario #2, the user utterances (user turn 1, user turn 2, and user turn 3) do not provide enough information for the model to extract "going to location" and "from location". In scenario #2, the dialogue system must also use system or bot responses (bot turn 1 and bot turn 2) to correctly label Seattle as "going to location" and Bremerton as "from location". However, even if the multi-turn dialogue system utilizes bot responses, the developer or builder of the multi-turn dialogue system must build a potential set of n! language understanding models or rules (where n is the number of slots / entities required by the application). Furthermore, the developer or builder of the multi-turn dialogue system must build a potential set of n! for each different and possible dialogue state. Each language understanding model or rule set presents a language complexity and resource requirement for building multi-turn dialogue systems, which is a bottleneck for the wider adoption of conversational interfaces. Consequently, the vast majority of developers are limited to single-turn dialogue systems. Without a precise multi-turn dialogue system, the dialogue between the application and the client may fail because the system cannot correctly interpret user queries. Therefore, builders of dialogue systems still need significant domain expert knowledge, expertise, time, and resources to leverage these prior systems and methodologies to build functional multi-turn dialogue systems. Currently, there is no simple or scalable way to create or build multi-turn dialogue systems.
[0083] The systems and methods disclosed in this paper relate to dialogue building platforms for constructing dialogue-state-specific multi-turn contextual language understanding (LU) systems, requiring only the builder to input the information necessary to build a single-slot language understanding (LU) system. This system can be rule-based or machine learning-based. More specifically, the systems and methods disclosed in this paper infer / derive, or are configured to infer / derive, state-specific decoding configurations, such as dialogue-state-dependent semantic patterns and / or dialogue-state-dependent rules. Accordingly, the systems and methods disclosed in this paper allow third-party developers to build dialogue-state-specific multi-turn contextual LU systems for digital agents, bots, messaging applications, voice agents, or any other application type, without requiring significant domain expert knowledge, time, and other resources.
[0084] The system and method described in this paper provide a dialogue building platform for constructing dialogue-state-specific multi-turn contextual LU systems. This requires no further input from the builder beyond the input necessary to form a single-slot LU model and / or single-slot rules, offering an easy-to-use and efficient service for building multi-turn dialogue systems.
[0085] Figure 1A and Figure 1B The illustrations depict various examples of a dialogue building platform 100 used by a builder 102 (or a user 102 of the dialogue building platform 100) via a client computing device 104, according to various aspects of this disclosure. The dialogue building platform 100 allows the builder 102 (or a user 102 of the dialogue building platform 100) to develop, build, or create a dialogue state-specific multi-turn contextual LU system 108, while only needing to provide the information 106 necessary to build a single-slot language understanding LU system. The single-slot LU system utilizes a single-slot LU model and / or a single-slot rule set. The single-slot LU system requires the user to provide all parameters, slots, and / or entities necessary to perform the desired task in one turn of a dialogue or session with the single-slot LU system. In contrast, the dialogue state-specific multi-turn contextual LU system 108 can collect the parameters, slots, and / or entities necessary to perform the desired task over any number of turns in a session with the user.
[0086] In some aspects, such as Figure 1A As illustrated, the dialogue building platform 100 is implemented on a client computing device 104. In a basic configuration, the client computing device 104 is a computer with both input and output elements. The client computing device 104 can be any suitable computing device for implementing the dialogue building platform 100. For example, the client computing device 104 can be a mobile phone, smartphone, tablet, phablet, smartwatch, wearable computer, personal computer, gaming system, desktop computer, laptop computer, etc. This list is merely exemplary and should not be considered limiting. Any suitable client computing device 104 for implementing the dialogue building platform 100 can be utilized.
[0087] In other aspects, such as Figure 1BAs illustrated, the dialogue building platform 100 is implemented on a server computing device 105. The server computing device 105 can provide data to and / or receive data from the client computing device 104 via a network 113. In some aspects, the network 113 is a distributed computing network, such as the Internet. In a further aspect, the aforementioned dialogue building platform 100 is implemented on more than one server computing device 105, such as multiple server computing devices 105 or a network of server computing devices 105. In some aspects, the dialogue building platform 100 is a hybrid system having a portion of the dialogue building platform 100 on the client computing device 104 and a portion of the dialogue building platform 100 on the server computing device 105.
[0088] The dialogue building platform 100 includes a user interface for building a dialogue-state-specific multi-turn contextual LU system 108. The user interface is generated by the dialogue building platform 100 and presented to the builder 102 via a client computing device 104. The user interface of the dialogue building platform 100 allows the builder 102 to select or provide task definitions 202 and / or necessary information 106 for building a single-slot LU model to the dialogue building platform 100. The client computing device 104 may have one or more input devices, such as a keyboard, mouse, pen, microphone or other sound or voice input device, touch or swipe input device, etc., to allow the builder 102 to provide task definitions 202 and / or information 106 via the user interface. The aforementioned devices are examples, and other devices may be used.
[0089] Figure 4A and Figure 4B Example user interfaces 400 for dialogue building platform 100 are shown during different stages of the construction process of building a rule-based, dialogue-state-specific multi-turn contextual LU system. The same or different interfaces 400 may be provided for building machine-learned, dialogue-state-specific multi-turn contextual LU systems by dialogue building platform 100, or for building composite, dialogue-state-specific multi-turn contextual LU systems. Dialogue building platform 100 may provide any suitable interface 400 for building a dialogue-state-specific multi-turn contextual LU system 108.
[0090] Figure 4AThe diagram illustrates a user interface 400A generated by the dialogue creation platform 100 at the beginning of the build phase for task or dialogue creation. In this task or dialogue creation interface 400A, the builder 102 has selected the "Dialogue" option or button 402 and is asked to name the dialogue that the builder 102 is building 404. In this example, the builder 102 enters the name 404 "Football Timetable" 406. In this example, the builder 102 is building or creating a dialogue for retrieving MLS team timetables. In some aspects, the user interface 400A provides the last time the named dialogue was modified under the "Last Modified" heading 405. In this example, "Football Timetable" 406 was last modified "Today" 407. Once the new dialogue has been named, the builder 102 can select the "Build Dialogue" button or option 408 to begin building the new dialogue. In a further aspect, the user interface 400A provides a simulation option or button 409 (shown as "Run Test" 409) for simulating the performance of a dialogue-state-specific multi-turn contextual LU system 108, which is formed based on the current constraints and / or a tuned single-slot LU model established using the dialogue building platform 100.
[0091] Figure 4B A user interface 400B generated by the dialogue building platform 100 for rule creation is shown. The rule creation interface 400B and / or model creation interface are provided to builder 102 after the “Establish Dialogue” button 408 has been selected by builder 102. In this interface 400B, builder 102 can enter a tag statement 412. In this example, builder 102 enters the tag statement “Get Timetable” 414. Builder 102 can add additional tag statements as needed by selecting the “Add New Statement” button or option 428. Furthermore, on this rule creation interface 400B, builder 102 can add a rule 416 for the tag statement 412 via the “Add Rule” button or option 426, edit the rule 416 via the “Edit” button or option 422, or delete the rule 416 via the “Delete” button or option 424. In this example, builder 102 has created two different rules, such as "Give me (MLS team) timetable" 418 and "Show me this (Microsoft.Date) match" 420. The rule creation interface 400B can also provide some rule formatting examples when selecting the "Rule Format Example" button or option 430.
[0092] In some aspects, the builder 102 provides a task definition 202 or a dialogue definition 202 to the dialogue building platform 100 via a computing device 104 and a provided user interface 400. The task definition 202 may include any data, modules, or systems necessary to perform the desired task. For example, the task definition 202 may include a task triggering module, a verification module, a language generation module, and a final action module. Modules, as used herein, can run one or more software applications of a program. In other words, a module can be a separate, interchangeable component of a program that contains one or more program functions and everything necessary to perform those functions. In some aspects, a module includes memory and one or more processors.
[0093] Builder 102 provides a user interface 400 presented via client computing device 104 to the dialogue building platform 100 with information 106 necessary for building a single-slot LU system. In some aspects, information 106 is a task defined based on a selected task. The single-slot LU system can be a rule-based single-slot LU system, a machine learning-based single-slot LU system, or a combined single-slot LU system. Information 106 can be data necessary for building a machine learning-based single-slot LU system, and / or data necessary for building a rule-based single-slot LU system. Alternatively, builder 102 provides the necessary information 106 to build a combined single-slot LU system. The combined single-slot LU system utilizes both a rule-based LU system and a machine learning-based LU system to decode the dialogue from the user. For example, information 106 can include any parameters, slots, entities, bindings, or mappings necessary for creating the single-slot LU system. Although builder 102 must only provide the information necessary to form the single-slot LU system, the builder may provide additional information as needed.
[0094] Dialogue building platform 100 receives information 106. Using this information, dialogue building platform 100 forms or configures a single-slot LU system. Once the single-slot LU system has been formed or configured, dialogue building platform 100 provides the single-slot LU system with a state-specific decoding configuration. The state-specific decoding configuration varies depending on the type of single-slot LU system. The single-slot LU system can be a machine learning-based LU system, a rule-based LU system, or a combination of machine learning and rule-based LU systems. In some aspects, components necessary for dialogue state-dependent decoding are inferred and / or derived by dialogue building platform 100. In alternative aspects, dialogue building platform 100 provides the ability to infer and / or derive components. This capability can be provided as a program module suitable for running software applications. Components can be dialogue state-dependent semantic patterns and / or dialogue state-dependent rules.
[0095] The dialogue building platform 100 utilizes an LU system with state-specific decoding configurations to form a dialogue-state-specific multi-turn context LU system 108. The dialogue building platform 100 provides the dialogue-state-specific multi-turn context LU system 108 to the builder 102. The builder 102 can then add the dialogue-state-specific multi-turn context LU system 108 to any desired digital agent, bot, messaging application, voice agent, or any other application type.
[0096] In some aspects, the dialogue building platform 100 is used to form, for example Figure 2A The diagram illustrates a machine learning-based, dialogue-state-specific, multi-turn contextual LU system 200A. In these aspects, a builder 102 may provide a task definition 202 (or dialogue definition 202) to a dialogue building platform 100. In these aspects, the builder 102 provides the dialogue building platform 100 with information 106 necessary for building a single-slot, machine learning-based model. The dialogue building platform 100 forms or configures a machine learning-based single-slot LU model 204A based on the information 106 from the builder 102. In these aspects, the dialogue building platform 100 configures the model 204A for dialogue-state-dependent decoding by constraining the decoding 206 of the machine learning-based single-slot LU model based on a dialogue-state-dependent semantic pattern, so as to output only dialogue-state-specific slots and dialogue-state-specific entities for each determined dialogue state, thereby forming a constrained machine learning-based single-slot model. The constrained machine learning-based single-slot model is implemented to form the machine learning-based, dialogue-state-specific, multi-turn contextual LU system 200A. In other words, when the machine-learned, dialogue-state-specific multi-turn contextual LU system 200A is used by the user or receives user input (such as utterances) from the user, the system 200A removes any provided parameters, slots, and / or entities from the single-slot LU model and then responds using a constrained single-slot LU model. This process continues until all the necessary data for performing the task or dialogue is provided by the user to the machine-learned, dialogue-state-specific multi-turn contextual LU system 200A.
[0097] In some aspects, the dialogue building platform 100 infers dialogue-state-dependent semantic patterns for different dialogue states based on a machine learning-learned single-slot LU model 204A. The dialogue building platform 100 infers dialogue-state-dependent semantic patterns before the formation of a machine learning-learned, dialogue-state-specific multi-turn contextual LU system 200A.
[0098] In an alternative, the dialogue construction platform 100 provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, based on a machine learning-learned single-slot LU model 204A and user input received by the system 200A from the dialogue with the user during decoding. This capability can be provided as a program module suitable for running software applications. In other words, the machine learning-learned, dialogue-state-specific, multi-turn contextual LU system 200A is configured to dynamically infer dialogue-state-dependent semantic patterns during the decoding of the user's dialogue.
[0099] In some aspects, the dialogue building platform 100 is used to form, for example Figure 2B The diagram illustrates a rule-based, dialogue-state-specific multi-turn contextual LU system 200B. In these aspects, a builder 102 may provide a task definition 202 (or dialogue definition 202) to a dialogue building platform 100. In these aspects, the builder 102 provides the dialogue building platform 100 with information 106 necessary for constructing a single-slot, rule-based system 204B. The dialogue building platform 100 forms a single-slot rule set based on the information 106 from the builder. In these aspects, the dialogue building platform 100 provides the system 204B with a state-specific decoding configuration by using derived dialogue-state-dependent rules 208 and inferred dialogue-state-dependent semantic patterns to label dialogue-state-specific slots and dialogue-state-specific entities for the determined dialogue state. The decoding configuration system is implemented by the dialogue building platform 100 to form the rule-based, dialogue-state-specific multi-turn contextual LU system 200B. In other words, when the rule-based, dialogue-state-specific multi-turn context LU system 200B receives user input from the user, the system 200B only labels the dialogue-state-specific slots and dialogue-state-specific entities for the determined dialogue state. This process continues until all the necessary data for performing the task is provided by the user to the rule-based, dialogue-state-specific multi-turn context LU system 200B.
[0100] In some aspects, the dialogue building platform 100 infers dialogue-state-dependent semantic patterns for different dialogue states based on the single-slot rules of the rule-based single-slot LU system 204B. In these aspects, the dialogue building platform 100 infers dialogue-state-dependent semantic patterns before the formation of the rule-based, dialogue-state-specific multi-turn context LU system 200B.
[0101] In an alternative aspect, the dialogue construction platform 100 provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, based on the single-slot rules of the rule-based single-slot LU system 204 and on user input received by the system 200B from the dialogue with the user during decoding. This capability can be provided as a program module suitable for running software applications. In other words, in these aspects, the rule-based, dialogue-state-specific multi-turn contextual LU system 200B dynamically infers dialogue-state-dependent semantic patterns during the decoding of the user dialogue.
[0102] In some aspects, the dialogue building platform 100 derives dialogue-state-dependent rules 208 for different dialogue states based on single-slot rules and dialogue-state-dependent semantic patterns. In these aspects, the dialogue building platform 100 infers the dialogue-state-dependent rules 208 before the formation of a rule-based, dialogue-state-specific multi-turn context LU system 200B.
[0103] In an alternative, the dialogue construction platform 100 provides the ability to infer or derive dialogue-state-dependent rules 208 for different dialogue states. This ability is based on single-slot rules, dialogue-state-dependent semantic patterns, and user input received by the system 200B from the dialogue with the user during decoding. In other words, the rule-based, dialogue-state-specific, multi-turn context LU system 200B dynamically derives dialogue-state-dependent rules 208 during the decoding of the user dialogue. This capability can be provided as a program module suitable for running software applications.
[0104] In some aspects, the dialogue building platform 100 is used to form, for example Figure 2C The diagram illustrates a combined, dialogue-state-specific, multi-turn contextual LU system 200C. In these aspects, the builder 102 may provide a task definition 202 (or dialogue definition 202) to the dialogue building platform 100. In these aspects, the builder 102 provides the dialogue building platform 100 with information 106 necessary for building a machine learning-based single-slot LU system 204A and a rule-based single-slot LU system 204B. Based on the information 106 from the builder 102, the dialogue building platform 100 forms the rule-based single-slot LU system 204B and the machine learning-based single-slot LU system 204A. In these aspects, the dialogue building platform 100 adjusts systems 204A & 204B for dialogue-state-dependent decoding 210 by:
[0105] • Based on the inferred dialogue-state-dependent semantic patterns, the decoding 206 of the machine-learned single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for at least one determined dialogue state, thereby forming a constrained, machine-learned single-slot language understanding model; and
[0106] • By utilizing derived dialogue-state-dependent rules 208 and inferred dialogue-state-dependent semantic patterns to label dialogue-state-specific slots and dialogue-state-specific entities for one or more determined dialogue states, a state-specific decoding configuration is provided to system 204B.
[0107] This results in an adjusted, combined single-slot system. The dialogue building platform 100 implements the adjusted, combined single-slot system to form a combined, dialogue-state-specific, multi-turn contextual LU system 200C.
[0108] In some aspects, the dialogue building platform 100 infers dialogue-state-dependent semantic patterns for different dialogue states based on a single-slot language understanding model and / or single-slot rules. In some aspects, the dialogue building platform 100 infers dialogue-state-dependent semantic patterns before the formation of a combined, dialogue-state-specific multi-turn contextual LU system 200C.
[0109] In an alternative aspect, the dialogue construction platform 100 provides the ability to infer dialogue-state-dependent semantic patterns for one or more dialogue states, based on a single-slot language understanding model and / or single-slot rules, and on user input received by the system 200C from the dialogue with the user during decoding. In other words, the dialogue-state-specific multi-turn contextual LU system 200C dynamically infers dialogue-state-dependent semantic patterns in response to the decoded dialogue from the user. This capability can be provided as a program module suitable for running software applications.
[0110] In some aspects, the dialogue building platform 100 derives dialogue-state-dependent rules 208 for one or more dialogue states based on single-slot rules and inferred dialogue-state-dependent semantic patterns. In some aspects, the dialogue building platform 100 derives the dialogue-state-dependent rules 208 prior to the formation of a combined, dialogue-state-specific multi-turn context LU system 200C.
[0111] In an alternative, the dialogue construction platform 100 provides the ability to derive dialogue-state-dependent rules 208 for at least one dialogue state. This capability is based on single-slot rules, dialogue-state-dependent semantic patterns, and user input received by the system 200C from the dialogue with the user during decoding. In other words, the dialogue-state-specific multi-turn context LU system 200C dynamically derives dialogue-state-dependent rules 208 during the decoding of the user dialogue. This capability can be provided as a program module suitable for running software applications.
[0112] Figure 3 A flowchart is illustrated, conceptually illustrating an example of a method 300 for constructing a dialogue-state-specific multi-turn contextual LU system. In some aspects, method 300 is performed by a construction platform 100 as described above. Method 300 provides a method for constructing a dialogue-state-specific multi-turn contextual LU system that does not require the builder to provide any further input beyond that necessary for constructing a single-slot LU system. More specifically, method 300 infers state-specific patterns or provides the ability to infer state-specific patterns, and / or derives state-specific rules from the formed single-slot language understanding system or provides the ability to derive state-specific rules. Thus, method 300 provides a method for constructing a dialogue-state-specific multi-turn contextual LU system that is easier to use and requires less expert knowledge, less time, and fewer resources than previously used methods for constructing dialogue-state-specific multi-turn contextual LU systems. The dialogue-state-specific multi-turn contextual LU system can be a machine learning-based dialogue-state-specific multi-turn contextual LU system, a rule-based dialogue-state-specific multi-turn contextual LU system, or a combined dialogue-state-specific multi-turn contextual LU system.
[0113] In some aspects, method 300 includes operation 302. In operation 302, a task definition or dialogue definition for the application is received from the builder. As discussed above, the task definition may include any data, modules, or systems necessary to perform the desired task. For example, the task definition may include a task triggering module, a verification module, a language generation module, and / or a final action module.
[0114] In operation 304, information for constructing a single-slot LU system is received. This information may be necessary for constructing the single-slot LU model and / or the single-slot rule set. The information is provided by the builder 102. In some aspects, the information includes parameters, entities, slots, tags, bindings, or mappings, etc. In some aspects, the received information is based on the task defined by the provided task.
[0115] In some aspects (during the construction of a rule-based, dialogue-state-specific multi-turn context LU system or a combined dialogue-state-specific multi-turn context LU system), method 300 includes operation 305. In operation 305, a single-slot rule set is formed based on the information received in operation 304.
[0116] In some aspects (during the construction of a dialogue-state-specific multi-turn contextual LU system learned by machine learning, or a combination of dialogue-state-specific multi-turn contextual LU systems), method 300 includes operation 306. In operation 306, a single-slot LU model is configured based on the information received at operation 304.
[0117] In operation 307, a dialogue-state-dependent semantic pattern is inferred, or the ability to infer a dialogue-state-dependent semantic pattern is provided. In some aspects, operation 307 infers dialogue-state-dependent semantic patterns for different dialogue states based on the single-slot LU model configured during operation 306 and / or the single-slot rule set formed at operation 305. The dialogue-state-dependent semantic pattern can be inferred at operation 307 before the formation of the dialogue-state-dependent multi-turn contextual LU system at operation 316, and before decoding by the formed system.
[0118] In an alternative aspect, the ability to infer dialogue-state-dependent semantic patterns for different dialogue states is provided at operation 307. This capability can be provided as a program module suitable for running software applications. In these aspects, the ability to infer dialogue-state-dependent semantic patterns for different dialogue states is based on a single-slot language understanding model configured during operation 306 and / or a single-slot rule set formed at operation 305, and on user input received during the dialogue with the user by a dialogue-state-dependent multi-turn contextual LU system. Thus, in these aspects, the dialogue-state-dependent multi-turn contextual LU system formed during operation 316 is able to dynamically infer dialogue-state-dependent semantic patterns in response to user input received from the dialogue with the user during decoding.
[0119] At operation 308, a state-specific decoding configuration is provided for the single-slot LU module and / or the single-slot rule set. The state-specific decoding configuration varies depending on the type of dialogue-state-specific multi-turn context LU system method 300 being built. In some aspects (during the construction of a machine-learned dialogue-state-specific multi-turn context LU system or a combined dialogue-state-specific multi-turn context LU system), the single-slot LU model is constrained at operation 308. In other aspects (during the construction of a rule-based dialogue-state-specific multi-turn context LU system or a combined dialogue-state-specific multi-turn context LU system), a dialogue-state-dependent rule set is derived at operation 308, or the ability to derive a dialogue-state-dependent rule set is provided.
[0120] In some aspects, at operation 308, dialogue-state-dependent rules are derived for different dialogue states, based on the single-slot rules formed during operation 305 and on the dialogue-state-dependent semantic patterns inferred at operation 307. In these aspects, the dialogue-state-dependent rules are derived prior to the formation of the dialogue-state-specific multi-turn contextual LU system at operation 316.
[0121] In an alternative aspect, the ability to derive dialogue-state-dependent rules for different dialogue states is provided at operation 308. This capability can be provided as a program module suitable for running software applications. In these aspects, the ability to derive dialogue-state-dependent rules for one or more different dialogue states is based on the single-slot rules received during operation 304, the dialogue-state-dependent semantic patterns inferred at operation 307, and the user input received by a dialogue-state-specific multi-turn contextual LU system during the decoding of the user's dialogue. In other words, the dialogue-state-specific multi-turn contextual LU system formed during operation 316 can dynamically derive dialogue-state-dependent rules in response to user input received during the decoding of the user's dialogue.
[0122] The derived set of rules, dependent on the dialogue state, labels dialogue-state-specific slots and dialogue-state-specific entities for one or more defined dialogue states during decoding. In these respects, the process continues during dialogue decoding until all necessary parameters are provided to the dialogue-state-specific multi-turn context LU system formed during operation 316.
[0123] In a further aspect, at operation 308, constraints are based on a dialogue-state-dependent semantic pattern and a machine-learned single-slot language understanding model to output only dialogue-state-specific slots and dialogue-state-specific entities for one or more determined dialogue states, thereby forming a constrained machine-learned single-slot language understanding model. In these aspects, when the machine-learned, dialogue-state-specific multi-turn contextual LU system formed by method 300 is used to decode received user input (such as utterances), the system removes any parameters, slots, and / or entities provided in the user input from the single-slot LU model and then responds to the user by utilizing the modified or constrained single-slot LU model. This process continues until all necessary data for performing the task or dialogue is provided by the user to the machine-learned, dialogue-state-specific multi-turn contextual LU system.
[0124] In one example, the combined single-slot LU model is adjusted as follows:
[0125] • Based on the inferred, dialogue-state-dependent semantic patterns, constrain the decoding of the machine-learned single-slot dialogue understanding model to output only dialogue-state-specific slots and dialogue-state-specific entities for at least one determined dialogue state; and
[0126] • By leveraging derived dialogue-state-dependent rules and inferred dialogue-state-dependent semantic patterns to label dialogue-state-specific slots and entities for one or more identified dialogue states, a state-specific decoding configuration is provided to a rule-based system.
[0127] This results in an adjusted combined single-slot LU system being formed at operation 308.
[0128] In some aspects, method 300 includes operations 310 and 312. At operation 310, a simulation request is received. The simulation request may be received from the builder via a user interface. The user interface may provide and / or display simulation options to the builder. At operation 312, in response to the simulation request received at operation 310, a simulation of the constrained single-slot LU model and / or the dialogue-state-dependent rule set is run. This simulation allows the builder to simulate or test how the currently constructed constrained single-slot LU model and / or the currently derived dialogue-state-dependent rule set will perform when implemented into a dialogue-state-specific multi-turn contextual LU system. For example, the simulation at operation 312 may consist of a set of utterances provided by the builder, or it may consist of automatically generated utterances from a simulated user and system responses, to test the implemented constrained single-slot LU model and / or the currently derived dialogue-state-dependent rule set.
[0129] In a further aspect, method 300 includes operation 314. In operation 314, an implementation request is received. The implementation request may be received from the builder via a user interface. The user interface may provide and / or display implementation simulation options to the builder.
[0130] In operation 316, a dialogue-state-specific multi-turn contextual LU system is formed. In some aspects, the dialogue-state-specific multi-turn contextual LU system is formed by implementing a constrained single-slot LU model at operation 316. In other aspects, the dialogue-state-specific multi-turn contextual LU system is formed by implementing a dialogue-state-dependent rule set at operation 316. In a further aspect, the dialogue-state-specific multi-turn contextual LU system is formed by implementing the ability to derive a dialogue-state-dependent rule set and / or state-dependent semantic patterns.
[0131] In some aspects, operation 316 is executed in response to the implementation request received at operation 314. In other aspects, operation 316 is executed automatically after operation 308. The established dialogue-state-specific multi-turn context LU system is provided to the builder. The builder can then add the established dialogue-state-specific multi-turn context LU system to any desired digital agent, bot, messaging application, voice agent, and / or any other application type.
[0132] Although the dialogue-state-specific multi-turn contextual LU system formed by method 300 does not require any further input from the builder at operation 304 other than the information necessary to form a single-slot LU system, the builder may provide additional information as needed.
[0133] Figures 5-8 The associated description provides a discussion of various operating environments in which various aspects of this disclosure can be practiced. However, regarding Figures 5-8 The devices and systems discussed and illustrated are for illustrative purposes only and are not intended to limit the wide range of computing device configurations that can be used to practice the various aspects of this disclosure described herein.
[0134] Figure 5This is a block diagram illustrating the physical components (e.g., hardware) of a computing device 500 that can be implemented using various aspects of this disclosure. For example, a dialogue building platform 100 can be implemented by the computing device 500. In some aspects, the computing device 500 is a mobile phone, smartphone, tablet computer, phablet, smartwatch, wearable computer, personal computer, desktop computer, gaming system, laptop computer, etc. The computing device components described below may include computer-executable instructions for the dialogue building platform 100, which can be executed to employ method 300 to establish a dialogue state-specific multi-turn context LU system 108 as disclosed herein. In a basic configuration, the computing device 500 may include at least one processing unit 502 and system memory 504. Depending on the configuration and type of the computing device, the system memory 504 may include, but is not limited to, volatile storage devices (e.g., random access memory), non-volatile storage devices (e.g., read-only memory), flash memory, or any combination of such memories. The system memory 504 may include an operating system 505 and one or more program modules 506 suitable for running software application 520. For example, operating system 505 can be used to control the operation of computing device 500. Furthermore, various aspects of this disclosure can be practiced in conjunction with graphics libraries, other operating systems, or any other application, and are not limited to any particular application or system. This basic configuration is... Figure 5 These components are illustrated within the dashed lines 508. The computing device 500 may have additional features or functions. For example, the computing device 500 may also include additional data storage devices (removable and / or non-removable), such as, for example, a disk, optical disc, or magnetic tape. These additional storage devices... Figure 5 The diagram is illustrated using removable storage device 509 and non-removable storage device 510. For example, any inferred dialogue-state-dependent rules of the inferential dialogue building platform 100, the configuration for inferring dialogue-state-dependent rules, any inferred dialogue-state-dependent semantic patterns, and / or the configuration for dialogue-state-dependent semantic patterns can be stored on any of the illustrated storage devices.
[0135] As described above, multiple program modules and data files can be stored in system memory 504. When executed on processing unit 502, program module 506 (e.g., dialogue building platform 100) can perform processes including, but not limited to, performing the methods described herein 300. For example, processing unit 502 can implement dialogue building platform 100. Other program modules that can be used according to various aspects of this disclosure, particularly program modules that generate screen content, may include: digital assistant applications, voice recognition applications, email applications, social networking applications, collaboration applications, enterprise management applications, messaging applications, word processing applications, spreadsheet applications, database applications, presentation applications, contact applications, game applications, e-commerce applications, e-commerce applications, transaction applications, exchange applications, device control applications, website interface applications, calendar applications, etc. In some aspects, dialogue building platform 100 allows the builder to construct a dialogue-state-specific multi-turn context LU system 108 for one or more of the applications mentioned above.
[0136] Furthermore, various aspects of this disclosure can be implemented in electrical circuits, including discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing microprocessors, and single chips containing electronic components or microprocessors. For example, various aspects of this disclosure can be implemented via a system-on-a-chip (SOC), wherein... Figure 5 Each or many of the components illustrated in the diagram can be integrated onto a single integrated circuit. Such a System-on-a-Chip (SoC) device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or “burned in”) onto a chip substrate as a single integrated circuit. When operating via the SoC, the capabilities described herein regarding the client switching protocol can be operated via dedicated logic integrated with other components of the computing device 500 on the single integrated circuit (chip).
[0137] Many aspects of this disclosure can also be practiced using other techniques capable of performing logical operations (such as, for example, AND, OR, and NOT), including but not limited to mechanical, optical, fluid, and quantum technologies. Furthermore, many aspects of this disclosure can be practiced within a general-purpose computer or in any other circuit or system.
[0138] The computing device 500 may also have one or more input devices 512, such as a keyboard, mouse, pen, microphone, or other sound or voice input device, touch or swipe input device, etc. It may also include output devices(s) 514, such as a display, speaker, printer, etc. The foregoing devices are examples, and other devices may also be used. The computing device 500 may include one or more communication connections 516 that allow communication with other computing devices 550. Examples of suitable communication connections 516 include, but are not limited to, RF transmitters, receivers, and / or transceiver circuitry, universal serial buses (USB), parallel and / or serial ports.
[0139] As used herein, the terms "computer-readable medium" or "storage medium" can include computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media that store information (such as computer-readable instructions, data structures, or program modules) implemented in any method or technology. System memory 504, removable storage device 509, and non-removable storage device 510 are examples of computer storage media (e.g., memory storage devices). Computer storage media can include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other storage technologies, CD-ROM, digital universal disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage devices or other magnetic storage devices, or any other article of manufacture that can be used to store information and can be accessed by computing device 500. Any such computer storage medium may be part of computing device 500. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0140] Communication media can be embodied in computer-readable instructions, data structures, program modules, or other data in modulated data signals (such as carrier waves or other transmission mechanisms), and include any information transmission medium. The term "modulated data signal" can describe a signal whose one or more characteristics are set or modified in such a way as to encode information in the signal. By way of example and without limitation, communication media can include wired media and wireless media, such as wired networks or direct wired connections, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0141] Figure 6A and Figure 6B The illustration depicts a mobile computing device 600, such as a mobile phone, smartphone, tablet computer, phablet, smartwatch, wearable computer, personal computer, desktop computer, gaming system, laptop computer, etc., from which various aspects of this disclosure can be practiced. Reference Figure 6AThe illustration depicts one aspect of a mobile computing device 600 suitable for implementation. In a basic configuration, the mobile computing device 600 is a handheld computer with both input and output elements. The mobile computing device 600 typically includes a display 605 and one or more input buttons 610, which allow the user to input information into the mobile computing device 600. The display 605 of the mobile computing device 600 can also function as an input device (e.g., a touchscreen display).
[0142] When included, the optional side input element 615 allows for further user input. The side input element 615 can be a rotary switch, a button, or any other type of manual input element. Alternatively, the mobile computing device 600 can include more or fewer input elements. For example, in some aspects, the display 605 may not be a touchscreen. In yet another alternative aspect, the mobile computing device 600 is a portable telephone system, such as a cellular phone. The mobile computing device 600 may also include an optional keypad 635. The optional keypad 635 can be a physical keypad or a “soft” keypad generated on a touchscreen display.
[0143] In addition to or as an alternative to a touchscreen input device associated with display 605 and / or keypad 635, a Natural User Interface (NUI) may be included in mobile computing device 600. As used herein, NUI includes any interface technology that enables a user to interact with the device in a “natural” manner, free from the artificial constraints imposed by input devices such as a mouse, keyboard, remote control, etc. Examples of NUI methods include those relying on: on-screen and near-screen speech recognition, touch and stylus recognition, gesture recognition, air gestures, head and eye tracking, voice and speech, vision, touch, gestures, and machine intelligence.
[0144] In various aspects, the output elements include a display 605 for displaying a graphical user interface (GUI). In the aspects disclosed herein, various user information collections may be displayed on the display 605. Further output elements may include a visual indicator 620 (e.g., a light-emitting diode) and / or an audio transducer 625 (e.g., a speaker). In some aspects, the mobile computing device 600 includes a vibration transducer for providing haptic feedback to a user. In yet another aspect, the mobile computing device 600 includes input and / or output ports for sending signals to or receiving signals from external devices, such as audio input (e.g., a microphone jack), audio output (e.g., a headphone jack), and video output (e.g., an HDMI port).
[0145] Figure 6BThis is a block diagram illustrating the architecture of one aspect of a mobile computing device. That is, the mobile computing device 600 may include a system (e.g., architecture) 602 to implement several aspects. In one aspect, the system 602 is implemented as a "smartphone" capable of running one or more applications (e.g., browser, email, calendar, contact manager, messaging client, games, and media clients / players). In some aspects, the system 602 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and wireless phone.
[0146] One or more applications 666 and / or conversation building platform 100 run on or are associated with operating system 664. Examples of applications include: telephone dialer, email program, personal information management (PIM) program, word processing program, spreadsheet program, internet browser program, messaging program, etc. System 602 also includes a non-volatile storage area 668 within memory 662. Non-volatile storage area 668 can be used to store persistent information that will not be lost in the event of a power outage of system 602. Application 666 can use and store information in non-volatile storage area 668, such as emails or other messages used by an email application, etc. A synchronization application (not shown) also resides on system 602 and is programmed to interact with a corresponding synchronization application residing on a hosting computer to keep the information stored in non-volatile storage area 668 synchronized with the corresponding information stored on the hosting computer. As should be understood, other applications can be loaded into memory 662 and run on mobile computing device 600.
[0147] System 602 has a power supply 670, which can be implemented as one or more batteries. The power supply 670 may further include an external power source, such as an AC adapter or electrical docking bracket for replenishing or recharging the batteries.
[0148] System 602 may also include a radio 672 that performs the functions of transmitting and receiving radio frequency communications. Radio 672 facilitates a wireless connection between system 602 and the "outside world" via a communication carrier or service provider. Transmissions to and from radio 672 are conducted under the control of operating system 664. In other words, communications received by radio 672 can be propagated to application 666 via operating system 664, and vice versa.
[0149] A visual indicator 620 can be used to provide visual notifications, and / or an audio interface 674 can be used to generate audio notifications via an audio transducer 625. In the illustrated aspect, the visual indicator 620 is a light-emitting diode (LED), and the audio transducer 625 is a speaker. These devices can be directly coupled to a power supply 670, such that when activated, these devices remain on for the duration indicated by the notification mechanism, even if the processor 660 and other components may be shut down to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes an action to indicate the device's power-on status. The audio interface 674 is used to provide auditory signals to and receive auditory signals from the user. For example, in addition to being coupled to the audio transducer 625, the audio interface 674 can also be coupled to a microphone to receive audio input. System 602 may further include a video interface 676, which enables the operation of the vehicle camera 630 to record still images, video streams, etc.
[0150] The system 602 implemented by the mobile computing device 600 may have additional features or functions. For example, the mobile computing device 600 may also include additional data storage devices (removable and / or non-removable), such as disks, optical discs, or magnetic tapes. Such additional storage devices... Figure 6B The diagram shows a non-volatile storage region 668.
[0151] Data / information generated or captured by mobile computing device 600 and stored via system 602 can be locally stored on mobile computing device 600 as described above, or the data can be stored on any number of storage media that can be accessed by the device via radio 672 or via a wired connection between mobile computing device 600 and a separate computing device associated with mobile computing device 600 (e.g., a server computer in a distributed computing network such as the Internet). It should be understood that such data / information can be accessed via mobile computing device 600 and via radio 672, or via a distributed computing network. Similarly, according to known data / information transfer and storage components, including email and collaborative data / information sharing systems, such data / information can be easily transferred between computing devices for storage and use.
[0152] Figure 7The illustration depicts one aspect of the architecture of a system for processing data received at a computing system from a remote source (such as the general-purpose computing device 704, tablet computer 706, or mobile device 708 described above). Content displayed and / or utilized at server device 702 can be stored in different communication channels or other storage types. For example, various files can be stored using directory services 722, web portals 724, email services 726, instant messaging storage 728, and / or social networking sites 730. As an example, a conversation building platform 100 can be implemented in the general-purpose computing device 704, tablet computing device 706, and / or mobile computing device 708 (e.g., a smartphone). In some aspects, such as... Figure 7 As shown in the diagram, server 702 is configured to enable dialogue building platform 100 via network 715.
[0153] Figure 8 An exemplary tablet computing device 800 capable of performing one or more aspects of the invention disclosed herein is illustrated. Furthermore, the aspects and functions described herein can operate on a distributed system (e.g., a cloud-based computing system), wherein application functions, memory, data storage and retrieval, and various processing functions can be remotely operated on each other over a distributed computing network (such as the Internet or an intranet). User interfaces and various types of information can be displayed via an in-vehicle computing device display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information can be projected onto a wall, and various types of user interfaces and information can be displayed and interacted with on the wall. Interaction with multiple computing systems in which aspects of the invention can be practiced includes keystroke input, touchscreen input, voice or other audio input, gesture input, wherein the associated computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures for controlling functions of the computing device, etc.
[0154] In some aspects, a system with a platform is provided for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0155] Receive information from the builder to construct a single-slot language understanding model;
[0156] Configure a single-slot language understanding model based on information;
[0157] Based on semantic patterns that depend on dialogue state, the decoding of the single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for each determined dialogue state; and
[0158] Implement a constrained single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0159] At least one processor may further operate to infer dialogue-state-dependent semantic patterns for different dialogue states based on a single-slot language understanding model and patterns for the single-slot language understanding model. At least one processor may further operate to provide the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, this ability being based on the single-slot language understanding model and on user input received from the dialogue with the user by a dialogue-state-specific multi-turn contextual language understanding system, as well as system prompts given for the dialogue or task. At least one processor may further operate to receive a task definition for the application from a builder. The single-slot language understanding model can be used to construct a task for the task definition. Information may include parameters, slots, and entities necessary for defining the task. In these aspects, the single-slot language understanding model is a machine-learned model. The system may be a server.
[0160] In other aspects, a system with a platform is provided for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0161] Receive information from the builder to create single-slot rules;
[0162] Single-slot rules are formed based on information;
[0163] Based on single-slot rules, infer dialogue-state-dependent semantic patterns for different dialogue states;
[0164] Based on single-slot rules and dialogue-state-dependent semantic patterns, dialogue-state-dependent rules are derived for different dialogue states; and
[0165] Implement dialogue-state-dependent rules to form a dialogue-state-specific multi-turn contextual language understanding system.
[0166] At least one processor may be further operable to: receive a task definition for the application from the builder. Single-slot rules may be for tasks defined in the task definition. Information may include parameters, slots, and / or entities necessary for defining the task. At least one processor may be further operable to: implement dialogue-state-dependent rules in response to receiving an implementation request from the builder to form a dialogue-state-specific multi-turn contextual language understanding system. The application may be:
[0167] Digital assistant applications;
[0168] Voice recognition applications;
[0169] Email applications;
[0170] Social networking applications;
[0171] Collaborative applications;
[0172] Enterprise management applications;
[0173] Messaging applications;
[0174] Word processing applications;
[0175] Spreadsheet applications;
[0176] Database applications;
[0177] Demonstration application;
[0178] Contacts application;
[0179] Game applications;
[0180] E-commerce applications;
[0181] E-commerce applications;
[0182] Trading applications;
[0183] Equipment control applications;
[0184] Website interface application;
[0185] Switching applications; and / or
[0186] Calendar app.
[0187] At least one processor may be further operable to: in response to receiving a simulation request from the builder, run a simulation based on rules that depend on the dialogue state.
[0188] In other aspects, a system with a platform is provided for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0189] Receive information from the builder to create single-slot rules;
[0190] Single-slot rules are formed based on information;
[0191] Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, forming a first provisioning capability based on single-slot rules and user input from the dialogue with the user during decoding;
[0192] Provides the ability to derive dialogue-state-dependent rules to form a second provisioning capability based on dialogue-state-dependent semantic patterns, single-slot rules, and user input from the dialogue with the user during decoding; and
[0193] Implement single-slot rules, first provisioning capabilities, and second provisioning capabilities to form a multi-turn contextual language understanding system specific to the dialogue state.
[0194] At least one processor may be further operable to receive a task definition for the application from the builder. Single-slot rules may be applied to tasks within the task definition. Information may include parameters, slots, and / or entities.
[0195] In other aspects, a system with a platform is provided for building an application-specific, dialogue-state-based, multi-turn contextual language understanding system. The system includes at least one processor and memory. The memory encodes computer-executable instructions that, when executed by the at least one processor, are operable as follows:
[0196] Receive information from the builder to create a composite single-slot language understanding system based on a combination of single-slot rules and single-slot models learned through machine learning;
[0197] Information is used to form a combined single-slot language understanding system, which includes single-slot rules and single-slot language understanding models learned by machine learning;
[0198] For decoding that depends on dialogue state, the combined single-slot language understanding system is adjusted to form an adjusted combined single-slot language understanding model; and
[0199] Implement a modified combined single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0200] A combined single-slot language understanding system for decoding that depends on dialogue state can be adapted as follows:
[0201] Based on a semantic pattern dependent on the dialogue state, the decoding of a machine-learned single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for at least one determined dialogue state; and
[0202] Processing single-slot rules and dialogue-state-dependent semantic patterns to derive dialogue-state-dependent rules or providing the ability to derive dialogue-state-dependent rules for tagging dialogue-state-specific slots and dialogue-state-specific entities for at least one determined dialogue state or another determined dialogue state.
[0203] At least one processor may be further operable to: infer dialogue-state-dependent semantic patterns for different dialogue states based on information, and derive dialogue-state-dependent rules. At least one processor may be further operable to: provide the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, based on single-slot rules or a single-slot language understanding model learned through machine learning, and based on user input received from the dialogue with the user during decoding by a dialogue-state-specific multi-turn contextual language understanding system. At least one processor may be further operable to: provide the ability to derive dialogue-state-dependent rules.
[0204] Another aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0205] Receive information from the builder to construct a single-slot language understanding model;
[0206] Configure a single-slot language understanding model based on information;
[0207] Based on semantic patterns that depend on dialogue state, the decoding of the single-slot language understanding model is constrained to output only dialogue-state-specific slots and dialogue-state-specific entities for each determined dialogue state; and
[0208] Implement a constrained single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0209] Another aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0210] Receive information from the builder to create single-slot rules;
[0211] Single-slot rules are formed based on information;
[0212] Based on single-slot rules, infer dialogue-state-dependent semantic patterns for different dialogue states;
[0213] Based on single-slot rules and dialogue-state-dependent semantic patterns, dialogue-state-dependent rules are derived for different dialogue states; and
[0214] Implement dialogue-state-dependent rules to form a dialogue-state-specific multi-turn contextual language understanding system.
[0215] A further aspect of this disclosure includes a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. The method includes:
[0216] Receive information from the builder to create single-slot rules;
[0217] Single-slot rules are formed based on information;
[0218] Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states, forming a first provisioning capability based on single-slot rules and user input from the dialogue with the user during decoding;
[0219] Provides the ability to derive dialogue-state-dependent rules to form a second provisioning capability based on dialogue-state-dependent semantic patterns, single-slot rules, and user input from the dialogue with the user during decoding; and
[0220] Implement single-slot rules, first provisioning capabilities, and second provisioning capabilities to form a multi-turn contextual language understanding system specific to the dialogue state.
[0221] Additional aspects of this disclosure include a method for an application-specific, dialogue-state-based multi-turn contextual language understanding system. This method includes:
[0222] Receive information from the builder to create a composite single-slot language understanding system based on a combination of single-slot rules and single-slot models learned through machine learning;
[0223] Information is used to form a combined single-slot language understanding system, which includes single-slot rules and single-slot language understanding models learned by machine learning;
[0224] For decoding that depends on dialogue state, the combined single-slot language understanding system is adjusted to form an adjusted combined single-slot language understanding model; and
[0225] Implement a modified combined single-slot language understanding model to form a dialogue-state-specific multi-turn contextual language understanding system.
[0226] For example, embodiments of the present disclosure are described above with reference to the block diagrams and / or operational illustrations of methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions described in the blocks may not occur in the order shown in any flowchart. For example, depending on the functions / actions involved, two blocks shown consecutively may actually be performed substantially simultaneously, or these blocks may sometimes be performed in reverse order.
[0227] This disclosure describes some embodiments of the present technology with reference to the accompanying drawings, which depict only a few of the possible aspects. However, other aspects may be embodied in many different forms, and the specific embodiments disclosed herein should not be construed as limiting the various aspects of the disclosure set forth herein. Rather, these exemplary aspects are provided to make this disclosure thorough and complete, and to fully convey the scope of other possible aspects to those skilled in the art. For example, aspects of the various embodiments disclosed herein may be modified and / or combined without departing from the scope of this disclosure.
[0228] Although specific aspects are described herein, the scope of this technology is not limited to these specific aspects. Those skilled in the art will recognize other aspects or improvements within the scope and spirit of this technology. Therefore, specific structures, actions, or media are disclosed only as illustrative aspects. The scope of this technology is defined by the claims and any of their equivalents.
Claims
1. A system with a platform, the platform being used to build an application-specific, dialogue-state-based, multi-turn contextual language understanding system, the system comprising: At least one processor; as well as The memory, which stores and encodes computer-executable instructions, is operable, when executed by the at least one processor, as follows: Receive information from the builder for creating a single-slot rule, wherein the information includes any slots and entities necessary for performing tasks in one round of the dialogue; The single-slot rule is formed based on the information; Based on the single-slot rule, infer the dialogue state-dependent semantic patterns for different dialogue states; Based on the single-slot rule and the dialogue state-dependent semantic pattern, derive dialogue state-dependent rules for different dialogue states. Implement the dialogue state-dependent rules to form the dialogue state-specific multi-turn contextual language understanding system; as well as The multi-turn contextual language understanding system is used to analyze the statements received from the user.
2. The system of claim 1, wherein the at least one processor is further operable to receive a task definition for the application from the builder, wherein the single-slot rule is for the task.
3. The system according to claim 1, wherein the information further includes parameters.
4. The system of claim 1, wherein the at least one processor is further operable to: in response to receiving an implementation request from the builder, implement the dialogue state-dependent rules to form the dialogue state-specific multi-turn contextual language understanding system.
5. The system of claim 1, wherein the application is at least one of the following: Digital assistant applications; Voice recognition applications; Email applications; Social networking applications; Collaborative applications; Enterprise management applications; Messaging applications; Word processing applications; Spreadsheet applications; Database applications; Demonstration application; Contacts application; Game applications; E-commerce applications; E-commerce applications; Trading applications; Equipment control applications; Website interface application; Switching applications; or Calendar app.
6. The system of claim 1, wherein the at least one processor is further operable to: in response to receiving a simulation request from the builder, run the simulation of the rules that depend on the dialogue state.
7. The system according to claim 1, wherein the system is a server.
8. The system of claim 1, wherein the system is rule-based.
9. A system with a platform, the platform being used to build an application-specific, dialogue-state-based, multi-turn contextual language understanding system, the system comprising: At least one processor; as well as The memory, which stores and encodes computer-executable instructions, is operable, when executed by the at least one processor, as follows: Receive information from the builder for creating a single-slot rule, wherein the information includes any slots and entities necessary for performing tasks in one round of the dialogue; The single-slot rule is formed based on the information; Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states based on the single-slot rules and based on user input from the dialogue with the user during decoding, to form a first provisioning capability; Providing the ability to derive dialogue-state-dependent rules based on the dialogue-state-dependent semantic patterns, the single-slot rules, and the user input from the dialogue with the user during the decoding process, to form a second provisioning capability; as well as The single-slot rule, the first provisioning capability, and the second provisioning capability are implemented to form the dialogue state-specific multi-turn contextual language understanding system.
10. The system of claim 9, wherein the at least one processor is further operable to receive a task definition for the application from the builder, wherein the single-slot rule is for the task.
11. The system of claim 9, wherein the information further includes parameters.
12. The system of claim 9, wherein the at least one processor is further operable to: in response to receiving an implementation request from the builder, implement the single-slot rule, the first provisioning capability, and the second provisioning capability to form the dialogue state-specific multi-turn contextual language understanding system.
13. The system of claim 9, wherein the application is at least one of the following: Digital assistant applications; Voice recognition applications; Email applications; Social networking applications; Collaborative applications; Enterprise management applications; Messaging applications; Word processing applications; Spreadsheet applications; Database applications; Demonstration application; Contacts application; Game applications; E-commerce applications; E-commerce applications; Trading applications; Equipment control applications; Website interface application; Switching applications; or Calendar app.
14. The system of claim 9, wherein the at least one processor is further operable to: in response to receiving a simulation request from the builder, run a simulation of the single-slot rule, the first provisioning capability, and the second provisioning capability.
15. The system of claim 9, wherein the system is a server.
16. The system of claim 9, wherein the system is rule-based.
17. A method for constructing an application-specific, dialogue-state-based multi-turn contextual language understanding system, the method comprising: Receive information from the builder for creating a single-slot rule, wherein the information includes any slots and entities necessary for performing tasks in one round of the dialogue; The single-slot rule is formed based on the information; Provides the ability to infer dialogue-state-dependent semantic patterns for different dialogue states based on the single-slot rules and based on user input from the dialogue with the user during decoding, to form a first provisioning capability; Providing the ability to derive dialogue-state-dependent rules based on the dialogue-state-dependent semantic patterns, the single-slot rules, and the user input from the dialogue with the user during the decoding process, to form a second provisioning capability; as well as The single-slot rule, the first provisioning capability, and the second provisioning capability are implemented to form the dialogue state-specific multi-turn contextual language understanding system.
18. The method of claim 17, wherein the information further includes parameters.
19. The method of claim 17, further comprising: Receive the implementation request from the builder. The implementation of the single-slot rule, the first provisioning capability, and the second provisioning capability to form the dialogue-state-specific multi-turn contextual language understanding system is executed in response to the implementation request.
20. The method of claim 17, further comprising: Receive simulation requests from the builder; as well as In response to the simulation request, the simulation of the single-slot rule, the first provisioning capability, and the second provisioning capability is run.