Error correction and extraction in request dialogue

JP7904786B2Active Publication Date: 2026-08-13ZOOM COMMUNICATIONS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-14
Publication Date
2026-08-13

AI Technical Summary

Benefits of technology

【0006】 【0006】本発明の実装により実現可能であるこれら及び他の利点は、以下の説明から明らかになるであろう。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007904786000001
    Figure 0007904786000001
  • Figure 0007904786000002
    Figure 0007904786000002
  • Figure 0007904786000003
    Figure 0007904786000003
Patent Text Reader

Abstract

The system includes a machine configured to operate in response to requests from a user and a detection means for detecting an operation mode dialogue stream from the user directed to the machine. The system also includes a computing system configured to train a neural network through machine learning to output a machine-directed correction request for each training example in the training dialogue stream dataset. The computing system is also configured, in the operation mode, to use the trained neural network to generate an operation mode correction request for the machine based on the operation mode dialogue stream from the user directed to the machine that has been detected by the detection means.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001]

[0001] (Priority Claim) This application claims the priority of U.S. Provisional Application No. 62 / 947,946, filed on December 13, 2019, with the same title and by the same inventors as above, and the entire disclosure thereof is incorporated herein by reference in its entirety as part of this specification.

[0002]

[0002] Errors and ambiguities are difficult to avoid during a conversation. It is possible to recover from an error by correction and resolve the ambiguity. For example, a household robot receives a request such as "Put the washed knife in the drawer for cutlery," but the robot does not identify which of the drawers is the drawer for cutlery. The robot selects one of the drawers and puts the knife in it. If this selection is incorrect, the user has to correct the robot, for example, by saying "No, put it in the drawer to the right of the sink." Instead, the robot might ask which drawer is the drawer for cutlery. A human explanation response such as "It's the drawer to the right of the sink" is also a correction in order to resolve the ambiguity in the response. Another type of correction occurs when the user's mind changes, for example, by saying "I've changed my mind, put the fork in instead."

Summary of the Invention

Problems to be Solved by the Invention

[0003]

[0003] All of these types of corrections can be processed in a similar manner. Therefore, in a general embodiment, the present invention relates to a software-based machine learning component that receives requests and corrections and outputs correction requests. In order to receive these correction requests, in one implementation, the correction phrase is replaced with the corresponding expression in the request. The request "Put the washed knife in the cutlery drawer" is translated to "Put the washed knife in the drawer to the right of the sink," along with its correction, "No, put it in the drawer to the right of the sink." Such components have two advantages compared to processing corrections in actual dialogue components. First, if an open-domain correction component exists, there is no need to learn corrections, which reduces the amount of training data required for actual dialogue components. Second, such types of correction components can be extended to output pairs of corrected items and corrected expressions. In this example, there is one pair, such as the cutlery drawer and the drawer to the right of the sink. These expression pairs can be used for learning, for example, in a lifelong learning component of a dialogue system, which can reduce the need for corrections in future dialogues. For example, a robot can learn which of the drawers is the cutlery drawer. [Means for solving the problem]

[0004]

[0004] In a general embodiment, the present invention relates to a system comprising a machine, a sensing means, and a computer system. For example, the machine, which may be a robot or a computer, is configured to operate in response to a request from a user. The sensing means is for detecting a stream of user-generated mode dialogues directed to the machine. The computer system comprises a neural network trained through machine learning to output a correction request for the machine for each training example in a training dialogue stream dataset. The computer system also uses the trained neural network to generate a corrected mode request for the machine based on the user-generated mode dialogues directed to the machine.

[0005]

[0005] In a more general embodiment, the present invention relates to a method for training a neural network through machine learning and for each training example in a dataset of training dialogue streams, outputting a machine-oriented correction request configured to operate in response to a user request. The method also includes, in the operating mode of the neural network after the neural network has been trained, (i) a sensing means detecting a machine-oriented user dialogue stream of the operating mode, and (ii) a computer system communicating with the sensing means using the trained neural network to generate a machine-oriented corrected request for the operating mode based on the operating mode dialogue stream.

[0006]

[0006] These and other advantages that can be realized by implementing the present invention will become apparent from the following description.

[0007]

[0007] Various implementations and embodiments of the present invention are described herein by reference with the following drawings. [Brief explanation of the drawing]

[0008] [Figure 1] This represents non-fluent speech annotated using restoration terminology. [Figure 2] These represent non-fluent utterances labeled with copy (C) and delete (D). [Figure 3] This represents the request and correction clauses, annotated using restoration terminology. [Figure 4] This represents a system according to various embodiments of the present invention. [Figure 5] This shows an exemplary template for generating a training dataset for the error correction module neural network shown in Figure 4, according to various embodiments of the present invention. [Figure 6] Figure 4 shows the size of the verification and test dataset for an example of training the error correction module neural network according to various embodiments of the present invention. [Figure 7] This describes sequence labeling approaches and inter-sequence approaches according to various embodiments of the present invention. [Figure 8] The evaluation results for the error correction module neural network shown in Figure 4, according to various embodiments of the present invention, are shown. [Figure 9] Figure 4 shows a computer system according to various embodiments of the present invention. [Modes for carrying out the invention]

[0009]

[0017] The task of request correction is related to the task of removing disfluency. In disfluency removal, there are the target of repair (which expression should be replaced), the interruption point (where the correction begins), the additional transition words (which phrases are signal phrases for correction), and the repair phrase (the corrected expression). Figure 1 shows a disfluent utterance annotated using this terminology.

[0010]

[0018] Much work has been done to remove unfluent speech. These efforts are assumed to be sufficient to remove tokens from unfluent speech and obtain fluent speech. Figure 2 shows unfluent speech with copy and delete labels. However, long-range substitutions may occur during the correction task. These long-range substitutions are shown in Figure 3.

[0011]

[0019] In a general embodiment, the present invention relates to a system for performing error correction in a request or dialogue stream. Figure 4 shows an example of the system according to various embodiments. As shown in Figure 4, the system 10 may comprise a user and a machine. The user outputs communication or exchange, such as a dialogue stream, including, for example, requests 11A to the machine and corrections 11B as necessary. Corrections 11B may be issued by the user, for example, in response to a question (audible or textual) from the machine (directly or indirectly), or in response to the user's observation of, or otherwise detection of, inappropriate behavior of the machine. The machine may be any machine that operates in response to commands or requests from the user in the dialogue stream, such as, for example, a mobile electromechanical robot 12 shown in Figure 4. In other embodiments, the machine may be a computer device such as, for example, a laptop computer, PC, server, workstation, mainframe, mobile terminal (e.g., smartphone or tablet computer), wearable computer device, or digital personal assistant. In further implementations, the machine may be a device or tool equipped with one or more processors, such as a kitchen or other household appliance, a power tool, a medical device, a medical diagnostic device, or an automotive system (e.g., an electronic or engine control unit (ECU) of a vehicle). The machine may also be an autonomous driving mobile terminal, such as an autonomous vehicle, or a part thereof.

[0012]

[0020] As shown in Figure 4, in various implementations, system 10 includes a computer system 14 that receives a user dialogue stream (e.g., including requests 11A and corrections 11B) and generates correction requests for the machine from this system. The computer system 14 may also include detection means for detecting user requests and corrections. For example, as shown in the example in Figure 4 where the user is making an audible request, the detection means may include one or more microphones 16 and a natural language processing (NLP) module 18 for processing the audible utterances from the user acquired by the microphones 16. The computer system 14 also includes an error correction module 20, which preferably implements a machine learning neural network, trained to generate correction requests for the machine based on the received dialogue stream. The computer system 14 may also include an output device 19 that outputs correction requests to the robot / machine.

[0013]

[0021] The detection means may comprise a microphone 16 and an NLP module 18, as shown in Figure 4. In this context, the user's response (e.g., a correction) may be a verbal response. In other implementations, other types of detection means may be used (in addition to or instead of the microphone), depending on the context and purpose of the system 10. For example, the detection means may also comprise a motion sensor, such as a motion sensor on the robot 12. This detects corrected motions being given to the robot 12 by the user (or a third party or something else). Using the example of a kitchen robot, the user can push the robot 12 toward a cutlery drawer, and the computer system 14 can use the detected motion input from the user to the robot 12 as a correction to generate a correction request. This motion sensor may be an accelerator and / or a gyroscope.

[0014]

[0022] Depending on the circumstances, the detection means may also include (1) a camera for capturing motion such as a physical response, motion, or gesture by the user, and (2) a module (e.g., a neural network) trained to interpret the user's physical response, motion, or gesture as a request 11A or description 11B. In addition, the detection means may include, for example, (1) a touch-sensitive display on a computer system 14 for receiving touch input from the user, and (2) a module for interpreting the user's touch input. In addition, the detection means may include a software program for receiving and processing text from the user. For example, the user may use an "app" (e.g., a software application for a mobile device) or another type of computer program (e.g., a browser) to generate a text-based dialogue stream including text-based requests and corrections for the robot 12.

[0015]

[0023] In this regard, as used herein, the term "dialogue stream" is not limited to a dialogue that includes only spoken words or text. Rather, a dialogue stream can be a sequence of requests and subsequent explanations in a format or modality suitable for creating requests and explanations, including spoken words, text, gestures recognized by a camera system, inputs received via a touch-based user interface, detected motions imparted to a robot by a user, and the like. Additionally, the original requests and explanations can use the same or different modalities. For example, the original request can include spoken words or text, while the subsequent explanation can include gestures recognized by a user, inputs received from a user via a touch-based user interface, detected motions imparted to a robot by a user, and the like. Additionally, the user does not have to be a single person. One human can create an initial request, while another human can create an explanation. Additionally, the user does not even have to be human, but instead, the user can be an intelligent system such as a virtual assistant (e.g., Siri by Apple, Google Assistant, Alexa by Amazon, etc.), or other types of intelligent systems.

[0016]

[0024] A user request for a machine can be, for example, a direct request or command for the machine. For a kitchen robot device, the request or command can be something like "Make me tea." The user request can also be less direct, such as the perceived or detected intention of the user for the machine. Continuing with the example of a kitchen robot, the user might say "I would like some tea." In such a case, the machine can be trained to translate the user's intention for "drinking tea" into a request to have the kitchen robot make tea. Thus, as used herein, the term "user intention" refers to the intention of the user for the machine. The user intention can be a direct request or explanation for the machine, but can also be the intention perceived by the user of the machine.

[0017]

[0025] In addition, although various embodiments and implementations of the system are described herein using the term "dialogue stream" between user and machine, it is clear from this description that the "dialogue stream" does not necessarily have to include utterances, nor does it necessarily have to include subsequent dialogue, nor does it have to be limited to dialogue between two parties (e.g., user and machine). For example, as described herein, in addition to or instead of spoken language, the user may use text or motion to express user intent to the machine. This motion may be a gesture, a nod, or a movement given to the machine by the user. The sensing means must be configured to detect any form of user intent expressed by the user. In this regard, the sensing means may include a microphone, NLP for text or spoken language, a motion sensor, a camera, a pressure sensor, a proximity sensor, a humidity sensor, an ambient light sensor, a GPS receiver, and / or a touch-sensitive display (e.g., for receiving input from the user via a touch-sensitive display). The dialogue stream or communication exchange between user and machine does not have to be continuous. All or part of the stream or exchange may be parallel or simultaneous.

[0018]

[0026] Figure 4 represents a computer system 14 as being unrelated to and separate from the user and machine. In this represented embodiment, the computer system 14 issues an audible correction request. In this case, the output device 19 is a speaker near the robot 12, which is picked up by the robot 12's microphone and processed accordingly by the robot. In other implementations, the output device 19 is a wireless communication circuit that transmits electronic radio communications to the robot 12 over a wireless network. The wireless data network may include ad-hoc and / or infrastructure wireless networks such as a Bluetooth network, a Zigbee network, a Wi-Fi network, a wireless mesh network, or other suitable wireless networks.

[0019]

[0027] In other implementations, computer system 14 can be part of or integrated with robot / machine 12. That is, for example, sensing means and error correction module 20 can be part of or integrated with robot / machine 12. In that case, output device 19 can issue commands to a controller of robot / machine 12 that controls the operation of robot / machine 12 via a data bus.

[0020]

[0028] Computer system 14 can also include a distributed system. For example, microphone 16 or other input device can be part of a device that wirelessly communicates with remote NLP module 18 and error correction module 20 and is carried, used, or transported by a user. For example, NLP module 18 and error correction module 20 can be part of a cloud computing system or other computing system that is remote from the microphone (or other input device). In that case, the input device can be in wireless communication with NLP module 18 (e.g., an ad hoc and / or infrastructure wireless network such as Wi-Fi).

[0021]

[0029] Error correction module 20 can be implemented using one or more machine learning networks (in their entirety), such as deep neural networks, that are trained with training examples sufficient to generate correction requests for the robot / machine based on the dialogue stream received from the user. In various implementations, the training can take into account and leverage the intended context or domain for the robot / machine. For example, for the kitchen robot example, error correction module 20 can leverage the dialogue stream and / or the correction requests are likely to include terms related to the kitchen (e.g., drawer, knife, fork, etc.). Similarly, for a medical diagnosis setting, error correction module 20 can leverage the dialogue stream and / or the correction requests are likely to include medical terms, etc.

[0022]

[0030] The following describes one method for training a neural network. A dataset can be created that includes requests and corresponding corrections. Robot / machine outputs, such as erroneous actions and questions for explanation, are not part of the dataset. For example, the training dataset could be a specific domain, such as a kitchen robot. One version of the dataset could focus on tasks such as moving a specified object to a specified location and cooking a specified recipe. For the task of moving a specified object to a specified location, many (e.g., 20) utterances and corresponding many (e.g., 20) corrective responses could be collected from each of several (e.g., 7) participants. For the task of cooking a specified recipe, several more (e.g., 10) utterances and corresponding several (e.g., 10) corrective responses could be collected from each participant. Examples of collected data include a request such as "Put the washed knife in the cutlery drawer," a correction for an inappropriate robot action such as "No, put it in the drawer to the right of the sink," an explanation such as "The drawer to the right of the sink," or a change of heart from the user such as "I've changed my mind, put the fork in." Correction responses were collected to correct or clearly describe objects, locations, objects and their locations, or recipes. Using the collected data, templates were constructed to cover the diversity of natural language. To enhance this diversity, synonyms, new objects, new locations, and / or new recipes can be added to the collected data. Exemplary templates are shown in Figure 5. For example, 15 templates for requests and 45 templates for correction responses can be created. As a result, 74,309 request-correction pairs are available to train the error correction module 20.

[0023]

[0031] A dataset in this context can have two targets. The first target is correction requests. The second target is pairs of items to be corrected and the corrected representations. Validation and test datasets can be used to test four capabilities: handling unknown representations, unknown templates, unknown representations and templates, and out-of-domain templates and representations. Regarding testing the handling of unknown representations, templates from the training dataset can be used, and all representations can be replaced with representations that do not occur within the training dataset. To test generalization with respect to natural language diversity, all templates in the training dataset can be replaced with unknown templates that have the same meaning. To combine both tests described, only unknown representations and templates may be used. For example, to test whether a trained model can be used in other domains or tasks without retraining, tasks such as purchasing a product and attaching a specified object to a specified location can be used. One validation used 400 request-correction pairs for validation and 1727 pairs for testing. The table in Figure 6 shows the number of pairs for each test.

[0024]

[0032] For correction and extraction, a sequence labeling approach can be used in various embodiments. In such embodiments, a neural network is trained to label all word tokens in requests and corrections for a particular token, such as C (copy), D (delete), R1 (representation 1 that may be replaced), R2 (representation 2 that may be replaced), S1 (representation to replace representation 1), or S2 (representation to replace representation 2). For correction targets, the labeled representations of S1 and S2 may be used to replace the labeled representations of R1 and R2. For extraction targets, the output may be pairs of R1 and S1, and pairs of R2 and S2. Figure 7 shows labeling to exemplary request and correction pairs, with both targets given. For sequence labeling, a fine-tuned BERT-based model in a transfer learning version with named entity recognition tools (12 Transformer blocks, 768 hidden sizes, 12 self-attention heads, and 110M parameters) may be used. Details about the fine-tuned BERT (a bidirectional encoder using Transformers) based model can be found in J. Devlin, et al., "BERT: Pre-training of deep bidirectional transformers for language understanding," Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NACL), 2019, and the entire work is incorporated herein as part of this specification. The three epochs of the training dataset are, for example, 2e -5 The initial fine-tuning learning rate can be adjusted. In this context, a Transformer is a deep learning model that processes continuous data, but does not need to process this continuous data sequentially.

[0025]

[0033] As an alternative to the sequence labeling approach, an inter-sequence approach can be used. In this case, the correction stream is output directly from the neural network. For the inter-sequence approach, a Transformer model can be trained, or a pre-trained Transformer model can be fine-tuned. Details on training the Transformer are provided in Vaswani et al., "Attention Is All You Need," arXiv 1706.03762 (2017), and details on pre-trained and fine-tuned Transformer models are provided in Raffel et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," arXiv 1910.10683 (2019).

[0026]

[0034] The architecture can be evaluated using the accuracy metric. In such evaluations, a request-correction pair can be accurately transformed if both the request and the target, such as the correction request and the pair of the target to be corrected and the corrected representation, meet their respective criteria. Figure 8 shows the results of a dataset for an exemplary example of the present invention. In this evaluated embodiment, substitution of unknown representations with trained representations, or substitution of unknown templates with trained templates, did not cause problems with the trained model. Accuracy rates of 98.54% and 97.03% were obtained, respectively. Combining both reduced performance to 90.24%. For this example, an accuracy rate of 88.35% was obtained for out-of-domain representations and templates.

[0027]

[0035] Therefore, the error correction module may include a machine learning neural network (or a collection of neural networks) trained to output correction streams or correction requests based on a training dataset of dialogue streams for controlling a robot / machine. In this case, the training dataset stream may include the original request and subsequent explanation. Once trained and in operating mode, the error correction module can generate correction streams for the robot / machine based on the dialogue stream detected by a microphone or other sensing means. For example, the neural network of the error correction module may be trained, using appropriate training examples, to identify the objects to be corrected during the correction of requests and corrections to those requests, and to generate correction streams for the machine based on the identified objects to be corrected and the corrections. The error correction module may include another machine learning system, such as a neural network, to further learn new information from the identified objects to be corrected and the corrections. Using the kitchen example again, if the original request was "Put the fork in the cutlery drawer" and the correction is "It must be the drawer to the right of the refrigerator, not the drawer to the left of the refrigerator," then the error correction module 20 can learn that the cutlery drawer is to the right of the refrigerator so that the computer system 14 can create correction requests without correction for future use of this system. For example, if the original request was "Put the fork in the cutlery drawer," then using the new knowledge, the error correction module 20 can automatically create the correction request "Put the fork in the cutlery drawer, and this drawer is to the right of the refrigerator."

[0028]

[0036] One or more neural networks in an error correction module may have a fixed-size or non-fixed-size output vocabulary. In particular, a sequence labeling approach to a neural network may have a fixed-size output vocabulary such as C, R1, R2, D, S1, S2. In an inter-sequence approach (where the correct dialogue is inferred directly from the corrected dialogue), a non-fixed-size vocabulary approach (e.g., all tokens of the relevant dialogue including the correction) may be used in addition to a fixed-size vocabulary approach. Word tokens in these approaches may be, for example, entire words or sub-parts of words.

[0029]

[0037] The detection means may include a combination of detection modalities. Continuing with the kitchen example, the user's correction may begin with the word "here," but simultaneously point to a location. This is an example of a deictic response / correction. The microphone / NLP can detect the user's utterance of "here," and the camera can detect the location pointed to by the user. Therefore, if the user's original request was "Put the knife in the cutlery drawer," and the correction is to point to the right drawer of the refrigerator and say "No, it's here," the error correction module 20 can generate the corrected request "Put the knife in the cutlery drawer, to the right of the refrigerator."

[0030]

[0038] Figure 9 shows computer systems 14 according to various embodiments. In the illustrated embodiments, the illustrated computer system 14 comprises a plurality of processor units 2402A-B, such that it comprises a plurality (N) sets of processor cores 2404A-N. Each processor unit 2402A-B may include onboard memory (ROM or RAM) (not shown) and offboard memory 2406A-B. The onboard memory may include a volatile memory device which is the main memory, and / or a non-volatile memory device (e.g., a memory device that can be directly accessed by the processor cores 2404A-N). The offboard memory 2406A-B may include a non-volatile memory device which is the auxiliary memory device (e.g., a memory device that cannot be directly accessed by the processor cores 2404A-N), such as ROM, HDD, SSD, or flash memory. The processor cores 2404A-N may be CPU cores, GPU cores, and / or AI accelerator cores. GPU cores operate in parallel (e.g., GPU-for-Purpose (GPGPU) pipelines) and therefore typically process data more efficiently than a collection of CPU cores, although all cores of a GPU execute the same code at once. AI accelerators are a type of microprocessor designed to accelerate artificial neural networks. They are typically used as coprocessors in devices that also have a host processor. AI accelerators typically have tens of thousands of matrix multiplication units that operate with lower precision than CPU cores, such as 8-bit precision for the AI ​​accelerator compared to 64-bit precision for the CPU core.

[0031]

[0039] In various embodiments, different processor cores 2404 may be trained and / or implement different components of the NLP module 18 and / or the error correction module 20. For example, in one embodiment, the core of the first processor unit 2402A may implement the NLP module 18, and the second processor unit 2402B may implement the error correction module 20. One or more host processors 2410 may coordinate and control the processor units 2402A-B. In other embodiments, the system 2400 may be implemented using a single processor unit 2402. In embodiments with multiple processor units, the processor units may be located in the same location or distributed. For example, the processor units 2402 may be interconnected by a data network such as a LAN, WAN, or the Internet, using suitable wired and / or wireless data communication links. Data may be shared among the various processing units 2402 using suitable data links such as a data bus (preferably a high-speed data bus) or a network link (e.g., Ethernet).

[0032]

[0040] The software for the NLP module 18 and the error correction module 20, and other computer functions described herein, may be implemented in computer software using .NET, C, C++, or Python, and any suitable computer programming language such as conventional functional or object-oriented techniques. For example, the error correction module 20 may be implemented using a stored software module, or otherwise maintained in a computer-readable medium such as RAM, ROM, or auxiliary storage. One or more processing cores of the machine learning system (e.g., CPU or GPU cores) may then execute software modules to implement the respective functions of the machine learning system (e.g., students, coaches, etc.). Programming languages ​​and other computer implementation instructions for computer software may be translated into machine language by a compiler or assembler before execution, and / or directly translated at runtime by an interpreter. Examples of assembly languages ​​include ARM, MIPS, and x86; examples of high-level languages ​​include Ada, BASIC, C, C++, C#, COBOL, Fortran, Java, Lisp, Pascal, Object Pascal, Haskell, and ML; and examples of scripting languages ​​include Bourne scripts, JavaScript, Python, Ruby, Lua, PHP, and Perl.

[0033]

[0041] In a general embodiment, the present invention thus relates to intelligent computer-based systems and methods. Various implementations of the system include a machine configured to operate in response to user intent (e.g., requests) from a user, and sensing means for detecting user communication (e.g., a dialogue stream) of operating modes to the machine. The system also includes a computer system communicating with the sensing means. The computer system is configured to train a neural network through machine learning, which outputs corrected user intent (requests) for the machine for each training example in a training dataset. The computer system is also configured to, in operating modes, use the trained neural network to generate corrected user intent (requests) for the machine based on user communication (dialogue streams) of operating modes detected by the sensing means.

[0034]

[0042] The method according to the present invention may include the steps of training a neural network through machine learning and outputting a machine-oriented correction request configured to operate in response to a user request for each training example in a dataset of training dialogue streams. The method also includes, in the operating mode of the neural network after the neural network has been trained, a sensing means detecting a machine-oriented user dialogue stream for the operating mode, and a computer system communicating with the sensing means using the trained neural network to generate a machine-oriented corrected request for the operating mode based on the dialogue stream for the operating mode.

[0035]

[0043] Depending on the implementation, the training dialogue stream dataset includes a training dialogue stream containing training requests for the machine and training corrections for the training requests, while the operation mode dialogue stream includes operation mode requests for the machine and operation mode corrections for the requests.

[0036]

[0044] Through various implementations, neural networks are trained to identify repair targets in training requests within the training dialogue stream dataset, identify training repairs in corrections to training requests within the training dialogue stream dataset, and generate machine-ready correction requests based on the repair targets and repairs within the training dialogue stream dataset. The neural network is also configured to generate machine-ready correction operation mode requests in operation modes, based on the repair targets of operation modes identified in operation mode requests and the repairs of operation modes identified in operation mode corrections.

[0037]

[0045] In various implementations, the neural network is configured to generate correction requests for the second operation mode for the machine, based on the operation mode to be corrected identified in the machine-oriented second operation mode request and the operation mode to be corrected identified in the machine-oriented pre-operation mode dialogue stream.

[0038]

[0046] In various implementations, the neural network is trained to assign labels to word tokens in the dialogue streams within the training dialogue stream dataset, and then, based on the assigned labels, to determine correction requests for each training example in the training dialogue stream dataset.

[0039]

[0047] In various implementations, the machine includes a robot. In such implementations, user correction of the operating mode may include the user's response to inappropriate actions by the robot. In addition, the detection means may include means for detecting the user's response to inappropriate actions by the robot. Furthermore, the response may include a response selected from the group consisting of a physical response by the user, a verbal response by the user, an imitative response by the user, and a gesture by the user.

[0040]

[0048] In various implementations, the machine includes processor-based devices such as computers, mobile terminals, devices, home entertainment systems, personal assistants, automotive systems, health management systems, or medical devices.

[0041]

[0049] In various implementations, the detected operation mode request may include voice from the user, and the detection means may include a microphone and a natural language processor (NLP). In addition, the detected operation mode request may include an electronic message containing text, and the detection means may include a natural language processor (NLP) for processing the text in the electronic message. In addition, the detection means may include a motion sensor, a camera, and a touch-sensitive display. The detection means may be part of a machine, and the computer system may be part of a machine.

[0042]

[0050] The embodiments presented herein are intended to illustrate potential and specific implementations of the invention. Those skilled in the art will understand that the embodiments are intended primarily to illustrate the invention. The specific aspects of the embodiments are not necessarily intended to limit the scope of the invention. Furthermore, the drawings and specification of the invention have been simplified to illustrate relevant elements for a clear understanding of the invention, omitting other elements for the purpose of clarity. While various embodiments are described herein, it is evident that various modifications, changes, and adaptations to these embodiments can be conceived by those skilled in the art, achieving at least some advantages. Therefore, the disclosed embodiments are intended to include all such modifications, changes, and adaptations without departing from the scope of the embodiments described herein.

Claims

1. It is a system, A machine configured to operate according to the user's intent, A detection means for detecting communication of the operating mode of the machine from the user, The system comprises a computer system that communicates with the detection means, and the computer system In an operating mode, a trained neural network is used and configured to generate one or more corrected user intents for the operating mode directed to the machine, based on the communication of the operating mode from the user to the machine. The communication of the aforementioned operating mode includes the user intent of the operating mode for the machine, and the correction of the operating mode in relation to the user intent of the aforementioned operating mode. The user intent for the one or more corrected operating modes for the machine is generated based on the one or more operating modes to be corrected identified in the user intent for the operating modes, and the one or more operating modes corrected identified in the correction of the operating modes. The communication in the aforementioned operating mode is detected by the detection means. The aforementioned neural network is Identify the target to be repaired in the user's intent for the aforementioned operating mode, Identify the correction in the user intent of the aforementioned operating mode, The machine is trained to output one or more corrected user intents to the machine based on the communication of the operating mode from the user to the machine. system.

2. The system according to claim 1, wherein the neural network is configured to generate a corrected user intent for the second operating mode for the machine, based on the target operating mode to be repaired identified in the user intent for the second operating mode for the machine and the repair of the operating mode identified in the communication for the pre-operating mode for the machine.

3. The system according to claim 1, further comprising a second neural network trained to learn the relationship between a repair target identified in the user intent of the operation mode and one or more corrective user intents for the identified repair target.

4. The system according to claim 1, wherein the neural network has a fixed-size output vocabulary.

5. The system according to claim 1, wherein the neural network is trained to assign labels to communication word tokens in a training dataset and to determine one or more corrective user intentions for each training example in the training dataset based on the assigned labels.

6. The system according to claim 1, wherein the neural network does not have a fixed-size output vocabulary.

7. The system according to claim 1, wherein the machine includes a robot.

8. The correction of the operation mode by the user includes the user's response to an inappropriate action by the robot. The detection means includes means for detecting the user's response to the inappropriate action by the robot, The system according to claim 7.

9. The system according to claim 8, wherein the response includes a response selected from the group consisting of a physical response by the user, a verbal response by the user, an imitative response by the user, and a gesture by the user.

10. The system according to claim 1, wherein the machine includes a processor-based device selected from the group consisting of computers, mobile terminals, devices, home entertainment systems, personal assistants, automotive systems, health management systems, and medical devices.

11. The user intent of the detected operating mode includes voice from the user, The detection means includes a microphone and a natural language processor (NLP), The system according to claim 1.

12. The user intent of the detected operating mode includes an electronic message containing text, The detection means includes a natural language processor (NLP) for processing the text in the electronic message. The system according to claim 1.

13. The system according to claim 1, wherein the detection means includes a sensor selected from the group consisting of a motion sensor, a camera, a pressure sensor, a proximity sensor, a humidity sensor, an ambient light sensor, a GPS receiver, and a touch sensor display.

14. The system according to claim 1, wherein the detection means is part of the machine.

15. The system according to claim 1, wherein the computer system is part of the machine.

16. The system according to claim 1, wherein the user intent from the user to the machine includes a command request from the user to the machine.

17. The system according to claim 1, wherein the communication of the operation mode detected by the detection means includes a communication modality selected from the group consisting of text, spoken language, physical things such as gestures, head movements, and actions.

18. The system according to claim 1, wherein the communication of the operation mode detected by the detection means includes a dialogue stream, and the dialogue stream includes dialogue from the user.

19. In the operation mode of the neural network after training the neural network, The detection means detects communication of an operating mode from a user to a machine, wherein the communication of the operating mode includes the user's intent for the operating mode to the machine, and corrections to the operating mode in relation to the user's intent for the operating mode. Identify one or more targets for repair of the operation mode in the user intent of the said operation mode, and identify one or more repairs in the correction of the user intent of the said operation mode, A computer system communicating with the detection means uses the trained neural network to generate one or more corrected user intents for the machine based on the communication of the operation mode detected by the detection means, wherein the one or more corrected user intents for the operation mode are generated based on one or more operation mode repair targets in the user intent of the operation mode, and the one or more repairs. Methods that include...

20. The method according to claim 19, further comprising generating a corrected user intent for the second operating mode for the machine by the neural network in the operating mode, based on the target operating mode to be repaired identified in the user intent for the second operating mode for the machine and the repair of the operating mode identified in the communication for the pre-operating mode for the machine.

21. The machine includes a robot, The correction of the operation mode by the user includes the user's response to an inappropriate action by the robot. Detecting communication in the aforementioned operating mode includes, by the detection means, detecting the user's response to the inappropriate action by the robot. The method according to claim 19.

22. The method according to claim 21, wherein detecting the response includes detecting by the detection means a response selected from the group consisting of a physical response by the user, a verbal response by the user, an imitative response by the user, and a gesture by the user.

Citation Information

Patent Citations

  • spoken language understanding system

    JP2018513405A

  • Disfluency detection for a speech-to-speech translation system using phrase-level machine translation with weighted finite state transducers

    US20080046229A1

  • Determining and utilizing corrections to robot actions

    US20190001489A1

  • Information processing device, information processing system, information processing method, and program

    WO2019142427A1