Learning methods for personalized intent
By combining a universal parser and a personal parser in an intelligent personal assistant, and using machine learning technology to train a personal parser, the problem of errors in user's natural language input intention analysis is solved, and personalized language comprehension ability and user experience are improved.
Patent Information
- Application Number
- CN201980013513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-23
- Filing Date
- 2019-02-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-11-19
AI Technical Summary
When existing smart personal assistants process user natural language input, it is difficult to correctly infer user intentions, especially because each user's language usage habits are different, resulting in parsing errors and unsatisfactory user experience.
By retrieving the user's natural language input on an electronic device, using a combination of a universal parser and a personal parser, the intent of the input is determined and a new personal intent is generated. At the same time, machine learning technology is used to train personal parsers to adapt to user personalized expressions.
Improves the personalized ability of intelligent personal assistants in natural language understanding, reduces parsing errors, enhances user experience, and does not need to access the training data of existing intention parsers, reducing the overhead of retraining.
Smart Images

Figure CN111819553B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments relate generally to virtual assistants, and more particularly to personalized intent learning for personalized expressions. Background Art
[0002] Customers can use voice-based personal assistants, e.g. and To answer questions, solve problems, perform tasks to save time, effort, and make their lives more convenient. Intent parsers are at the core of artificial intelligence (AI) technology, which converts a user's natural language (NL) query into an intent class, which is then executed by calling a predefined action routine. These intent parsers are trained using machine learning techniques on large labeled datasets with the most common user queries / expressions and their corresponding intents. However, these datasets are never exhausted because there are potentially multiple ways to paraphrase a sentence expressing a specific intent. Users often encounter situations where these AI assistants are unable to correctly infer their desired intent. This can be amplified because language usage varies from person to person, and everyone has their own speaking style and preferences. Summary of the invention BRIEF DESCRIPTION OF THE DRAWINGS
[0003] For a fuller understanding of the nature and advantages of the embodiments and the preferred mode of use, reference should be made to the following detailed description read in conjunction with the accompanying drawings, wherein:
[0004] Figure 1 shows a schematic diagram of a communication system according to some embodiments;
[0005] Figure 2 A block diagram illustrating the architecture of a system including an electronic device having a personal intent learning application according to some embodiments;
[0006] Figure 3 An example usage of a virtual personal assistant (PA) for correct and incorrect natural language (NL) expression understanding is shown;
[0007] Figure 4 shows an example PA scenario for NL expression intent learning according to some embodiments;
[0008] Figure 5 shows a block diagram for reactive intent parsing for a PA according to some embodiments;
[0009] Figure 6 shows a block diagram for customizing interpretation of a PA according to some embodiments;
[0010] Figure 7shows a block diagram of active intent parsing for a PA according to some embodiments;
[0011] Figure 8 shows a block diagram of recursive reactive intent parsing for PA according to some embodiments;
[0012] Fig. 9 shows a block diagram of recursive proactive intent parsing for PA according to some embodiments;
[0013] Fig.10 A block diagram illustrating a process for personalized expressions of intent learning for a virtual personal assistant according to some embodiments; and
[0014] Fig.11 is a high-level block diagram illustrating an information processing system including a computing system that implements one or more embodiments.
[0015] Specific implementation method
[0016] One or more embodiments generally relate to personalized expression intent learning for intelligent personal assistants. In one embodiment, a method includes: retrieving a first natural language (NL) input on an electronic device, wherein neither a universal parser nor a personal parser can determine the intent of the first NL input; retrieving a paraphrase of the first NL input on the electronic device; determining the intent of the paraphrase of the first NL input using at least one of the universal parser, the personal parser, or a combination thereof; generating a new personal intent for the first NL input based on the determined intent; and training the personal parser using the existing personal intent and the new personal intent.
[0017] In another embodiment, an electronic device includes: a memory storing instructions; and at least one processor executing instructions including a process, the process being configured to: retrieve a first natural language (NL) input, wherein neither a universal parser nor a personal parser can determine the intent of the first NL input; retrieve a paraphrase of the first NL input; determine the intent of the paraphrase of the first NL input using at least one of the universal parser, the personal parser, or a combination thereof; generate a new personal intent for the first NL input based on the determined intent; and train the personal parser using the existing personal intent and the new personal intent.
[0018] In one embodiment, a non-transitory processor-readable medium includes a program that, when executed by a processor, performs a method comprising: retrieving a first natural language (NL) input on an electronic device, wherein neither a universal parser nor a personal parser can determine the intent of the first NL input; retrieving a paraphrase of the first NL input on the electronic device; determining the intent of the paraphrase of the first NL input using at least one of the universal parser, the personal parser, or a combination thereof; generating a new personal intent for the first NL input based on the determined intent; and training the personal parser using the existing personal intent and the new personal intent.
[0019] These and other aspects and advantages of one or more embodiments will become apparent from the following detailed description, which, when taken in conjunction with the accompanying drawings, illustrates, by way of example, the principles of the one or more embodiments.
[0020] The following description is intended to illustrate the general principles of one or more embodiments and is not meant to limit the inventive concepts claimed herein. In addition, the specific features described herein may be used in combination with other described features in each of the various possible combinations and arrangements. Unless otherwise expressly defined herein, all terms should be given the broadest interpretation, including the meaning implied from the specification and the meaning understood by those skilled in the art and / or as defined in dictionaries, monographs, etc.
[0021] It should be noted that the term "at least one of..." refers to the subsequent one or more elements. For example, "at least one of a, b, c or a combination thereof" can be interpreted as "a", "b" or "c", respectively, or as "a" and "b" combined together, or as "b" and "c" combined together; or as "a" and "c" combined together; or as "a", "b" and "c" combined together.
[0022] One or more embodiments provide personalized expressed intent learning for an intelligent personal assistant. Some embodiments include a method comprising retrieving a first natural language (NL) input on an electronic device. Neither a universal parser nor a personal parser can determine the intent of the first NL input. Retrieving a paraphrase of the first NL input on the electronic device. Determine the intent of the paraphrase of the first NL input using at least one of the universal parser, the personal parser, or a combination thereof. Generate a new personal intent for the first NL input based on the determined intent. Train the personal parser using the existing personal intent and the new personal intent.
[0023] In some embodiments, a personal assistant (PA) NL understanding (NLU) system includes two parsers, a "general intent parser" that is the same for each user, and a "personal paraphrase retriever" (personal parser) that is private and different for each user (i.e., personalized). When both the general parser and the personal paraphrase retriever cannot determine the intent of the user's NL input X (e.g., "find me a ride to the airport"), a "learn new intent" process will provide the user with the opportunity to define a new personalized intent. The user can define any new intent, which can be a combination of default intents (e.g., common intents such as dial a phone call, launch a website, send an email or text message, etc.) and personalized intents. In one embodiment, the personal paraphrase retriever does not need access to the training data (dataset) used to train the general intent parser, and can therefore be used with third-party parsers.
[0024] One or more embodiments can be easily scaled to support millions of users and provide an interface to allow users to re-paraphrase / interpret personalized expressions (e.g., convert speech to text, text input, etc.). Some embodiments integrate a personalized intent parser (personal paraphrase retriever) with an existing intent parser, thereby enhancing the intent parsing capabilities of the combined system and tailoring it to the end user. One or more embodiments can be integrated into existing parsers and integrated into off-the-shelf PAs. The personalized intent parser can understand complex user expressions and map them to a possible series of simpler expressions, and provide a scalable process involving a personalized parser that learns to understand increasingly personalized expressions over time.
[0025] Some advantages of one or more embodiments over conventional PA are that the process does not require access to training data for existing intent parsers. Some embodiments do not require modifying parameters of existing intent parsers in order to learn new intents. For one or more embodiments, no separate vocabulary summarization algorithm is required. Additionally, some embodiments are more scalable, practical, and have less retraining overhead than conventional PA.
[0026] Some embodiments improve the personalized language understanding capabilities of the smart PA. In addition, some embodiments can be easily integrated into any existing intent parser (without actually modifying it). When an expression is encountered that the intent parser cannot parse, the user will be provided with the opportunity to paraphrase the expression using a single or multiple simpler expressions that can be parsed. Using this paraphrase example provided by the user, machine learning techniques are then used to train a customized user-specific personalized intent parser so that the next time the user uses a similar (but not necessarily identical) expression to express the same intent, the entire process through the PA will now automatically parse the expression into the user's desired intent (e.g., for performing the desired action).
[0027] Figure 1 1 is a schematic diagram of a communication system 10 according to one embodiment. The communication system 10 may include a communication device (sending device 12) that initiates an outgoing communication operation and a communication network 110 that the sending device 12 may use to initiate and conduct communication operations with other communication devices within the communication network 110. For example, the system 10 may include a communication device that receives communication operations from the sending device 12 (receiving device 11). Although the communication system 10 may include multiple sending devices 12 and receiving devices 11, in Figure 1 Only one of each device is shown in order to simplify the drawing.
[0028] Any suitable circuits, devices, systems, or combinations of these that are operable to create a communication network (e.g., wireless communication infrastructure including communication towers and telecommunication servers) may be used to create the communication network 110. The communication network 110 can provide communications using any suitable communication protocol. In some embodiments, the communication network 110 may support, for example, traditional telephone lines, cable television, Wi-Fi (e.g., IEEE 802.11 protocol), High frequency systems (e.g., 900 MHz, 2.4 GHz, and 5.6 GHz communication systems), infrared, other relatively localized wireless communication protocols, or any combination thereof. In some embodiments, the communication network 110 may support wireless and cellular telephones and personal email devices (e.g., ) protocol used. Such protocols may include, for example, GSM, GSM plus EDGE, CDMA, quadband, and other cellular protocols. In another example, the remote communication protocol may include Wi-Fi and protocols for making or receiving calls using VOIP, LAN, WAN, or other TCP-IP based communication protocols. When located within the communication network 110, the sending device 12 and the receiving device 11 can communicate on a two-way communication path such as path 13 or on two unidirectional communication paths. Both the sending device 12 and the receiving device 11 are capable of initiating communication operations and receiving initiated communication operations.
[0029] The sending device 12 and the receiving device 11 may include any suitable device for sending and receiving communication operations. For example, the sending device 12 and the receiving device 11 may include, but are not limited to, mobile phone devices, television systems, cameras, camcorders, devices with audio and video capabilities, tablet computers, wearable devices, smart devices, smart photo frames, and any other device capable of wireless communication (with or without the help of wirelessly supported accessory systems) or through a wired path (e.g., using traditional telephone lines). The communication operations may include any appropriate form of communication, including, for example, data and control information, voice communication (e.g., telephone calls), data communication (e.g., emails, text messages, media messages), video communication, or a combination of these (e.g., video conferencing).
[0030] Figure 2 A functional block diagram of an architecture system 100 that can be used for a PA to enhance the natural language understanding capabilities and personalization of the PA, for example, using an electronic device 120 (e.g., a mobile phone device, a television (TV) system, a camera, a camcorder, a device with audio and video capabilities, a tablet computer, a tablet device, a wearable device, a smart device, a smart photo frame, a smart lighting, etc.) Transmitting device 12( Figure 1 ) and receiving device 11 may include some or all of the features of electronic device 120. In one embodiment, electronic device 120 may include display 121, microphone 122, audio output 123, input mechanism 124, communication circuit 125, control circuit 126, camera 128, personal intent learning (or PA) application 129 (including personal interpretation retriever 520 ( Figure 5 and 7 -9) and machine learning (ML) for learning personal intent from personal expressions), and communicate with communication circuit 125 to obtain / provide information from cloud or server 130; and may include, but is not limited to, any of the processes of the examples and embodiments described below and any other suitable components. In one embodiment, applications 1-N 127 are provided, and applications 1-N 127 may be obtained from a cloud or server 130, a communication network 110, etc., where N is a positive integer equal to or greater than 1.
[0031] In one embodiment, all applications employed by audio output 123, display 121, input mechanism 124, communication circuitry 125, and microphone 122 may be interconnected and managed by control circuitry 126. In one example, a handheld music player capable of sending music to other tuned devices may be incorporated into electronic device 120.
[0032] In one embodiment, audio output 123 may include any suitable audio components for providing audio to a user of electronic device 120. For example, audio output 123 may include one or more speakers (e.g., mono or stereo speakers) built into electronic device 120. In some embodiments, audio output 123 may include an audio component that is remotely coupled to electronic device 120. For example, audio output 123 may include an audio component that can be coupled to electronic device 120 via a wire (e.g., via a jack) or wirelessly (e.g., Headphones or Headset) is an earphone, headphone or earbud coupled to a communications device.
[0033] In one embodiment, display 121 may include any suitable screen or projection system for providing a visible display to a user. For example, display 121 may include a screen (e.g., an LCD screen, an LED screen, an OLED screen, etc.) incorporated into electronic device 120. As another example, display 121 may include a removable display or projection system (e.g., a video projector) for providing a display of content on a surface remote from electronic device 120. Display 121 may be operable to display content (e.g., information about communication operations or information about available media selections) under the direction of control circuitry 126.
[0034] In one embodiment, input mechanism 124 may be any suitable mechanism or user interface for providing user input or instructions to electronic device 120. Input mechanism 124 may take a variety of forms, such as buttons, a keypad, a dial, a click wheel, a mouse, a visual indicator, a remote control, one or more sensors (e.g., a camera or visual sensor, a light sensor, a proximity sensor, etc.), or a touch screen. Input mechanism 124 may include a multi-touch screen.
[0035] In one embodiment, the communication circuit 125 may be operable to connect to a communication network (e.g., Figure 1 The communication circuit 125 may be configured to communicate with other devices within the communication network 110 and transmit communication operations and media from the electronic device 120 to other devices within the communication network. The communication circuit 125 may be configured to use any suitable communication protocol, such as Wi-Fi (e.g., IEEE 802.11 protocol), High frequency systems (e.g., 900 MHz, 2.4 GHz, and 5.6 GHz communication systems), infrared, GSM, GSM plus EDGE, CDMA, quad-band and other cellular protocols, VOIP, TCP-IP, or any other suitable protocol.
[0036] In some embodiments, the communication circuit 125 may be operated to create a communication network using any suitable communication protocol. For example, the communication circuit 125 may use a short-range communication protocol to create a short-range communication network to connect to other communication devices. For example, the communication circuit 125 may be operated to use a short-range communication protocol to create a short-range communication network to connect to other communication devices. The protocol creates a local communication network to connect the electronic device 120 with Headphone coupling.
[0037] In one embodiment, the control circuit 126 can be operated to control the operation and performance of the electronic device 120. The control circuit 126 may include, for example, a processor, a bus (e.g., for sending instructions to other components of the electronic device 120), a memory, a storage, or any other suitable component for controlling the operation of the electronic device 120. In some embodiments, the processor can drive the display and process inputs received from the user interface. The memory and storage may include, for example, cache, flash memory, ROM and / or RAM / DRAM. In some embodiments, the memory can be dedicated to storing firmware (e.g., for device applications such as operating systems, user interface functions, and processor functions). In some embodiments, the memory can be operated to store information related to other devices with which the electronic device 120 performs communication operations (e.g., storing contact information related to communication operations or storing information related to different media types and media items selected by the user).
[0038] In one embodiment, the control circuit 126 may be operable to perform operations of one or more applications implemented on the electronic device 120. Any suitable number or type of applications may be implemented. Although the following discussion will list different applications, it should be understood that some or all of the applications may be combined into one or more applications. For example, the electronic device 120 may include applications 1-N 127, and the applications 1-N 127 include, but are not limited to: automatic speech recognition (ASR) applications, OCR applications, conversation applications, mapping applications, media applications (e.g., QuickTime, MobileMusic.app, or MobileVideo.app), social networking applications (e.g., The electronic device 120 may include one or more applications operable to perform communication operations. For example, the electronic device 120 may include a messaging application, an email application, a voicemail application, an instant messaging application (e.g., for chatting), a video conferencing application, a fax application, or any other suitable application for performing any suitable communication operation.
[0039] In some embodiments, the electronic device 120 may include a microphone 122. For example, the electronic device 120 may include a microphone 122 to allow a user to send audio (e.g., voice audio) for voice control and navigation of the applications 1-N 127 during performance of a communication operation, as a means of establishing a communication operation, or as an alternative to using a physical user interface. The microphone 122 may be incorporated into the electronic device 120, or may be remotely coupled to the electronic device 120. For example, the microphone 122 may be incorporated into a wired headset, the microphone 122 may be incorporated into a wireless headset, the microphone 122 may be incorporated into a remote control, etc.
[0040] In one embodiment, camera module 128 includes one or more camera devices that include functionality for capturing still and video images, editing functionality, communication interoperability for sending, sharing photos / videos, etc.
[0041] In one embodiment, the electronic device 120 may include any other components suitable for performing communication operations. For example, the electronic device 120 may include a power supply, a port or interface for coupling to a host device, an auxiliary input mechanism (e.g., an ON / OFF switch), or any other suitable component.
[0042] Figure 3 Example usage of PA for correct and incorrect NL expression understanding is shown. A conventional smart PA has a pre-trained intent parser (generic intent parser) that can map NL queries to intent classes and take appropriate actions based on the classes. The conventional generic intent parser is the same for each user and is not personalized for the individual. In a first scenario 310, the generic intent parser of PA 330 receives a command X 320 NL expression. In the first scenario 310, the generic intent parser of PA 330 understands command X 320 and PA 330 issues a correct action A 340.
[0043] In the second scenario 311, the generic intent parser of the PA 330 receives the command X 321 NL utterance, but the generic intent parser of the PA 330 cannot determine / understand the intent from the NL utterance (command X 321) with sufficient confidence, and then gives an output 322 (e.g., simulated speech) of "Sorry, I can't understand." Therefore, no action is taken 323. This is called failure case 1. In failure case 1, the conventional PA will wait for another utterance that the generic intent parser can understand.
[0044] In the third scenario 312, the generic intent parser of the PA 330 receives the command X 322 NL utterance. The generic intent parser of the PA 330 misunderstands the intent and issues an incorrect action B 324. This is referred to as failure case 2. For failure case 2, the user may become frustrated or may have to undo the incorrect action B 324 and repeat another / new NL utterance until the generic intent parser of the PA 330 can understand the intent.
[0045] Figure 4 Example PA scenarios for NL expression intent learning according to some embodiments are shown. In a first scenario 410, the general intent parser and the personal intent parser of PA 430 receive the command X 420 (X = find me a ride to the airport) NL expression. In the first scenario 410, neither the general intent parser nor the personal intent parser of PA 430 can understand the command X 420, and PA 430 outputs a response 422 "Sorry, I don't understand. Can you rephrase it?". Since PA 430 cannot understand its intent, no action is taken 423.
[0046] In some embodiments, in the second scenario 411, the universal intent parser of PA 430 receives the paraphrased / re-expressed command Y(s) 421NL expression. The universal parser of PA 430 understands the intent of the paraphrased / re-expressed command Y 421NL expression and issues the correct action A 440. The personal intent parser of PA 430 uses machine learning techniques to learn the intent of command X 420 to obtain the paraphrased / re-expressed command Y (421) (e.g., the intent dataset is updated based on the user's understood personal intent). In the third scenario 412, the universal intent parser of PA 430, which has access to the updated dataset, receives the command X′ 422NL expression of “I want to go to the airport.” PA 430 understands command X′ as being equal to command X 420 (i.e., X′=X at 450) and issues the correct action A 440. The embodiments described below provide more details on personalized PA learning and use and personal intent parsers (e.g., personal paraphrase retriever 520, Figure 5-9 )’s use.
[0047] Figure 5 1. The PA (eg, PA application 129 ( Figure 2 )) is a block diagram of reactive intent parsing. In block 501, an electronic device (e.g., Figure 2The electronic device 120 of the PA parses the NL (user) expression X′501 using the PA's general intent parser 510. The general parser 510 determines in box 515 whether the intent of the expression X′501 is found. If the intent is found in box 515, the PA issues a corresponding action in box 525, proceeds to box 550, and then stops (e.g., waiting for the next NL expression command). Otherwise, if the general intent parser 510 does not find the intent in box 515, the expression X′501 is sent to the PA's personal interpretation retriever (personalized intent parser) 520 in box 520. In box 530, it is determined whether a personalized interpretation is found in the PA dataset to determine the intent. If it is a personalized interpretation found in the PA (e.g., represented by the expression sequence Y(s)), the expression sequence Y(s) is sent to the general intent parser 510 again for parsing. If the intent is found in box 540, the action 526 is issued and executed by the PA. Then, PA proceeds to block 550 and stops (e.g., waiting for the next NL expression command). If no personalized paraphrase is found in block 530, then at block 535, machine learning techniques are used to learn the expression X′ (learn new intent algorithm), and the intent and paraphrase are added to the dataset (for following / future NL expressions), and PA stops at block 550 (e.g., waiting for the next NL expression with an updated dataset).
[0048] Using reactive intent parsing for PA, failure case 1 can be mitigated if a specific case X only occurs once ( Figure 3 ). If the user chooses to define the expression X as a new paraphrase / personalized intent, then the next time the generic intent parser 510 receives the NL expression X, the generic intent parser 510 queries for X′ that is functionally similar to X to correctly parse the expression X to obtain the correct intent.
[0049] Figure 6 FIG. 6 shows a block diagram for customizing the interpretation of a PA according to some embodiments. In block 610, in the general intent parser 510 ( Figure 5 ) and the personalized paraphrase retriever 520 failed to resolve the intent of the NL expression X, and the PA hopes to learn the intent. The PA prompts the electronic device (e.g., Figure 2The user of the electronic device 120) asks whether he / she wants to define the failed NL expression X as a new intent. In box 620, it is determined whether the user wants PA to learn a new command. If PA receives a "no" for the NL expression, PA proceeds to box 670 and stops (e.g., waiting for the next NL expression). Otherwise, if the user wants PA to learn a new command, the user is prompted to use an NL expression or a series of NL paraphrase expressions Y(s) to interpret X. In box 630, PA enters the paraphrase expression Y(s) of X. Then, PA checks whether each expression Y in Y(s) can be parsed by the general intent parser 510 or the personal paraphrase retriever 520 (personalized intent parser), but no action is taken at this time. If any expression in Y(s) fails to be parsed, PA again requests the user to enter a simpler paraphrase. After the user enters the paraphrase expression Y(s) of expression X, PA creates a new customized user-specific intent "I" in box 640, represented as a series of paraphrase expressions P. In box 650, PA adds the custom paraphrase expression P and intent I to personal paraphrase retriever 520. In box 660, personal paraphrase retriever 520 is retrained using all previous and new personalized intents. PA then proceeds to box 670 and stops (e.g., waiting for the next NL expression).
[0050] In some embodiments, the personal paraphrase retriever 520 can be created as follows. In one or more embodiments, the personal paraphrase retriever 520 is required to be able to map the NL expression X to a sequence of one or more expressions Y(s). In essence, the personal paraphrase retriever 520 is a paraphrase generator. A paraphrase generation model (e.g., a machine learning model) can be trained that can map a single NL expression X to a sequence of expressions Y(s), which together form the paraphrase of X. The paraphrase generation model is used as a personal paraphrase retriever (personalized intent parser) 520. When a new custom intent "I" and paraphrase P are added, the PA first checks whether the general intent parser 510 can parse each individual expression Y in the sequence Y(s). If so, the PA only adds this new example {X, Y(s)} to the personal paraphrase retriever 520 and retrains the personal paraphrase retriever (personalized intent parser) 520.
[0051] Figure 7A block diagram for proactive intent resolution for PA according to some embodiments is shown. In some embodiments, in block 501, a user NL expression X′ 501 is received and then parsed using a personal paraphrase retriever (personalized intent parser) 520. In block 530, it is determined whether the personal paraphrase retriever 520 has found a paraphrase (e.g., represented as a sequence of expressions Y(s)). If the personal paraphrase retriever 520 has found a paraphrase, the expression sequence Y(s) is sent to the general intent parser 510 for parsing. In block 515, it is determined whether an intent is found in the general intent parser 510. If an intent is found, in block 525, a corresponding action is performed, and PA processing stops at block 720 (e.g., PA waits for the next NL expression to act). If the personal paraphrase retriever 520 has not found a paraphrase in block 530, the NL expression X′ 501 (instead of Y(s)) is sent to be parsed by the general intent parser 510. In box 540, if an intent is found, then the corresponding action is performed in box 526, and the PA processing stops at box 720. If no intent is found in box 540, the NL expression X' is sent to box 535 to use machine learning (e.g., learn intent process, Figure 6 ) and user input to learn new intents. The PA process then proceeds to block 720 and stops (e.g., waiting for the next NL expression). In some embodiments, for block 710, the PA process also has a user intervention 710 input process, where the user can define new intents at any time, not necessarily after a parsing failure. This helps to solve failure cases 2 (which may be difficult for PA to detect without any user feedback) Figure 3 ). If the user chooses to define X′ as a new paraphrased / personalized intent, proactive intent parsing for PA processing could potentially mitigate failure cases 1 and 2.
[0052] Figure 8 A block diagram of recursive reactive intent resolution for PA according to some embodiments is shown. In some embodiments, recursive reactive intent resolution for PA combines features from proactive and reactive processing (e.g., Figure 5 and 7 ). In the learn new intent box 535, when the user interprets X′ with a sequence of expressions Y(s), the PA process also allows one or more of these interpreted expressions to become personalized expressions of the user, which can be mapped to the personalized intent in the current interpretation / personal intent retriever (personalized intent parser) 520.
[0053] In one or more embodiments, in box 501, a user NL expression X′ is received and sent to the general intent parser 510. In box 515, if an intent is found for the user NL expression X′ 501, the PA processing proceeds to box 840, where the intent actions are queued. The PA processing proceeds to box 820, where it is determined whether all intents are found in the paraphrase sequence Y(s). If all intents are found in box 820, the PA processing performs the actions in the queued actions in box 830, and the PA processing stops at box 850 (e.g., PA waits for the next NL expression). If all intents are not found, the PA processing proceeds to box 810, where an unresolved expression Y in Y(s) is selected and sent to the general intent parser 510 for recursive processing.
[0054] In one or more embodiments, in block 515, if no intent is found for user NL expression X'501, the PA process proceeds to send user NL expression X'501 to personal paraphrase retriever 520. In block 530, the PA process determines whether a personalized paraphrase is found. If a personalized paraphrase is found, the PA process proceeds to block 835, where Y'(s) is appended to Y(s), and the PA process proceeds to block 820 for recursive processing. Otherwise, if no personalized paraphrase is found, the PA process proceeds to block 535 to learn a new intent using machine learning (learn intent process, Figure 6 ). PA processing then proceeds to block 850 and stops (eg, waiting for the next NL expression).
[0055] Fig. 9 A block diagram of a recursive proactive intent parsing for PA according to some embodiments is shown. In one or more embodiments, the recursive proactive intent parsing automatically detects an erroneous action (or failure case 2, Figure 3 ). When the user Figure X When it is incorrectly parsed by the general intent parser 510 and mapped to the wrong action ( Figure 3 In the case of failure 2), the user can provide a paraphrase X' in the next utterance to correct the failed action. In some embodiments, the PA process trains and uses a paraphrase detector (PD) to detect whether two consecutive user utterances (X and X') are paraphrases. If the PD detects that X and X' are paraphrases, the PA process calls the learn intention process ( Figure 6 ) to learn the personalized intent when the user wants to define it.
[0056] In some embodiments, in box 501, a user NL expression X' is received and sent to a personal paraphrase retriever (personal intent resolver) 520. In box 530, if a personalized paraphrase is found for the user NL expression X' 501, the PA process proceeds to box 835, where Y'(s) is appended to Y(s). The PA process proceeds to box 810, where an unresolved expression Y in Y(s) is selected and sent to the general intent resolver 510. The PA process then continues to determine whether the general intent resolver 510 has found an intent. If it is determined that one or more intents are found in box 515, the PA process queues the intent action and proceeds to box 820. In box 820, the PA process determines whether all intents are found in Y(s). If it is determined that all intents are found in Y(s), the PA process executes the action from the queued actions in box 830, and the PA process stops at box 920 (e.g., PA waits for the next NL expression). If not all intents are found in Y(s) in block 820, the PA process proceeds to block 810 where unresolved expressions Y in Y(s) are selected and sent to the personal paraphrase retriever 520 for recursive processing.
[0057] In one or more embodiments, in block 515, if the intent of the user NL expression X′501 is not found, the PA process continues to send the user NL expression X′501 to block 910, where it is determined whether the parsing fails in both the general intent parser 515 and the personal paraphrase retriever 520. If it is determined that the parsing in both parsers fails, the PA process proceeds to block 535 to learn a new intent using machine learning (learn intent process, Figure 6 ). The PA process then proceeds to block 920 and stops (eg, waiting for the next NL expression). If it is determined in block 910 that both parsers have not failed, the PA process proceeds to block 820 for recursive processing.
[0058] In block 530, if no personalized paraphrase is found for the user NL expression X' 501, the PA process proceeds to the generic intent resolver 510 and (e.g., in parallel) to block 810, where an unresolved expression Y in Y(s) is selected and then sent to the generic intent resolver 510 (while also sending the expression X' to the generic intent resolver 510). The PA process then proceeds to block 515 and proceeds as described above.
[0059] Fig.10 1000 is a block diagram of a process 1000 for learning a personalized expression of a virtual PA according to some embodiments. In block 1010, the process 1000 is performed on an electronic device (e.g., electronic device 120, Figure 2) retrieves a first NL input (e.g., a user NL expression), where a universal parser (e.g., Figure 5 and 7 -9's general intent parser 510) and personal parsers (e.g., Figure 5-9 In box 1020, process 1000 retrieves a paraphrase of the first NL input at the electronic device. In box 1030, process 1000 uses a universal parser, a personal parser, or a combination thereof to determine the intent of the paraphrase of the first NL input. In box 1040, process 1000 generates a new personal intent for the first NL input based on the determined intent. In box 1050, process 1000 trains the personal parser using the existing personal intent and the new personal intent. In one or more embodiments, machine learning techniques are used to train the personal parser (e.g., see Figure 6 ). In block 1060, process 1000 updates the training data set with the new personal intent as a result of training the personal parser, where only the universal parser has access to the training data set.
[0060] In some embodiments, process 1000 may further include processing the second NL input on the electronic device using a personal parser for determining the personal intent of the second NL input (the second NL input may be similar to or different from the first NL input. When the personal intent cannot be determined, using a universal parser to determine a universal intent for the second NL input.
[0061] In one or more embodiments, process 1000 may include processing a third NL input on an electronic device using a personal resolver for determining a personal intent of the third NL input (the third NL input may be similar to or different from the first NL input. The personal resolver then determines whether the third NL input is a paraphrase of the first NL and determines a new personal intent for the third NL input.
[0062] In some embodiments, process 1000 may further include processing a fourth NL input on the electronic device using a personal parser for determining a personal intent of the fourth NL input (the fourth NL input may be similar to or different from the first NL input. The fourth NL input is parsed into a sequence of known inputs, and one or more intents of the fourth NL input are determined based on processing the sequence of known inputs. In one or more embodiments, the personal parser is private and personalized (e.g., to a user of the electronic device), and the new personal intent includes a combination of a default intent and a personalized intent. In process 1000, the personal parser improves the personalized NL understanding of the smart PA.
[0063] Fig.111100 includes one or more processors 1111 (e.g., ASICs, CPUs, etc.), and may further include: an electronic display device 1112 (for displaying graphics, text, and other data), a main memory 1113 (e.g., random access memory (RAM), cache devices, etc.), a storage device 1114 (e.g., a hard drive), a removable storage device 1115 (e.g., a removable storage drive, a removable memory, a tape drive, an optical drive, a computer-readable medium having stored therein computer software and / or data), a user interface device 1116 (e.g., a keyboard, a touch screen, a keypad, a pointing device), and a communication interface 1117 (e.g., a modem, a wireless transceiver (e.g., Wi-Fi, a cellular network)), a network interface (e.g., an Ethernet card), a communication port or a PCMCIA slot and card).
[0064] The communication interface 1117 allows software and data to be transferred between the computer system and external devices via the Internet 1150, mobile electronic devices 1151, servers 1152, networks 1153, etc. The system 1100 also includes a communication infrastructure 1118 (e.g., a communication bus, cross bar, or network) connected to the above-mentioned devices 1111 to 1117.
[0065] Information transmitted via communication interface 1117 may be in the form of signals, such as electronic, electromagnetic, optical, or other signals capable of being received by communication interface 1117 via a communication link carrying the signals, and may be implemented using wire or cable, optical fiber, a telephone line, a cellular telephone link, a radio frequency (RF) link, and / or other communication channels.
[0066] In one implementation of one or more embodiments in a mobile wireless device (e.g., a mobile phone, a tablet computer, a wearable device, etc.), the system 1100 further includes: an image capture device 1120, such as a camera 128 ( Figure 2 ); audio capture device 1119, such as microphone 122 ( Figure 2 ). The system 1100 may further include application processing or processors, such as MMS 1121, SMS 1122, email 1123, social network interface (SNI) 1124, audio / video (AV) player 1125, web browser 1126, image capture 1127, etc.
[0067] In one embodiment, the system 1100 includes a personal intent learning process 1130 that can implement the personal intent learning application 129 ( Figure 2 ) and for the similar processing described above with respect to Figure 5-10 In one embodiment, the personal intent learning process 1130 may be implemented with the operating system 1129 as executable code residing in a memory of the system 1100. In another embodiment, the personal intent learning process 1130 may be provided in hardware, firmware, or the like.
[0068] In one embodiment, the main memory 1113 , the storage device 1114 , and the removable storage device 1115 may each store, alone or in any combination, instructions of the above-described embodiments that may be executed by one or more processors 1111 .
[0069] As known to those skilled in the art, according to the above architecture, the above example architecture described above can be implemented in a variety of ways, such as program instructions executed by a processor, as a software module, microcode, as a computer program product on a computer-readable medium, as an analog / logic circuit, as an application-specific integrated circuit, as firmware, as a consumer electronic device, AV device, wireless / wired transmitter, wireless / wired receiver, network, multimedia device, etc. Embodiments of the described architecture can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment containing both hardware and software elements.
[0070] One or more embodiments have been described with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to one or more embodiments. Each frame or combination thereof of such diagrams / charts can be implemented by computer program instructions. When the computer program instructions are provided to the processor, the machine program produces a machine so that the instructions executed by the processor create a device for implementing the function / operation specified in the flowchart and / or block diagram. Each frame in the flowchart / block diagram can represent the hardware and / or software module or logic that implements one or more embodiments. In alternative embodiments, the functions indicated in the frame may not occur in the order indicated in the figure, may occur simultaneously, etc.
[0071] The terms "computer program medium", "computer usable medium", "computer readable medium" and "computer program product" are generally used to refer to media such as main memory, secondary memory, removable storage drives, hard disks installed on hard drives, and the like. These computer program products are devices for providing software to computer systems. The computer readable medium allows a computer system to read data, instructions, messages or message packets, and other computer readable information from the computer readable medium. For example, computer readable media may include non-volatile memory such as floppy disks, ROMs, flash memory, disk drive memory, CD-ROMs, and other permanent memory. For example, it is useful for transferring information (e.g., data and computer instructions) between computer systems. Computer program instructions may be stored in a computer readable medium that can direct a computer, other programmable data processing device, or other device to operate in a specific manner so that the instructions stored in the computer readable medium produce an article of manufacture including the instructions. They implement the functions / actions specified in the flowchart and / or block diagram blocks.
[0072] The computer program instructions representing the block diagrams and / or flow charts of this article can be loaded onto a computer, a programmable data processing device, or a processing device so that a series of operations performed thereon produce a computer-implemented process. The computer program (i.e., computer control logic) is stored in a main memory and / or an auxiliary memory. The computer program can also be received via a communication interface. Such a computer program enables a computer system to perform the features of the embodiments discussed herein when executed. In particular, the computer program enables a processor and / or a multi-core processor to perform the features of a computer system when executed. Such a computer program represents a controller of a computer system. A computer program product includes a tangible storage medium that can be read by a computer system and stores instructions executed by a computer system to perform the method of one or more embodiments.
[0073] Although embodiments have been described with reference to particular versions thereof, other versions are possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the preferred versions contained herein.
Claims
1. A method comprising: Retrieving a first natural language NL input of a user on the electronic device, wherein neither the universal parser nor the personal parser can determine the intent of the first NL input; Receiving a first paraphrased NL expression of a first NL input on an electronic device; Using the universal parser to determine whether the intent expressed by the first paraphrase NL is found; When the universal parser cannot determine the intention of the first paraphrased NL expression, using the personal parser to determine the intention of the first paraphrased NL expression; When the personal parser successfully determines the intent of the first paraphrased NL expression, performing an action corresponding to the intent of the first paraphrased NL expression; Determine the intent of the first paraphrase NL expression using a personal parser; When both the universal parser and the personal parser are unable to determine the intent of the first paraphrased NL expression, determining the intent of the first paraphrased NL expression by using a machine learning model; generating a new personal intention for the first NL input based on the determined intention; updating the training dataset by adding new personal intents and first paraphrased NL expressions to the training dataset; and Use the updated training dataset to train your personal parser.
2. The method according to claim 1, wherein: Only the universal parser has access to the training dataset.
3. The method according to claim 1, further comprising: processing a second NL input on the electronic device using the personal resolver to determine a personal intent of the second NL input, wherein the second NL input is different from the first NL input; and When the individual intent cannot be determined, a universal parser is used to determine the universal intent of the second NL input.
4. The method according to claim 1, further comprising: processing a third NL input on the electronic device using the personal resolver to determine the personal intent of the third NL input, wherein the third NL input is different from the first NL input; determining, by the personal parser, that the third NL input is a paraphrase of the first NL; and Determine the new personal intention for the third NL input.
5. The method according to claim 1, further comprising: processing a fourth NL input on the electronic device using the personal resolver to determine a personal intent of the fourth NL input, wherein the fourth NL input is different from the first NL input; Parsing the fourth NL input as a known input sequence; and One or more intentions of a fourth NL input are determined based on processing of the known input sequence.
6. The method according to claim 1, wherein: The new personal intent includes a combination of a default intent and a personalized intent.
7. The method according to claim 1, wherein: The personal parser improves the personalized NL understanding of the intelligent personal assistant PA, and training the personal parser includes a machine learning model training process.
8. An electronic device comprising: Memory, which stores instructions; and at least one processor executing instructions comprising a process configured to: Retrieving a first natural language NL input of the user, wherein neither the universal parser nor the personal parser can determine the intent of the first NL input, receiving a first paraphrased NL expression of a first NL input, Use the generic parser to determine if the intent of the first paraphrased NL expression is found, When the universal parser cannot determine the intention of the first paraphrased NL expression, using the personal parser to determine the intention of the first paraphrased NL expression; When the personal parser successfully determines the intent of the first paraphrased NL expression, performing an action corresponding to the intent of the first paraphrased NL expression; Determine the intent of the first paraphrase NL expression using a personal parser; When both the universal parser and the personal parser are unable to determine the intent of the first paraphrased NL expression, determining the intent of the first paraphrased NL expression by using a machine learning model, Generate a new personal intent for the first NL input based on the determined intent, updating the training dataset by adding new personal intents and first paraphrased NL expressions to the training dataset; and Use the updated training dataset to train your personal parser.
9. The electronic device according to claim 8, wherein: The process is also configured to: Processing a second NL input using the personal parser to determine a personal intent of the second NL input, wherein the second NL input is different from the first NL input; and When the individual intent cannot be determined, a universal parser is used to determine the universal intent of the second NL input.
10. The electronic device according to claim 8, wherein: The process is also configured to: processing a third NL input using the personal parser to determine the personal intent of the third NL input, wherein the third NL input is different from the first NL input; determining, by the personal parser, that the third NL input is a paraphrase of the first NL; and Determine the new personal intention for the third NL input.
11. The electronic device according to claim 8, wherein: The process is also configured to: processing a fourth NL input using the personal parser to determine a personal intent of the fourth NL input, wherein the fourth NL input is different from the first NL input; Parsing the fourth NL input as a known input sequence; and One or more intentions of a fourth NL input are determined based on processing of the known input sequence.
12. The electronic device according to claim 8, wherein: The new personal intent includes a combination of a default intent and a personalized intent; The personal parser improves the personalized NL understanding of the intelligent personal assistant PA, and Training the personal parser includes a machine learning model training process.
13. A computer-readable medium comprising a program which, when executed by a processor, performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Dialogue apparatus and method
US20170140754A1