CONTEXT-DRIVEN MESSAGE TRANSMISSION SYSTEM

DE112016002588B4Active Publication Date: 2026-07-09GOOGLE LLC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2016-05-04
Publication Date
2026-07-09

AI Technical Summary

Technical Problem

Existing computing devices require user intervention for sending and receiving text-based messages, including confirming recipients and message content, which is cumbersome and time-consuming.

Method used

A computing device equipped with speech recognition and synthesis capabilities automatically determines the intended recipient and content of messages based on context information, allowing for seamless text-based conversations without additional user input.

Benefits of technology

This solution reduces the need for manual prompts and confirmations, enabling efficient and natural conversational exchanges by automatically processing incoming and outgoing messages.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method for automatically detecting a conversation and determining the intended recipient of an outgoing communication, comprising: receiving, by a computer device associated with a first user, a message from a device associated with a second user; receiving, by the computer device, an audio input; determining, by the computer device and without issuing a request for instructions to the first user, and based at least in part on the message from the device associated with the second user, the audio input, and contextual information, whether the first and second users are conducting the conversation; in response to a determination that the first and second users are conducting the conversation: determining, by the computer device, that the second user is the intended recipient of the outgoing communication;Generating, by the computer device and based on the audio input, the outgoing communication as a response to the message received by the device associated with the second user; and sending, by the computer device and without receiving any user acknowledgment of the contents of the outgoing communication or any user instruction to send the outgoing communication, the outgoing communication to the second user.
Need to check novelty before this filing date? Find Prior Art

Description

GENERAL STATE OF THE ART

[0001] Some computer devices are equipped with speech recognition functionality to convert spoken language into text. For example, a computer device might have speech recognition functionality that can receive audio input (such as a user's voice) and determine written content (such as a text message, email, search query, device command, etc.) based on the audio input. Some computer devices are equipped with speech synthesis functionality to convert written text into spoken words. For example, a computer device might have speech synthesis functionality that can receive text content and output audio data that specifies the text content.

[0002] A user can instruct a computer to search for audio input, so that the computer receives the audio and converts it to text. The user may need to confirm the message contents and then instruct the computer to send the message. The user may need to repeat these steps each time they want to send a message.

[0003] Similarly, a computer device can receive text communication and ask the user if they wish to listen to the content. The computer device can display a prompt to the user each time text communication is received before converting the text to speech. SUMMARY OF THE REVELATION

[0004] In one example, a procedure might involve a computer device associated with a user receiving a message from a source and the computer device receiving an audio input. The procedure might involve the computer device determining, based at least partially on the audio input and contextual information, the probability that the user intends to send a reply message to the source. Furthermore, in response to a determination that the probability of the user intending to send the reply message to the source reaches a certain probability threshold, the procedure might involve the computer device determining that the user intends to send the reply message to the source.The procedure may also involve, in response to a determination that the user intends to send the reply message to the originating source, the generation of the reply message by the computer device based on the audio input, and the sending of the reply message to the originating source by the computer device.

[0005] In another example, a device can include an audio output device, an audio input device, a communication unit, and a message management module, operable by at least one processor. The message management module can receive a message from a source via the communication unit. The message management module can also receive audio input via the audio input device. Furthermore, the message management module can determine, at least partially based on the audio input and contextual information, a probability that a user associated with the device intends to send a reply message to the source.The message management module can determine, in response to a determination that the probability of the user intending to send the reply message to the originating source has reached a certain probability threshold, that the user intends to send the reply message to the originating source. Furthermore, in response to a determination that the user intends to send the reply message to the originating source, the message management module can generate the reply message based on the audio input and send the reply message to the originating source via the communication unit.

[0006] In another example, a computer-readable storage medium may contain instructions which, when executed, configure one or more processors of a computer system to receive a message from a source source, receive audio input, determine, at least in part based on the audio input and context information, a probability that a user associated with the computer system intends to send a reply message to the source source, and, in response to determining that the probability that the user intends to send a reply message to the source source reaches a probability threshold, determine that the user intends to send the reply message to the source source.The instructions, when executed, further configure the one or more processors, in response to a determination that the user intends to send the response message to the originating source, to generate the response message based on the audio input and to send the response message to the originating source.

[0007] In another example, a method may involve the output of an audio signal by a computer device associated with a user, representing a text message from a source source. The method may involve the computer device receiving audio data representing a speech utterance from the user. The method may also involve the computer device determining, without additional input from the user, a probability that the user intends to send a response, based at least partially on the audio data and one or more of the frequency of incoming messages from the source source, the frequency of outgoing messages to the source source, the time since the last message was received from the source source, or the time since the last message was sent to the source source.Furthermore, in response to a determination that the probability has reached a threshold and without additional input from the user, the procedure can involve transmitting a transcription of at least part of the audio data to the source.

[0008] The details of one or more examples of the disclosure are set forth below in the accompanying drawings and description. Other features, subject matter, and advantages of the invention will become apparent from the description, the drawings, and the claims. List of characters Fig. Figure 1 shows a conceptual diagram illustrating an exemplary system for sending and receiving text-based messages according to one or more aspects of the present disclosure. Fig. Figure 2 shows a block diagram illustrating an exemplary computer device configured to send and receive text-based messages according to one or more aspects of the present disclosure. FIGS. 3A-3H show conceptual diagrams illustrating an exemplary operation of the computer device. Fig. Figure 4 shows a flowchart illustrating an example of the operation of the computer device. Fig. Figure 5 shows a flowchart illustrating an example of the operation of the computer device. DETAILED DESCRIPTION

[0009] In general, techniques from this disclosure can enable a computer device to automatically determine that a user is engaged in a text-mediated conversation and to facilitate such conversations. In some examples, the computer device can automatically perform speech synthesis conversions on incoming communications and automatically perform speech recognition conversions on outgoing communications. In several instances, techniques from this disclosure can enable a computer device to intelligently determine an intended recipient of an outgoing communication. A computer device can determine the probability that the user intends to send a message to a particular recipient and, based on that probability, determine whether a message is sent to that particular recipient.In this way, techniques from this revelation, instead of requiring the user to instruct the computer device to send a message, confirm the recipient and the content of the message, can enable the computer device to automatically recognize the conversation and automatically determine an intended recipient of an outgoing communication, which can reduce the amount of user interaction required for the user to participate in the conversation.

[0010] Fig. Figure 1 shows a conceptual diagram that, according to one or more aspects of the present disclosure, represents a system 100 illustrated as an exemplary system for sending and receiving text-based messages. System 100 includes computer equipment 110 , Information Server System (“ISS”) 160 and message transmission devices 115A - 115N(collectively referred to as "Message Transmission Devices 115"), which operate via network 130 are communicatively coupled.

[0011] Message transmission devices 115 Each represents a computer device, such as a mobile phone, a laptop computer, a desktop computer, or another type of computer device that is configured to transmit information over a network, such as a network. 130 , to send and receive. Messaging devices 115 These include text-based messaging applications for sending and receiving text-based messages, such as email, short message service (SMS), multimedia messaging service (MMS), instant messaging (IM), or other types of text-based messages. The messaging devices 115 form a group of message transmission devices, from which the respective ones communicate with the message transmission devices 115A - 115Nassociated user text-based messages to computer device 110 send and text-based messages from computer device 110 can receive them.

[0012] computer device 110 This can be a mobile device, such as a mobile phone, tablet computer, laptop computer, smartwatch, smart glasses, smart gloves, or any other type of wearable computing device. Additional examples of computing devices 110 This includes desktop computers, televisions, personal digital assistants (PDAs), portable gaming systems, media players, e-book readers, mobile television platforms, automotive navigation and entertainment systems, or any other types of portable and non-portable computing devices configured to transmit information over a network, such as the network. 130 , to send and receive.

[0013] computer device 110includes a user interface device 112 , a user interface (UI) module 111 and a message management module (MMM) 120 Module 111 , 120 can perform the described operations using software, hardware, firmware, or a combination of software, hardware, and firmware installed in the computer device 110 is resident and / or is executed on it. Computer device 110 can the modules 111 , 120 Run with multiple processors or multiple devices. Computer device 110 can the modules 111 , 120 as virtual machines running on underlying hardware. The modules 111 , 120 They can be executed as one or more services of an operating system or computer platform. The modules 111 , 120can be executed as one or more executable programs at an application level of a computer platform.

[0014] UID 112 of the computer device 110 can be used as a suitable input and / or output device for the computer device 110 function. UID 112 It can be implemented using various technologies. UID 112 For example, it can function as an input device using presence-sensitive input screens, such as resistive touchscreens, SAW (Surface Acoustic Wave) touchscreens, capacitive touchscreens, projective capacitive touchscreens, pressure-sensitive screens, APR (Acoustic Pulse Recognition) touchscreens, or other presence-sensitive display technologies. Additionally, UID can 112This includes microphone technologies, infrared sensor technologies, or other input device technologies for use in receiving user input.

[0015] UID 112 It may also use one or more display devices, such as LCDs (Liquid Crystal Displays), dot matrix displays, LEDs (Light Emitting Diodes), OLEDs (Organic Light Emitting Diodes), e-paper displays, or similar monochrome or color displays that provide visible information to a user of a computer device. 110 They can be output and function as an output device (e.g., display device). Additionally, UID can 112 This includes loudspeaker technologies, haptic feedback technologies, or other output device technologies for use in delivering information to a user.

[0016] UID 112Each can include presence-sensitive displays that require tactile input from a user of the respective computer device. 110 can receive UID 112 It can receive information from tactile input by recognizing one or more gestures from a user (e.g., from the user who uses a finger or a stylus to touch one or more digits of the UID). 112 touches or points to). UID 112 It can present output to a user, e.g., on relevant presence-sensitive displays. UID 112 The output can be in the form of the respective graphical user interfaces (e.g., user interface). 114 ) represent the computer device 110 The provided functionality can be assigned to it. For example, UID can be... 112 various user interfaces (e.g. user interface) 114) presenting which relate to text-based messages or other features of computer platforms, operating systems, applications and / or services that are connected to or from computer equipment 110 are executed or accessible (e.g., electronic messaging applications, internet browser applications, mobile or desktop operating systems, etc.). UID 112 It can output audio signals to a user, for example using a speaker. For example, UID can 112 Output audio signals that indicate the content of a text-based message.

[0017] UI module 111 manages user interactions with UID 112 and other components of the computer device 110 UI module 111 can UID 112 cause a user interface, such as a user interface, to appear. 114 (or other exemplary user interfaces) to be displayed when a user of a computer device 110Outputs considered and / or inputs to UID 112 performs. UI module 111 and UID 112 can receive one or more pieces of information from user input at different times when users interact with the graphical user interface and when the user and computer device 110 located at different locations. UI module 111 and UID 112 can be linked to UIDs 112 Interpret recognized inputs and provide information about the UIDs. 112 forward recognized inputs to one or more linked platforms, operating systems, applications and / or services located on the computer device 110 to be executed, for example, to operate a computer device 110 to cause functions to be executed.

[0018] UI module 111can receive information and instructions from one or more linked platforms, operating systems, applications and / or services connected to the computer device 110 and / or one or more remote computer systems, such as ISS 160 , can be executed. Additionally, the UI module can be used. 111 act as an intermediary between one or more linked platforms, operating systems, applications and / or services that are connected to the computer device 110 and the different output devices of the computer device 110 (e.g., loudspeakers, LED displays, audio or electrostatic haptic output devices, etc.) are executed to produce an output (e.g., a graphic, a flash of light, a sound, a haptic response, etc.) to the computer device 110 to produce.

[0019] The ISS 160represents any remote computer system, such as one or more desktop computers, laptops, mainframes, servers, cloud computing systems, etc., that is used to send and receive information to and from a network, such as a network. 130 are able to. ISS 160 It hosts (or at least provides access to) speech recognition services for converting speech into text-based messages and speech synthesis services for converting text-based messages into audio data. In some examples, ISS represents 160 a cloud computing system that provides speech recognition and speech synthesis services through a network 130 for one or more computer devices 110 provides access to the resources provided by the ISS 160 Provided cloud access to speech recognition and speech synthesis services.

[0020] The network 130Represents any public or private communications network, such as a mobile network, Wi-Fi, and / or any other network type used to transmit data between computer systems, servers, and computer devices. Network 130 can include one or more network hubs, network switches, network routers, or any other network devices that are operationally coupled together, thereby enabling the exchange of information between the ISS 160 , computer device 110 and the messaging devices 115 will be provided. Computer device 110 , message transmission devices 115 and ISS 160 Data can be transferred over the network using any suitable communication techniques. 130 send and receive.

[0021] ISS 160 , computer device 110 and communication devices 115Each can be operated operationally using appropriate network connections with the network 130 be docked. ISS 160 , computer device 110 and communication devices 115 can operate using various network connections with the network 130 be coupled. The connections that the ISS 160 , computer device 110 and messaging devices connected to the network 130 The connections can be Ethernet, ATM or other types of network connections, and these connections can be wireless and / or wired.

[0022] According to techniques of the present disclosure, system 100 Automatically recognize the conversation and automatically determine an intended recipient of outgoing communication. For example, one or more messaging devices can be used. 115 via network 130a message to computer device 110 send. Computer device 110 The computer device receives the message and can, in response, output a message description. 110 It can determine whether it outputs a visual (e.g., graphic) or audible indication of the message. Computer device 110 can determine whether to output a message without additional input from the user (e.g., acoustic or gesture-based input).

[0023] In response to a request to output an audio message, the computer device can 110 Convert the text-based message into audio data that specifies the message by performing speech synthesis processing on the message. In some examples, computer equipment 110 , in order to convert the text-based message into audio data, at least part of the message is sent to the ISS for speech synthesis processing 160 send. Speech synthesis module 164 from the ISS160 can convert at least part of the message into audio data, while the ISS 160 the audio data to computer device 110 can send. In several cases, computer devices 110 and ISS 160 Each computer device performs speech synthesis processing on at least part of the message to convert the text-based message into audio data that indicates the message content. (Computer device) 110 Can the audio data be accessed via UID? 112 spend.

[0024] After computer device 110 The computer device outputs audio data indicating the received message. 110 Detect that a user is speaking (e.g., in a conversation with another person, when providing audio input to a computer device). 110 , when singing along to a song on the radio, etc.). Computer device 110 audio data of the language can be transmitted via UID 112receive the audio data and determine whether to send a text-based reply message. Computer device 110 can determine whether to send a text-based response message without additional input from the user (e.g., acoustic or gesture-based input).

[0025] If computer device 110 Determined that the user intended to send a reply message, the computer device 110 Convert the audio data into text data that specifies the audio data by performing speech recognition processing on the audio data. In some examples, computer equipment 110 at least part of the audio data for speech recognition processing at ISS 160 send. Speech recognition module 162 can convert at least some of the audio data into text data, while ISS 160 the text data to computer device 110 can send. In some examples, both computer devices can be used. 110as well as the ISS 160 Perform speech recognition processing on at least a portion of the audio data and convert the audio data into text data that specifies the audio data. Computer device 110 can generate a text-based reply message using the text data. Computer device 110 Can the reply message be sent to a specific messaging device? 115 send.

[0026] computer device 110 Can a text-based message be sent from a messaging device? 115A Received. Message transmission device 115A can be used with a contact in the contact list of a computer device 115S (e.g., Aaron) be associated with. Computer device 110 can be done via UI 114 Display a graphical representation of the message. For example, a computer device can 110 UI 114cause the message to be displayed: "Incoming message from Aaron: 'Are you coming to Jimmy's tonight?'" Similarly, computer equipment 110 a text-based message from a second messaging device 115B received, which are associated with a contact in the contact list for computer device 110 (e.g., Jimmy) may be associated with a computer device. 110 can UI 114 cause the message to be displayed: "Incoming message from Jimmy: 'Are you coming tonight?'"

[0027] In some examples, MMM can 120 determine to output an audio description of the received messages. In some examples, MMM determines 120 , whether it outputs an audio summary of the message without additional user input. In response to a request to output an audio summary of the initial message, the computer device can 110 the UID 112(e.g., a speaker) to output the audio data: "Incoming message from Aaron: 'Are you coming to Jimmy's tonight?'" In response to a request to output an audio sample of the second message, the computer device can 110 UID 112 (e.g., a loudspeaker) to output the audio data: "Incoming message from Jimmy: 'Are you coming tonight?'"

[0028] After computer device 110 The audio data that indicates the first received message and / or the second received message can be accessed by a user of a computer device. 110 Speak a response. For example, the user can reply to the first message by saying "Yes." Computer device 110 can recognize the user's response and can be accessed via UID. 112 (e.g., a microphone) receives audio data that indicates the answer. MMM 120can determine whether to send a text-based reply message to the messaging device 115A sends. In some examples, MMM can 120 Without additional user input, the device determines whether to send a reply message. In response to a user's instruction, a reply message is sent to the messaging device. 115A to send, computer device 110 Generate a text-based response message based on audio data. Computer device 110 Can the reply message be sent to the messaging device? 115A send. In some examples, computer device 110 Provide a visual or audible indication that the reply message has been sent. For example, the computer device can play the audio cue "Message sent to Aaron".

[0029] A user of computer equipment 110 can reply to the second received message, for example by saying "Yes." Computer device 110can recognize the user's response, and MMM 120 can determine whether to send a text-based reply message to one or both of the message delivery devices. 115A , 115B sends. In some examples, MMM can 120 The determination can be made without additional user input. In some examples, MMM can 120 specify that a reply message should be sent to only one messaging device (e.g., messaging device). 115B to send. Computer device 110 can generate a text-based response message based on the audio data. Computer device 110 Can the reply message be sent to the messaging device? 115B send. In some examples, computer device 110 Provide a visual or audible indication that the reply message has been sent. For example, a computer device can 110 output the audio data "Message sent to Jimmy."

[0030] Techniques of this revelation can simplify and accelerate the exchange of text-based messages. By automatically determining whether a user is engaged in a text-based conversation, techniques of this revelation can reduce or eliminate cumbersome and time-consuming prompts, voice acknowledgments, and touch inputs that would otherwise be required to send or hear a received text-based message. Techniques of this revelation can enable a computer device to efficiently process communications by shifting the conversation from a cumbersome transaction-oriented approach to a more natural, dialogue-oriented one.

[0031] Fig. Figure 2 shows a conceptual diagram illustrating an exemplary computer device configured to send and receive text-based messages. 210 out of Fig. 2 will be discussed below within the context of Fig. 1 described. Fig. 2 illustrates only one specific example of computer equipment 210 , while many other examples of computer equipment 210 can be used in other cases. Other examples of the computer device 210 may include a subset of the components found in the exemplary computer device 210 are included, or may include additional components not included. Fig. 2 will be shown.

[0032] As in the example from Fig. 2 shown, includes computer device 210 a user interface device (UID) 212 , one or more processors 240 , one or more input devices 242 , one or more communication units 244 , one or more output devices 246 and one or more storage devices 248 Storage device 248 of computer device210 It also includes a message management module 220 . MMM 220 application modules 222A - 222N (collectively referred to as "Application Module 222"), speech recognition module 224 , speech synthesis module 226 and Conversation Management Module (CMM) 228 include one or more communication channels. 250 can each of the components 212 , 240 , 242 , 244 , 246 and 248 to connect components (physical, communicative, and / or operational) for communication purposes. In some examples, the communication channels can 250 include a system bus, a network connection, a cross-process communication data structure, or another technology for communicating data.

[0033] One or more input devices 242 of the computer device 210They can receive input. Examples of input include tactile, motion, audio, and video input. The input devices 242 of the computer device 210 In one example, a presence-sensitive display could be used. 213 , a touch-sensitive screen, a mouse, a keyboard, a voice response system, a video camera, a microphone (e.g. microphone) 243 ) or include another type of device for detecting human or machine input.

[0034] One or more output devices 246 of computer device 210 They can produce outputs. Examples of outputs are tactile, electromagnetic, audio, and video outputs. The output devices 246 of computer device 210 One example includes a presence-sensitive display, loudspeakers (e.g., speakers). 247), a cathode ray tube (CRT) monitor, a liquid crystal display (LCD), or any other type of device for generating output to humans or machines. The output devices 246 can use one or more from a sound card or a video graphics adapter card to produce either acoustic or visual output.

[0035] One or more communication units 244 of computer device 210 Communication units can communicate with external devices over one or more networks by sending and / or receiving network signals over those networks. 244 They can connect to any public or private communication network. For example, the computer device 210 Communication unit 244They are used to send and / or receive radio signals in a wireless network, such as a mobile network. Similarly, communication units can be used. 244 Transmit and / or receive satellite signals in a global navigation satellite system (GNSS) network, such as the global positioning system (GPS). Examples of the communication unit 244 They can include a network interface card (e.g., an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, or any other type of device capable of sending or receiving information. Other examples of communication units 244 These can include shortwave radios, mobile data radios, wireless Ethernet network radios (e.g. WLAN), and Universal Serial Bus (USB) interfaces.

[0036] One or more storage devices 248 within computer device 210Information may be processed during the operation of the computer device. 210 save. In some examples, the storage device functions as 248 as a temporary storage device, which means that the storage device 248 not used for long-term storage. The storage devices 248 of the computer device 210 Memory devices can be configured as volatile storage for short-term information storage, which means that when they are switched off, the stored contents are lost. Examples of volatile memory include random access memory (RAM), dynamic memory (DRAM), static memory (SRAM), and other forms of volatile memory known in the field.

[0037] The storage devices 248 In some examples, they also include one or more computer-readable storage media. The storage devices 248They can store larger amounts of information than volatile storage devices. 248 They can also be configured for long-term information storage as non-volatile memory and to retain information after power cycles. Examples of non-volatile memory include magnetic hard disks, optical hard disks, floppy disks, flash memory, or forms of electrically programmable memory (EPROM) or electrically rewritable and programmable memory (EEPROM). These storage devices 248 Program instructions and / or data can be used in conjunction with the modules. 220 , 222 , 224 , 226 and 228 save.

[0038] One or more processors 240 can implement functions and / or instructions within the computer device 210 execute. The processors 240 of computer device 210For example, they can receive and execute instructions from the storage devices. 248 were saved, which enabled the functionality of the message management module 220 , application modules 222 , speech recognition module 224 , speech synthesis module 226 and CMM 228 execute these through the processors. 240 Instructions followed can be used with computer equipment 210 cause the program to run in the storage devices during program execution. 248 To store information. The processors 240 can be in modules 220 , 222 , 224 , 226 and 228 Execute instructions to convert audio input to text and send a text-based message based on the audio input, or to convert a text-based message to speech and output audio based on the text message. This means that modules 220 , 222 , 224, 226 and 228 through the processors 240 They can be operated to perform multiple actions, including converting received audio data and sending the transcribed data to a remote device, as well as converting received text data into audio data and outputting the audio data.

[0039] Application modules 222 can include any other application that uses a computer device 210 in addition to the other modules specifically described in this disclosure, the application modules can execute. For example, the application modules 222 Messaging applications (e.g., email, SMS, MMS, IM, or other text-based messaging applications), a web browser, a media player, a file system, a mapping program, or any other number of applications or features that include computer equipment 210 may include.

[0040] According to the techniques of this revelation, computer equipment 210 Determine the probability that the user intends to hear an audio message indicating a received text-based message. Computer device 210 Can a text-based message be sent via a communication unit? 244 received CMM 228 It can determine whether to output an audio version of the message based on the probability that the user intends to hear an audio version.

[0041] CMM 228 can determine the probability that a user of a computer device will 110 Intended to hear an audio version of a text-based message. Contextual information can include, as non-restrictive examples, the frequency of incoming messages from a particular messaging device. 115 (e.g. message transmission device) 115A), the frequency of outgoing messages to the message transmission device 115A , elapsed time since the last message from the communication device 115A received message, elapsed time since the last message was sent to the messaging device 115A the sent message. For example, the user of a computer device 210 frequently send SMS messages over a predetermined period of time using a messaging device 115A exchange. Due to the frequency of SMS messages between the user and the messaging device. 115A can CMM 228 Determine that there is a high probability that the user intends to hear an audio version of the message. The user of a computer device 210 can sporadically send SMS messages to another of the messaging devices over a predetermined period of time. 115 (e.g. message transmission device) 115Nexchange. Based on sporadic message exchange with a message transmission device. 115N can CMM 228 determine that there is a low probability that the user intends to receive an audio version of the message from the messaging device 115N listen.

[0042] The contextual information used to determine the probability may also include one or more of the following: a user's location, a time of day, calendar entries from the user's calendar, information on whether a message is being sent to (or received from) a contact in the user's contact list, or whether the user recently had a telephone conversation with a user of a particular messaging device. 115The context information may include the actions taken by the user. In some examples, the context information may also include one or more actions performed by the user, such as using an application (e.g., using a web browser, playing music, using navigation programs, taking a photo, etc.), muting a computer device. 210 Sending or receiving a voice message (e.g., a phone call or video chat), sending or receiving a text-based message, speaking a command to a computer device 210 or any other action that can indicate whether a user of computer device 220 intended to listen to an audio version of a received text-based message.

[0043] CMM 228It can determine, based on some kind of contextual information, the probability that the user intends to listen to a received message. For example, if the user starts playing music on a computer device... 210 to play CMM 228 Determine that the probability that the user intends to hear a message is low. In some examples, CMM can 228 Based on several types of contextual information, determine the probability that the user intends to hear a message. For example, CMM can 228 Based on whether the sender is in the user's contact list, and that the user has exchanged a certain number of messages with them within a given time period, determine the probability that the user intends to listen to a received message.

[0044] In some examples, CMM can 228Types of contextual information can be considered independently of each other. For example, CMM can 228 , provided CMM 228 based on the frequency of incoming messages from the message transmission device 115 and based on whether one is connected to the messaging device 115 The probability that the user intends to hear an audio version of the message is determined by whether an associated third party is in the user's contact list. This probability is increased if the frequency of incoming messages reaches a threshold or if the third party is in the user's contact list. CMM 228 However, in some examples, the probability can be determined using a weighting. For example, CMM can 228A high probability can be determined, even though the message frequency is low, if the third party sending and / or receiving messages is in the user's contact list. In contrast, CMM 228 Despite a high frequency of messages, a low probability is determined if the third party sending and / or receiving messages is not in the user's contact list.

[0045] CMM 228 It can determine whether the user intends to hear an audio version of a received message by comparing the probability that a user intends to hear the message to a probability threshold. In some examples, CMM can 228The probability that the user intends to hear the message is compared using different probability thresholds. Each of the different probability thresholds can correspond to a different conversation state, and CMM 228 It can perform different actions depending on the conversation status.

[0046] CMM 228 can establish a conversation status between a user of a computer device 210 and message transmission device 115 Determine based on the probability that the user intends to hear a text-based message. For example, CMM can 228 determine that the user cannot have a conversation with a user of a messaging device 115 This results in what is subsequently referred to as the "idle state". In some examples, CMM can 228determine that a user can have a conversation with another user of a messaging device to a limited extent 115 This results in what is subsequently referred to as the "recent state". Furthermore, CMM can 228 In some examples, determine that a user is having an intensive conversation with another user of a messaging device. 115 This leads to what is subsequently referred to as the "active state".

[0047] In some examples, CMM can 228 a conversation status between the user of computer device 210 and a specific messaging device 115 determined on an individual basis. In other words, the conversation status between the user and an initial messaging device can be determined. 115 from the conversation status between the user and a second messaging device 115 differ. For example, CMM can 228determine that a conversation takes place between the user and a specific messaging device. 115 is in a recent state, and that a conversation is taking place between the user and another messaging device. 115 found in an active state. In some examples, CMM can 228 the conversation status between the user of the computer device 210 and a specific group of messaging devices 115 Determine on a group basis. For example, computer equipment 210 the conversation status between the user of the computer device 210 and a group of communication devices 115(e.g., contacts participating in a group message) determine the conversation status so that the conversation status is the same for all group members. In some examples, the conversation management module can determine a conversation status between the user of a computer device. 210 and determine all contacts on a global basis. For example, CMM can 228 Specify that the conversation state is a sleep state for all conversations (e.g., the user can turn off the computer device). 210 put it into a "Do Not Disturb" mode).

[0048] CMM 228CMM can determine that the conversation status is in an active state when the probability that the user intends to hear an audio version of a received message reaches a first probability threshold and a second probability threshold (e.g., the probability is higher than both the first and second probability thresholds). 228 Determined that the conversation is in an active state, it's possible that the user doesn't need to issue any commands to send or listen to a message. For example, in an active state, the user can receive a message from a specific messaging device. 115 receive, and computer device 210 It can perform speech synthesis processing on the received message without requesting instructions from the user. Speech synthesis (TTS) module226 can convert the message into audio data, whereupon the computer device 210 the audio data via loudspeakers 247 can spend.

[0049] CMM 228 The conversation state can be determined to be in a recent past state if the probability that the user intends to hear an audio version of a received message does not reach a first probability threshold but does reach a second probability threshold (e.g., the probability lies between a first and a second probability threshold). When the conversation state is in a recent past state, it may only require minor commands from the user to send or hear a message. In some examples, the TTS module 226In a recent state, perform speech synthesis processing on the message to convert the audio data into text data. Computer device 210 It can output audio data with minimal message context, such as the sender's name. (If computer device) 210 For example, if it receives an SMS, the TTS module can 226 Convert the text-based message into an audio output so that the computer device outputs the message context "Jimmy said" and the audio data "Hey. Colleague, where are you going tonight?"

[0050] CMM 228CMM can determine that the conversation state is in a dormant state if the probability that the user intends to hear an audio version of a received message does not reach either probability threshold (e.g., the probability is less than both the first and second probability thresholds). 228 If the conversation state is set to idle, the user may be required to perform an action to send a message or listen to a received message. Computer device 210 can issue a request for additional instructions from the user. For example, a computer device can 210 in a sleep state, a text-based message is sent from a specific messaging device. 115It can receive and output audio data that asks if the user wants to hear the message. For example, a computer device 210 Play the audio message: "Message received from Jimmy. Would you like to listen to the message?"

[0051] CMM 228 It can determine the conversation status based on the probability that the user intends to listen to a received text-based message. In some examples, CMM can 228 Determine the conversation state based on the probability that the user intends to send a message.

[0052] In some examples, a user of computer equipment 210 Send a message. Computer device 210 Can audio input be received from the user via microphone? 243 received CMM 228can determine a probability that the user intends to send a text-based message to a specific messaging device. 115 to send. The likelihood that the user intends to send a message can be based on contextual information, such as the contextual information used to determine whether the user intends to hear an audio version of a received text-based message. As an additional example, contextual information can include the positive connotation or strength of a command given by the user to the computer device. 210 This includes, for example, the command "Text Jimmy" may have a less positive connotation than the command "Talk to Jimmy", so the first command may indicate a lower probability than the second command.

[0053] CMM 228It can determine whether the user intends to send a message by comparing the probability of the user intending to send a message with a probability threshold. In some examples, CMM can 228 The probability that the user intends to send the message is compared using different probability thresholds. Each of these thresholds can correspond to a different conversation state, and CMM can perform different actions depending on the conversation state.

[0054] In some examples, CMM can 228Determine that the conversation is in an active state. In an active state, the user can send a message by speaking aloud the message they wish to send, without any command such as "say," "type," "send," or other commands. For example, the user can say "I'll be there in five minutes" without specifically stating, "Send a message to Jimmy." Computer device 210 Can receive audio data from a user. Speech recognition module. 224 It can perform speech recognition (STT) processing on audio data and convert the audio data into text data. CMM 228 can generate a text-based message based on the text data, whereupon the computer device 210 the message automatically to a specific messaging device 115 (e.g. a messaging device associated with Jimmy) 115 can send.

[0055] If CMM 228Determined that the conversation status is in a recent state, the user may be able to send a message to a specific messaging device with minimal commands. 115 to send. For example, the user can speak a message that includes a message command (e.g., "say," "type," "send") and the message content ("I'll be there in five minutes."). Computer device 210 can transmit the message command and message content via microphone 243 received. STT module 224 CMM can convert audio input into text data. 228 can generate a text-based message based on the text data in such a way that the communication module 244 It may send a text-based message (where the message reads "I'll be there in five minutes.") without requiring the user to confirm the content of the message or the user's intention to send the message.

[0056] In some examples, CMM can 228 Determine that the conversation is in a sleep state. (If computer device) 210 If a computer device receives audio input from a user while a conversation is in a sleep state, it may 210 issue a request for additional information from the user. For example, a computer device 210 In a sleep state, display a message prompting the user to confirm whether they wish to send a message. Computer device 210 It can receive audio input confirming the user's intention to send a message and can receive audio input indicating a message to be sent. STT module 224 CMM can perform speech recognition processing on audio input and convert the audio data into text data. 228 can generate a text-based message based on the text data, whereupon the computer device 210the message to a specific messaging device 115 can send.

[0057] computer device 210 It can provide the user with a visual or audible indication of the conversation status. For example, a computer device can 210 The user is notified of the conversation status via sound signals (e.g., a series of beeps or speech synthesis notifications). In some examples, the computer device can 210 the user via a visual notification (e.g., an on-screen message) 114 The displayed status symbol indicates the conversation status.

[0058] CMM 228 CMM can determine different conversation statuses for incoming messages compared to outgoing messages. For example, CMM can 228 determine a high probability that the user intends to receive messages from a specific messaging device 115to hear. CMM 228 However, it can determine that the probability that the user intends to send a message is lower than the probability that the user intends to hear a received message. Consequently, the computer device 210 In some examples, it automatically outputs an audio version of a received message, but it can also request additional instructions from the user before sending an outgoing message.

[0059] FIGS. 3A-3H show conceptual diagrams illustrating an exemplary operation of computer equipment 210 illustrate. Computer device 210 can receive a text-based message from a source. CMM 228 can determine a probability that the user of computer equipment 210 Intended to listen to the received message. CMM 228CMM can determine the probability based on one or more types of contextual information. For example, if the contextual information includes the frequency of incoming messages from the source and the frequency of incoming messages is low, CMM can 228 determine that the probability that the user intends to listen does not reach a probability threshold. Consequently, CMM can 228 Determine that the conversation is in a sleep state. Computer device 210 can output a message to notify the user about the incoming message ( Fig. 3A). Computer device 210For example, it can display a message asking if the user wants to hear it. In some examples, the user confirms their intention to hear the message by saying "yes," "read message," "ok," or any other response indicating that they want to hear the message.

[0060] computer device 210 can receive audio data from the user via microphone 243 Received signals indicating that the user wishes to hear the message content. TTS module 226 It can perform speech synthesis processing on the received text-based message and convert the text data into audio data. In response to receiving a command from the user, the computer device can 210 output the audio data that indicates the content of the text-based message ( Fig. 3B). Since CMM 228 has determined that the conversation status is in a sleep state, computer device 210Output message context, such as the name of the contact who sent the message. For example, a computer device 210 Output the message context (e.g., "Jimmy said") followed by the audio data (e.g., "Hey mate! Where are you going tonight?"). In some examples, the computer device can 210 After outputting the audio data, issue a request for additional commands from the user.

[0061] In some examples, the user can access the computer device 210 Command to send a reply message to the originating source. For example, the user can reply, say "tell Jimmy," or any other words indicating that the user wants to send a reply message to the originating source. Microphone 243 of computer device 210 It can receive audio input spoken by the user. CMM 228can determine the probability that the user intends to send a reply message to the originating source. In some examples, when using a computer device 210 If CMM has only received a text-based message from the originating source and the user gives a command to reply to the message, 228 Determine that the probability that the user intends to send a reply message has not reached a probability threshold and that the conversation is still in a dormant state. As a result, the computer device may 210 issue a request for a response message ( Fig. 3C). Computer device 210 Can the reply message be sent via microphone? 243 Received as audio input. STT module 224It can perform speech recognition processing on the audio input and convert the audio data into text data. Since the conversation state is still in a sleep state, the computer device 210 Issue a request to the user to confirm whether the reply message should be sent ( Fig. 3D). In some examples, computer equipment can 210 Send the reply message to the originating source and display a message to confirm to the user that the reply message has been sent ( Fig. 3E).

[0062] As in Fig. As shown in 3F, the originating source can send the user a second text-based message. CMM 228 Based on contextual information (e.g., an increase in message frequency between the user and the originating source), it can determine that the likelihood of the user intending to listen to the received message has increased. CMM 228It can determine that the probability that the user intends to hear the message reaches a first probability threshold but does not reach a second probability threshold (e.g., the probability that the user intends to hear the message lies between a first probability threshold and a second probability threshold). Consequently, CMM can 228 Determine that the conversation status between the user and the originating source is in a recent state. In a recent state, the TTS module can 226 Perform speech synthesis processing on the received message and convert the text data into audio data. Computer device 210 It can automatically output the text data. For example, a computer device 210Output the message context (e.g., "Jimmy said") followed by the audio data (e.g., "Are you bringing snacks?").

[0063] In some examples, the user can reply to the message from the originating source by speaking a reply message. Computer device 210 can transmit audio data corresponding to the user's reply message via microphone 243 Received. For example, the user can say: "Tell Jimmy I'm bringing cookies." CMM 228 can specify that the response includes a command to send a message (e.g., "Tell Jimmy"). CMM 228 CMM can determine that the probability that the user intends to send a message reaches a first probability threshold because the message contains a command, but does not reach a second probability threshold. Consequently, CMM can 228Determine that the conversation between the user and the originating source is in a recent state. STT module 226 CMM can perform speech recognition processing on the received audio data and convert the audio data into text data. 228 can generate a text-based response message based on the text data, whereupon the computer device 210 can send the text-based reply message to the originating source.

[0064] computer device 210 can receive a third incoming message from the originating source, and CMM 228 It can determine the probability that the user intends to listen to the received message. For example, CMM can 228Based on the frequency of messages exchanged between the user and the originating source, a threshold is determined for the probability that the user intends to hear the message to be reached, thus placing the conversation in an active state. TTS module 226 It can convert text data into audio data. Computer device 210 can automatically output the audio data (e.g., "Great, see you soon!"). Fig. 3H).

[0065] For subsequent messages between the user and the originating source, CMM can 228 Determine the probability that the user intends to send a message or listen to a received message. If CMM 228 Determined that the conversation status has changed, computer device 210 Display prompts and message context according to the respective conversation status, as described above.

[0066] In some examples, a user can have a text-based conversation using a specific messaging device. 115 Initiate the conversation. The user can start the conversation by physically entering the computer device. 210 (e.g. by pressing the presence-sensitive display) 5 ) initiate or by speaking a voice command. Computer device 210 can accept the voice command in the form of an audio input via microphone 243 received CMM 228 Based on the voice command and other contextual information, CMM can determine the likelihood that the user intends to send a text-based message. For example, the user might say, "Text Jimmy," so CMM 228The probability that the user intends to send a message to the recipient (e.g., Jimmy) and the corresponding conversation status can be determined. The probability that the user intends to send a message and the corresponding conversation status can depend on the positive connotation of the original command. For example, if the user says, "Text Jimmy," CMM can 228 Determine a probability, but if the user says "Talk to Jimmy", CMM can 228 determine a different probability. CMM 228 CMM can determine that the probability that a user intends to send a message to the recipient when the user says "text Jimmy" is higher than a first probability threshold, but lower than a second probability threshold. 228However, it can determine that the probability that the user intends to send a message to the recipient is higher than both the first probability threshold and the second probability threshold when the user says, "talk to Jimmy". Consequently, CMM can 228 Depending on the positive connotation of the received command, different conversation statuses are determined.

[0067] In some examples, CMM can 228 Determine that a conversation has ended based on explicit actions or commands by the user (e.g., the probability that the user intends to listen to a message is very low). For example, the user might press a button on a computer device. 210 (e.g. on the presence-sensitive display) 5 Press ) to end the conversation. In some examples, CMM may 228Based on more than one type of contextual information, such as the content of a message, determine that a conversation has ended. For example, the user might say "Goodbye" or "End conversation." If CMM 228 Determined that the conversation is over, computer device 210 Require a full set of commands and acknowledgments from the user to send additional messages or listen to received messages.

[0068] In some examples, CMM can 228Determine whether the probability that the user intends to send or receive a message has transitioned from low (i.e., the conversation state is idle) to high (i.e., the conversation state is active) or vice versa, without passing through an intermediate range. In other words, the conversation state can skip the recent state if the probability suddenly increases or decreases significantly.

[0069] CMM 228 It can determine a temporary or interim conversation status. For example, the user can initiate a short conversation with a specific messaging device. 115 to initiate (i.e., the conversation is temporarily in an active state) by waiting for a certain period of time on the message-transmitting device 115associated contact information or by displaying the contact information associated with the messaging device on the screen 114 This can be represented in some examples. 228 Specify that the conversation remains in a temporary state for a certain period of time or as long as the contact information is displayed.

[0070] In some examples, a user can have multiple conversations using different messaging devices. 115 lead to. For example, computer equipment 210 a message from a first message transmission device 115 and a message from a second messaging device 115 received. In response to receiving audio data from the user, CMM can 228 the probability that the user intends to send a message to a first message transmission device 115to send, and the probability that the user intends to send a message to a second messaging device. 115 to send, determine. In some examples, CMM can 228 Analyze the content of the audio data and determine whether the content of the audio data is more relevant for the conversation with the first messaging device or the second messaging device.

[0071] CMM 228 can determine which message transmission device 115 the message is to be received by increasing the probability that the user intends to send a message to the first message transmission device 115 to send, compared with the probability that the user intends to send a message to a second messaging device. 115to send, and determines which probability is higher. If the probability that the user intends to send the message to the first message-transmitting device is higher, then the probability is higher. 115 to send is higher than the probability that the user intends to send the message to the second messaging device. 115 to send CMM 228 determine that the user intends to send the message to the first messaging device 115 to send.

[0072] In some examples, CMM can 228 determine which message transmission device 115 the message is to be received by increasing the probability that the user intends to send a message to the first message transmission device. 115 to send, compared with the probability that the user intends to send a message to the second messaging device. 115to send, and compares each of the probabilities with a probability threshold. For example, if the probability that the user intends to send a message to the first message delivery device 115 to send, a probability threshold has been reached, and the probability that the user intends to send a message to the second message transmission device 115 To send, once a probability threshold has been reached, CMM can 228 determine that the user intends to send the message to the message delivery device associated with the higher probability 115 to send.

[0073] In some examples, CMM can 228 the probability that the user intends to send the message to the first messaging device 115to send, compare with a probability threshold and can determine the probability that the user intends to send the message to the second message transmission device. 115 to send, compare with the probability threshold. If CMM 228 For example, it determines the probability that the user intends to send a message to the first message transmission device. 115 to send, a probability threshold has been reached and that the probability that the user intends to send a message to the second message transmission device 115 To send, once a probability threshold has been reached, CMM can 228 determine that computer device 210 the message to both the first and second messaging devices 115should send. However, if the probability is that the user intends to send a message to the first messaging device... 115 to send, a probability threshold has been reached and the probability that the user intends to send a message to the second message transmission device 115 To send, a probability threshold is reached, computer device 210 Issue a request to the user to confirm which messaging device is being used. 115 should receive the outgoing message.

[0074] If the probability is that the user intends to send the message to the first messaging device 115 to send, a probability threshold is reached and the probability that the user intends to send the message to the second message transmission device 115If a probability threshold is not reached, CMM can send a message. 228 In some examples, a prompt will appear to the user to confirm whether a message should be sent. CMM 228 It can also issue a request to the user to confirm which messaging device is being used. 115 to receive the message.

[0075] In some examples, the probability threshold for sending a message can change if the user is conducting multiple conversations. For instance, the probability threshold for sending a message in the active state might be a first probability threshold if the user is only conducting one conversation. However, the probability threshold for sending a message in the active state might increase to a second probability threshold if the user is conducting more than one conversation (i.e., if there is at least one other conversation that is not in a resting state).

[0076] Fig. Figure 4 shows a flowchart illustrating an exemplary operation of a computer device. 210 Illustrated. In some examples, computer equipment can be used. 210 receive a text-based message from a source ( 400The text-based message can include an email, instant message, SMS, or any other type of text-based message. CMM 228 It can determine the probability that the user intends to hear an audio version of the message and can compare this probability to a probability threshold. In some examples, the TTS module can 226 Perform speech synthesis processing on the received message and convert the text data into audio data. Computer device 210 can output the audio data.

[0077] computer device 210 can receive audio input ( 410 For example, the user can speak a message that is picked up through the microphone. 243 of computer device 210 is received. CMM 228 can determine the probability that the user intends to send a reply message to the originating source ( 420The likelihood that the user intends to send a reply message to the originating source can be based on explicit commands or contextual information. For example, contextual information can determine the frequency of messages sent to and received by the originating source. CMM 228 It can determine whether the probability that the user intends to send a reply message to the originating source reaches a certain probability threshold. CMM 228 can determine a conversation state (e.g., idle, recent, or active) by comparing the probability that the user intends to send a reply message with a first probability threshold and a second probability threshold.

[0078] In response to a determination that the probability that the user intends to send a reply message has reached a probability threshold, the computer device may 210 generate the reply message based on the audio input ( 430 ). Computer device 210 For example, it can perform speech recognition processing on audio input and convert the audio data into text data. CMM 228 can generate a text-based reply message based on the text data. Computer device 210 can send the reply message to the originating source ( 440 ).

[0079] In some examples, a method may involve a computer device associated with a user outputting an audio signal representing a text message from a source. The method may involve the computer device receiving audio data representing a speech utterance from the user. The method may also involve the computer device determining, without additional input (such as acoustic or gesture-based input) from the user, a probability that the user intends to send a response, at least partially based on the audio data and one or more factors from the frequency of incoming messages from the source, the frequency of outgoing messages to the source, the time elapsed since the last message received from the source, or the time elapsed since the last message sent to the source.The procedure may also involve transmitting, in response to a determination that the probability has reached a probability threshold and without additional input (e.g. acoustic or gesture-based input) from the user, a transcription of at least part of the audio data to the source.

[0080] Fig. Figure 5 shows a flowchart illustrating an exemplary operation of a computer device. 210 Illustrated. In some examples, a user can conduct multiple conversations with different sources. Computer device. 210 For example, a text-based message can originate from an initial source (i.e., an initial message transmission device). 115 ) received ( 500 The computer device can receive a text-based message from a second source (i.e., a message transmission device). 115 ) received ( 510The text-based message from the first source and the text-based message from the second source can contain different types of messages. For example, the text-based message from the first source can be an SMS message, and the text-based message from the second source can be an instant message. CMM 228 CMM can determine the probability that the user intends to hear an audio version of the message from the original source. 228 The TTS module can compare the probability that a user intends to hear the audio version of the message to a probability threshold. In response to a determination that the probability of the user intending to hear an audio version of the message reaches a probability threshold, the TTS module can 226 convert the text data into audio data, whereupon the computer device210 which can output audio data. Likewise, computer devices can 210 Determine the probability that the user intends to hear an audio version of the message from the second source and compare this probability to a probability threshold. In response to determining that the probability reaches a probability threshold, the computer device can 210 Convert the text data into audio data and output the audio data.

[0081] computer device 210 can receive audio input ( 510 ). After computer device 210 For example, if the computer receives a message from the first source and a message from the second source, the user can speak a message. The computer can then hear the message from the user via the microphone. 243 received as audio input.

[0082] CMM 228can determine a probability that the user intends to send a reply message to the initial source ( 530 The likelihood that the user intends to send a reply message to the initial source can be based on an explicit command and / or contextual information. An explicit command might include a statement, such as "Tell Aaron." Contextual information might include the frequency of communication between the computer device. 210 and the first source of exchanged messages, the elapsed time since the last exchange between computer devices 210 and the message exchanged from the original source, or any other type of contextual information.

[0083] CMM 228 can determine a probability that the user intends to send a reply message to the second source ( 540The likelihood that the user intends to send a reply message to the second source can be based on an explicit command and / or contextual information. An explicit command might include a statement, such as "Tell Jimmy." Contextual information might include the frequency of communication between the computer device. 210 and the second source of exchanged messages, the elapsed time since the last exchange between computer devices 210 and the second source of the exchanged message or other type of contextual information.

[0084] CMM 228 can determine whether the user intends to send the reply message to the first source, the second source, both the first and second sources, or neither of the sources ( 550 In some examples, CMM can 228Compare the probability that the user intends to send a message to the first source with the probability that the user intends to send a message to the second source, determine which probability is higher, and computer device 210 cause the reply message to be sent to the originating source with a higher probability.

[0085] In some examples, CMM can 228 Compare the probability that the user intends to send the reply message to the first source with a probability threshold, and compare the probability that the user intends to send the reply message to the second source with the probability threshold. If CMM 228For example, it determines that the probability that the user intends to send the reply message to the first source reaches a probability threshold, and that the probability that the user intends to send the reply message to the second messaging device 115 To send, once a probability threshold has been reached, CMM can 228 Determine that the user intends to send the reply message to both the first and second source. If the probability that the user intends to send a message to a first source reaches a certain probability threshold, and the probability that the user intends to send a message to the second source reaches a certain probability threshold, the computer device can 210Issue a request to the user to confirm which origin source should receive the response message. If the probability that the user intends to send the response message to the first origin source does not reach a certain probability threshold, and the probability that the user intends to send the response message to the second message delivery device is higher, the system will return a message to the second source. 115 To send, a probability threshold is not reached, computer device 210 In some examples, a prompt will appear to the user to confirm whether a message should be sent. Computer device 210 It can also issue a request to the user to confirm which source the message should be received from.

[0086] In some examples, CMM can 228Determine whether the user intends to send the reply message to the first or second source by comparing the probability that the user intends to send a message to the first source with the probability that the user intends to send a message to the second source, and comparing the respective probabilities to a probability threshold. For example, if the probability that the user intends to send the reply message to the first source reaches a probability threshold, and the probability that the user intends to send the reply message to the second source reaches a probability threshold, CMM can 228 determine that the user intends to send the message to the source associated with the higher probability.

[0087] computer device 210 can generate the reply message based on the audio input ( 560 For example, the STT module 224 Convert the audio data into text data that specifies the audio data received from the user. In some examples, the computer device 210 at least part of the audio data for speech recognition processing at ISS 160 send so that the ISS 160 generate text data and send the text data to a computer device 210 can send CMM 228 can generate a text-based response message based on the text data.

[0088] After computer device 210 The reply message generated can be used by a computer device 210 send the reply message ( 570 The reply message can be sent to the address provided by CMM. 228 specific source(s) will be sent.

[0089] Appended to this description are numerous claims directed to several embodiments of the disclosed subject matter. It is understood that embodiments of the disclosed subject matter may also fall within the scope of several combinations of said claims, such as dependencies and multiple dependencies among them. Thus, all dependencies and multiple dependencies, whether explicitly referenced or otherwise incorporated, form part of this description.

[0090] In one or more examples, the described functions can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on or transmitted via a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which correspond to physical media, such as data storage media, or communication media, including media that facilitate the transmission of a computer program from one location to another, for example, according to a communication protocol. In this way, computer-readable media can generally be considered physical computer-readable ( 1 ) Storage media that are non-volatile or ( 2) a communication medium, such as a signal or a carrier wave. Data storage media can be any available medium accessible to one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product can include a computer-readable medium.

[0091] For example, but not limited to, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures accessible by a computer. Furthermore, any connection is referred to as a computer-readable medium.For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted-pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then these technologies are included in the definition of medium. However, it should be clear that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other physical media, but instead refer to non-volatile, physical storage media.Hard disks and floppy disks, as used herein, include Compact Disc (CD), Laserdisc, optical disc, Digital Versatile Disc (DVD), floppy disk, and Blu-ray Disc, with floppy disks typically reproducing data magnetically, while discs reproducing data optically using lasers. Combinations of the foregoing storage media should also be included in the scope of computer-readable media.

[0092] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-oriented integrated circuits (ASICs), field-programmable general-purpose circuits (FPGAs), or any other equivalent integrated or discrete logic circuits. Accordingly, the term "processor," as used herein, can refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. Furthermore, some aspects of the functionality described herein can be provided within dedicated hardware and / or software modules. Additionally, the techniques could be implemented entirely within one or more circuits or logic elements.

[0093] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). This disclosure describes various components, modules, or units to emphasize functional aspects of devices configured to perform the disclosed techniques, but these do not necessarily require implementation by different hardware units. Rather, as described above, different units can be combined in a single hardware unit or provided by a collection of interoperable hardware units, including one or more processors as described above, in conjunction with suitable software and / or firmware.

[0094] Several examples have been described. These and other examples fall within the scope of the following claims.

Claims

[1] Procedure, encompassing: Receiving, by a computer device associated with a user, a message from a source; Received, by the computer device, an audio input; Determine, through the computer device and based at least partially on audio input and contextual information, a probability that the user intends to send a reply message to the originating source; Determine, by the computer device, in response to a determination that the probability that the user intends to send the reply message to the originating source reaches a probability threshold that the user intends to send the reply message to the originating source; and in response to a determination that the user intends to send the reply message to the originating source: Generate, by the computer device and based on the audio input, the response message; and Sending, via the computer device, the reply message to the originating source. [2] Method according to claim 1, further comprising: in response to a determination that the probability that the user intends to send the reply message to the originating source does not reach the probability threshold: Output, by the computer device, of a request for additional actions by a user; Received by the computer device and by the user, a second audio input indicating the user's intention to send a message; and Sending, through the computer device and at least partially based on the second audio input, the reply message to the originating source. [3] Method according to one of claims 1-2, wherein the source of origin is a first source of origin, the method further comprising: before receiving the audio input, receiving, through the computer device, a message from a second source; Determine, through the computer device and at least partially based on the audio input and contextual information, a probability that the user intends to send the response message to the second source; Determine, using the computer device and based on the probability that the user intends to send the reply message to the first source of origin and the probability that the user intends to send the reply message to the second source of origin, whether the user intends to send the reply message to the first source of origin or to the second source of origin; and in response to a determination that the user intends to send the reply message to the second source: Generating the reply message by the computer device based on the audio input; and Sending the reply message to the second source by the computer device. [4] Method according to claim 3, where a determination that the user intends to send the reply message to the first source is further made in response to a determination that the probability that the user intends to send the reply message to the first source reaches the probability threshold, but that the probability that the user intends to send the reply message to the second source does not reach the probability threshold, and where a determination that the user intends to send the reply message to the second source is further made in response to a determination that the probability that the user intends to send the reply message to the second source reaches the probability threshold, but that the probability that the user intends to send the reply message to the first source does not reach the probability threshold. [5] Method according to claim 3, wherein a determination that the user intends to send the reply message to the first source is further made in response to a determination that the probability that the user intends to send the reply message to the first source reaches the probability threshold and is higher than the probability that the user intends to send the reply message to the second source, and where determining that the user intends to send the reply message to the second source is further performed in response to determining that the probability that the user intends to send the reply message to the second source reaches the probability threshold and is higher than the probability that the user intends to send the reply message to the first source. [6] Method according to any one of claims 1-5, further comprising: Determine, through the computer device, a probability that the user intends to hear the message from the original source; and in response to a determination that the probability that the user intends to hear the message from the source has reached a probability threshold for hearing the message: Generated by the computer device and based on the message from the source, from audio data; and Output, through the computer device, of the audio data. [7] Method according to any one of claims 1-6, wherein the context information includes one or more of the following: Frequency of incoming messages from the source, frequency of outgoing messages to the source, time elapsed since the last message received from the source or time elapsed since the last message sent to the source. [8] Method according to any one of claims 1-7, wherein the probability that the user intends to send the reply message to the originating source is not based on a user command. [9] Device, comprising: an audio output device; an audio input device; a communication unit; a message management module that can be operated by at least one processor to: to receive a message from a source via the communication unit; to receive audio input via the audio input device; to determine, at least partially based on audio input and contextual information, a probability that a user associated with the device intends to send a reply message to the originating source; in response to a determination that the probability that the user intends to send the reply message to the originating source has reached a probability threshold, to determine that the user intends to send the reply message to the originating source; and in response to a determination that the user intends to send the reply message to the originating source: to generate the reply message based on the audio input; and to send the reply message to the originating source via the communication unit. [10] Device according to claim 9, wherein the message management module is further operable by the at least one processor to: in response to a determination that the probability that the user intends to send the reply message to the originating source does not reach the probability threshold: to issue a request for additional action by a user via the audio output device; to receive a second audio input via the audio input device, indicating the user's intention to send a message; and to send the reply message to the originating source via the communication unit and based at least partially on the second audio input. [11] Device according to one of claims 9-10, wherein the source of origin is a first source of origin, wherein the message management module is further operable by the at least one processor to: to receive a message from a second source before receiving the audio input via the communication unit; to determine, at least partially based on the audio input and contextual information, a probability that the user intends to send the reply message to the second source; based on the probability that the user intends to send the reply message to the first source and the probability that the user intends to send the reply message to the second source, to determine whether the user intends to send the reply message to the first or the second source; and in response to a determination that the user intends to send the reply message to the second source: to generate the reply message based on the audio input; and to send the reply message via the communication network to the second source. [12] Device according to claim 11, where a determination that the user intends to send the reply message to the first source is further made in response to a determination that the probability that the user intends to send the reply message to the first source reaches the probability threshold, but the probability that the user intends to send the reply message to the second source does not reach the probability threshold; and where determining that the user intends to send the reply message to the second source is further performed in response to determining that the probability that the user intends to send the reply message to the second source reaches the probability threshold, but the probability that the user intends to send the reply message to the first source does not reach the probability threshold. [13] Device according to claim 11, wherein a determination that the user intends to send the reply message to the first source is further made in response to a determination that the probability that the user intends to send the reply message to the first source reaches the probability threshold and is higher than the probability that the user intends to send the reply message to the second source, and where determining that the user intends to send the reply message to the second source is further performed in response to determining that the probability that the user intends to send the reply message to the second source reaches the probability threshold and is higher than the probability that the user intends to send the reply message to the first source. [14] Device according to one of claims 9-13, wherein the message management module is furthermore operable by the at least one processor to: to determine a probability that the user intends to hear the message from the original source; and in response to a determination that the probability that the user intends to hear the message from the source has reached a probability threshold for hearing the message: to generate audio data based on the message from the source; and to output the audio data via the audio output device. [15] Computer-readable storage medium comprising instructions which, when executed, configure one or more processors of a computer system to perform one of the methods of claims 1-8.

Citation Information

Patent Citations

  • automatically activating intelligent responses based on activities of remote devices

    DE112014003653T5

  • Message Processing Method, Terminal and System

    US20140189027A1

  • System and method for inferring user intent from speech inputs

    US20140365209A1

  • Evaluating transcriptions with a semantic parser

    US8868409B1

  • Ordering of conversations based on monitored recipient user interaction with corresponding electronic messages

    WO2007044806A2