Systems and methods for efficient multimodal input collection using mobile devices
By integrating input collection elements and machine learning models into the lock screen interface of mobile devices, the problem of time-consuming information capture in traditional devices is solved, enabling fast and effective multi-modal input and improving device efficiency and power utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2021-04-28
- Publication Date
- 2026-05-26
Smart Images

Figure CN116888586B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to multimodal input collection. More specifically, this disclosure relates to efficient and intuitive multimodal input collection for mobile devices. Background Technology
[0002] Mobile communication devices (e.g., smartphones) have become increasingly important tools for consumers in both business and personal settings. Users typically expect to capture and record information in response to time-sensitive events (e.g., quickly capturing images of restaurant menus, capturing video of events occurring, etc.). However, to do this, users traditionally need to navigate a series of time-consuming user interfaces and applications to initiate information capture (e.g., navigating through the lock screen, opening the camera app, switching to video recording mode within the camera app, etc.). Therefore, systems and methods for efficiently collecting multimodal input from mobile devices are desired. Summary of the Invention
[0003] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.
[0004] One example aspect of this disclosure relates to a computer-implemented method for contextualized input collection and intent determination for a mobile device. The method includes providing a lock screen interface associated with a mobile computing system comprising one or more computing devices, wherein the lock screen interface includes input collection elements configured to induce the capture of sensor data when selected. The method includes obtaining an input signal from a user of the mobile computing system indicating that the input collection element has been selected by the mobile computing system. The method includes capturing sensor data from a plurality of sensors of the mobile computing system in response to obtaining the input signal, wherein the plurality of sensors include an audio sensor and one or both of a front-facing image sensor or a rear-facing image sensor. The method includes determining one or more proposed action elements from a plurality of predefined action elements by the mobile computing system, at least in part, based on the sensor data from the plurality of sensors, wherein the one or more proposed action elements respectively instruct one or more device actions. The method includes displaying the one or more proposed action elements within the lock screen interface at a display device associated with the mobile computing system.
[0005] Another example aspect of this disclosure relates to a mobile computing system. The mobile computing system includes one or more processors. The mobile computing system includes multiple sensors, including: one or more image sensors, including one or more of a front-facing image sensor, a rear-facing image sensor, or a peripheral image sensor; and an audio sensor. The mobile computing system includes a display device. The mobile computing system includes one or more tangible, non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. Operations include providing a lock screen interface associated with the mobile computing system, wherein the lock screen interface includes an input collection element configured to cause the capture of sensor data when selected. Operations include obtaining an input signal from a user of the mobile computing system for selecting the input collection element. Operations include capturing sensor data from the multiple sensors of the mobile computing system in response to obtaining the input signal. Operations include determining one or more suggested action elements from a plurality of predefined action elements based at least in part on the sensor data from the multiple sensors, wherein the one or more suggested action elements respectively instruct one or more device actions. Operations include providing the one or more suggested action elements at the display device for display within the lock screen interface.
[0006] Another example aspect of this disclosure relates to one or more tangible, non-transitory computer-readable media that collectively store instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations include providing a lock screen interface associated with a mobile computing system, wherein the lock screen interface includes an input collection element configured to cause the capture of sensor data when selected. The operations include obtaining an input signal from a user of the mobile computing system to select the input collection element. The operations include capturing sensor data from a plurality of sensors of the mobile computing system in response to obtaining the input signal, wherein the plurality of sensors include an audio sensor and one or both of a front-facing or rear-facing sensor. The operations include determining one or more proposed action elements from a plurality of predefined action elements by the mobile computing system, at least in part, based on the sensor data from the plurality of sensors, wherein the one or more proposed action elements respectively instruct one or more device actions. The operations include providing one or more proposed action elements at a display device associated with the mobile computing system to display within the lock screen interface.
[0007] Other aspects of this disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0008] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description
[0009] A detailed discussion of embodiments for those skilled in the art is set forth in the specification with reference to the accompanying drawings, in which:
[0010] Figure 1A A block diagram of an example computing system for performing efficient input signal collection according to an example embodiment of the present disclosure is depicted.
[0011] Figure 1B A block diagram of an example computing device for performing efficient input signal collection according to an example embodiment of the present disclosure is depicted.
[0012] Figure 1C A block diagram of an example computing device for performing machine learning actions to determine the training and utilization of a model according to an example embodiment of the present disclosure is depicted.
[0013] Figure 2 A block diagram depicts an example machine learning action for determining a model according to an example embodiment of the present disclosure.
[0014] Figure 3 An example lock screen interface including an input collection element is depicted according to an example embodiment of the present disclosure.
[0015] Figure 4 An example lock screen interface including a live preview input collection element is depicted according to an example embodiment of the present disclosure.
[0016] Figure 5 A graphical view is depicted for displaying one or more suggested action elements within a lock screen interface, according to an example embodiment of the present disclosure.
[0017] Figure 6 A graphical view is depicted for capturing sensor data over a period of time within a lock screen interface, according to an example embodiment of the present disclosure.
[0018] Figure 7 A graphical view is depicted according to an example embodiment of the present disclosure for displaying one or more suggested action elements in response to an input signal received from a user.
[0019] Figure 8 A flowchart is depicted for an example method for performing effective multimodal input collection according to an example embodiment of the present disclosure.
[0020] The repeated reference numerals across multiple figures are intended to identify the same features in various embodiments. Detailed Implementation
[0021] Overview
[0022] Generally, this disclosure relates to multimodal input collection. More specifically, this disclosure relates to efficient and intuitive multimodal input collection for mobile devices. As an example, a mobile computing system (e.g., a smartphone, tablet, wearable device, etc.) may display a lock screen interface (e.g., an initial interface requesting user interaction and / or authentication before granting access to an application) on a display device associated with the mobile computing system. The lock screen interface may include an input collection element that, once selected, is configured to begin capturing sensor data from the mobile computing system's sensors. The mobile computing system may receive input signals from the user who selected the input collection element (e.g., providing a touch gesture at the location of the input collection element, providing a voice command, moving the mobile computing system in a certain manner, etc.). In response to receiving the input signals, the mobile computing system may capture sensor data from multiple sensors (e.g., an audio sensor, a front-facing image sensor, a rear-facing image sensor, a peripheral image sensor, etc.). Based on the sensor data, one or more suggested action elements may be determined from multiple action elements (e.g., using a machine learning model, etc.). Each of the suggested action elements can instruct a corresponding device action (e.g., delete recorded data, save recorded data, open the app, share recorded data, view recorded data, etc.). Once determined, the suggested action element can be displayed within the lock screen interface. In this way, users can record multimodal input data almost instantly, thus allowing users to react quickly and effectively to real-time events as they unfold (e.g., record video of the situation in real time, quickly generate reminders, etc.).
[0023] More specifically, the lock screen interface of a mobile computing system can be displayed on a display device (e.g., a touchscreen display device, etc.) associated with the mobile computing system. The lock screen interface can be, or otherwise includes, an interface that requests interaction and / or authentication from the user before granting access to a separate interface of the mobile computing system. More specifically, the lock screen interface can generally be understood as the interface first presented to the user after the mobile computing system exits a resting state. As an example, the mobile computing system can be in a "resting" state, wherein the display device associated with the mobile computing system is not in an active state (e.g., "sleep" mode, etc.). The mobile computing system can exit the resting state based on a specific stimulus (e.g., a specific movement of the mobile computing system, pressing a button on the mobile computing system, touching the display device of the mobile computing system, receiving notification data for an application running on the mobile computing system, etc.) and can activate the display device associated with the mobile computing system. Once the display device has been activated, the lock screen interface can be displayed on the display device.
[0024] In other implementations, the lock screen interface may include a blank or non-lit screen. As an example, in the case of, for instance, an OLED screen, the device may open the gate to input signals directly from the resting state. This reduces the need to "wake up" to illuminate the lock screen interface. In other words, the device can begin input, for example, by touching the lower third of the device's display and immediately starting to speak or record with the camera, or both.
[0025] In some implementations, the lock screen interface may request authentication data (e.g., fingerprint data, facial recognition data, password data, etc.) from the user. Additionally or alternatively, in some implementations, the lock screen interface may request input signals from the user (e.g., a swipe gesture on a "turn on screen" action element, movement of the mobile computing system, pressing a physical button on the mobile computing system, a voice command, etc.). As an example, the lock screen interface may include an authentication collection element. The authentication collection element may indicate a particular type of authentication data to be collected and may also indicate the status of authentication data collection. For example, the authentication collection element may be an icon representing a fingerprint (e.g., overlaid on a location part of a display device configured to collect fingerprint data, etc.), and the authentication collection element may be modified to indicate whether authentication data has been collected and / or accepted (e.g., changing the icon from red to green after successful authentication of biometric data, etc.).
[0026] In some implementations, user authentication may be required to access the suggested action. In other implementations, user authentication may not be required to access the suggested action. Providing the suggested feature as part of the lock screen interface can have the advantage that the lock screen is the first screen presented to the user when they pick up their device, and therefore having this feature in the lock screen means one less screen to pass through in order to access the desired action / application.
[0027] The lock screen interface may include an input collection element. This input collection element can be configured to capture multi-mode sensor data when selected. The input collection element can be an icon or other representation selectable by the user. For example, the input collection element could be an icon indicating the initiation of multi-mode input collection when selected. In some implementations, the input collection element may include, or otherwise represent, a preview of the multi-mode sensor data to be collected. For example, the input collection element may include, or otherwise represent, a preview of what the camera's image sensor is currently capturing. For instance, the rear image sensor of a mobile computing system may periodically capture image data (e.g., every two seconds, periodically when the mobile computing system is moved, etc.). The captured image data may be included as a thumbnail within the input collection element. As another example, the input collection element may include an audio waveform symbol indicating audio data currently captured by the audio sensor of the mobile computing system. In this way, the input collection element can indicate to the user that interaction with the input collection element can initiate multi-mode data collection, and may also indicate a preview of what data will be collected when the input collection element is selected (e.g., a preview of what the image sensor is currently capturing, etc.).
[0028] In some implementations, the input collection element can be selected for a time period and configured to capture sensor data during that period. As an example, the input collection element can be configured to be selected via a touch gesture or a touch-and-hold gesture (e.g., placing a finger on the input collection element and holding it there for a period of time). If selected via a touch gesture, the input collection element can capture sensor data corresponding to that moment of the touch gesture, or it can capture sensor data for a predetermined amount of time (e.g., two seconds, three seconds, etc.). If selected via a touch-and-hold gesture, the input collection element can capture sensor data for the duration of the touch-and-hold gesture (e.g., recording video and audio data whenever the user touches the input collection element).
[0029] A mobile computing system can acquire input signals from a user of the mobile computing system who selects an input collection element. As an example, the input signal can be a touch gesture or a touch-and-hold gesture at the location of the input collection element displayed on a display device. It should be noted that the input signal does not necessarily have to be a touch gesture. Instead, the input signal can be any type or manner of input signal from the user who selects the input collection element. As an example, the input signal can be a voice command from the user. As another example, the input signal can be movement by the user on the mobile computing system. As yet another example, the input signal can be a gesture performed by the user and captured by the image sensor of the mobile computing system (e.g., a gesture performed in front of the front-facing image sensor of the mobile computing system). As yet another example, the input signal can be or otherwise include multiple input signals from the user. For example, the input signal can include movement of the mobile computing system and voice commands from the user.
[0030] In some implementations, the type of sensor data collected can be based at least in part on the type of input signal obtained from the user. As an example, if the input signal is or otherwise includes a touch gesture (e.g., briefly touching the location of an input collection element on a display device), the input collection element can use an image sensor (e.g., a front image sensor, a rear image sensor, a peripheral image sensor, etc.) to collect a single image. If the input signal is or otherwise includes a touch and hold gesture (e.g., placing a finger at the location of the input collection element and holding the finger in that position for a period of time), the input collection element can collect video data (e.g., multiple image frames) as long as the touch and input signal are maintained. In this way, the input collection element can provide the user with precise control over the type and / or duration of input signals collected from the sensors of the mobile computing system.
[0031] In some implementations, preliminary sensor data can be collected for a period of time before receiving an input signal to select an input collection element. More specifically, preliminary sensor data (e.g., image data, etc.) can be continuously collected and updated, such that the preliminary sensor data collected within a time period can be appended to the sensor data collected after selecting the input collection element. As an example, a mobile computing system can continuously capture preliminary image data for the last five seconds before selecting an input collection element. The five seconds of preliminary image data can be appended to the sensor data collected after selecting the input collection element. In this way, in the case of real-time, rapidly occurring events, the user can access the sensor data even if they are not fast enough to select the input collection element in time to capture the real-time event. Additionally, in some implementations, the input collection element may include at least a portion of the preliminary sensor data (e.g., image data, etc.). As an example, the preliminary sensor data may include image data. The image data may be presented within the input collection element (e.g., depicted in the center of the input collection element, etc.). In this way, the input collection element can act as a preview to the user of what image data will be collected if the user selects the input collection element.
[0032] In response to receiving an input signal, the mobile computing system can capture sensor data from multiple sensors within the mobile computing system. These multiple sensors can include any conventional or future sensor devices included within the mobile computing system (e.g., image sensors, audio sensors, accelerometers, GPS sensors, LiDAR sensors, infrared sensors, ambient light sensors, proximity sensors, biometric sensors, barometers, gyroscopes, NFC sensors, ultrasonic sensors, etc.). As an example, the mobile computing system may include a front-facing image sensor, a rear-facing image sensor, peripheral image sensors (e.g., image sensors positioned around the edge of the mobile computing system and perpendicular to the front and rear image sensors), and an audio sensor. In response to receiving an input signal, the mobile computing system can capture sensor data from the front-facing image sensor and the audio sensor. As another example, the mobile computing system may include an audio sensor, a rear-facing image sensor, and a LiDAR sensor. In response to receiving an input signal, the mobile computing system can capture sensor data from the rear-facing image sensor, the audio sensor, and the LiDAR sensor.
[0033] Based at least in part on sensor data from multiple sensors, one or more suggested action elements can be determined from a plurality of predefined action elements. The one or more action elements may or may include elements that can be selected by a user of the mobile computing system. More specifically, each of the one or more suggested action elements may indicate a corresponding device action and may be configured to perform the corresponding device action when selected.
[0034] The suggested action element can be an element selectable by the user of the mobile computing system. As an example, the suggested action element could be an interface element displayed on a display device (e.g., a touch icon, etc.) and could be selected by the user using a touch gesture. As another example, the suggested action element could be or otherwise include descriptive text and could be selected by the user using a voice command. As yet another example, the suggested action element could be an icon indicating a movement pattern and could be selected by the user by replicating the movement pattern using the mobile computing system. For example, the suggested action element could be configured to share data with a separate user nearby, and the suggested action element could indicate a "shaking" motion (e.g., gripping the mobile computing system, indicating that the mobile computing system is being shaken, etc.). If the user replicates the movement pattern (e.g., shaking the mobile computing system, etc.), the data can be shared with the separate user.
[0035] One or more suggested action elements can instruct one or more corresponding device actions. Device actions can include actions that can be performed by a mobile computing system using captured sensor data (e.g., storing data, displaying data, editing data, sharing data, deleting data, transcribing data, providing sensor data to an application, opening an application associated with the data, generating instructions for an application based on the data, etc.). As an example, the captured data can include image data. Based on the image data, suggested action elements instructing device actions that share the image data with a second user can be determined. Following the previous example, a second suggested action element can be determined that instructs device actions that open an application that can utilize the image data (e.g., a social networking application, a photo editing application, a messaging application, a cloud storage application, etc.). As another example, the captured data can include image data depicting a scene and audio data. Audio data can include a user's voice recording of an application name. Suggested action elements instructing device actions that perform the voice recording of the application name can be determined. As another example, the captured data can include image data depicting a scene and audio data. Audio data can include a user's voice recording of a virtual assistant command. Suggested action elements instructing device actions that provide sensor data to a virtual assistant can be determined. In response, visual assistant applications can provide users with additional suggested action elements based on sensor data (e.g., search results from captured image data, search results from queries included in captured audio data, etc.).
[0036] One or more suggested action elements can be selected from a plurality of predefined action elements. Following the previous example, each of the previously described action elements can be included among a plurality of predefined action elements (e.g., copying data, sharing data, opening a virtual assistant application, etc.). Based on sensor data, one or more suggested action elements can be selected from a plurality of predefined action elements.
[0037] One or more suggested action elements can be displayed on the lock screen interface of a mobile computing system's display device. As an example, suggested action elements may or may not include icons selectable by the user (e.g., touch gestures on the display device). Icons can be displayed on the lock screen interface of the display device (e.g., above, below, or around an input collection element).
[0038] In some implementations, input signals can be obtained from a user who selects one or more suggested action elements. In response, the mobile computing system can execute a device action indicated by the suggested action element. As an example, suggested action elements instructing a virtual assistant application can be displayed within the lock screen interface on the display device. Input signals from the user can select suggested action elements (e.g., through touch gestures, voice commands, motion input, etc.). The mobile computing system can execute the virtual assistant application and provide sensor data to the virtual assistant application. It should be noted that suggested action elements can be selected in the same or substantially similar manner as previously described regarding input collection elements.
[0039] In some implementations, the mobile computing system may stop displaying the lock screen interface and instead display an interface corresponding to the virtual assistant application on the display device. Alternatively, in some implementations, the mobile computing system may determine and display additional suggested action elements in response to providing sensor data to the virtual assistant application. As an example, sensor data may be provided to the virtual assistant application in response to a user selecting a suggested action element that instructs the virtual assistant application. The virtual assistant application may process the sensor data and generate output (e.g., process image data depicting text content and generate search results based on the text content). One or more additional suggested action elements may be displayed within the lock screen interface based on the output of the virtual assistant application (e.g., providing suggested action elements that instruct mapping data associated with the sensor data, etc.). For example, if the sensor data includes a query, the suggested action element based on the output data may or may not include the result in response to that query. For another example, if the sensor data includes audio data, the suggested action element based on the output data may or may not depict a transcription of the audio data (e.g., a text transcription displayed on the lock screen interface). Therefore, it should be broadly understood that in some implementations, the suggested action element may not instruct a device action that can be performed by the mobile computing system. Instead, the suggested action element may or may depict information intended to be conveyed to the user.
[0040] In some implementations, to determine one or more suggested action elements, a machine learning action determination model (e.g., a neural network, recurrent neural network, convolutional neural network, one or more multilayer perceptrons, etc.) can be used to process sensor data. The machine learning action determination model can be configured to determine one or more suggested action elements from a plurality of predefined action elements. In some implementations, the machine learning action determination model can be a personalized model configured to determine one or more suggested action elements that the user is most likely to expect by training the model at least in part based on user-associated data (e.g., training the model in an unsupervised manner at least in part based on the user's historical choices of suggested action elements, etc.). As an example, one or more parameters of the machine learning action determination model can be adjusted at least in part based on the suggested action elements.
[0041] In some implementations, sensor data may include image data depicting one or more objects (e.g., from a front image sensor, a rear image sensor, a peripheral image sensor, etc.). One or more proposed action elements may be at least partially based on one or more objects, which may be determined by a mobile computing system (e.g., using one or more machine learning object recognition models, etc.). As an example, the object depicted in the image data may be a fast-food restaurant sign. Based at least partially on this object, the proposed action elements may instruct a food delivery application to perform device actions to deliver food to the restaurant. Alternatively, sensor data (e.g., annotations of the sensor data, etc.) may be provided to the food delivery application, providing the application with information about the identified fast-food restaurant sign. In this way, the mobile computing system may analyze image data and / or audio data (e.g., using one or more machine learning models, etc.) to determine the proposed action elements.
[0042] In some implementations, determining one or more suggested action elements may include displaying text content describing at least a portion of the audio data within the lock screen interface. As a more specific example, sensor data captured by the mobile computing system may include image data and audio data including speech from the user. The mobile computing system may determine one or more suggested action elements and may also display text content describing at least a portion of the audio data. For example, the mobile computing system may display a transcription of at least a portion of the audio data within the lock screen interface. In some implementations, the transcription may be displayed in real-time within the lock screen interface whenever the user selects an input collection element.
[0043] The systems and methods disclosed herein offer several technical effects and benefits. As an example, current user interface implementations of mobile devices typically require a series of cumbersome and time-consuming interactions from the user. Furthermore, the complexity and time required to navigate to certain applications can prevent users from capturing real-time events. For instance, users wishing to quickly capture real-time events (e.g., a moment in a sporting event) often need to navigate through a series of interfaces (e.g., lock screen, home screen, etc.) before opening the camera app to initiate data recording, making it exceptionally difficult to capture sensor data when the event occurs. As another example, the need to navigate multiple user interfaces before opening an application can often cause users to forget the initial purpose of operating their mobile device before they can access the originally intended application, significantly reducing user efficiency when using the device. In response, the systems and methods of this disclosure allow users to initiate sensor data capture almost instantaneously, thus significantly improving the speed and efficiency of user navigation of their mobile device's interface. Additionally, by significantly reducing the time required to operate the device before capturing sensor data, the systems and methods of this disclosure can significantly reduce the use of power, computing resources (e.g., CPU cycles, memory, etc.), and battery life involved in initiating sensor data capture.
[0044] Exemplary embodiments of this disclosure will now be discussed in more detail with reference to the accompanying drawings.
[0045] Example devices and systems
[0046] Figure 1A A block diagram of an example computing system 100 performing efficient input signal collection according to an exemplary embodiment of the present disclosure is depicted. The computing system 100 includes a mobile computing system 102 and a server computing system 130 communicatively coupled via a network 180.
[0047] Mobile computing system 102 can be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing system (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0048] Mobile computing system 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118 executed by processor 112 to cause mobile computing system 102 to perform operations.
[0049] In some implementations, the mobile computing system 102 may store or include one or more machine learning action determination models 120. For example, the machine learning action determination model 120 may be, or may otherwise include, various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models). Reference Figure 2 The discussion example is machine learning action determination model 120.
[0050] In some implementations, one or more machine learning action determination models 120 may be received from server computing system 130 via network 180, stored in mobile computing system memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, mobile computing system 102 may implement multiple parallel instances of a single machine learning action determination model 120 (e.g., performing parallel action determination across multiple instances of the machine learning action determination model).
[0051] More specifically, in some embodiments, to determine one or more suggested action elements, the mobile computing system 102 may utilize a machine learning action determination model 120 (e.g., a neural network, recurrent neural network, convolutional neural network, one or more multilayer perceptrons, etc.) to process sensor data. The machine learning action determination model 120 may be configured to determine one or more suggested action elements from a plurality of predefined action elements. In some embodiments, the machine learning action determination model 120 may be a personalized model configured to determine one or more suggested action elements that the user is most likely to expect by training the model 120 at least in part based on data associated with the user (e.g., training the model in an unsupervised manner at least in part based on the user's historical choices of suggested action elements, etc.). As an example, one or more parameters of the machine learning action determination model 120 may be adjusted at least in part based on the suggested action elements.
[0052] The mobile computing system 102 may also include one or more user input components 122 for receiving user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices through which the user can provide input.
[0053] Mobile computing system 102 may include multiple sensors 124. As an example, sensor 124 may include multiple image sensors (e.g., front image sensor 124A, rear image sensor 124B, peripheral image sensor 124C, etc.). For example, mobile computing system 102 may be or otherwise include a smartphone device, and the peripheral image sensor may be or otherwise include one or more cameras positioned around the periphery of the smartphone (e.g., the front image sensor 124A positioned perpendicular to the edge of the smartphone). As another example, sensor 124 may include an audio sensor 124D (e.g., a microphone, etc.). As yet another example, sensor 124 may include one or more of various sensors 124E (e.g., image sensors, audio sensors, accelerometers, GPS sensors, LiDAR sensors, infrared sensors, ambient light sensors, proximity sensors, biometric sensors, barometers, gyroscopes, NFC sensors, ultrasonic sensors, heartbeat sensors, etc.). It should be noted that mobile computing system 102 may include any sensors conventionally utilized in or by mobile computing devices.
[0054] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 that are executed by processor 132 to cause server computing system 130 to perform operations.
[0055] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by one or more server computing devices. In instances where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0056] Mobile computing system 102 and / or server computing system 130 can train model 120 via interaction with training computing system 150, which is communicatively coupled through network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.
[0057] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.
[0058] The training computation system 150 may include a model trainer 160, which uses various training or learning techniques—such as, for example, backpropagation of error—to train a machine learning model 120 stored at the mobile computation system 102. For example, a loss function can be used to update one or more parameters of the model (e.g., gradients based on the loss function) through backpropagation of the model. Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update parameters across multiple training iterations.
[0059] In some implementations, backpropagation of the error may include performing truncated backpropagation over time. The model trainer 160 may perform multiple generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0060] Specifically, model trainer 160 may train model 120 based on a set of training data 162. Training data 162 may include, for example, historical user data indicating past suggested action elements selected by a user. Additionally or alternatively, in some embodiments, training data 162 may include historical user data indicating past suggested action elements selected by multiple users. In this way, machine learning action determination model 120 may be trained to generate suggested action elements most likely to be served for a particular user and / or all user preferences. In some embodiments, training data may also include records of sensor data captured prior to a user making a particular choice. For example, in addition to knowing that a user selected a microphone to record audio data in a particular instance, training data may also indicate that the choice corresponds to the presence of sound in the environment at the time the choice was made.
[0061] In some implementations, training examples may be provided by mobile computing system 102 if the user has provided consent. Therefore, in such an implementation, the model 120 provided to mobile computing system 102 can be trained by training computing system 150 on user-specific data received from mobile computing system 102. In some instances, this process may be referred to as a personalized model.
[0062] Model trainer 160 includes computer logic for providing desired functionality. Model trainer 160 can be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some embodiments, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium such as RAM, a hard disk, or optical or magnetic media.
[0063] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication on network 180 can be carried over any type of wired and / or wireless connection using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0064] Figure 1A An example computing system that can be used to implement this disclosure is illustrated. Other computing systems may also be used. For example, in some embodiments, mobile computing system 102 may include a model trainer 160 and a training dataset 162. In such embodiments, model 120 can be trained and used locally at mobile computing system 102. In some of such embodiments, mobile computing system 102 may implement model trainer 160 to personalize model 120 based on user-specific data. Similarly, the proposed system can be implemented even when the mobile or user device is not connected to a network. For example, the device may activate all sensors and route the user to the most useful options via services on the device.
[0065] Figure 1B A block diagram of an example computing device 10 performing efficient input signal collection according to an exemplary embodiment of the present disclosure is depicted. The computing device 10 may be a mobile computing device or a server computing device.
[0066] The computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.
[0067] like Figure 1B As shown, each application can communicate with multiple other components of the computing device—such as, for example, one or more sensors, a scene manager, a device status component, and / or additional components. In some implementations, each application may communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is application-specific.
[0068] Figure 1C A block diagram of an example computing device 50, which performs machine learning actions to determine the training and utilization of a model according to an example embodiment of the present disclosure, is depicted. The computing device 50 may be a mobile computing device or a server computing device.
[0069] Computing device 50 includes multiple applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc. In some implementations, each application may communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0070] The central intelligence layer comprises multiple machine learning models. For example, such as... Figure 1C As shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model is managed by a central intelligent layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligent layer can provide a single model for all applications. In some embodiments, the central intelligent layer is included within or otherwise implemented by the operating system of the computing device 50.
[0071] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for computing device 50. For example... Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device—such as, for example, one or more sensors, a field manager, a device status component, and / or additional components. In some implementations, the central device data layer may communicate with each device component using an API (e.g., a proprietary API).
[0072] Example model layout
[0073] Figure 2 A block diagram of an example machine learning action determination model 200 according to an exemplary embodiment of the present disclosure is depicted. In some embodiments, the machine learning action determination model 200 is trained to receive a set of input data 204 describing a plurality of input signals, and as a result of receiving the input data 204, provides output data 206 including one or more suggested action elements. More specifically, the input data 204 may include, or otherwise, sensor data describing a plurality of sensors (e.g., camera, microphone, accelerometer, etc.) from a mobile device. The input data 204 may be processed by the machine learning action determination model 200 (e.g., neural network, recurrent neural network, convolutional neural network, one or more multilayer perceptrons, etc.) to obtain the output data 206. The output data 206 may include one or more suggested action elements from a plurality of predefined action elements. In some embodiments, the machine learning action determination model 200 may be a personalized model configured to determine one or more suggested action elements that a user is most likely to expect.
[0074] Figure 3An example lock screen interface 301, including an input collection element 302, is depicted according to an example embodiment of the present disclosure. More specifically, the lock screen interface 301 may be displayed on a display device (e.g., a touchscreen display device, etc.) associated with the mobile computing system 300. The lock screen interface 301 can generally be understood as an interface 301 first presented to the user after the mobile computing system 300 exits a sleep state. As an example, the mobile computing system 300 may be in a "resting" state, wherein the display device associated with the mobile computing system 300 is not in an active state (e.g., "sleep" mode, etc.). The mobile computing system 300 may exit the resting state based on a specific stimulus (e.g., a specific movement of the mobile computing system 300, pressing a button on the mobile computing system 300, touching the display device of the mobile computing system 300, receiving notification data for an application executed on the mobile computing system 300, etc.) and may activate the display device associated with the mobile computing system 300. Once the display device has been activated, the lock screen interface 301 may be displayed on the display device. In one example, in certain situations, such as displaying an OLED screen with low power availability, the device can switch directly from a resting state to capturing a signal.
[0075] The lock screen interface 301 may include an input collection element 302. The input collection element 302 may be configured to capture multi-mode sensor data when selected. The input collection element 302 may be an icon or another representation selectable by the user. Following the illustrated example, the input collection element 302 may depict an icon indicating the collection of input signals.
[0076] Figure 4 An example lock screen interface 401, including a real-time preview input collection element 402, is depicted according to an example embodiment of the present disclosure. More specifically, the lock screen 401 may be displayed on a display device of a mobile computing system 400. The lock screen 401 may include an input collection element 402. The input collection element 402 may be similar to... Figure 3The input collection element 302 may include, or otherwise depict a preview of the sensor data 405 to be collected (e.g., image data, etc.). For example, the rear image sensor 405 of the mobile computing system 403 may periodically (e.g., every two seconds, periodically when the mobile computing system 403 moves, etc.) capture sensor data including image data 406. The captured image data 406 may be included within the input collection element 402 or depicted as a thumbnail within the input collection element 402. As another example, the input collection element 402 may include an audio waveform symbol indicating the audio data currently captured by the audio sensor of the mobile computing system 403. In this way, the input collection element 402 may indicate to the user that interaction with the input collection element can initiate multi-mode data collection, and may also indicate a preview of what data will be collected when the input collection element is selected (e.g., a preview of what the image sensor is currently capturing, etc.).
[0077] Additionally, the lock screen interface 401 may include an authentication collection element 404. The authentication collection element 404 may indicate a certain type of authentication data to be collected and may also indicate the status of authentication data collection. Following the illustrated example, the authentication collection element 404 may be an icon representing a fingerprint. In some embodiments, the authentication collection element 404 may overlay the location of a biometric sensor of the mobile computing system 403 (e.g., a fingerprint sensor located below the display device of the mobile computing system 403). Alternatively, in some embodiments, the authentication collection element 404 may indicate that authentication data from the user is required. For example, when the user has not yet provided authentication data (e.g., facial recognition data, biometric data, etc.), the authentication collection element 404 may be depicted in a first color. Once the user provides authentication data, the authentication collection element 404 may be modified. For example, if the authentication data provided by the user is accepted, the authentication collection element 404 may change to a second color, and if the authentication data provided by the user is rejected, the authentication collection element 404 may change to a third color.
[0078] Figure 5 A graphical view depicting a lock screen interface 501 for displaying one or more suggested action elements according to an example embodiment of the present disclosure is provided. More specifically, a mobile computing system 500 may include a lock screen interface 501, which includes an input collection element 502. The mobile computing system 500 may obtain an input signal 504 for selecting the input collection element 502 from a user of the mobile computing system 500. To follow the depicted example, a user may provide a touch gesture input signal 504 by touching a display device (e.g., a touchscreen) of the mobile computing system 500 where the input collection element 502 is located.
[0079] When the mobile computing system 500 receives an input signal 504, it can collect sensor data from multiple sensors. Following the illustrated example, image data 506 can be captured from the rear image sensor of the mobile computing system 500, and audio data 507 (e.g., voice from a user) can be captured from the audio sensor of the mobile computing system 500. After capturing sensor data 506 / 507, a suggested action element can be determined from a plurality of predefined action elements, and the suggested action element can be displayed within the lock screen interface 501. Furthermore, the input collection element 502 can be modified to display at least a portion of the captured sensor data. For example, the input collection element can be expanded and can depict a portion of the captured image data 506.
[0080] One or more suggested action elements (e.g., 508, 510, 512, etc.) can be displayed within the lock screen interface. These suggested action elements can be determined at least in part based on captured sensor data 506 / 507. As an example, image data 506 can depict a parking lot. Image data 506 can be processed (e.g., using...) Figure 2 The image data 506 can be processed using machine learning action determination models, etc., and based on the content depicted in the image data 506, the suggested action element 508 for "drop pin" can be determined and displayed within the lock screen interface 501. As another example, the audio data 507 may include a voice command from the user instructing the mobile computing system 500 (e.g., a virtual assistant application of the mobile computing system) to save the reminder. Based on the audio data 507, the suggested action element 510 for "create reminder" can be determined and displayed within the lock screen interface 501. As yet another example, the image data 506 (e.g., using machine learning action determination models, etc.) can be determined... Figure 2 Machine learning action determination models (such as those used for motion determination) are insufficient to determine suggested action elements (e.g., based on blurred, cropped, distorted, etc. image data). Due to the insufficiency of the captured image data 506, suggested action elements 512 can be displayed, which are configured to capture additional sensor data when selected by the user. In this way, suggested device action indicators can be determined and displayed according to the context based on the content of the captured sensor data 506 / 507.
[0081] Additionally, the suggested action element 514 can be determined based on the captured sensor data 506 / 507 and displayed within the lock screen interface 501. The suggested action element 514 can open an application based at least in part on the sensor data 506 / 507. As an example, the captured sensor data may or may not include image data 506. Based on the capture of image data 506, a suggested action element 514 can be determined to open an application that can utilize image data (e.g., a social networking application, a photo sharing application, a video editing application, a cloud storage application, etc.). As another example, the captured image data 506 may depict a storefront. The image data 506 can be processed (e.g., using...). Figure 2 The system processes the image data 506 (using machine learning action determination models, etc.) and, based on the image data 506, can determine a suggested action element 514 to open an application associated with the store (e.g., an online shopping application corresponding to the store). As another example, the captured audio data 507 may include voice from a user instructing the mobile computing system 500 (e.g., a virtual assistant application of the mobile computing system 500, etc.) to open a virtual assistant application. Based on the audio data 507, the suggested action element 514 to open the virtual assistant application can be determined.
[0082] It should be noted that, in addition to opening the application upon selection, the suggested action element 514 may additionally or alternatively provide the application with captured sensor data 506 / 507, and / or generate instructions for the application. As an example, when selected, the suggested action element 514 may generate instructions for the application to share image data 506 (e.g., as specified by voice captured from the user in audio data 507). Therefore, it should be understood that selection of the suggested action element 514 can facilitate the application's functionality (e.g., by generating instructions, etc.) without requiring navigation out of the lock screen interface 501 or otherwise opening the application.
[0083] In some implementations, the user interface may include a "Save All" option, which allows the user to postpone the decision. For example, the user can postpone multiple options that allow the user to save as a reminder, video, still image, and / or other options. This can be referred to as a global save postponement.
[0084] Figure 6 A graphical view depicting a lock screen interface 601 for capturing sensor data over a time period is illustrated according to an example embodiment of the present disclosure. More specifically, the mobile computing system 600 may display the lock screen interface 601, which may include an input collection element 602. The mobile computing system 600 may receive an input signal 604 from a user who selects the input collection element 602.
[0085] In response, the mobile computing system 600 can acquire sensor data, including audio data 603 and image data 605. This can be similar to... Figure 5 Sensor data 506 / 507 is used to collect sensor data, except that input signal 604 can be a touch gesture performed for a period of time 607 at a display device (e.g., a touchscreen) of the mobile computing system 600. For example, input signal 604 can be a user touching the location of input collection element 602 for a period of time 607. During this period of time, audio data 603 and image data 605 can be collected.
[0086] While capturing audio data 603 and image data 605 within the time period 607, suggested action elements are displayed on the lock screen interface 601. As an example, when the user provides input signal 604 within a time period, a suggested action element 612 for "transcription" can be determined and displayed on the lock screen interface 601. The suggested action element 612 may or may not include real-time text transcription of speech captured within the audio data 603 (e.g., using a speech recognition model employing machine learning).
[0087] As another example, when a user provides input signal 604 within a time period, a suggested action element 606 for "recording capacity" can be determined and displayed within the lock screen interface 601. The suggested action element 606 may depict or otherwise indicate the maximum amount of time the mobile computing system 600 can capture image data 605. Simultaneously, image data 605 may be displayed within the lock screen interface 601. In this way, the mobile computing system 600 can indicate to the user the maximum value of the time period 607 during which the user provides input signal 604. After the user provides input signal 604 to the mobile computing system 600 for the time period 607, a suggested action element 610 may be displayed, allowing the user to manipulate the captured sensor data 603 / 605 (e.g., store at least a portion of sensor data 603 / 605, delete at least a portion of sensor data 603 / 605, display at least a portion of sensor data 603 / 605, provide sensor data 603 / 605 to the application, etc.).
[0088] Figure 7 A flowchart illustrating an example embodiment of the present disclosure for displaying one or more suggested action elements in response to receiving an input signal 703 from a user is depicted. More specifically, the mobile computing system 700 may display a lock screen interface 701 within a display device (e.g., a touchscreen device, etc.) of the mobile computing system. The lock screen interface 701 may include, as previously described... Figure 5The input collection element 702 is described. The mobile computing system 700 may receive an input signal 703. To follow the illustrated example, the input signal 703 may be or otherwise include audio data that selects the input collection element 702. For example, the input signal 703 may be audio data including voice describing a command from a user programmed to activate the input collection element 702 (e.g., the user says "starting recording," etc.). In some embodiments, the input signal 703 may be audio data including voice describing a command from a user programmed to activate the input collection element 702 for a period of time (e.g., the user says "record for 15 seconds," etc.). It should be noted that the input signal 703 is depicted as audio data merely to illustrate exemplary embodiments of this disclosure. Rather, the input signal 703 may be any kind of signal or set of signals. As an example, the input signal 703 may be accelerometer data indicating a user's movement mode of the mobile computing system 700, which is pre-programmed to select the input collection element 702. Therefore, it should be understood broadly that the input signal 703 may or may not include any type of signal data from any sensor and / or sensor set of the mobile computing system 700.
[0089] In response to receiving input signal 703, mobile computing system 700 can capture sensor data 705. Following the depicted example, sensor data 705 can be audio data, including voice from a user describing a command to open virtual assistant application 704. Based on sensor data 705, mobile computing system 700 can determine one or more suggested action elements and display them within lock screen interface 701. As an example, mobile computing system 700 can determine a suggested action element 718 for an "application window" corresponding to virtual assistant application 704, and can display the suggested action element 718 within the lock screen interface. It should be noted that in some embodiments, the suggested action element can be or otherwise include a window displayed within the lock screen in which the application is executed. Following the depicted example, virtual assistant application 704 can be executed within suggested action element 718 (e.g., as an application window, etc.). In this way, the user can interact directly with the application from the lock screen, eliminating the need to navigate a series of user interfaces to directly open the application. Additionally, the action elements 718 suggested by the "Application Window" may include additional suggested action elements that the user can select to interact with the application 704 (e.g., 708, 710, 712, etc.).
[0090] The suggested action element 718 may include multiple additional suggested action elements, which may be determined at least in part based on sensor data 705. As an example, sensor data 705 may correspond to a specific music artist (e.g., image data depicting album covers, audio data including a portion of music from the artist, etc.). In response, suggested action element 708 may instruct a device action to be performed by a music application separate from the virtual assistant application 704. As another example, sensor data 705 may include geolocation data indicating the user's location at an airport. In response, suggested action element 710 may instruct a device action prompting the virtual assistant application to interact with a ride-sharing application. As yet another example, sensor data 750 may include user history data indicating the user's preference to call a family member at the current time. In response, suggested action element 712 may correspond to an action to initiate a call to a family member when selected.
[0091] The proposed action element 718 may include an input collection element 702. As previously discussed... Figure 4 As described, the input collection element 702 may include, or otherwise depict, a preview of the sensor data to be collected. Additionally, the suggested action element 718 may include suggested action elements corresponding to actions of the input control device. As an example, the suggested action element 718 may include a suggested action element 714 instructing the mobile computing system 700 to collect sensor data from a front-facing image sensor instead of a rear-facing image sensor, or vice versa. As another example, the suggested action element 718 may include a suggested action element 716 instructing the mobile computing system 700 to collect additional user input from a virtual keyboard application. In this way, the suggested action element 718 may include additional suggested action elements allowing the user to control various settings or functions of the mobile computing system 700 (e.g., various power operation modes, input modes, sensor modes, etc. of the mobile computing system 700, such as a Wi-Fi switch, airplane mode switch, etc.).
[0092] Example Method
[0093] Figure 8 A flowchart depicting an example method 800 performed according to an example embodiment of the present disclosure is provided. Although Figure 8 The steps performed in a specific order are depicted for illustrative and discussion purposes, but the method of this disclosure is not limited to the specifically shown order or arrangement. The various steps of method 800 may be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of this disclosure.
[0094] At point 802, the computing system (e.g., a mobile computing system, a smartphone device, etc.) may display a lock screen interface on a display device of the mobile computing system, including input collection elements. More specifically, the mobile computing system may display a lock screen interface on a display device (e.g., a touchscreen display device, etc.) associated with the mobile computing system. The lock screen interface may be or otherwise include an interface that requests interaction and / or authentication from the user before granting access to a separate interface of the mobile computing system. More specifically, the lock screen interface can generally be understood as the interface first presented to the user after the mobile computing system exits a resting state. As an example, the mobile computing system may be in a "resting" state, wherein the display device associated with the mobile computing system is not in an active state (e.g., "sleep" mode, etc.). The mobile computing system may exit the resting state based on a specific stimulus (e.g., a specific movement of the mobile computing system, pressing a button on the mobile computing system, touching the display device of the mobile computing system, receiving notification data for an application running on the mobile computing system, etc.) and may activate the display device associated with the mobile computing system. Once the display device is activated, the lock screen interface may be displayed on the display device.
[0095] In some implementations, the lock screen interface may request authentication data (e.g., fingerprint data, facial recognition data, password data, etc.) from the user. Additionally or alternatively, in some implementations, the lock screen interface may request input signals from the user (e.g., a swipe gesture on a "turn on screen" action element, movement of the mobile computing system, pressing a physical button on the mobile computing system, a voice command, etc.). As an example, the lock screen interface may include an authentication collection element. The authentication collection element may indicate a particular type of authentication data that needs to be collected, or it may indicate the status of authentication data collection. For example, the authentication collection element may be an icon representing a fingerprint (e.g., overlaid on a location part of a display device configured to collect fingerprint data, etc.), and the authentication collection element may be modified to indicate whether authentication data has been collected and / or accepted (e.g., changing the icon from red to green after successful authentication of biometric data, etc.).
[0096] The lock screen interface may include an input collection element. This input collection element can be configured to capture multi-mode sensor data when selected. The input collection element can be an icon or other representation selectable by the user. For example, the input collection element could be an icon indicating the initiation of multi-mode input collection when selected. In some implementations, the input collection element may include, or otherwise represent, a preview of the multi-mode sensor data to be collected. For example, the input collection element may include, or otherwise represent a preview of what the camera's image sensor is currently capturing. For instance, the rear image sensor of a mobile computing system may periodically (e.g., every two seconds, periodically when the mobile computing system is moving, etc.) capture image data. The captured image data may be included as a thumbnail within the input collection element. As another example, the input collection element may include an audio waveform symbol indicating audio data currently captured by the audio sensor of the mobile computing system. In this way, the input collection element can indicate to the user that interaction with the input collection element can initiate multi-mode data collection, and may also indicate a preview of what data will be collected when the input collection element is selected (e.g., a preview of what the image sensor is currently capturing, etc.).
[0097] In some implementations, the input collection element can be selected for a time period and configured to capture sensor data during that period. As an example, the input collection element can be configured to be selected via a touch gesture or a touch-and-hold gesture (e.g., placing a finger on the input collection element and holding it there for a period of time). If selected via a touch gesture, the input collection element can capture sensor data corresponding to that moment of the touch gesture or it can capture sensor data for a predetermined amount of time (e.g., two seconds, three seconds, etc.). If selected via a touch-and-hold gesture, the input collection element can capture sensor data for the duration of the touch-and-hold gesture (e.g., recording video and audio data whenever the user touches the input collection element).
[0098] At 804, the computing system can acquire an input signal that selects an input collection element. More specifically, the computing system can acquire an input signal from a user of the mobile computing system that selects the input collection element. As an example, the input signal can be a touch gesture or a touch and hold gesture at the location of the input collection element displayed on a display device. It should be noted that the input signal does not necessarily have to be a touch gesture. Instead, the input signal can be any type or manner of input signal from the user who selects the input collection element. As an example, the input signal can be a voice command from the user. As another example, the input signal can be movement by the user on the mobile computing system. As another example, the input signal can be a gesture performed by the user and captured by the image sensor of the mobile computing system (e.g., a gesture performed in front of the front image sensor of the mobile computing system). As another example, the input signal can be or otherwise include multiple input signals from the user. For example, the input signal can include movement of the mobile computing system and voice commands from the user.
[0099] In some implementations, the type of sensor data collected can be based at least in part on the type of input signal obtained from the user. As an example, if the input signal is or otherwise includes a touch gesture (e.g., briefly touching the location of an input collection element on a display device), the input collection element can use an image sensor (e.g., a front image sensor, a rear image sensor, a peripheral image sensor, etc.) to collect a single image. If the input signal is or otherwise includes a touch and hold gesture (e.g., placing a finger at the location of the input collection element and holding the finger in that position for a period of time), the input collection element can collect video data (e.g., multiple image frames) for as long as the touch and input signal are held. In this way, the input collection element can provide the user with precise control over the type and / or duration of input signals collected from the sensors of the mobile computing system.
[0100] In some implementations, preliminary sensor data can be collected for a period of time before receiving an input signal to select an input collection element. More specifically, preliminary sensor data (e.g., image data, etc.) can be continuously collected and updated, such that the preliminary sensor data collected within a time period can be appended to the sensor data collected after selecting the input collection element. As an example, a mobile computing system can continuously capture the last five seconds of preliminary image data before selecting an input collection element. The five seconds of preliminary image data can be appended to the sensor data collected after selecting the input collection element. In this way, in the case of real-time, rapidly occurring events, the user can access the sensor data even if the user is not fast enough to select the input collection element in time to capture the real-time event. Additionally, in some implementations, the input collection element may include at least a portion of the preliminary sensor data (e.g., image data, etc.). As an example, the preliminary sensor data may include image data. The image data may be presented within the input collection element (e.g., depicted in the center of the input collection element, etc.). In this way, the input collection element can act as a preview for the user of what image data will be collected if the user selects the input collection element.
[0101] At point 806, the computing system can capture sensor data from multiple sensors of the mobile computing system. More specifically, the computing system can capture sensor data from multiple sensors of the mobile computing system in response to receiving an input signal. The multiple sensors can include any conventional or future sensor devices included within the mobile computing system (e.g., image sensors, audio sensors, accelerometers, GPS sensors, LiDAR sensors, infrared sensors, ambient light sensors, proximity sensors, biometric sensors, barometers, gyroscopes, NFC sensors, ultrasonic sensors, etc.). As an example, the mobile computing system may include a front-facing image sensor, a rear-facing image sensor, peripheral image sensors (e.g., image sensors positioned around the edge of the mobile computing system and perpendicular to the front and rear image sensors), and an audio sensor. In response to receiving an input signal, the mobile computing system can capture sensor data from the front-facing image sensor and the audio sensor. As another example, the mobile computing system may include an audio sensor, a rear-facing image sensor, and a LiDAR sensor. In response to receiving an input signal, the mobile computing system can capture sensor data from the rear-facing image sensor, the audio sensor, and the LiDAR sensor.
[0102] At point 808, the computing system can determine a proposed action element from a plurality of predefined action elements. More specifically, the computing system can determine one or more proposed action elements from the plurality of predefined action elements based at least in part on sensor data from a plurality of sensors. The one or more action elements may be or otherwise include elements that can be selected by a user of the mobile computing system. More specifically, each of the one or more proposed action elements may indicate a corresponding device action and may be configured to perform the corresponding device action when selected.
[0103] The suggested action element can be an element selectable by the user of the mobile computing system. As an example, the suggested action element could be an interface element displayed on a display device (e.g., a touch icon, etc.) and could be selected by the user using a touch gesture. As another example, the suggested action element could be or otherwise include descriptive text and could be selected by the user using a voice command. As yet another example, the suggested action element could be an icon indicating a movement pattern and could be selected by the user by replicating the movement pattern using the mobile computing system. For example, the suggested action element could be configured to share data with a separate user nearby, and the suggested action element could indicate a "shaking" motion (e.g., gripping the mobile computing system, indicating that the mobile computing system is being shaken, etc.). If the user replicates the movement pattern (e.g., shaking the mobile computing system, etc.), the data can be shared with the separate user.
[0104] One or more suggested action elements can instruct one or more corresponding device actions. Device actions can include actions that can be performed by a mobile computing system using captured sensor data (e.g., storing data, displaying data, editing data, sharing data, deleting data, transcribing data, providing sensor data to an application, opening an application associated with the data, generating instructions for an application based on the data, etc.). As an example, the captured data can include image data. Based on the image data, suggested action elements instructing a device action to share the image data with a second user can be determined. Following the previous example, a second suggested action element can be determined instructing a device action to open an application that can utilize the image data (e.g., a social networking application, a photo editing application, a messaging application, a cloud storage application, etc.). As another example, the captured data can include image data depicting a scene and audio data. Audio data can include a user's voice recording of an application name. Suggested action elements instructing a device action to perform the voice recording of the application name can be determined. As another example, the captured data can include image data depicting a scene and audio data. Audio data can include a user's voice recording of a virtual assistant command. Suggested action elements instructing a device action to provide sensor data to a virtual assistant can be determined. In response, visual assistant applications can provide users with additional suggested action elements based on sensor data (e.g., search results for captured image data, search results for queries included in captured audio data, etc.).
[0105] One or more suggested action elements can be selected from a plurality of predefined action elements. Following the previous example, each of the previously described action elements can be included among a plurality of predefined action elements (e.g., copying data, sharing data, opening a virtual assistant application, etc.). Based on sensor data, one or more suggested action elements can be selected from a plurality of predefined action elements.
[0106] At point 810, the computing system may display suggested action elements within the lock screen interface. More specifically, the computing system may display suggested action elements within the lock screen interface of the computing system's display device. As an example, the suggested action elements may or may include icons that can be selected by the user (e.g., via touch gestures at the display device). The icons may be displayed in the lock screen interface of the display device (e.g., above, below, or around the input collection element).
[0107] In some implementations, input signals can be obtained from a user who selects one or more suggested action elements. In response, the mobile computing system can perform a device action indicated by the suggested action element. As an example, suggested action elements instructing a virtual assistant application can be displayed within the lock screen interface of a display device. Input signals from the user can select suggested action elements (e.g., via touch gestures, voice commands, mobile input, etc.). The mobile computing system can execute the virtual assistant application and provide sensor data to the virtual assistant application. It should be noted that suggested action elements can be selected in the same or substantially similar manner as previously described regarding input collection elements.
[0108] In some implementations, the mobile computing system may stop displaying the lock screen interface and instead display an interface corresponding to the virtual assistant application on the display device. Alternatively, in some implementations, the mobile computing system may determine and display additional suggested action elements in response to providing sensor data to the virtual assistant application. As an example, sensor data may be provided to the virtual assistant application in response to a user selecting a suggested action element that instructs the virtual assistant application. The virtual assistant application may process the sensor data and generate output (e.g., process image data depicting text content and generate search results based on the text content). One or more additional suggested action elements may be displayed within the lock screen interface based on the output of the virtual assistant application (e.g., providing suggested action elements that instruct mapping data associated with the sensor data, etc.). For example, if the sensor data includes a query, the suggested action element based on the output data may or may include the result in response to that query. For another example, if the sensor data includes audio data, the suggested action element based on the output data may or may depict a transcription of the audio data (e.g., a text transcription displayed on the lock screen interface). Therefore, it should be broadly understood that in some implementations, the suggested action elements may not instruct device actions that can be performed by the mobile computing system. Instead, the suggested action elements may or may depict information intended to be conveyed to the user.
[0109] In some implementations, to determine one or more suggested action elements, a machine learning action determination model (e.g., a neural network, recurrent neural network, convolutional neural network, one or more multilayer perceptrons, etc.) can be used to process sensor data. The machine learning action determination model can be configured to determine one or more suggested action elements from a plurality of predefined action elements. In some implementations, the machine learning action determination model can be a personalized model configured to determine one or more suggested action elements that the user is most likely to expect by training the model at least in part based on user-associated data (e.g., training the model in an unsupervised manner at least in part based on the user's historical choices of suggested action elements, etc.). As an example, one or more parameters of the machine learning action determination model can be adjusted at least in part based on the suggested action elements.
[0110] In some implementations, sensor data may include image data depicting one or more objects (e.g., from a front image sensor, a rear image sensor, a peripheral image sensor, etc.). One or more proposed action elements may be at least partially based on one or more objects, which may be determined by the mobile computing system (e.g., using one or more machine learning object recognition models, etc.). As an example, the object depicted in the image data may be a fast-food restaurant sign. Based at least partially on this object, the proposed action elements may instruct a device to perform actions that would allow a food delivery application to deliver food to the restaurant. Alternatively, sensor data (e.g., annotations of the sensor data, etc.) may be provided to the food delivery application, providing the application with information about the identified fast-food restaurant sign. In this way, the mobile computing system may analyze the image data and / or audio data (e.g., using one or more machine learning models, etc.) to determine the proposed action elements.
[0111] In some implementations, determining one or more suggested action elements may include displaying text content describing at least a portion of the audio data within the lock screen interface. As a more specific example, sensor data captured by the mobile computing system may include image data and audio data including speech from the user. The mobile computing system may determine one or more suggested action elements and may also display text content describing at least a portion of the audio data. For example, the mobile computing system may display a transcript of at least a portion of the audio data within the lock screen interface. In some implementations, the transcript may be displayed in real-time within the lock screen interface whenever the user selects an input collection element.
[0112] Additional Disclosure
[0113] This article discusses technical reference servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0114] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation rather than limitation. Those skilled in the art will readily generate substitutions, variations, and equivalents for such embodiments upon understanding the foregoing. Therefore, this disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter, which will be apparent to those skilled in the art. For example, features shown or described as part of an embodiment may be used with another embodiment to produce yet another embodiment. Therefore, this disclosure is intended to cover such substitutions, variations, and equivalents.
Claims
1. A computer-implemented method for contextualized input collection and intent determination for mobile devices, comprising: A lock screen interface associated with a mobile computing system comprising one or more computing devices is provided, wherein the lock screen interface includes an input collection element configured to trigger the capture of sensor data when selected; The mobile computing system obtains an input signal from the user of the mobile computing system to select the input collection element; In response to receiving the input signal, the mobile computing system captures sensor data from a plurality of sensors of the mobile computing system, wherein the plurality of sensors include an audio sensor and one or both of a front-facing image sensor or a rear-facing image sensor; The mobile computing system determines one or more proposed action elements from a plurality of predefined action elements, at least in part, based on sensor data from the plurality of sensors, wherein each of the one or more proposed action elements indicates one or more device actions, and wherein the device actions are actions that can be performed by the mobile computing system using the sensor data; and The mobile computing system provides one or more suggested action elements within the lock screen interface at a display device associated with the mobile computing system for display.
2. The computer-implemented method according to claim 1, wherein, The method further includes: The mobile computing system obtains an input signal from the user of the mobile computing system, and the input signal selects a suggested action element from the one or more suggested action elements.
3. The computer-implemented method according to claim 2, wherein, The method further includes the mobile computing system performing the device action indicated by the proposed action element.
4. The computer-implemented method according to claim 2, wherein, Determining the action elements of the one or more proposed actions includes: The mobile computing system processes the sensor data using a machine learning action determination model to determine one or more suggested action elements from the plurality of predefined action elements.
5. The computer-implemented method according to claim 4, wherein: The machine learning action determination model is trained, at least in part, based on data associated with the user; and The method further includes adjusting one or more parameters of the machine learning action determination model by the mobile computing system, at least in part, based on the proposed action elements.
6. The computer-implemented method according to claim 1, wherein: The sensor data includes image data depicting one or more objects from one or more of the front image sensor or the rear image sensor; as well as The one or more suggested action elements are at least partially based on the one or more objects.
7. The computer-implemented method according to claim 6, wherein: The display device associated with the mobile computing system includes a touchscreen display; The input collection elements include touch elements; and The input collection element includes an authentication element configured to evaluate the user's fingerprint.
8. The computer-implemented method according to claim 1, wherein: The sensor data includes image data from one or more of the front image sensor or the rear image sensor; The input signal for selecting the input collection element includes performing a touch gesture at the input collection element for a period of time; as well as The sensor data is captured at least during the stated time period.
9. The computer-implemented method according to claim 1, wherein: The sensor data includes audio data from the audio sensor; as well as The determination of the one or more suggested action elements that respectively instruct the one or more devices to act further includes: the mobile computing system providing text content describing at least a portion of the audio data for display within the lock screen interface.
10. The computer-implemented method according to claim 1, wherein, The input signal used to select the input collection element includes one or more of the following: Touch gesture at the location of the input collection element on the display device associated with the mobile computing system; Voice commands; Gestures performed by the user; or The movement of the mobile computing system performed by the user.
11. The computer-implemented method according to claim 1, wherein, The one or more device actions include one or more of the following: Store at least a portion of the sensor data; Delete at least a portion of the sensor data; Display at least a portion of the sensor data; Provide the sensor data to the application; Open the application; or Generate instructions for one or more applications.
12. The computer-implemented method according to claim 1, wherein: Before obtaining the input signal for selecting the input collection element, the method includes: capturing preliminary image data from the front image sensor or the rear image sensor by the mobile computing system; and The input collection elements include at least a portion of the preliminary image data.
13. The computer-implemented method according to claim 1, wherein: Before obtaining the input signal that selects the input collection element, the method includes: capturing preliminary sensor data from the plurality of sensors by the mobile computing system over a time period; and The capture of sensor data from the plurality of sensors of the mobile computing system in response to obtaining the input signal further includes: the mobile computing system attaching the preliminary sensor data and the sensor data.
14. The computer-implemented method according to claim 1, wherein: The sensor data further includes geographic location data, and the plurality of sensors further includes location sensors; and The sensor data further includes accelerometer data, and the plurality of sensors further include accelerometers.
15. A mobile computing system, comprising: One or more processors; Multiple sensors, the multiple sensors including: One or more image sensors, including one or more of a front image sensor, a rear image sensor, or a peripheral image sensor; and Audio sensor; Display devices; and One or more tangible, non-transitory computer-readable media, the one or more tangible, non-transitory computer-readable media jointly storing instructions, the instructions, when executed by the one or more processors, causing the one or more processors to perform operations, the operations including: A lock screen interface associated with the mobile computing system is provided, wherein the lock screen interface includes an input collection element configured to trigger the capture of sensor data when selected; The user of the mobile computing system obtains an input signal to select the input collection element; In response to receiving the input signal, sensor data is captured from the plurality of sensors of the mobile computing system; At least in part, one or more suggested action elements are determined from a plurality of predefined action elements based on sensor data from the plurality of sensors, wherein the one or more suggested action elements respectively instruct one or more device actions, and wherein the device actions are actions that can be performed by the mobile computing system using the sensor data; and The suggested action elements within the lock screen interface are provided and displayed on the display device.
16. The mobile computing system according to claim 15, wherein, The operation further includes: Obtaining an input signal from the user of the mobile computing system, the input signal selecting a suggested action element from the one or more suggested action elements; and Perform the device action indicated by the suggested action element.
17. The mobile computing system according to claim 15, wherein, Determining the action elements of the one or more proposed actions includes: The sensor data is processed using a machine learning action determination model to determine one or more suggested action elements from the plurality of predefined action elements.
18. The mobile computing system according to claim 17, wherein: The machine learning action determination model is trained, at least in part, based on data associated with the user; and The operation further includes adjusting one or more parameters of the machine learning action determination model based at least in part on the proposed action elements.
19. One or more tangible, non-transitory computer-readable media that commonly stores instructions, which, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: A lock screen interface associated with a mobile computing system is provided, wherein the lock screen interface includes an input collection element configured to trigger the capture of sensor data when selected; The user of the mobile computing system obtains an input signal to select the input collection element; In response to receiving the input signal, sensor data is captured from a plurality of sensors of the mobile computing system, wherein the plurality of sensors include an audio sensor and one or both of a front-facing image sensor or a rear-facing image sensor; The mobile computing system determines one or more proposed action elements from a plurality of predefined action elements, at least in part, based on sensor data from the plurality of sensors, wherein each of the one or more proposed action elements indicates one or more device actions, and wherein the device actions are actions that can be performed by the mobile computing system using the sensor data; and The suggested action elements within the lock screen interface are provided and displayed at the display device associated with the mobile computing system.
20. One or more tangible, non-transitory computer-readable media according to claim 19, wherein, Determining the action elements of the one or more proposed actions includes: The mobile computing system processes the sensor data using a machine learning action determination model to determine one or more suggested action elements from the plurality of predefined action elements.