Gesture recognition based on machine learning

By using machine learning models combined with sensor data on wireless audio output devices, the problem of touchless sensor devices being unable to detect gestures is solved, accurate recognition of user gestures and corresponding operations are achieved, and the convenience of the device is improved.

CN112868029BActive Publication Date: 2025-10-03APPLE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080001961.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-25
Filing Date
2020-07-23
Publication Date
2025-10-03
Estimated Expiration
2041-04-02

AI Technical Summary

Technical Problem

Existing wearable devices such as earbuds have difficulty detecting user touch input and gestures without touch sensors.

Method used

Using machine learning models, specifically convolutional neural networks, combined with output data from accelerometers, optical sensors, and microphones, the model is trained to detect user gestures, such as taps and swipes, and this data is processed in real time on the wireless audio output device through a dedicated processor.

Benefits of technology

It achieves accurate detection of user gestures and performs corresponding operations, such as adjusting the volume, without relying on touch sensors, improving the convenience and functionality of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112868029B_ABST
    Figure CN112868029B_ABST
Patent Text Reader

Abstract

The subject technology receives a first sensor output of a first type from a first sensor of a device. The subject technology receives a second sensor output of a second type from a second sensor of the device, wherein the first sensor and the second sensor are non-touch sensors. The subject technology provides the first sensor output and the second sensor output as input to a machine learning model that has been trained to output a predicted touch-based gesture based on the first type of sensor output and the second type of sensor output. The subject technology provides the predicted touch-based gesture based on the output from the machine learning model. In addition, the subject technology adjusts an audio output level of the device based on the predicted gesture, wherein the device is an audio output device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification generally relates to gesture recognition, including gesture recognition based on machine learning. Background Art

[0002] The present disclosure relates generally to electronic devices and, in particular, to detecting gestures made by a user wearing or otherwise operating an electronic device. Wearable technology is receiving considerable attention, including audio accessory devices (e.g., earbuds), where users can potentially enjoy the benefits of mobile technology with increased convenience. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Certain features of the subject technology are set forth in the appended claims.For purposes of explanation, however, several embodiments of the subject technology are set forth in the following figures.

[0004] Figure 1 An exemplary network environment for providing machine learning-based gesture recognition in accordance with one or more implementations is illustrated.

[0005] Figure 2 An exemplary network environment including an exemplary electronic device and an exemplary wireless audio output device is shown in accordance with one or more implementations.

[0006] Figure 3 An exemplary architecture is shown that may be implemented by a wireless audio output device for detecting gestures based on other sensor output data without using a touch sensor, in accordance with one or more implementations.

[0007] Figures 4A to 4B An example timing diagram of sensor output from a wireless audio output device that may indicate corresponding gestures is shown according to one or more implementations.

[0008] Figure 5 A flowchart illustrating an example process for machine learning-based gesture recognition according to one or more implementations is shown.

[0009] Figure 6 Exemplary electronic systems are shown that can be used to implement various aspects of the subject technology according to one or more implementations. DETAILED DESCRIPTION

[0010] The specific embodiments shown below are intended to be descriptions of various configurations of the subject technology and are not intended to represent the only configuration in which the subject technology can be put into practice. The accompanying drawings are incorporated herein and constitute a part of the specific embodiments. The specific embodiments include specific details intended to provide a thorough understanding of the subject technology. However, the subject technology is not limited to the specific details described herein and can be put into practice using one or more other specific implementations. In one or more specific implementations, structures and components are shown in block diagram form to avoid making the concept of the subject technology vague.

[0011] Wearable devices such as earphones / headphones / earphones and / or one or more earbuds can be configured to include various sensors. For example, the earbuds can be equipped with various sensors, such as optical sensors (e.g., photoplethysmography (PPG) sensors), motion sensors, proximity sensors, and / or temperature sensors, which can work independently and / or in conjunction to perform one or more tasks, such as detecting when the earbud is placed in the user's ear, detecting when the earbud is placed in the housing, etc. The earbuds and / or earphones that include one or more of the aforementioned sensors may also include one or more additional sensors, such as a microphone or microphone array. However, given size / space constraints, power constraints, and / or manufacturing costs, the earbuds and / or earphones may not include a touch sensor for detecting touch input and / or touch gestures.

[0012] Nevertheless, it may be desirable to allow devices that do not include touch sensors, such as earbuds, to detect touch input and / or touch gestures from a user. The subject technology enables devices that do not include touch sensors to detect touch input and / or touch gestures from a user by utilizing input received via one or more non-touch sensors included in the device. For example, input received from one or more non-touch sensors can be applied to a machine learning model to detect whether the input corresponds to a touch input and / or gesture. In this way, the subject technology can enable detection of touch input and / or touch gestures, such as taps, swipes, etc., on a device surface without the use of a touch sensor.

[0013] Figure 1 An exemplary network environment for providing machine learning-based gesture recognition according to one or more implementations is shown. However, not all depicted components may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Variations in the arrangement and types of these components may be made without departing from the spirit or scope of the claims set forth herein. Additional, different, or fewer components may be provided.

[0014] The network environment 100 includes an electronic device 102, a wireless audio output device 104, a network 106, and a server 108. The network 106 may communicatively couple (directly or indirectly) the electronic device 102 and / or the server 108, for example. Figure 1 , the wireless audio output device 104 is shown as not being directly coupled to the network 106 ; however, in one or more implementations, the wireless audio output device 104 can be directly coupled to the network 106 .

[0015] The network 106 may be an interconnected network that may include the Internet or devices communicatively coupled to the Internet. In one or more specific implementations, the connection through the network 106 may be referred to as a wide area network connection, and the connection between the electronic device 102 and the wireless audio output device 104 may be referred to as a peer-to-peer connection. For purposes of explanation, Figure 1 The network environment 100 shown in FIG. 1 includes electronic devices 102 - 105 and a single server 108 ; however, the network environment 100 may include any number of electronic devices and any number of servers.

[0016] The server 108 may be and / or may include the following with respect to Figure 6 All or part of the electronic system discussed. Server 108 may include one or more servers, such as a server cloud. For purposes of explanation, a single server 108 is shown and discussed with respect to various operations. However, these operations and other operations discussed herein may be performed by one or more servers, and each different operation may be performed by the same or different servers.

[0017] The electronic device may be, for example, a portable computing device such as a laptop, a smartphone, a peripheral device (e.g., a digital camera, a headset), a tablet device, a smart speaker, a set-top box, a content streaming device, a wearable device such as a watch, a band, etc., or any other suitable device including one or more wireless interfaces (such as one or more near field communication (NFC) radios, WLAN radios, Bluetooth radios, Zigbee radios, cellular radios, and / or other wireless radios). Figure 1 In the embodiment, the electronic device 102 is depicted as a smart phone by way of example. The electronic device 102 may be and / or may include the following with respect to Figure 2 The electronic devices discussed and / or below with respect to Figure 6 All or part of the electronic system in question.

[0018] The wireless audio output device 104 may be, for example, a wireless headset, a wireless headphone, one or more wireless earbuds, a smart speaker, or generally any device that includes audio output circuitry and one or more wireless interfaces, such as a near field communication (NFC) radio, a WLAN radio, a Bluetooth radio, a Zigbee radio, and / or other wireless radios. Figure 1 , by way of example, the wireless audio output device 104 is depicted as a set of wireless earbuds. As discussed further below, the wireless audio output device 104 may include one or more sensors that can be used and / or repurposed for detecting input received from a user; however, in one or more specific implementations, the wireless audio output device 104 may not include a touch sensor. As referred to herein, a touch sensor may refer to a sensor that measures information directly resulting from a physical interaction corresponding to a physical touch (e.g., when a touch sensor receives touch input from a user). Some examples of touch sensors include surface capacitance sensors, projected capacitance sensors, line resistance sensors, surface acoustic wave sensors, and the like. The wireless audio output device 104 may be and / or may include the following references to Figure 2 The wireless audio output devices discussed and / or referenced below Figure 6 All or part of the electronic system in question.

[0019] The wireless audio output device 104 can be paired with the electronic device 102, such as via Bluetooth. After the two devices 102 and 104 are paired together, when the devices 102 and 104 are located in close proximity to each other, such as within the Bluetooth communication range of each other, they can automatically form a secure peer-to-peer connection. The electronic device 102 can stream audio, such as music, phone calls, etc., to the wireless audio output device 104.

[0020] For purposes of explanation, the subject technology is described herein with respect to a wireless audio output device 104. However, the subject technology may also be applied to wired devices that do not include a touch sensor, such as a wired audio output device. Further for purposes of explanation, the subject technology is discussed with respect to a device that does not include a touch sensor. However, in one or more specific implementations, the subject technology may be used in conjunction with a touch sensor, such as to enhance and / or improve detection of touch input to the touch sensor. For example, a device may include a low-cost and / or low-power touch sensor that can coarsely detect touch input and / or touch gestures, and the subject technology may be used to refine the coarsely detected touch input / gesture.

[0021] Figure 2 An exemplary network environment 200 including an exemplary electronic device 102 and an exemplary wireless audio output device 104 is shown according to one or more implementations. For purposes of explanation, Figure 2However, one or more components of the electronic device 102 may also be implemented by other electronic devices. Similarly, for the purpose of explanation, Figure 2 A wireless audio output device 104 is shown in the figure; however, one or more components of the wireless audio output device 104 may also be implemented by other devices. However, not all depicted components may be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Variations in the arrangement and types of these components may be made without departing from the spirit or scope of the claims set forth herein. Additional components, different components, or fewer components may be provided.

[0022] The electronic device 102 may include a host processor 202A, memory 204A, and radio frequency (RF) circuitry 206A. The wireless audio output device 104 may include a host processor 202B, memory 204A, RF circuitry 206B, a digital signal processor (DSP) 208, one or more sensors 210, a dedicated processor 212, and a speaker 214. In one embodiment, the sensor 210 may include one or more of a motion sensor (such as an accelerometer), an optical sensor, and an acoustic sensor (such as a microphone). It should be understood that the aforementioned sensors do not include capacitive or resistive touch sensor hardware (or any of the aforementioned examples of touch sensors).

[0023] The RF circuitry 206A-B may include one or more antennas and one or more transceivers for transmitting / receiving RF communications, such as WiFi, Bluetooth, cellular, etc. In one or more implementations, the RF circuitry 206A of the electronic device 102 may include circuitry for forming a wide area network connection and a peer-to-peer connection, such as WiFi, Bluetooth, and / or cellular circuitry, while the RF circuitry 206B of the wireless audio output device 104 may include Bluetooth, WiFi, and / or other circuitry for forming a peer-to-peer connection.

[0024] In one embodiment, the RF circuit 206B may be used to receive audio content that may be processed by the host processor 202B and sent to the speaker 214, and / or may also be used to receive signals from the RF circuit 206A of the electronic device 102 for performing tasks, such as adjusting the volume output of the speaker 214, among other types of tasks.

[0025] The host processors 202A-202B may include suitable logic, circuitry, and / or code to enable processing data and / or controlling the operation of the electronic device 102 and the wireless audio output device 104, respectively. In this regard, the host processors 202A-202B may be enabled to provide control signals to various other components of the electronic device 102 and the wireless audio output device 104, respectively. Additionally, the host processors 202A-202B may be enabled to implement an operating system or otherwise execute code to manage the operation of the electronic device 102 and the wireless audio output device 104, respectively. The memories 204A-204B may include suitable logic, circuitry, and / or code to enable storage of various types of information, such as received data, generated data, code, and / or configuration information. The memories 204A-204B may include, for example, random access memory (RAM), read-only memory (ROM), flash memory, and / or magnetic storage. The DSP 208 of the wireless audio output device 104 may include suitable logic, circuitry, and / or code to enable specific processing.

[0026] As described herein, a given electronic device, such as the wireless audio output device 104, may include a dedicated processor (e.g., dedicated processor 212) that is always powered on and / or in an active mode, for example, even when the device's host processor / application processor (e.g., host processor 202B) is in a low-power mode or when such an electronic device does not include a host processor / application processor (e.g., a CPU and / or GPU). Such a dedicated processor may be a low-computing-power processor designed to also utilize less energy than a CPU or GPU, and in one example, is also designed to continuously run on the electronic device to collect audio and / or sensor data. In one example, such a dedicated processor may be an always-powered processor (AOP), which is a small, low-power auxiliary processor implemented as an embedded motion coprocessor. In one or more specific implementations, the DSP 208 may be and / or may include all or part of the dedicated processor 212.

[0027] The dedicated processor 212 may be implemented as dedicated, customized, and / or proprietary hardware, such as a low-power processor that may be always powered on (e.g., to detect audio triggers, collect and process sensor data from sensors such as accelerometers, optical sensors, etc.) and continuously running on the wireless audio output device 104. The dedicated processor 212 may be used to perform certain operations in a more computationally efficient and / or power efficient manner. In one example, to enable deployment of a neural network model on a dedicated processor 212 that has less computational power than a host processor (e.g., host processor 202B), modifications to the neural network model are performed during the compilation process of the neural network model to make it compatible with the architecture of the dedicated processor 212. In one example, operations from such a compiled neural network model may be executed using the dedicated processor 212, which is described below. Figure 3 In one or more implementations, the wireless audio output device 104 may include only the dedicated processor 212 (eg, without the host processor 202B and / or the DSP 208).

[0028] In one or more specific implementations, the electronic device 102 can be paired with the wireless audio output device 104 to generate pairing information that can be used to form a connection (such as a peer-to-peer connection between the devices 102 and 104). The pairing can include, for example, exchanging communication addresses such as Bluetooth addresses. After pairing, the devices 102 and 104 can store the generated and / or exchanged pairing information (e.g., communication addresses) in the corresponding memories 204A-204B. In this way, when within the communication range of the corresponding RF circuits 206A-206B, the devices 102 and 104 can automatically and without user input connect to each other using the corresponding pairing information.

[0029] Sensors 210 may include one or more sensors for detecting device motion, user biometric information (e.g., heart rate), sound, light, wind, and / or generally any environmental input. For example, sensors 210 may include one or more of an accelerometer for detecting device acceleration, one or more microphones for detecting sound, and / or an optical sensor for detecting light. As described below with respect to Figures 3 to 5 Further described, the wireless audio output device 104 may be configured to output a predicted gesture based on output provided by the one or more sensors 210 (eg, corresponding to input detected by the one or more sensors 210 ).

[0030] In one or more implementations, one or more of the host processors 202A-202B, memories 204A-204B, RF circuits 206A-206B, DSP 208, and / or special purpose processor 212, and / or one or more portions thereof, may be implemented in software (e.g., subroutines and code), in hardware (e.g., an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gated logic components, discrete hardware components, or any other suitable device), and / or a combination of both.

[0031] Figure 3 An exemplary architecture 300 is shown that can be implemented by the wireless audio output device 104 according to one or more implementations for detecting gestures based on other sensor output data without using a touch sensor. However, not all depicted components can be used in all implementations, and one or more implementations may include additional or different components than those shown in the figures. Variations in the arrangement and types of these components may be made without departing from the spirit or scope of the claims set forth herein. Additional components, different components, or fewer components may be provided.

[0032] In one or more specific implementations, the architecture 300 can provide gesture detection without using information from the touch sensor based on sensor output fed into the gesture prediction engine 302, which can be executed by the dedicated processor 212. An example of a dedicated processor is detailed in U.S. Provisional Patent Application No. 62 / 855,840, filed on May 31, 2019, entitled "Compiling Code For A Machine Learning Model For Execution On A Specialized Processor," which is hereby incorporated by reference in its entirety for all purposes. As shown, the gesture prediction engine 302 includes a machine learning model 304. In one example, the machine learning model 304 is implemented as a convolutional neural network model that is configured to detect gestures over time using such sensor input. Furthermore, before deploying the machine learning model 304 on the wireless audio output device 104, the machine learning model can be pre-trained on a different device (e.g., the electronic device 102 and / or the server 108) based on sensor output data.

[0033] As described herein, a neural network (NN) is a computational model that uses a collection of connected nodes to process input data based on machine learning techniques. A neural network can be represented by connecting different operations together, hence the name "network." A model of a NN (e.g., a feedforward neural network) can be represented as a graph that shows how these operations are connected together from an input layer through one or more hidden layers, and finally to an output layer, where each layer includes one or more nodes, and where different layers perform different types of operations on corresponding inputs.

[0034] A convolutional neural network (CNN), as described above, is a type of neural network. As described herein, a CNN refers to a specific type of neural network, but uses different types of layers consisting of nodes that exist in three dimensions, and these dimensions can vary between layers. In a CNN, a node in a layer can only be connected to a subset of nodes in the previous layer. The final output layer can be fully connected and can be sized according to the number of classifiers. A CNN can include various combinations, and in some cases multiple types and orders of the following types of layers: an input layer, a convolutional layer, a pooling layer, a rectified linear unit layer (ReLU), and a fully connected layer. Part of the operations performed by a convolutional neural network includes obtaining a set of filters (or kernels) that iterate over the input data based on one or more parameters.

[0035] In one example, a convolutional layer reads input data (e.g., a 3D input volume corresponding to sensor output data, a 2D representation of sensor output data, or a 1D representation of sensor output data) using a kernel that reads in small segments at a time and steps across the entire input field. Each read can cause an input to be projected onto a filter graph and represent an internal interpretation of the input. As described herein, a CNN such as machine learning model 304 can be applied to human activity recognition data (e.g., sensor data corresponding to motion or movement), where the CNN model learns to map a given window of signal data to an activity (e.g., a gesture and / or a portion of a gesture) in which the model reads across each window of data and prepares an internal representation of the window.

[0036] like Figure 3 As shown, the Figure 2 A set of sensor outputs from the aforementioned sensors (e.g., sensor 210), from a first sensor output 306 to an Nth sensor output 308, are provided as input to the machine learning model 304. In one embodiment, each sensor output may correspond to a time window, such as 0.5 seconds, during which sensor data was collected by the corresponding sensor. In one or more embodiments, the sensor outputs 306-308 may be filtered and / or pre-processed, such as normalized, before being provided as input to the machine learning model 304.

[0037] In one specific implementation, the first sensor output 306 to the Nth sensor output 308 include data from an accelerometer, data from an optical sensor, and audio data from a microphone of the wireless audio output device 104. Based on the collected sensor data, a machine learning model (e.g., a CNN) can be trained. In a first example, the model is trained using data from the accelerometer and data from the optical sensor. In a second example, the model is trained using data from the accelerometer, data from the optical sensor, and audio data from the microphone.

[0038] In one example, input detected by an optical sensor, an accelerometer, and a microphone may individually and / or collectively indicate touch input. For example, a touch input / gesture along and / or on the body / housing of the wireless audio output device 104, such as a swipe up or down, a tap, etc., may result in a specific change in light detected by an optical sensor disposed on and / or along the body / housing. Similarly, such touch input / gesture may result in a specific sound detectable by the microphone and / or a specific vibration detectable by the accelerometer. Furthermore, if the wireless audio output device 104 includes multiple microphones disposed at different locations, the sound detected at each microphone may vary based on the location along the body / housing where the touch input / gesture is received.

[0039] After training, the machine learning model 304 generates a set of output predictions corresponding to the predicted gestures 310. After the predictions are generated, a policy may be applied to the predictions to determine whether to indicate an action to be performed by the wireless audio output device 104, as described below. Figures 4A to 4B This is discussed in more detail in the description of .

[0040] Figures 4A to 4B An example timing diagram of sensor output from the wireless audio output device 104 that may be indicative of corresponding touch gestures is shown in accordance with one or more implementations.

[0041] like Figures 4A to 4B As shown, graph 402 and graph 450 include respective timing graphs of sensor output data from respective motion sensors of wireless audio output device 104. The x-axes of graph 402 and graph 450 correspond to motion values ​​from an accelerometer or other motion sensor, and the y-axes of graph 402 and graph 450 correspond to time.

[0042] As also shown, graphs 404 and 452 include corresponding timing diagrams for the optical sensor of wireless audio output device 104. The x-axis of graphs 404 and 452 corresponds to the value of luminosity from the optical sensor (e.g., the amount of light reflected from the user's skin), and the y-axis of graphs 402 and 450 corresponds to time. Segments 410 and 412 correspond to time periods during which specific gestures occurred based on the sensor output data shown in graphs 402 and 404. Similarly, segments 460 and 462 correspond to time periods during which specific gestures occurred based on the sensor output data shown in graphs 450 and 452.

[0043] The corresponding predictions of the machine learning model 304 are shown in graph 406 and graph 454. In one embodiment, the machine learning model 304 provides ten predictions (or some other quantity of predictions) per second based on the aforementioned sensor output data, which is visually shown in graph 406 and graph 454. The machine learning model 304 provides prediction outputs that fall into one of seven different categories: no data ("none" or no gesture), the beginning of an upward swipe ("before swipe up"), the middle of an upward swipe ("swipe up"), or the end of an upward swipe ("after swipe up"), the beginning of a downward swipe ("before swipe down"), the middle of a downward swipe ("swipe down"), or the end of a downward swipe ("after swipe down").

[0044] The different categories of predicted gestures make the boundary conditions more robust, so that the machine learning model 304 does not need to provide a "hard" classification between no gesture and swipe up. Therefore, there is a transition period in which the sensor output data can go between gesture stages corresponding to no (e.g., no gesture), before swipe up, swipe up, then after swipe up, and then back to swipe down to no gesture. In graphs 406 and 454, the x-axis indicates the number of frames, where each frame can correspond to a single data point. In one example, multiple frames (e.g., 125 frames) can be used for each prediction provided by the machine learning model 304.

[0045] In one embodiment, as described above, the machine learning model 304 further utilizes a strategy to determine the predicted output. As mentioned herein, a strategy can correspond to a function that determines a mapping of a specific input (e.g., sensor output data) to a corresponding action (e.g., providing a corresponding prediction). In one example, the machine learning model 304 utilizes only sensor output data corresponding to an up swipe or a down swipe for classification, and the strategy can determine an average of a number of previous predictions (e.g., 5 previous predictions). The machine learning model 304 uses previous predictions within a specific time window, and when the average of these predictions exceeds a specific threshold, the machine learning model 304 can instruct the wireless audio output device 104 to initiate a specific action (e.g., adjust the volume of the speaker 214). In one or more embodiments, a state machine can be utilized to further refine the prediction, for example, based on previous predictions within a time window.

[0046] In one embodiment, for a pair of wireless audio output devices, such as a pair of earbuds, a corresponding machine learning model may be run independently on each wireless audio output device in the pair. The pair of wireless audio output devices may communicate with each other to determine whether the two wireless audio output devices detected a swipe gesture at approximately the same time. If the two wireless audio output devices detected a swipe gesture at approximately the same time, regardless of whether the swipe gestures were in the same direction or in opposite directions, the policy may determine to suppress detection of the swipe gesture to help prevent swipes in different directions from being detected on the two wireless audio output devices, or causing false triggering that is not caused by touch input on one of the wireless audio output devices but by general motion (e.g., an action that does not correspond to a specific action to be taken by the wireless audio output device).

[0047] Figure 5 A flowchart of an exemplary process for machine learning-based gesture recognition according to one or more specific implementations is shown. For the purpose of explanation, this document is primarily intended to Figure 1 The process 500 is described with reference to the wireless audio output device 104. However, the process 500 is not limited to Figure 1 The wireless audio output device 104 of the process 500 may be used, and one or more blocks (or operations) of the process 500 may be performed by one or more other components and other suitable devices. Further, for the purpose of explanation, the blocks of the process 500 are described herein as occurring sequentially or linearly. However, multiple blocks of the process 500 may occur in parallel. In addition, the blocks of the process 500 do not need to be performed in the order shown, and / or one or more blocks of the process 500 do not need to be performed and / or may be replaced by other operations.

[0048] The wireless audio output device 104 receives a first sensor output of a first type from a first sensor (eg, one of the sensors 210) (502). The wireless audio output device 104 receives a second sensor output of a second type from a second sensor (eg, another of the sensors 210) (504).

[0049] The wireless audio output device 104 provides the first sensor output and the second sensor output as input to a machine learning model that has been trained to output a predicted gesture based on the first type of sensor output and the second type of sensor output (506).

[0050] The wireless audio output device 104 provides a predicted gesture based on the output from the machine learning model (508). In one embodiment, the first sensor and the second sensor may be non-touch sensors. In one or more embodiments, the first sensor and / or the second sensor may include an accelerometer, a microphone, or an optical sensor. The first sensor output and the second sensor output may correspond to sensor input detected from a touch gesture provided by the user relative to the wireless audio output device 104, such as a tap, an upward swipe, a downward swipe, etc. In one embodiment, the predicted gesture includes at least one of the following: a start upward swipe, an intermediate upward swipe, an end upward swipe, a start downward swipe, an intermediate downward swipe, an end downward swipe, or a non-swipe. The corresponding sensor input for each of the aforementioned gestures may be a different input. For example, the sensor input for starting a swipe down or starting a swipe up may correspond to an accelerometer input indicating acceleration or movement in a first direction or a second direction, respectively, and / or whether a sensor input is received at a specific optical sensor located on the wireless audio output device 104 (e.g., if the wireless audio output device 104 includes at least two optical sensors), and / or a sensor input at a specific microphone (e.g., if the wireless audio output device 104 may include one or more microphones). In another example, the sensor input for ending a swipe up or ending a swipe down may correspond to an accelerometer input indicating that acceleration or movement has ended, and / or a sensor input to a second specific microphone, and / or whether a sensor input is received at a specific optical sensor located on the wireless audio output device 104.

[0051] The output level of the wireless audio output device 104, such as the audio output level or volume, can be adjusted based at least on the predicted gesture. In one example, sensor input corresponding to a swipe-down gesture can be determined based on a combination of predicted gestures including a start swipe-down, a middle swipe-down, and an end swipe-down. Furthermore, the swipe-down gesture can be predicted based on a specific order of the aforementioned predicted gestures, such as where the start swipe-down is predicted first, the middle swipe-down is predicted second, and the end swipe-down is predicted last. The system then determines a specific action for the wireless audio output device 104 based on the predicted gestures, such as adjusting the audio output level of the wireless audio output device 104.

[0052] In one example, sensor input corresponding to a middle-up swipe or a middle-down swipe results in a corresponding predicted gesture for performing an action of increasing or decreasing the audio output level of the wireless audio output device 104. In one example, the predicted gesture is based at least in part on the middle-up swipe or the middle-down swipe, and adjusting the audio output level includes increasing or decreasing the audio output level of the audio output device by a specific increment.

[0053] For another example, various combinations of gestures and / or a specific order of gesture combinations may correspond to specific actions based on continuous gestures (i.e., controls that are proportional to the distance traveled by the finger). In one example, adjusting the audio output level is based at least in part on: 1) a first distance based on the start of the upward swipe, the middle of the upward swipe, and the end of the upward swipe, or 2) a second distance based on the start of the downward swipe, the middle of the downward swipe, and the end of the downward swipe. In this example, adjusting the audio output level includes increasing or decreasing the audio output level of the audio output device 104 in proportion to the first distance or the second distance.

[0054] As described above, one aspect of the present technology is the collection and use of data that can be obtained from specific and legitimate sources to provide user information associated with messaging. The present disclosure contemplates that, in some instances, the collected data may include personal information data that uniquely identifies or can be used to identify a specific person. Such personal information data may include demographic data, location-based data, online identifiers, phone numbers, email addresses, home addresses, data or records related to the user's health or fitness level (e.g., vital signs measurements, medication information, exercise information), date of birth, or any other personal information.

[0055] The present disclosure recognizes that the use of such personal information data in the present technology can be used to benefit users. For example, personal information data can be used to provide information associated with messaging corresponding to the user. Therefore, the use of such personal information data can facilitate transaction processing (e.g., online transaction processing). In addition, the present disclosure also anticipates other uses of personal information data that benefit users. For example, health and fitness data can be used according to the user's preferences to provide insights into their overall health status, or can be used as positive feedback to individuals who use technology to pursue health goals.

[0056] This disclosure contemplates that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information will adhere to established privacy policies and / or practices. Specifically, such entities are expected to implement and consistently apply privacy practices generally recognized as meeting or exceeding industry or government requirements for maintaining user privacy. Such information regarding the use of personal data should be prominently displayed and easily accessible to users and updated as the collection and / or use of data changes. Users' personal information should be collected only for lawful uses. Furthermore, such collection / sharing should occur only after receiving user consent or other lawful basis as provided in applicable law. Furthermore, such entities should consider taking any necessary steps to safeguard and secure access to such personal information and ensure that others with access to such personal information adhere to their privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific types of personal information collected and / or accessed and adapted to applicable laws and standards, including specific considerations specific to jurisdictions that may impose higher standards. For example, in the United States, the collection or access of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while health data in other countries may be subject to other regulations and policies and should be handled accordingly.

[0057] Regardless of the foregoing, the present disclosure also contemplates implementation schemes in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates providing hardware elements and / or software elements to prevent or block access to such personal information data. For example, in the case of providing information corresponding to a user associated with messaging, the present technology may be configured to allow a user to choose to "opt in" or "opt out" to participate in the collection of personal information data at any time during or after registration for a service. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notifications related to access or use of personal information. For example, a user may be notified that their personal information data will be accessed when downloading an application, and then be reminded again just before the personal information data is accessed by the application.

[0058] Furthermore, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use. Risk can be minimized by limiting data collection and deleting data once it is no longer needed. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated where appropriate by removing identifiers, controlling the amount or specificity of stored data (e.g., collecting location data at a city level rather than an address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods such as differential privacy.

[0059] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, this disclosure also contemplates that various embodiments may be implemented without access to such personal information data. That is, various embodiments of the present technology will not be unable to function properly due to the lack of all or part of such personal information data.

[0060] Figure 6 An electronic system 600 is shown that can be utilized to implement one or more implementations of the subject technology. The electronic system 600 can be Figure 1 One or more of the electronic devices 102-106 and / or server 108 shown, and / or may be a portion thereof. The electronic system 600 may include various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 600 includes a bus 608, one or more processing units 612, a system memory 604 (and / or cache), a ROM 610, a permanent storage device 602, an input device interface 614, an output device interface 606, and one or more network interfaces 616, or subsets and variations thereof.

[0061] The bus 608 generally represents the entire system bus, peripheral bus, and chipset bus that communicatively connect the many internal devices of the electronic system 600. In one or more implementations, the bus 608 communicatively connects one or more processing units 612 with the ROM 610, the system memory 604, and the permanent storage device 602. The one or more processing units 612 retrieve instructions to be executed and data to be processed from these various memory units in order to perform the processes disclosed herein. In different implementations, the one or more processing units 612 can be a single processor or a multi-core processor.

[0062] ROM 610 stores static data and instructions required by one or more processing units 612 and other modules of electronic system 600. On the other hand, permanent storage device 602 can be a read-write memory device. Permanent storage device 602 can be a non-volatile memory unit that stores instructions and data even when electronic system 600 is turned off. In one or more specific implementations, a mass storage device (such as a magnetic or optical disk and its corresponding disk drive) can be used as permanent storage device 602.

[0063] In one or more implementations, a removable storage device (such as a floppy disk, a flash drive, and its corresponding disk drive) can be used as the permanent storage device 602. Like the permanent storage device 602, the system memory 604 can be a read-write memory device. However, unlike the permanent storage device 602, the system memory 604 can be a volatile read-write memory, such as random access memory. The system memory 604 can store any of the instructions and data that one or more processing units 612 may need during execution. In one or more implementations, the processes disclosed herein are stored in the system memory 604, the permanent storage device 602, and / or the ROM 610. The one or more processing units 612 retrieve instructions to be executed and data to be processed from these various memory units in order to perform the processes of one or more implementations.

[0064] The bus 608 is also connected to an input device interface 614 and an output device interface 606. The input device interface 614 enables a user to transmit information and select commands to the electronic system 600. Input devices that can be used with the input device interface 614 may include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 606 may, for example, enable the display of images generated by the electronic system 600. Output devices that can be used with the output device interface 606 may include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a flexible display, a flat panel display, a solid-state display, a projector, or any other device for outputting information. One or more specific implementations may include a device that acts as both an input device and an output device, such as a touch screen. In these specific implementations, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input.

[0065] Finally, if Figure 6 As shown, bus 608 also couples electronic system 600 to one or more networks and / or to one or more network nodes, such as a network interface 616. Figure 1. In this manner, electronic system 600 can be part of a computer network, such as a LAN, a wide area network ("WAN"), or an intranet, or can be part of a network of networks, such as the Internet. Any or all components of electronic system 600 can be used with the subject disclosure.

[0066] Implementations within the scope of the present disclosure may be implemented in part or in whole using a tangible computer-readable storage medium (or multiple tangible computer-readable storage media of one or more types) having one or more instructions programmed thereon. The tangible computer-readable storage medium may also be non-transitory in nature.

[0067] Computer-readable storage media can be any storage medium that can be read, written, or otherwise accessed by a general-purpose or special-purpose computing device, including any processing electronics and / or processing circuitry capable of executing instructions. For example, and without limitation, computer-readable media can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, and TTRAM. Computer-readable media can also include any non-volatile semiconductor memory, such as ROM, PROM, EPROM, EEPROM, NVRAM, flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, and Millipede memory.

[0068] Furthermore, the computer-readable storage medium may include any non-semiconductor memory, such as optical disk storage, magnetic disk storage, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions. In one or more implementations, the tangible computer-readable storage medium may be directly coupled to the computing device, while in other implementations, the tangible computer-readable storage medium may be indirectly coupled to the computing device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.

[0069] Instructions can be directly executable or can be used to develop executable instructions. For example, instructions can be implemented as executable or non-executable machine code, or can be implemented as high-level language instructions that can be compiled to produce executable or non-executable machine code. In addition, instructions can also be implemented as data, or can include data. Computer executable instructions can also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, etc. As those skilled in the art will appreciate, the details including but not limited to the number, structure, sequence and organization of instructions can be significantly different without changing the underlying logic, function, processing and output.

[0070] While the above discussion primarily relates to microprocessors or multi-core processors that execute software, one or more implementations are performed by one or more integrated circuits such as ASICs or FPGAs. In one or more implementations, such integrated circuits execute instructions stored on the circuits themselves.

[0071] Those skilled in the art will recognize that the various illustrative blocks, modules, elements, parts, methods and algorithms described herein can be implemented as electronic hardware, computer software or a combination of the two. In order to illustrate this interchangeability of hardware and software, various illustrative blocks, modules, elements, parts, methods and algorithms have been generally described above in terms of functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Technicians can implement the described functionality in different ways for each specific application. Various components and blocks can be arranged differently (e.g., arranged in different orders, or divided in different ways) without departing from the scope of the present subject technology.

[0072] Should be understood that the specific order or the hierarchical structure of the frames in the process disclosed by the present invention are illustrations of exemplary methods. Based on design preference requirements, it should be understood that the specific order or the hierarchical structure of the frames in the process can be rearranged or all the frames shown are executed. Any frame in these frames can be executed simultaneously. In one or more specific implementations, multitasking and parallel processing may be advantageous. In addition, the division of each system component in the above-mentioned specific implementation should not be understood as requiring this type of division in all specific implementations, and should be understood that program components and systems can generally be integrated together in a single software product or be packaged in multiple software products.

[0073] As used in this specification and any claims of this patent application, the terms "base station," "receiver," "computer," "server," "processor," and "memory" refer to electronic devices or other technical equipment. These terms exclude people or groups of people. For the purposes of this specification, the terms "display" or "displaying" mean displaying on an electronic device.

[0074] As used herein, the phrase "at least one of" following a list of items, any of which is separated by the terms "and" or "or," modifies the list as a whole, rather than each member (i.e., each item) of the list. The phrase "at least one of" does not require selection of at least one of each item listed; rather, the phrase allows for a meaning that includes at least one of any one item and / or at least one of any combination of items and / or at least one of each item. For example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0075] The predicate words "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of a subject matter and are intended to be used interchangeably. In one or more implementations, a processor configured to monitor and control an operation or component may also mean that the processor is programmed to monitor and control the operation or that the processor is operable to monitor and control the operation. Similarly, a processor configured to execute code may be interpreted as a processor that is programmed to execute code or operable to execute code.

[0076] Phrases such as aspect, this aspect, another aspect, some aspects, one or more aspects, an implementation, this implementation, another implementation, some implementations, one or more implementations, an embodiment, this embodiment, another embodiment, some embodiments, one or more embodiments, configuration, this configuration, other configurations, some configurations, one or more configurations, subject technology, disclosure, the present disclosure, other variations thereof, and the like are used for convenience and do not imply that disclosure involving such one or more phrases is essential to the subject technology or that such disclosure applies to all configurations of the subject technology. Disclosure involving such one or more phrases may apply to all configurations or one or more configurations. Disclosure involving such one or more phrases may provide one or more examples. Phrases such as aspect or some aspects may refer to one or more aspects and vice versa, and this applies similarly to the other aforementioned phrases.

[0077] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" or as an "example" is not necessarily to be construed as preferred or advantageous over other implementations. Furthermore, to the extent that the terms "including," "having," and the like are used in the specification or claims, such terms are intended to be inclusive, similar to the way the term "comprising" is interpreted when used as a transitional word in a claim.

[0078] All structural and functional equivalents to the elements of various aspects described throughout this disclosure that are known or later come to be known to one of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. In addition, nothing disclosed herein is intended to be made available to the public regardless of whether such disclosure is explicitly recited in the claims. No claim element shall be construed under 35 U.S.C. §112(f) unless the element is explicitly recited using the phrase “means for” or, in the case of a method claim, the phrase “step for”.

[0079] The previous description is provided to enable those skilled in the art to practice various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the present claims are not intended to be limited to the aspects shown herein, but are intended to make the full scope consistent with the language claims, wherein reference to elements in singular values ​​is not intended to mean "only one", but refers to "one or more", unless specifically noted. Unless otherwise specifically stated, the term "some" refers to one or more. Male pronouns (e.g., his) include female and neutral (e.g., her and its), and vice versa. Titles and subtitles (if any) are used only for convenience and do not limit this subject disclosure.

Claims

1. A method comprising: receiving a first sensor output of a first type from a first sensor of the device; receiving a second sensor output of a second type different from the first type from a second sensor of the device, wherein the first sensor and the second sensor do not include touch sensors, and wherein the first sensor output and the second sensor output correspond to sensor input detected in response to a touch-based gesture provided by a user on a surface of the device; providing the first sensor output and the second sensor output as input to a machine learning model, the machine learning model having been trained to output a predicted gesture based on the first type of sensor output and the second type of sensor output; determining a predicted touch-based gesture by combining a plurality of predicted gestures from the machine learning model; as well as An audio output level of the device is adjusted based on the predicted touch-based gesture, wherein the device comprises an audio output device. 2 . The method of claim 1 , wherein each of the first sensor and the second sensor comprises at least one of an accelerometer, a microphone, or an optical sensor. The method of claim 1 , wherein an action is performed when a specific prediction is received a specific number of times. 4 . The method of claim 1 , wherein the predicted gesture comprises at least one of: a start swipe up, a middle swipe up, a end swipe up, a start swipe down, a middle swipe down, a end swipe down, or no swipe.

5. The method of claim 4 , wherein the predicted gesture is based at least in part on the middle swipe up or the middle swipe down, and adjusting the audio output level comprises increasing or decreasing the audio output level of the audio output device by a specific increment, or wherein adjusting the audio output level is based at least in part on: a first distance based on the starting swipe up, the middle swipe up, and the ending swipe up, or a second distance based on the starting swipe down, the middle swipe down, and the ending swipe down, and adjusting the audio output level comprises increasing or decreasing the audio output level of the audio output device proportionally to the first distance or the second distance.

6. A device comprising: sensor; at least one processor; as well as a memory comprising instructions that, when executed by the at least one processor, cause the at least one processor to: receiving a sensor output of a predefined type from the sensor, the sensor being a non-touch sensor, wherein the sensor output corresponds to a sensor input detected in response to a touch-based gesture provided by a user on a surface of the device; providing the sensor output to a machine learning model that has been trained to output a predicted touch-based gesture based on previous sensor output of the predefined type; determining a predicted touch-based gesture by combining a plurality of predicted gestures from the machine learning model; as well as An audio output level of the device is adjusted based on the predicted touch-based gesture.

7. The apparatus of claim 6, wherein the sensor comprises at least one of: an accelerometer, a microphone, or an optical sensor.

8. The device of claim 6, wherein the predicted gesture comprises at least one of: a start swipe up, a middle swipe up, a end swipe up, a start swipe down, a middle swipe down, a end swipe down, or no swipe.

9. The device of claim 8, wherein the predicted gesture is based at least in part on the middle swipe up or the middle swipe down, and adjusting the audio output level comprises increasing or decreasing the audio output level by a specific increment, Or wherein adjusting the audio output level is based at least in part on: a first distance based on the starting upward swipe, the middle upward swipe and the ending upward swipe, or a second distance based on the starting downward swipe, the middle downward swipe and the ending downward swipe, and adjusting the audio output level includes increasing or decreasing the audio output level proportionally to the first distance or the second distance.

10. The device of claim 6, further comprising a second sensor, the instructions further causing the at least one processor to: receiving a second sensor output of a second predefined type from the second sensor, the second sensor being a non-touch sensor; and The second sensor output is provided along with the sensor output to the machine learning model, the machine learning model having been trained to output the predicted gesture based on previous sensor outputs of the predefined type and the second predefined type.

11. A computer program product comprising code stored in a non-transitory computer-readable storage medium, the code comprising: for receiving, from a first sensor of a device, a code of a first sensor output of a first type associated with a gesture provided by a user relative to the device; code for receiving, from a second sensor of the device, a second sensor output of a second type associated with the gesture provided by the user, wherein the first sensor and the second sensor do not comprise touch sensors, and wherein the first sensor output and the second sensor output correspond to sensor input detected in response to a touch-based gesture provided by the user on a surface of the device; code for providing the first sensor output and the second sensor output as input to a machine learning model, the machine learning model having been trained to output a predicted gesture based on the first type of sensor output and the second type of sensor output; code for determining a predicted touch-based gesture by combining a plurality of predicted gestures from the machine learning model; as well as Code for adjusting an audio output level of the device based on the predicted touch-based gesture.

12. The computer program product of claim 11, wherein each of the first sensor and the second sensor comprises an accelerometer, a microphone, or an optical sensor.

13. The computer program product of claim 11, wherein the predicted gesture comprises at least one of: a start swipe up, a middle swipe up, an end swipe up, a start swipe down, a middle swipe down, an end swipe down, or no swipe.

14. The computer program product of claim 13 , wherein the predicted gesture is based at least in part on the middle swipe up or the middle swipe down, and adjusting the audio output level comprises increasing or decreasing the audio output level by a specific increment, Or wherein adjusting the audio output level is based at least in part on: a first distance based on the starting upward swipe, the middle upward swipe and the ending upward swipe, or a second distance based on the starting downward swipe, the middle downward swipe and the ending downward swipe, and adjusting the audio output level includes increasing or decreasing the audio output level proportionally to the first distance or the second distance.

Citation Information

Patent Citations

  • Microelectromechanical systems (MEMS) acoustic sensor-based gesture recognition

    US20160091308A1