Determination of less destructive commands
By using the radar system in the computing device to detect and compare the radar signal characteristics of fuzzy gestures, the problem of inaccurate gesture recognition is solved, and a higher user satisfaction and interactive experience is achieved.
Patent Information
- Application Number
- CN202280100597.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to accurately identify fuzzy gestures in gesture recognition, causing the computing device to perform unintentional operations and affect user satisfaction.
By integrating the radar system in the computing device, the fuzzy gestures performed by the user are detected and compared with the stored radar signal characteristics to determine the corresponding command. The computing device may determine a less destructive command and perform operations associated with the command.
Improve the accuracy of gesture recognition, reduce user frustration caused by incorrect recognition, and improve the user's interactive experience with the device.
Smart Images

Figure CN120051745A_ABST
Abstract
Description
Background Art
[0001] Year after year, computing devices play an increasingly important role in people's lives. However, this more important role has a correspondingly greater demand for seamless and universal interaction between people and their devices. Gone are the days of interacting with a desktop computer solely through a physical keyboard. To meet this demand, new ways to interact with devices have been developed. However, many of these ways of interaction include attendant design difficulties. For example, some computing devices use gesture recognition to enable users to control their devices without the user physically contacting the device or the device's peripherals. Detecting or recognizing these types of gestures can be difficult to perform accurately, which can cause the computing device to perform differently than intended. Summary of the invention
[0002] This document describes techniques, devices, and systems for determining less destructive commands. A computing device may detect an ambiguous gesture performed by a user and compare a radar signal characteristic of the ambiguous gesture to one or more stored radar signal characteristics to correlate the ambiguous gesture with a first gesture and a second gesture. The first gesture and the second gesture may cause the computing device to execute a first command and a second command, respectively. The computing device may determine a less destructive command of the first command and the second command, and perform an operation associated with the less destructive command. In doing so, a device performing radar-based gesture detection may reduce the consequences of inaccurate gesture recognition, thereby improving user satisfaction.
[0003] Aspects described below include methods, systems, devices, and means for determining a less destructive command by a radar-based gesture detector. The method may include detecting an ambiguous gesture performed by a user at a computing device and using a radar system. The ambiguous gesture may be associated with a specific radar signal characteristic. The radar signal characteristic may be compared to one or more stored radar signal characteristics, the comparison effectively correlating the ambiguous gesture with a first gesture and a second gesture. The first gesture and the second gesture may be associated with a first command and a second command, respectively. The computing device may determine that the first command will be less destructive than the second command. In response to determining that the first command will be less destructive than the second command, the computing device, an application associated with the computing device, or another device associated with the computing device may be directed to execute the first command.
[0004] A device is also described, the device comprising a radar system capable of transmitting and receiving radar signals. The device also comprises at least one computer-readable storage medium storing instructions that, when executed by at least one processor, determine a less destructive command according to the above method. Means for performing the method are also described. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] This article describes devices and techniques for determining less destructive commands. Throughout the drawings, the same reference numerals are used to reference the same features and components:
[0006] Figure 1 An example environment with a radar-enabled computing device, a user, a proximity area, and a radar system is shown;
[0007] Figure 2 Shows Figure 1 Example implementations of a radar-enabled computing device;
[0008] Figure 3 An example environment is shown in which a plurality of radar-enabled computing devices are connected via a communication network to form a computing system;
[0009] Figure 4 An example environment is shown in which a radar system is used by a computing device to detect, distinguish, and / or recognize a user or a gesture being performed by the user;
[0010] Figure 5 An example implementation of an antenna, analog circuitry, and system processor for a radar system is shown;
[0011] Figure 6 shows an example implementation in which a user module can distinguish between users;
[0012] Figure 7 An example implementation of a machine learning (ML) model for distinguishing users of a computing device is shown;
[0013] Figure 8 An example implementation of a gesture module is shown that utilizes a spatiotemporal machine learning model to improve detection and recognition of gestures;
[0014] Fig. 9 An example implementation of deep learning techniques utilized by the frame model is shown;
[0015] Fig.10 An example implementation of deep learning techniques utilized by a temporal model is shown;
[0016] Fig.11 Experimental results indicating improved performance in gesture recognition when utilizing radar enhancement techniques are shown;
[0017] Fig.12 Experimental data showing a user performing a tap gesture;
[0018] Fig.13 Experimental data showing a user performing a tap gesture, a right swipe, a strong left swipe, and a weak left swipe;
[0019] Fig.14 Experimental data showing three sets of negative data that can be stored to improve detection and recognition of gestures from background motion;
[0020] Fig.15 Experimental results on the accuracy of gesture recognition in the presence of background motion are shown;
[0021] Fig.16 Experimental results on the accuracy of gesture detection and recognition when adversarial negative data is additionally used are shown;
[0022] Fig.17 The experimental results (confusion matrix) related to the accuracy of gesture recognition are shown;
[0023] Fig.18 Experimental results corresponding to the accuracy of gesture recognition at various linear and angular displacements of the antenna of the radar system are shown;
[0024] Fig.19 An example implementation of a radar-enabled computing device that uses additional sensors to improve the fidelity of user detection and differentiation is shown;
[0025] Fig. 20 An example environment is shown in which privacy settings are modified based on user presence;
[0026] Fig.21 An example environment is shown in which user differentiation techniques may be implemented using a plurality of computing devices forming a computing system;
[0027] Fig. 22 An example environment is shown in which operations are continuously performed across multiple computing devices of a computing system;
[0028] Fig.23 An example environment is shown in which a computing system implements continuity of operations across multiple computing devices;
[0029] Fig.24 A technique for radar-based ambiguous gesture determination using contextual information is shown;
[0030] Fig.25 An example implementation is shown in which a gesture module can recognize a gesture performed by a user;
[0031] Fig.26 An example environment is shown in which a computing device can utilize contextual information of a user's habits to improve gesture recognition;
[0032] Fig. 27 An anticipation of user presence based on a location of a computing device is shown;
[0033] Fig.28 Techniques for using room-dependent context to improve recognition of ambiguous gestures are shown;
[0034] Fig.29 A technique for improving recognition of ambiguous gestures using the state of an operation being performed at the current time is shown;
[0035] Fig.30 Techniques for distinguishing ambiguous gestures based on foreground operations and background operations being performed at the current time are shown;
[0036] Fig.31 Techniques for using contextual information including past operations and / or future operations of a computing device are shown;
[0037] Fig.32 Techniques for identifying ambiguous gestures based on less destructive operations are shown;
[0038] Fig.33 showing a user performing an ambiguous gesture that the user intended as a first gesture (e.g., a known gesture);
[0039] Fig.34 An example method for online learning based on user input to improve ambiguous gesture recognition is shown;
[0040] Fig.35 An example technique for online learning of new gestures for a radar-enabled computing device is shown;
[0041] Fig.36 An example technique for configuring a context-sensitive master sensor for a computing device is shown;
[0042] Fig.37 An environment is shown in which techniques for detecting user engagement with a device may be implemented;
[0043] Fig.38 An example method for determining the presence of a registered user is shown;
[0044] Fig.39 An example method for radar-based ambiguous gesture determination using contextual information is shown;
[0045] Fig.40 An example method for continuous online learning for radar-based gesture recognition is shown;
[0046] Fig.41 An example method for long-range radar-based gesture detection is shown;
[0047] Fig.42An example method for online learning based on user input is shown;
[0048] Fig.43 An example method for online learning of new gestures for a radar-enabled computing device is shown;
[0049] Fig.44 An example method for sensor capability determination is shown;
[0050] Fig.45 An example method of identifying ambiguous gestures based on less destructive operations is shown; and
[0051] Fig.46 An example method of detecting user engagement is shown. DETAILED DESCRIPTION
[0052] Overview
[0053] Computing devices are increasingly used in everyday life, and users choose to rely on these devices to support a wide variety of tasks. For example, home automation using virtual assistant (VA) technology is becoming increasingly popular as a way to increase residential security, comfort, and convenience. Residents can easily control lighting, climate, entertainment systems, appliances, alarm clocks, etc. using computing devices equipped with VA technology. In order to increase the usability and user satisfaction of these devices, manufacturers aim to provide users with convenient ways to interact with their computing devices to provide efficient and accurate control of the devices when performing these functions. As the functionality of computing devices expands to become useful in increasingly complex scenarios, additional interaction methods are developed to enable users to communicate with their devices.
[0054] One such form of interaction is touchless gesture recognition. This allows the device to be controlled without the user having to make physical contact with the device. Instead, the user can use their entire body or part of their body to perform a gesture (e.g., a hand motion made in the air), and the computing device can recognize this gesture as a specific command, in response to which the computing device performs a specific function associated with the command. This enables the user to control the computing device even when their hands are dirty, such as when cooking, or when they are located a distance away from the computing device. As such, enabling touchless gesture recognition on a computing device can improve the user's ability to interact with the device, and thereby improve user satisfaction.
[0055] While gesture recognition can enable users to more conveniently interact with their devices, the inability of these devices to accurately recognize gestures can frustrate users, which in some cases can cause them to avoid using gestures to control their devices. Specifically, a computing device may receive a gesture performed by a user that was a particular known gesture that was intended by the user. Due to inconsistencies in the performance or recognition of the gesture, the computing system may be unable to determine the gesture from one or more known gestures. Some of these known gestures may result in operations on the computing device that, if performed without the user's intent, may result in changes to the device that are not easily reversible or inconvenient (e.g., terminating an application). Given that these operations may not be easily reversible, a user may be frustrated when the computing device performs one of these harmful operations due to incorrect gesture recognition of a gesture performed by the user.
[0056] Typical gesture recognition devices may not be able to recognize differences in the consequences of different operations caused by different gestures. In contrast, these techniques enable the device to determine the destructiveness of each command associated with a set of known gestures that may be associated with a gesture performed by a user. The destructiveness of each command can represent the consequences of the computing device inaccurately recognizing the gesture as a gesture associated with the command. The consequences of determining a gesture performed by a user as a particular gesture can be incorporated into the decision to recognize the gesture, thereby reducing user frustration caused by incorrectly recognized gestures.
[0057] It should be noted that this is only one example of determination of a less destructive command, and other examples are described below. The present disclosure will now turn to a description of an example operating environment, followed by an example computing device, examples of radar-based gesture detection and recognition, and various techniques for utilizing and improving radar-based gesture recognition.
[0058] Example Environment
[0059] Figure 1An example environment 100 is shown in which a radar-enabled computing device 102 performs the techniques described in this document, such as detecting and distinguishing users and user engagement, detecting and recognizing gestures, and causing execution of commands, as well as improving any of these techniques. A radar-enabled computing device 102 (computing device 102) can be used to cause execution of a task (e.g., turning off lights, lowering music volume, starting an oven, changing a TV channel) by recognizing a gesture associated with a command. A swipe of a user's hand can indicate a command to change a song that is playing, while a push-pull gesture can indicate a command to check the status of a timer in the kitchen. Computing device 102 can allow multiple users 104, each of which can enjoy a customized experience by being distinguished by computing device 102 in some cases. In addition, multiple computing devices (e.g., computing devices 102-X, where X represents an integer value of 1, 2, 3, 4, ..., etc.) can be connected (e.g., wirelessly, such as through connection to one or more wireless networks and / or through direct wireless communication) to create a network such as about Figure 3 and Figure 4 An interconnected network of radar systems described. This network of computing devices can be arranged to detect and recognize gestures being performed by a user 104 (such as a first user 104-1, a second user 104-2, or a different user 104-X, where X represents an integer value of 3, 4, 5, ... etc.), for example, in any one or more rooms of a residence.
[0060] Specifically, these techniques may include: (1) detecting the presence of user 104 within proximity 106 of computing device 102; (2) distinguishing the user from other users to enable a customized experience; and then (3) directing computing device 102, an application associated with computing device 102, or another device to execute a command upon recognizing that a known gesture has been performed. Additionally, or in lieu of (3), computing device 102 may prompt user 104 to begin or continue gesture training based on the user's training history.
[0061] Computing device 102 may transmit radar transmit signals discretely or continuously (e.g., without a "wake-up trigger") over time to detect user presence and / or performance of gestures within proximity 106. Any one or more of these radar transmit signals may be reflected by objects in proximity 106 (e.g., user 104 making movement or stationary objects), thereby generating one or more radar receive signals. Computing device 102 may determine radar signal characteristics (e.g., temporal information or topological information of the object or movement) based on the radar receive signals. If the determined radar signal characteristics are correlated with one or more stored radar signal characteristics, computing device 102 may classify the object or movement. By doing so, computing device 102 may determine that the movement is likely to be a gesture, rather than some other moving or non-moving object. In this example, these techniques compare the radar signal characteristics with one or more stored radar signal characteristics associated with registered or unregistered users or known gestures, and thereby attempt to detect and distinguish users making movement and detect and identify gestures being performed.
[0062] For example, assume that user 104-1 makes a motion with his or her hand. Computing device 102 may detect that this motion is a gesture, rather than a non-gesture movement, based on radar signal characteristics of the radar receive signals, based on one or more radar receive signals reflected by the user's hand. Computing device 102 correlates one or more of the radar signal characteristics with one or more stored radar signal characteristics of a known gesture (e.g., a waving gesture associated with a command to turn on a light) simultaneously with or after detecting the gesture. If the determined radar signal characteristic correlation meets a desired confidence level (e.g., a threshold criterion), the device may determine that a waving gesture was performed and caused the light to turn on.
[0063] In another example, the computing device 102 may detect that the first user is in the neighboring area 106 (eg, Figure 1 The device may determine the presence of the first user 104-1 within the proximity area 106. The device may determine at least one radar signal characteristic (e.g., height, shape, movement) of the first user 104-1 and correlate it with one or more stored radar signal characteristics of "registered users". If the determined radar signal characteristic correlation reaches a desired confidence level, the device may determine that the first user 104-1 is a registered user and that the user is currently located within the proximity area 106.
[0064] In this disclosure, a "registered user" will generally refer to a user 104 that is associated with at least one stored radar signal characteristic and / or has an account or other registration item that is accessible by the computing device 102. An account can be manually set up (e.g., by a registered user or another user of the device) or automatically set up (e.g., when interacting with the device). An account can include or be associated with one or more stored radar signal characteristics, user settings, preferences, gesture training history, user habits, etc., which may or may not be dependent on access to personally identifiable information. Through an account, a registered user can have permissions to modify, store, or access a certain amount of information associated with the computing device (e.g., the user's settings or preferences). These permissions can be granted to registered users, but (by default) are not granted to users who are not registered on the device.
[0065] For some aspects, the account registered by the registered user corresponds to their account in such (It provides "Hey or OK "Voice Assistant Service) or (It provides An account on a cloud-based smart home service platform of a virtual assistant provider of a virtual assistant service (such as a voice assistant service). In such an example, a registered user may be the primary user of an account associated with a cloud-based service platform for a particular residence, sometimes referred to as a primary user, a billing user, a supervisory user, an administrator user, a super user, a master user, etc., depending on the nature of the platform. The role of the primary user is usually assumed by a dad, a mom, another family leader, a "family technology expert," or other designated person. A registered user may alternatively be a secondary user (or tertiary user, etc.) of an account associated with a cloud-based service platform, such as a teenager or non-primary user adult, who enjoys at least some of the rights known to the cloud-based service platform, but typically has a more limited set of permissions or capabilities. However, it should be understood that other types of user registration items that are different from the user registration items of the cloud-based service platform are also within the scope of the present teachings, including but not limited to user accounts established using independent, offline, or off-grid device groups that have their own local account establishment and registered user establishment types.
[0066] More specifically, computing device 102 uses radar system 108 (eg, Figure 1104) to detect the presence of users in the adjacent area 106. When an object (e.g., user 104) is detected within the adjacent area 106, the radar transmit signal may be reflected by the user 104 and modified (e.g., in amplitude, phase, and / or frequency) based on the terrain topography and / or the movement of the user 104. The modified radar transmit signal (e.g., radar receive signal) may be received by the radar system 108 and include information for distinguishing the user 104 from other users, such as, for example, Figure 6 and its accompanying description. The radar system 108 uses the radar received signals to determine the velocity, size, shape, surface smoothness or material of the user 104 because the received signals will be affected by factors such as the object's velocity (e.g., via Doppler), size (see Figure 6 ), shape (see Figure 6 ) etc. The radar system 108 may also determine the distance between the user 104 and the computing device 102 and / or the orientation of the user 104 relative to the computing device 102, such as through time-of-flight analysis.
[0067] Although the neighboring area 106 of the example environment 100 is depicted as a hemisphere, in general, the neighboring area 106 is not limited to the topography shown. The topography of the neighboring area 106 can also be affected by nearby obstacles (e.g., walls, large objects). In general, the computing device 102 can be located in a neighboring area 106 (e.g., a bedroom) in a larger physical area (e.g., a residence, an environment) than the neighboring area 106. In addition, the radar system 108 can detect and detect users and / or gestures outside the neighboring area 106 depicted in the example environment 100. The boundary of the neighboring area 106 corresponds to an accuracy threshold, wherein the user and / or gesture detected within this boundary is more likely to be accurately distinguished or recognized than the user or gesture detected outside this boundary. The example neighboring area 106 is one centimeter to eight meters, depending on the power usage of the radar system 108, the desired confidence, and whether the radar system 108 is configured to detect users, distinguish users, detect gestures, recognize gestures, and / or detect user engagement.
[0068] For example, computing device 102 transmits a first radar transmit signal to neighboring area 106, and then receives a first radar receive signal (e.g., a reflected radar transmit signal) associated with the presence of an object (e.g., first user 104-1). This first radar receive signal includes one or more radar signal characteristics (e.g., radar cross section (RCS) data, motion characteristics, gesture performance, etc.), which can be used to distinguish first user 104-1 from other users 104. Specifically, radar system 108 can compare the first radar receive signal with stored radar signal characteristics of registered users to determine whether first user 104-1 is a registered user who has previously interacted with the device and / or established an account. Example ways of doing this include using Figure 7 and Figure 8 , and corresponding topological distinction 600-1, time distinction 600-2, gesture distinction 600-3, and context distinction 600-4. In this example, the radar signal characteristics of the first radar reception signal are correlated with one or more stored radar signal characteristics of the registered users using the machine learning model 700 to achieve a desired confidence level, and the machine learning model is trained using the stored radar signal characteristics and the radar signal characteristics of the first radar reception signal reflected from one of the users 104 as input.
[0069] Generally speaking, a radar transmission signal may refer to a single (discrete) signal, a burst of signal pulses, or a continuous signal stream transmitted over time from one or more antennas of computing device 102. Figure 8 The radar receive signal may be transmitted at any time in the event of a "wake-up" trigger event (described in more detail below). The radar receive signal may be transmitted by the same transmit antenna, a different antenna of the computing device 102, or a different antenna of the computing device 102 (described in more detail below). Figure 3 and Figure 4 The antenna of another device of the computing system (described in more detail) receives the signal.
[0070] The computing device 102 in the example environment 100 may (but is not required to) forgo “personally identifying” the first user 104-1 (e.g., its private or personally identifiable information) to determine that the detected object is a registered user. For example, the computing device 102 may determine that the first user 104-1 is a registered user without the need for personally identifiable information, which may include legally identifiable information (e.g., legal name). In addition, the computing device 102 may forgo identifying the first user 104-1’s personal device (e.g., a mobile phone, a device equipped with an electronic tag), collecting facial recognition information, or performing speech-to-text of a potentially private conversation to determine that the first user 104-1 is a registered user based on the user’s preferences or settings. Instead of personally identifying the first user 104-1, the computing device 102 may use radar signal characteristics that do not contain personally identifiable information and / or confidential information to “distinguish” the first user 104-1 from another user (e.g., the second user 104-2), as described below with respect to Figure 4 As described.
[0071] Controls may be provided to the user 104 to allow the user 104 to make choices regarding whether and when the techniques described herein may enable the collection of user information (e.g., information about the user's social network, social behavior, social activities, or occupation; photos taken by the user; audio recordings made by the user; the user's preferences; the user's current location, etc.); and whether to send content or communications from the server to the user 104. Additionally, certain data may be processed in one or more ways before it is stored or used to remove personally identifiable information. For example, the user's identity may be processed so that no personally identifiable information can be determined for the user 104, or the user's geographic location may be generalized (e.g., to a city, ZIP code, or state level) when location information is obtained so that the specific location of the user 104 cannot be determined. Thus, the user 104 may control what information is collected about the user 104, how that information is used, and what information is provided to the user 104.
[0072] Gesture training
[0073] Gesture training involves an interactive user experience through which a radar-enabled computing device helps a user learn how to make gesture commands or inputs to the device. Gesture training may, for example, involve the device performing the following steps: (i) conveying information to the user about how to make one or more specific gesture commands or inputs; (ii) suggesting / proposing that the user attempt to make the gesture; (iii) monitoring the user while the attempt is being made; and (iv) providing the user with evaluative feedback about whether the attempt was successful and whether they could try again in a different manner. In one exemplary scenario, the device is a smart home display assistant (e.g., NEST HUB TM ), the smart home display assistant has touchless gesture (e.g., air hand gesture) recognition based on FMCW (frequency modulated continuous wave) radar using radar signals in the 60GHz range. For example, this smart home assistant is able to recognize gestures such as swipe left, swipe right, swipe up, swipe down, air knob turn, and push inward. The smart home assistant can provide gesture training for the left swipe gesture by: (i) showing the user a short animation or video of a left swipe motion while displaying or saying "This is a swipe-left"; (ii) displaying or saying "Now you try"; (iii) performing radar-based monitoring when the user tries the gesture; and if successful (iv) displaying or saying "Awesome, you've got it! (Great, you did it!)" For example, if the distance from left to right of the gesture is too short, the smart home assistant can say "Try again using a longer sweeping motion" and so on. Other gesture training methods may be similar in an overall manner, but expressed in more interesting activities, such as having the user control or direct an on-screen character or other on-screen object in a game-like setting. Other kinds of interactive user experiences for gesture training may be provided without departing from the scope of the present teachings. For the sake of user convenience, encouragement, and continued engagement with the device, it is generally desirable to avoid requiring a one-time single-event gesture training session in which all gesture training is completed at one time unless the user explicitly requests such a single-event gesture training session. Instead, it is generally desirable to suggest and offer smaller modular lessons at appropriate times over time (e.g., a first lesson on swiping left at an appropriate time on the first day, a first lesson on turning an air knob at an appropriate time on a later day, and so on).
[0074] Thus, according to one aspect, before, after, or simultaneously with determining that the first user 104-1 is a registered user (or unregistered but has stored radar signal characteristics), the computing device 102 may prompt the first user 104-1 to start or continue gesture training. For example, the first user 104-1 may be in the middle of training and has completed training for a first gesture (e.g., a left swipe in the aforementioned example of the previous paragraph). The computing device 102 may have stored information about the manner in which the first user 104-1 performed the first gesture during training in a training history. When the first user 104-1 is detected, the computing device 102 may access the training history and prompt the first user 104-1 to continue training for a second gesture (e.g., turning the knob in the air) instead of repeating training for the first gesture. Thus, determining that the first user 104-1 is a registered user or an unregistered person with stored radar signal characteristics may allow the computing device 102 to improve the efficiency of gesture training. The differentiation of the first user 104 - 1 may allow the computing device 102 to activate settings (eg, privacy settings, preferences) for a registered user or an unregistered person, thereby providing a customized experience for the first user 104 - 1 .
[0075] for Figure 1 , assume that second user 104-2 joins first user 104-1 on a couch in adjacent area 106. Computing device 102 may use radar system 108 to send a second radar transmit signal to detect the presence of another object (e.g., second user 104-2). Radar system 108 may then compare the second radar receive signal with stored radar signal characteristics of registered users to determine whether second user 104-2 is another registered user. In this example, the second radar receive signal is not found to be correlated with one or more stored radar signal characteristics of another registered user. Therefore, second user 104-2 is distinguished as an "unregistered person."
[0076] In the present disclosure, "unregistered personnel" will generally refer to users 104 who are not registered on the device and are therefore not associated with one or more accounts. Unlike registered users, unregistered personnel may not have the authority to modify, store or access information of the computing device. For example, a new visitor (e.g., a guest who is not associated with one or more accounts of the device) can be considered an unregistered person. This new visitor may not have interacted with the computing device 102 before and is therefore not associated with one or more stored radar signal characteristics, or the visitor may have done so and is associated with the stored radar signal characteristics, but does not have an account or other special permissions. Therefore, previous visitors (e.g., nannies, housekeepers, gardeners) may be associated with one or more stored radar signal characteristics, but do not have an account for the device. According to one or more aspects, the computing device 102 can still store the radar signal characteristics of this user to improve their user experience. However, the device can prevent previous visitors (unregistered personnel) from exercising the permissions enjoyed by registered users, such as modifying, accessing or storing information of the device.
[0077] After determining that the second user 104-2 is an unregistered person, the computing device 102 assigns an unregistered user identification (e.g., a simulated identity, a pseudo identity) to the unregistered person, which can be associated with one or more radar signal characteristics of the second radar reception signal, such as a unique random number used to later identify the unregistered user. The unregistered user identification can be stored to enable the second user 104-2 to be distinguished at a future time. Specifically, the unregistered user identification can be used to associate future received radar reception signals with one or more associated radar signal characteristics of the unregistered person. The computing device 102 can also prompt the second user 104-2 to register on the computing device 102.
[0078] The computing device 102 may or may not require personally identifiable information of the second user 104-2 to determine that the other object is an unregistered person. After having distinguished the second user 104-2 from the first user 104-1, the computing device 102 may determine that the privacy settings of the first user 104-1 need to be adapted (e.g., modified, restricted) to ensure that the first user's information remains private. For example, the first user 104-1 may want the device to avoid broadcasting calendar reminders (e.g., doctor's appointments) when the second user 104-2 is present. Additionally, the computing device 102 may prompt the second user 104-2 to start gesture training, which may be recorded in another training history of the second user 104-2 (e.g., associated with an unregistered user identification).
[0079] At a later time (not depicted), computing device 102 may again use radar system 108 to send a third radar transmit signal to detect whether user 104 is within proximity 106. If user 104 (e.g., first user 104-1, second user 104-2) is present at this time, the third radar transmit signal may be reflected by user 104, and computing device 102 may receive a third radar receive signal, which includes one or more radar signal characteristics. Radar system 108 may compare these radar signal characteristics with, for example, stored radar signal characteristics of first user 104-1 (a registered user) and second user 104-2 (an unregistered person associated with an unregistered user identifier) to determine whether first user 104-1 or second user 104-2 is present. Based on this determination, computing device 102 may customize settings and training prompts accordingly.
[0080] In the example, radar system 108 uses the third radar reception signal to determine that first user 104-1 (a registered user) is again present in proximity area 106 based on the stored radar signal characteristics associated with first user 104-1. Computing device 102 may then prompt first user 104-1 to complete their gesture training and / or activate their user settings based on their training history. Alternatively, if radar system 108 determines that second user 104-2 (an unregistered person) is again present in proximity area 106, computing device 102 may prompt second user 104-2 to continue their gesture training and / or activate predetermined user settings. Figure 2 Computing device 102 and radar system 108 are further described.
[0081] Example computing device
[0082] Figure 2 An example implementation 200 of a radar system 108 as part of a computing device 102 is shown. The computing device 102 is shown with various non-limiting example devices 202, including a home automation and control system 202-1, a smart display 202-2 associated with the home automation and control system, a desktop computer 202-3, a tablet computer 202-4, a laptop computer 202-5, a television 202-6, a computing watch 202-7, computing glasses 202-8, a gaming system 202-9, a microwave oven 202-10, a smart thermostat interface 202-11, and a car with computing capabilities 202-12. Other devices may also be used, such as security cameras, baby monitors, A router, a drone, a trackpad, a drawing tablet, a netbook, an e-reader, other forms of home automation and control systems, a wall display, a virtual reality headset, another vehicle (e.g., an electric bicycle or airplane), and other home appliances, to name a few examples. It should be noted that the computing device 102 can be wearable, non-wearable but mobile, or relatively non-mobile (e.g., a desktop and an appliance), all without departing from the scope of the present teachings.
[0083] The computing device 102 may include one or more processors 204 and one or more computer readable media (CRM) 206, which may include memory media and storage media. Applications and / or operating systems (not shown) embodied as computer readable instructions on the CRM 206 may be executed by the processor 204 to provide some of the functionality described herein. The CRM 206 may also include a radar-based application 208 that uses data generated by the radar system 108 to perform functions such as gesture-based control, human vital sign notification, collision avoidance for autonomous driving, and the like. For example, the radar system 108 may recognize a gesture performed by the user 104 that indicates a command to turn off the lights in the room. This command data may be used by the radar-based application 208 to send a control signal (e.g., a trigger) to turn off the lights in the room.
[0084] The computing device 102 may also include a network interface 210 for communicating data via a wired, wireless, or optical network. For an interconnected system of multiple computing devices 102-X, each computing device 102 may communicate with another computing device 102 via the network interface 210. For example, the network interface 210 may communicate data via a local area network (LAN), a wireless local area network (WLAN), a personal area network (PAN), a wide area network (WAN), an intranet, the Internet, a peer-to-peer network, a point-to-point network, a mesh network, etc. The multiple computing devices 102-X may use the following information about the interconnected system: Figure 3 The described communication networks are in communication with each other. The computing device 102 may also include a display.
[0085] Radar system 108 may be used as a stand-alone radar system, or may be used with or embedded within many different computing devices or peripherals, such as in a control panel for controlling home appliances and systems, in an automobile to control internal functions (e.g., volume, cruise control, or even sedan steering), or as an accessory to a laptop computer to control computing applications on the laptop.
[0086] Radar system 108 may include communication interface 212 to transmit radar data (e.g., radar signal characteristics) to a remote device, but this communication interface may not be used when radar system 108 is integrated within computing device 102. In general, radar data provided by communication interface 212 may be in a format that can be used to detect, distinguish, and / or identify a user, user engagement, or gesture, such as radar signal characteristics (e.g., values of a frame corresponding to a complex range-Doppler map, see Figure 8 and Figure 12 to Figure 14 ) or a determination of detection or recognition by computing device 102. Communication interface 212 may also or alternatively communicate with a remote instance of radar-based application 208, such as a command associated with the recognized gesture or an identification of the recognized gesture (e.g., indicating to radar-based application 208 on a remote computing device that a push or pull gesture has been performed).
[0087] The radar system 108 may also include at least one antenna 214 for transmitting and / or receiving radar signals. In some cases, the radar system 108 may include multiple antennas 214, which are implemented as antenna elements of an antenna array. The antenna array may include at least one transmitting antenna element and at least one receiving antenna element. In some scenarios, the antenna array may include multiple transmitting antenna elements to implement a multiple input multiple output (MIMO) radar capable of transmitting multiple different waveforms (e.g., different waveforms for each transmitting antenna element) at a given time. For implementations including three or more receiving antenna elements, the receiving antenna elements may be positioned in a one-dimensional shape (e.g., a line) or a two-dimensional shape (e.g., a triangle, a rectangle, or an L-shape). A one-dimensional shape may enable the radar system 108 to measure an angular dimension (e.g., an azimuth or an elevation), while a two-dimensional shape may enable the measurement of two angular dimensions (e.g., both an azimuth and an elevation). Each antenna 214 may alternatively be configured as a transducer or a transceiver. In addition, any one or more antennas 214 may be circularly polarized, horizontally polarized, or vertically polarized.
[0088] Using antenna arrays, the radar system 108 can form a steerable or non-steerable, wide or narrow beam (e.g., one degree to 45 degrees, 15 degrees to 90 degrees) or a shaped beam (e.g., hemispherical, cubic, fan-shaped, conical, or cylindrical). The one or more transmitting antenna elements may have a non-steerable omnidirectional radiation pattern, or may be capable of producing a wide steerable beam. Any of these techniques can enable the radar system 108 to radar illuminate a large volume of space. In order to achieve target angular accuracy and angular resolution, the receiving antenna elements can be used to generate thousands of narrow steerable beams (e.g., 2000 beams, 4000 beams, or 6000 beams) through digital beamforming. In this way, the radar system 108 can efficiently monitor users and gestures in the environment.
[0089] The radar system 108 may also include at least one analog circuit 216, which includes circuit systems and logic for transmitting and receiving radar signals using the at least one antenna 214. The components of the analog circuit 216 may include amplifiers, mixers, switches, analog-to-digital converters, filters, etc., for conditioning radar signals. The analog circuit 216 may also include logic for performing in-phase / quadrature (I / Q) operations such as modulation or demodulation. Various modulations may be used to generate radar signals, including linear frequency modulation, triangular frequency modulation, step frequency modulation, or phase frequency modulation. The analog circuit 216 may be configured to support continuous wave or pulse radar operations.
[0090] The analog circuit 216 may generate a radar signal (e.g., a radar transmit signal) in a spectrum (e.g., a frequency range) including frequencies between 1 gigahertz (GHz) and 400 GHz, between 1 GHz and 24 GHz, between 2 GHz and 6 GHz, between 4 GHz and 100 GHz, or between 57 GHz and 63 GHz. In some cases, the spectrum may be divided into a plurality of sub-spectra having similar or different bandwidths. Example bandwidths may be in the order of 500 megahertz (MHz), one GHz, two GHz, etc. Different sub-spectra may include, for example, frequencies between approximately 57 GHz and 59 GHz, between 59 GHz and 61 GHz, or between 61 GHz and 63 GHz. Although the example sub-spectra described above are continuous, other sub-spectra may not be continuous. In order to achieve coherence, multiple radar signals may be generated by the analog circuit 216 using multiple sub-spectra (continuous or discontinuous) having the same bandwidth, which may be transmitted simultaneously or separated in time. In some scenarios, a single radar signal may be transmitted using multiple continuous sub-rate spectra, thereby enabling the radar signal to have a wide bandwidth.
[0091] The radar system 108 may also include one or more system processors 218 and system media 220 (e.g., one or more computer-readable storage media). For example, the system processor 218 may be implemented within the analog circuit 216 as a digital signal processor or a low-power processor (or both). The system processor 218 may execute computer-readable instructions stored within the system media 220. Example digital operations performed by the system processor 218 may include fast Fourier transforms (FFTs), filtering, modulation or demodulation, digital signal generation, digital beamforming, etc.
[0092] System medium 220 may optionally include user module 222 and gesture module 224, which may be implemented using hardware, software, firmware, or a combination thereof. User module 222 and gesture module 224 may enable radar system 108 to process radar receive signals (e.g., electrical signals received at analog circuit 216) to detect the presence of user 104 and distinguish the user and detect and recognize gestures, as well as other functions such as object (non-user) detection and detection of user engagement.
[0093] The user module 222 and the gesture module 224 may include one or more machine learning algorithms and / or machine learning models, such as an artificial neural network (referred to herein as a neural network), to improve user differentiation and gesture recognition accordingly. A neural network may include a set of connected nodes (e.g., neurons or perceptrons) organized into one or more layers. As an example, the user module 222 and the gesture module 224 may include a deep neural network including an input layer, an output layer, and a plurality of hidden layers between the input layer and the output layer. The nodes of the deep neural network may be partially connected or fully connected between the layers.
[0094] In some cases, the deep neural network may be a recurrent deep neural network (e.g., a long short-term (LSTM) recurrent deep neural network), in which the connections between nodes form a cycle to retain information from a previous part of the input data sequence for use by a subsequent part of the input data sequence. In other cases, the deep neural network may be a feedforward deep neural network, in which the connections between nodes do not form a cycle. Figure 7 and Figure 8 Describe an example deep neural network. The user module 222 and the gesture module 224 may also include models capable of performing clustering (e.g., trained using unsupervised learning), anomaly detection, or regression, such as a single linear regression model, a multivariate linear regression model, a logistic regression model, a stepwise regression model, a multivariate adaptive regression spline, a locally estimated scatter point smoothing model, and the like.
[0095] In general, the machine learning architecture may be customized based on available power, available memory, or computing power. For user modules 222, the machine learning architecture may also be customized based on the number of radar signal characteristics that radar system 108 is designed to recognize. For gesture modules 224, the machine learning architecture may additionally be customized based on the number of gestures and / or various versions of gestures that radar system 108 is designed to recognize.
[0096] Computing device 102 may optionally (not depicted) include at least one additional sensor (other than antenna 214) to improve the fidelity of user module 222 and / or gesture module 224. In some cases, for example, user module 222 may detect the presence of user 104 with low confidence (e.g., the amount of confidence and / or accuracy is below a threshold). Such detection may occur, for example, when user 104 is far away from computing device 102 or when a large object (e.g., furniture) obscures user 104. In order to increase the accuracy of user detection and differentiation and / or gesture detection and recognition, computing device 102 may use one or more additional sensors (e.g., about Fig.19 The sensors may be passive, active, remote, and / or touch-based. Example sensors (some of which may sense in more than one of passive, active, remote, and touch-based ways) include microphones, ultrasonic sensors, ambient light sensors, cameras, health and / or biometric sensors, barometers, inertial measurement units (IMUs) and / or accelerometers, gyroscopes, magnetic sensors (e.g., magnetometers or Hall effect), proximity sensors, pressure sensors, touch sensors, thermostats / temperature sensors, optical sensors, and the like.
[0097] The user module 222 may also use context information to differentiate between the users 104 (eg, the first user 104-1 and the second user 104-2). Such context information may also improve understanding of, for example, Fig.24 Interpretation of ambiguous gestures (e.g., gestures that cannot be recognized with a desired confidence level) of those described. For example, each user 104 may perform their own version of a known gesture, typically based on their personality, physique, abilities / disabilities, mood, etc. Such contextual information may be accessed by the gesture module 224 to improve gesture recognition.
[0098] In the example, the first user 104-1 performs a gesture, and the gesture module 224 determines that the first user 104-1 accidentally performed an ambiguous gesture that is not correlated to a known gesture (not reaching the desired confidence level). However, if the user module 222 determines that the first user 104-1 (rather than the second user 104-2 or another user) performed the ambiguous gesture, the gesture module 224 may additionally access contextual information about stored radar signal characteristics of previous gesture performances by the user. Using this contextual information, the gesture module 224 may be able to correlate the ambiguous gesture with one or more stored radar signal characteristics associated with the distinguished users, thereby better identifying the ambiguous gesture as a known gesture. In this way, the user module 222 of the computing device 102 can improve the fidelity of gesture recognition.
[0099] Example computing system
[0100] Figure 3 An example environment 300 is shown in which multiple computing devices 102-1 and 102-2 are connected via a communication network 302 to form a computing system. The example environment 300 depicts a residence having a first room 304-1 (living room) and a second room 304-2 (kitchen). The first room 304-1 is equipped with a first computing device 102-1 including a first radar system 108-1, and the second room 304-2 is equipped with a second computing device 102-2 including a second radar system 108-2. In this example, the first room 304-1 is separate from the second room 304-2, but connected by a door in the residence. The first computing device 102-1 in the first room 304-1 can detect users 104 and gestures within the first adjacent area 106-1, and the second computing device 102-2 in the second room 304-2 can detect users 104 and gestures within the second adjacent area 106-2.
[0101] The residence of the example environment 300 may not be limited to the arrangement and number of computing devices 102 shown. In general, an environment (e.g., a residence, a building, a workplace, a car, an airplane, a public space) may include one or more computing devices 102 distributed across one or more different areas (e.g., rooms 304). For example, a room 304 may contain two or more computing devices 102 that are positioned close to or far away from each other (or the radar systems 108 associated with a single computing device 102). Although the first neighboring area 106-1 depicted in the example environment 300 does not spatially overlap with the second neighboring area 106-2, and therefore the computing devices in each area cannot sense the radar reception signals of the other area, in general, the neighboring areas 106 may also be positioned to partially overlap. Although Figure 3 The environment depicted in is a residence, but in general, the environment may include any private or public indoor and / or outdoor space, such as a library, an office, a workplace, a factory, a garden, a restaurant, a terrace, an airplane, or a car.
[0102] For environments with two or more computing devices 102, these devices can communicate with each other through one or more communication networks 302. The communication network 302 can be a LAN, a WAN, a mobile or cellular communication network such as a 4G or 5G network, an extranet, an intranet, the Internet, In some examples, computing device 102 may use technologies such as near field communication (NFC), radio frequency identification (RFID), Waiting for short-range communication to communicate.
[0103] In addition, the computing system may include one or more memories separate from or integrated into one or more of the computing devices 102-1 and 102-2. In one example, the first computing device 102-1 and the second computing device 102-2 may include a first memory and a second memory, respectively, wherein the contents of each memory are shared between devices using the communication network 302. In another example, the memory may be separate from the first computing device 102-1 and the second computing device 102-2 (e.g., cloud storage), but both devices may access the memory. The memory may be used to store, for example, radar signal characteristics of registered users, user preferences, security settings, training history, unregistered user identification, and radar signal characteristics.
[0104] In the example, the first computing device 102-1 can use the first network interface 210-1 (see Figure 2 ) is connected to a communication network 302 to exchange information with a second computing device 102-2. Using this communication network 302, the computing devices 102 can exchange stored information about one or more users 104, which can include radar signal characteristics, training history, user settings, etc. In addition, the computing devices 102 can exchange information about ongoing operations (e.g., timers, music being played) to maintain continuity of operations and / or information about operations across various rooms 304. These operations can be performed simultaneously or independently by one or more computing devices 102 based on, for example, detecting the presence of a user in a room 304. Each computing device 102 can also use the radar system 108 to associate (and store in memory) commonly detected users and commands associated with the device location. About Figure 4 Radar system 108 is further described.
[0105] Radar-enabled user detection and differentiation
[0106] Figure 4An example environment 400 is shown in which radar systems 108 are used by computing devices 102 to detect the presence of users 104 and distinguish the users. Example environment 400 depicts a first computing device 102-1 having a first radar system 108-1 and a second computing device 102-2 having a second radar system 108-2. The first radar system 108-1 and the second radar system 108-2 may transmit one or more radar transmit signals 402 (e.g., 402-Y, where Y represents an integer value of 1, 2, 3, etc.) to detect users (and / or gestures) in a first neighboring area 106-1 and a second neighboring area 106-2, respectively. Note that for simplicity, each of these areas is shown as a cone, but its contour is determined by the amplitude and quality of the radar field in which the radar receive signal can be received by the corresponding radar system 108 in that area. Each radar transmit signal 402-Y may be referred to as a composite radar transmit signal 402-Y, which represents the composite radar transmit signal transmitted from the corresponding antenna 214 (see FIG. 2 ) at a given time. Figure 2 ). Using radar transmit signal 402-1, first radar system 108-1 may illuminate an object (e.g., user 104) entering first proximity zone 106-1 with a wide 150° radar pulse beam (e.g., one or more radar transmit signals) operating at a frequency of 1 gigahertz to 100 gigahertz (GHz; e.g., 60 GHz). Although reference may be made to radar transmit signal 402 in the present disclosure, it should be understood that one or more radar transmit signals 402 may be transmitted and / or include one or more radar pulses within a period of time.
[0107] When encountering the user 104, a portion of the energy associated with the radar transmit signal 402-Y may be reflected back toward the first radar system 108-1 and / or the second radar system 108-2 in the form of one or more radar receive signals 404-Z (where Z may represent an integer value of 1, 2, 3, ...). Each radar receive signal 404-Z may be referred to as a composite radar receive signal 404-Z that represents a superposition of multiple reflections of the radar transmit signal 402-Y at the one or more antennas 214 at a given time. In the example environment 400, two radar receive signals 404-1 and 404-2 are depicted as being received by the radar systems 108-1 and 108-2, respectively. The radar receive signals 404-1 and 404-2 may be reflected from one or more discrete dynamic scattering centers of the user 104. Each radar receive signal 404 may represent a modified version of its corresponding radar transmit signal 402, wherein the amplitude, phase, and / or frequency are modified by the one or more dynamic scattering centers. These radar reception signals 404 may allow one or more of radar systems 108-1 and 108-2 to distinguish users 104 and / or recognize gestures, for example, using radial distance, geometry (e.g., size, shape, height), orientation, surface texture, material composition, etc. For additional details on how this is performed, see at least the Figures 7 to 18 .
[0108] Although the first computing device 102-1 and the second computing device 102-2 of the example environment 400 can independently detect and distinguish the user 104 and / or recognize the gesture performance, they can also work together (e.g., dependently, collaboratively). This can be particularly useful in situations where, for example, each device is unable to distinguish the user 104 and / or recognize the gesture performance to a desired confidence level alone. In this case, the two devices can exchange radar signal characteristics determined from radar receive signals received from each device. In the example, the first computing device 102-1 can detect an ambiguous user associated with the first radar signal characteristic within the first adjacent area 106-1. Simultaneously or at a separate time, the second computing device 102-2 can also detect this ambiguous user and determine a second radar signal characteristic associated with the presence of the user. If the first radar signal characteristic and the second radar signal characteristic are not sufficient alone to distinguish the ambiguous user to a desired level of accuracy, the devices can work together and / or exchange information to achieve user differentiation.
[0109] In the first scenario, first computing device 102-1 may access the second radar signal characteristic and then compare the first radar signal characteristic and the second radar signal characteristic with one or more stored radar signal characteristics to distinguish the ambiguous user. In the second scenario, second computing device 102-2 may access the first radar signal characteristic and then compare the first radar signal characteristic and the second radar signal characteristic with the one or more stored radar signal characteristics to distinguish the ambiguous user (e.g., by using Figure 7 or Figure 8 In a third scenario, first computing device 102-1 and second computing device 102-2 may work collaboratively, cooperatively, consistently, etc. to distinguish ambiguous users. This technology may also be applied to ambiguous gesture commands. Figure 5 1 and 2 will now describe radar system 108 in greater detail.
[0110] Figure 5 An example implementation 500 is shown that includes an antenna 214, analog circuitry 216, and a system processor 218 of a radar system 108. In the depicted configuration, the analog circuitry 216 may be coupled between the antenna 214 and the system processor 218 to implement techniques for both user detection and differentiation and gesture detection and recognition. The analog circuitry 216 may include a transmitter 502 equipped with a waveform generator 504 and a receiver 506 including at least one receive channel 508. The waveform generator 504 and the receive channel 508 may each be coupled between the antenna 214 and the system processor 218.
[0111] Although one antenna 214 is depicted in the example implementation 500, in general, the radar system 108 may include one or more antennas to form an antenna array. When utilizing an antenna array, the waveform generator 504 may generate similar or different waveforms for each antenna 214 to transmit into the adjacent area 106. Additionally, although one receive channel 508 is depicted in the example implementation 500, in general, the radar system 108 may include one or more receive channels. Each receive channel 508 may be configured to accept a single or multiple versions of the radar receive signal 404-Z at any given time.
[0112] During operation, the transmitter 502 may communicate an electrical signal to the antenna 214, which may transmit one or more radar transmit signals 402-Y to detect user presence and / or gestures in the adjacent area 106. Specifically, the waveform generator 504 may generate an electrical signal having a specified waveform (e.g., a specified amplitude, phase, frequency). The waveform generator 504 may additionally communicate information about the electrical signal to the system processor 218 for digital signal processing. If the radar transmit signal 402-Y interacts with the user 104, the radar system 108 may receive a radar receive signal 404-Z on the receive channel 508. The radar receive signal 404-Z (or multiple versions thereof) may be sent to the system processor 218 to enable user detection (using the user module 222 of the system medium 220) and / or gesture detection (using the gesture module 224). The user module 222 may determine whether the user 104 is located within the adjacent area 106, and then distinguish the user 104 from other users. The user 104 may be distinguished based on the one or more radar receive signals 404-Z, such as regarding Figure 6 Further described.
[0113] Figure 6 Example implementations 600-1 to 600-4 are shown in which the user module 222 can distinguish between users 104. The user module 222 can use one or more radar received signals 404 in part to distinguish, for example, the first user 104-1 from the second user 104-2 with or without personally identifying the first user 104-1 or the second user 104-2. By distinguishing between the users 104, the user module 222 can enable the computing device 102 to provide a customized experience for each user 104, such as invoking training history, preferences, privacy settings, etc. In this way, the computing device 102 can improve some virtual assistant (VA)-equipped devices by meeting each user's privacy and / or functionality expectations.
[0114] To distinguish users 104, user module 222 may analyze radar receive signal 404 to determine (1) topological distinction, (2) temporal distinction, (3) gesture distinction, and / or (4) contextual distinction of user 104. In the present disclosure, topological, temporal, gesture, and contextual distinctions may be determined based in part on one or more radar signal characteristics and, in some cases, non-radar data from non-radar sensors. User module 222 is not limited to Figure 6 The four distinguishing categories depicted in , and may include other categories not shown. In addition, these four distinguishing categories are shown as example categories, and may be combined and / or modified to include subcategories that implement the techniques described herein. Figure 6 The described techniques may also be applied to the gesture module 224 as described herein (e.g., in relation to Fig.25 ).
[0115] In the example implementation 600-1, the user module 222 may use topological information in part to distinguish users 104. This topological information may include radar cross section (RCS) data, such as the height, shape, or size of the user 104. For example, the first user 104-1 (e.g., father) may be significantly larger than the second user 104-2 (e.g., child). When the father and child enter the adjacent area 106, the radar system 108 may obtain a radar receive signal 404 indicating the presence of each user. These radar receive signals 404 may include, in part, radar signal characteristics associated with topological information, which indicates the height, shape, or size of each user. In this example, the father's radar signal characteristics may be different from the child's radar signal characteristics. Then, the user module 222 may compare the radar signal characteristics of each user with the stored radar signal characteristics of the registered user to determine whether the father and child are registered users (or unregistered persons with associated radar signal characteristics).
[0116] In this example, user module 222 can determine that first user 104-1 (father) is a registered user. Specifically, user module 222 can correlate the father's stored (e.g., saved to a memory shared by multiple computing devices 102-X) radar signal characteristics with the one or more radar reception signals 404 to determine that he is a registered user. Upon determining that first user 104-1 is a registered user, computing device 102 can activate the father's settings, prompt the father to continue gesture training (based on the father's training history), etc.
[0117] The user module 222 may also determine that the second user 104-2 (the child) is an unregistered person who does not have an account on the computing device 102. Specifically, the user module 222 may compare the stored radar signal characteristics of the registered users with the one or more radar received signals 404, which include the radar signal characteristics of the second user 104-2. Upon determining that the topological information associated with the radar signal characteristics of the second user 104-2 is not correlated with the one or more stored radar signal characteristics of the registered users (e.g., assuming a certain level of fidelity), the user module 222 determines that the child is an unregistered person. The radar system 108 may assign an unregistered user identifier to the child, which includes the radar signal characteristics of the child so that the child can be distinguished from other users (e.g., the father) using topological information at a future time. This unregistered user identifier may include information such as Figures 12 to 15The data shown or information determined from the data, such as the child's height range, movement data, or door, etc. The computing device 102 can also prompt the child to begin gesture training and / or implement predetermined settings (e.g., standard preferences, settings programmed by the owner of the computing device 102).
[0118] In general, the stored radar signal characteristics of registered users can be collected and saved once or more. For example, the radar system 108 can store the radar signal characteristics each time the user 104 interacts with the computing device 102 to improve user identification. The radar system 108 can also continuously store the radar signal characteristics over time to improve user and / or gesture detection. The stored radar signal characteristics of the user 104 may include topological, time, gesture and / or context information inferred from one or more radar received signals 404 associated with the user 104. Additionally, the radar system 108 can utilize one or more models used by the user module 222 to distinguish them based on the radar signal characteristics corresponding to each user 104. The one or more models may include machine learning (ML) models, predicate logic, hysteresis logic, etc. to improve user differentiation.
[0119] In the example implementation 600-2, the user module 222 may use temporal information in part to distinguish users 104. Unlike conventional radar detectors that may require high spatial resolution, the radar system 108 of the present disclosure may rely more on temporal resolution (rather than spatial resolution) to detect and distinguish users 104 and / or recognize gesture performance. In this way, the radar system 108 can distinguish users 104 moving into the adjacent area 106 by receiving motion characteristics (e.g., the unique way in which the user 104 typically moves). The motion characteristics may include gait (depicted in the drawing of the example implementation 600-2), limb movement (e.g., corresponding arm movement), weight distribution, breathing characteristics, unique habits, etc. The motion characteristics of a user may include a limp, an energetic step, pigeon-toed steps, knock-kneed, bowed legs, etc. Using this information, the user module 222 may be able to detect the user's motion (e.g., the movement of his hands) without identifying details that may be considered private (e.g., facial features). The detection of motion signatures by radar system 108 may enable a user to maintain greater anonymity than using, for example, devices that perform facial recognition or speech-to-text technology.
[0120] Although described in the context of distinguishing a user from one or more other users, or distinguishing a user as a specific registered user, these techniques can also be used in conjunction with detecting a user. Thus, the radar signal characteristics of the radar reception signal reflected from the user can be used to detect both the presence of a user (e.g., any person) and the presence of the detected user as a specific user. Thus, these operations of detecting the presence of a user and distinguishing a user can be performed separately or as one operation.
[0121] In example implementation 600-3, user module 222 may also use, in part, gesture performance information to distinguish users 104. User module 222 may communicate with gesture module 224 while utilizing radar signal characteristics associated with gesture performance. A user may perform a gesture (e.g., a push-pull gesture) in a unique manner (or a partially unique manner) that may help distinguish users 104 (while still sufficiently conforming to the push-pull gesture paradigm, of course, so as to be recognizable as a push-pull gesture). For example, a push-pull gesture may include a user's hand pushing in one direction, followed by their hand pulling in the opposite direction. Although radar system 108 may expect push and pull motions to be complementary (e.g., equal extent of motion, equal speed), user 104 may perform motions that are different from those expected. Each user may perform this gesture in a unique manner that is recorded on the device for user differentiation and gesture recognition.
[0122] As depicted in example implementation 600-3, a first user 104-1 (e.g., a father) may perform a push-pull gesture differently than a second user 104-2 (e.g., a child). For example, the first user 104-1 may push his hand to a first extent (e.g., a distance) at a first rate, but pull his hand back to a second extent at a second rate. The second extent may include a shorter distance than the first extent, and the second rate may be much slower than the first rate. The radar system 108 may be configured to recognize this unique or partially unique push-pull gesture based on the first user's training history (if available).
[0123] When distinguishing the first user 104-1, the radar system 108 can receive one or more radar reception signals 404, which include radar signal characteristics of the first user 104-1 associated with its performance of the push-pull gesture. The user module 222 can compare these radar signal characteristics with the stored radar signal characteristics of the registered users to determine whether there is a correlation (see Figures 7 to 17 and Fig.25and accompanying description for a way to do this). If there is a correlation (e.g., assuming a certain level of fidelity), the user module 222 can determine that the first user 104-1 is a registered user (the father) based on the performance of the push-pull gesture. Similar to the teachings above regarding the example implementation 600-1, the computing device 102 can then activate the father's settings, prompt the father to continue gesture training, etc. In this example, it is assumed that the father has performed the push-pull gesture at least once in the past, and the radar signal characteristics of that performance are recorded in the father's training history, thereby partially achieving the distinction between the father's presence and the presence of other users.
[0124] As also depicted in example implementation 600-3, second user 104-2 (child) may attempt to perform a push-pull gesture. Second user 104-2 may push its hand to a first degree at a first rate, but pull its hand back to a third greater degree and a third rate. Specifically, radar system 108 may receive one or more radar receive signals 404, which include radar signal characteristics of second user 104-2 associated with the execution of this push-pull gesture. User module 222 may compare these radar signal characteristics with the stored radar signal characteristics of registered users to determine whether there is a correlation. Similar to the teachings of example implementation 600-1 above, user module 222 may determine that the push-pull gesture of second user 104-2 is not related to the stored radar signal characteristics of registered users. Therefore, user module 222 may determine that the child is an unregistered person and assign an unregistered user identification to the child. However, the radar signal characteristics associated with the child's push-pull gesture may be included in the unregistered user identification to achieve future distinction.
[0125] The user module 222 may also use context information in part to distinguish users 104. The context information may be determined by the user module 222 using, for example, the antenna 214, another sensor of the computing device 102, data stored on a memory (e.g., user habits), local information (e.g., time, relative position), etc. In the example implementation 600-4, the user module 222 may use the local time as a background to achieve differentiation of a particular user. If the user 104 (e.g., father) consistently sits on the sofa in the living room at 5:30pm every day, the user module 222 may note this habit to improve user differentiation. Whenever the user 104 is detected on the sofa at 5:30pm, the radar system 108 may use this context information in part to distinguish the user 104 as the father. In another example, if the computing device 102 is located in a child's room, the radar system 108 may determine over time that the child is the most common user in the room. The context information may be used to achieve user differentiation. Similarly, if computing device 102 is located in a shared space (e.g., a backyard, an entryway), radar system 108 may determine over time that unregistered persons are common in that area (e.g., guests, nannies, housekeepers, gardeners, contractors, freelance help). It should be understood that while the scope of the present teachings is not necessarily limited to camera-less environments, and thus for some embodiments, cameras and facial recognition may be used to enhance contextual information, one advantageous feature provided by the camera-less embodiments described herein is that desired contextual information may indeed be derived without the use of cameras, as the presence of cameras in residential environments, particularly in more sensitive areas of a residence, may induce a sense of unease and invasion of privacy.
[0126] The context information collected by the user module 222 can be used alone to distinguish users 104, or in combination with topological information, time information, and / or gesture information. In general, the user module 222 can use any one or more of the depicted distinction categories in any combination at any time to distinguish users 104. For example, the radar system 108 can collect topological and time information about users 104 who have entered the adjacent area 106 but lack gesture and context information. In this case, the user module 222 can distinguish users 104 based on the analysis of the topological and time information. In another case, the radar system 108 can collect topological and time information, but determine that the information is not sufficient to correctly distinguish users 104 (e.g., to achieve a desired confidence level). If context information is available, the radar system 108 can use the context to distinguish users 104 (similar to the example implementation 600-4). Figure 6 Any one or more of the categories depicted in may take precedence over another category.
[0127] The user module 222 can utilize one or more logic systems (e.g., including predicate logic, hysteresis logic, etc.) to improve user differentiation. The logic system can be used to prioritize certain user differentiation techniques over other differentiation techniques (e.g., favoring time differentiation rather than contextual information), add weights (e.g., confidence) to certain results when relying on two or more differentiation categories, etc. For example, the user module 222 can determine with low confidence that the first user 104-1 may be a registered user. The logic system can determine that the low confidence is below an allowed threshold standard (e.g., a limit), and instead prompt the radar system 108 to send a second radar transmission signal 402-2 (or a set of signals transmitted within a certain time period) to re-detect the adjacent area 106. The user module 222 can also include one or more machine learning models to improve user differentiation, such as regarding Figure 7 Further described.
[0128] Figure 7 An example implementation of a machine learning model 700 for distinguishing between users 104 and / or recognizing gestures is shown. The machine learning model 700 can perform classification, wherein the machine learning model 700 provides a numerical value for each of one or more categories that describes the extent to which the input data is believed to be classified into the corresponding category. In some instances, the numerical value provided by the machine learning model 700 can be referred to as a probability or "confidence score" that indicates the corresponding confidence in classifying the input into the corresponding category. In some implementations, the confidence score can be compared to one or more threshold criteria to present discrete category predictions. In some implementations, only a specific number of categories (e.g., one) with relatively maximum confidence scores can be selected to present discrete category predictions.
[0129] In an example implementation, the machine learning model 700 can provide probabilistic classification. For example, given a sample input, the machine learning model 700 can predict a probability distribution of a set of categories. Therefore, rather than outputting only the most likely category to which the sample input should belong, the machine learning model 700 can output, for each category, the probability that the sample input belongs to that category. In some implementations, the sum of the probability distributions of all possible categories can be one.
[0130] The machine learning model 700 can be trained using supervised learning techniques. For example, the machine learning model 700 can be trained on a training data set that includes training examples that are labeled as belonging to (or not belonging to) one or more categories. At least a portion of the training can be performed before the user purchases the computing device 102 to initialize the machine learning model 700. This type of training is called offline training. During offline training, the training data set is not necessarily associated with the user. In some implementations, the computing device 102 enables the user to perform user gesture training. During user gesture training, the machine learning model 700 can collect new training data sets specific to the user and operate as a perpetual learning machine by performing instant training using training data associated with the user. In this way, the machine learning model 700 can adapt to the user's unique radar characteristics and the way the user performs gestures to improve performance. This type of training is called online or on-line training.
[0131] In the depicted configuration, the machine learning model 700 is implemented as a deep neural network and includes an input layer 702, a plurality of hidden layers 704, and an output layer 706. The input layer 702 includes a plurality of inputs 708-1, 708-2 ... 708-N, where N represents a positive integer equal to the number of radar signal characteristics 710 associated with one or more radar receive signals 404. The plurality of hidden layers 704 may include layers 704-1, 704-2 ... 704-M, where M represents a positive integer. Each hidden layer 704 may include a plurality of neurons, such as neurons 712-1, 712-2 ... 712-Q, where Q represents a positive integer. Each neuron 712 may be connected to at least one other neuron 712 in the previous hidden layer 704 or the next hidden layer 704. The number of neurons 712 between different hidden layers 704 may be similar or different. In some cases, the hidden layer 704 may be a copy of the previous layer (e.g., layer 704-2 may be a copy of layer 704-1). The output layer 706 may include outputs 714 - 1 , 714 - 2 . . . 714 -N associated with differentiated users 716 (eg, registered users, unregistered persons) that may have been detected within the proximity 106 .
[0132] In general, a variety of different deep neural networks may be implemented with various numbers of inputs 708, hidden layers 704, neurons 712, and outputs 714. The number of layers within the machine learning model 700 may be based on the number of radar signal characteristics and / or the ability to distinguish or identify categories (e.g., Figure 6As an example, the machine learning model 700 may include four layers (e.g., one input layer 702, one output layer 706, and two hidden layers 704) to distinguish the first user 104-1 from the second user 104-2, as described with respect to the example environment 100 and the example implementation 600. Alternatively, the number of hidden layers may be on the order of one hundred.
[0133] When utilized by the user module 222, the machine learning model 700 can improve the fidelity of user differentiation. The machine learning model 700 can collect multiple inputs 708 (e.g., radar signal characteristics 710 associated with one or more radar received signals 404) over time, the multiple inputs including topological, temporal, gesture, and / or contextual information about the user 104. For example, during a first interaction with the computing device 102, the second user 104-2 (e.g., a child) may be located away from the radar system 108, thereby generating a first set of inputs 708 for distinguishing the child as an unregistered person. The first set of inputs 708 can be included in the unregistered user identification assigned to the child. In a second interaction, the child may be seated close to the radar system 108, thereby generating a second set of inputs 708 for distinguishing the child, which is different from the first set and can also be included in the unregistered user identification of the child. This process can continue over time, providing more inputs 708 to the machine learning model 700 to better (e.g., with higher accuracy, faster rate) distinguish the child at a future time.
[0134] When utilized by user module 222, machine learning model 700 analyzes complex radar data (e.g., phase and / or amplitude data) and generates probabilities. Some of these probabilities are associated with various gestures that radar system 108 can recognize. Another of these probabilities may be associated with background tasks (e.g., background noise or gestures that cannot be recognized by radar system 108). Although described with respect to gestures, machine learning model 700 may be extended to indicate other events, such as whether a user is present within a given distance.
[0135] The gesture module 224 may also collect gesture execution information of the user 104 during the user gesture training as an input to the machine learning model 700 to enable the user module 222 to distinguish users based on gesture execution. If the first user 104-1 performs four push and pull gestures during the user gesture training, the machine learning model 700 may have at least four inputs 708-1, 708-2, 708-3, and 708-4. The user module 222 may partially utilize one or more outputs 714 of the machine learning model 700 to distinguish the first user 104-1 (e.g., father) when performing the push and pull gestures at a future time.
[0136] In general, the machine learning model 700 can be integrated into the user module 222, the radar system 108, or the computing device 102, or located separately from the computing device 102 (e.g., a shared server). The gesture module 224 can also include a similar machine learning model 700, which can improve the detection and recognition of gestures being performed by the user 104. For example, the gesture module 224 can detect one or more radar signal characteristics 710 associated with a gesture being performed by the first user 104-1. The gesture module 224 can use the output 714 of the machine learning model 700 to identify the gesture as a known gesture (e.g., a push-pull gesture) associated with a command (e.g., starting the oven). The operations of the gesture module 224 can be performed simultaneously with the operations performed by the user module 222, or at a separate time. The gesture module 224 can additionally include one or more deep learning algorithms, such as a convolutional neural network (CNN), to improve the detection and recognition of gestures. About Figure 8 An example of integrating a CNN into the gesture module 224 is further described.
[0137] Although the above Figure 7 and the following Figure 8 The machine learning model is described as distinguishing users and recognizing gestures, but detecting a user or a gesture can be performed as an operation of distinguishing the user or recognizing the gesture, respectively. However, in some cases, such as when attempts to both detect and recognize a gesture fail because the gesture is sufficient to detect but not recognize (e.g., the correlation with known gestures is too low to identify which gesture was performed, such as a low confidence level, but sufficient to determine that the gesture is a certain gesture rather than a non-gesture movement), multiple or more complex operations are used. Therefore, one or more radar signal characteristics of one or more radar received signals reflected from a user can be used to both detect that a gesture was performed and also to identify that the detected gesture is a known gesture.
[0138] Example spatiotemporal machine learning model
[0139] Figure 8 An example implementation 800 is shown including a gesture module 224 that utilizes a spatiotemporal machine learning model 802 (e.g., one or more CNNs) to improve detection and recognition of gestures. Such a spatiotemporal machine learning model 802 can enable the computing device 102 to detect and recognize gestures at a desired confidence level at long-range range distances (such as four meters) as well as close distances (such as a few centimeters). The gesture module 224 is depicted as having a signal processing module 804, a frame model 806, a temporal model 808, and a gesture debouncer 810.
[0140] The spatiotemporal machine learning model 802 has a multi-level architecture that includes a first level (e.g., a frame model 806) and a second level (e.g., a time model 808). In the first level, the spatiotemporal machine learning model 802 processes complex radar data (e.g., a complex range Doppler map) across the spatial domain, which involves processing the complex radar data pulse by pulse. In the second level, the spatiotemporal machine learning model 802 cascades the results of the frame model 806 across multiple pulses. By cascading these results, the second level processes the complex radar data across the time domain. Through the multi-level architecture, the overall size and inference time of the spatiotemporal machine learning model 802 can be significantly reduced compared to other types of machine learning models. This property can enable the spatiotemporal machine learning model 802 to run on a computing device 102 with limited computing resources.
[0141] The gesture module 224 is not limited to the arrangement depicted in the example implementation 800, and may include more or fewer components as shown. For example, the gesture module 224 may not have a gesture de-jitter 810, but include multiple signal processing modules 804 arranged before the frame model 806, before the time model 808, and / or after the time model 808. Additionally, any one or more of the depicted components of the gesture module 224 may be arranged separately from the gesture module 224. For example, the output of the time model 808 may be sent to the gesture de-jitter 810 that is separate from the gesture module 224. The spatiotemporal machine learning model 802 may also be separate from the gesture module 224. In one example, the spatiotemporal machine learning model 802 may be arranged within the radar system 108, but separate from the gesture module 224. In another example, the spatiotemporal machine learning model 802 may be separate from the computing device 102 (e.g., located on a remote server). In this document, such as with respect to Fig.25 To describe further details about gesture recognition.
[0142] In the example implementation 800, three radar transmit signals 812-1, 812-2, and 812-3 are transmitted using antennas 214-1, 214-2, and 214-3, respectively. The three radar transmit signals 812-1, 812-2, and 812-3 represent component signals that can be superimposed during propagation to form a composite transmit signal 402-Y. The composite transmit signal 402-Y propagates into the surrounding environment (e.g., a residence). The composite transmit signal 402-Y is reflected, for example, by the environment 814 and / or a gesture 816 being performed by the user 104. The environment 814 can include, for example, fixed environments such as stationary objects (e.g., furniture) and non-fixed environments such as the movement of objects not associated with the gesture (e.g., the user 104 walking and / or interacting with his environment 814, a ceiling fan, the movement of domestic animals, etc.). In Figure 88, composite transmit signal 402-Y is reflected by environment 814 and gesture 816 (e.g., a user's hand) to produce composite radar receive signal 404-Z. Antennas 214-1, 214-2, and 214-3 each receive a version of composite radar receive signal 404-Z represented by radar receive signals 818-1, 818-2, and 818-3. Radar receive signals 818-1, 818-2, and 818-3 correspond to at least three respective radar signal characteristics and are sent to analog circuitry 216 (see FIG. 1 ) before being sent to signal processing module 804. Figure 5 Analog circuitry 216 may modify (eg, digitize) radar receive signal 818 (associated with radar signal characteristics) to implement operations of signal processing module 804 .
[0143] exist Figure 8 In one example, computing device 102 transmits composite radar transmit signal 402-Y as a 16-chirped pulse train at a high pulse repetition rate (PRF) of 3 kilohertz (kHz). Each pulse train includes a wide 150 degree frequency modulated continuous wave radar beam to illuminate the surrounding environment of adjacent area 106 (e.g., environment 814 and gesture 816). Each pulse train is transmitted periodically over time (at a rate of 30 Hz) to enable unsegmented detection of gestures. Each antenna 214 captures a superposition of reflections from scattering surfaces within adjacent area 106 (corresponding to radar receive signal 404-Z) at a long range (e.g., four meters, although other ranges such as approximately two meters, six meters, or eight meters are also contemplated). In general, radar transmit signals can be transmitted at variable time periods that are non-fixed time periods (e.g., segmented detection periods).
[0144] Although three radar transmission signals 812 are depicted in the example implementation 800, in general, the computing device 102 may transmit one or more signals simultaneously from one or more antennas 214. In general, the computing device 102 may detect and recognize gestures at one or more locations within the proximity area 106, which extends to a long-range range, such as a linear range of one to four meters. The computing device 102 of the present disclosure does not require the user 104 to perform gestures at any particular location within the proximity area 106 (e.g., above the device interface), which may allow the user 104 to freely perform gesture commands from various locations in their home, for example, without having to walk to the device interface.
[0145] In the example implementation 800, radar receive signals 818-1, 818-2, and 818-3 are processed by a signal processing module 804, which applies a high pass filter to remove reflections from stationary objects. This high pass filter may include, for example, one or more resistors, capacitors, inductors, operational amplifiers (op amps), etc. Alternatively or additionally, each radar receive signal 818-1, 818-2, and 818-3 may be processed by a multi-stage fast Fourier transform (FFT) to generate one or more complex range Doppler maps 820-A (with respect to Fig.12 describes an example).
[0146] The complex range-Doppler map 820 may be a two-dimensional representation including a distance dimension (e.g., a slant distance dimension) and a Doppler dimension. The distance dimension may correspond to the displacement of a scattering surface (e.g., the surface of a user's hand) of gesture 816 from computing device 102, and the Doppler dimension may correspond to the range rate of change of the scattering surface relative to computing device 102. Thus, one or more complex range-Doppler maps 820 may enable radar system 108 to determine the relative position and motion of an object (e.g., gesture 816) within its vicinity 106. Radar system 108 may determine one or more complex range-Doppler maps 820 with similar or different FFT window sizes over time. For example, the FFT window size may be set to 128 by 16, corresponding to the bin size of the range and Doppler data, respectively. In this example, the range resolution Δr is 0.027 meters (m), while the Doppler resolution Δf d is 0.38 meters per second (m / s), defined by the following equation:
[0147]
[0148] Where c is the speed of light, B is the transmission bandwidth set to 5.5 GHz, PRF is the pulse repetition frequency set to 3 kHz, and f c is the center frequency set to 60.75 GHz, and l is the number of chirps per pulse train set to 16.
[0149] The signal processing module 804 then sends one or more complex range Doppler maps 820 to the frame model 806. As depicted in the example implementation 800, the signal processing module 804 sends three complex range Doppler maps 820-1, 820-2, and 820-3 to the frame model 806. Although one frame model 806 is depicted, the gesture module 224 may include one or more frame models 806, each of which utilizes CNN (convolutional neural network) technology. Referring to the previous example, the frame model 806 can receive the complex range Doppler maps 820-1 to 820-3, which are formatted as tensors of size 128 (number of range bins) by 16 (number of Doppler bins) by 6 (3 antennas × 2 values), where "2 values" corresponds to real and imaginary values as floating point representations. If the neighboring area 106 is reduced to a smaller size (e.g., 1.5m), the tensor size can be reduced. In this case, the number of distance bins may be truncated at the 64th bin, corresponding to user 104 standing 1.7 m from computing device 102 and performing gesture 816 with their hand at 1.5 m from the device.
[0150] The frame model 806 may output frame results 822-B that include a one-dimensional representation of the complex range Doppler map 820-A that has been processed for each pulse train. In this case, the frame results 822-1, 822-2, and 822-3 associated with different pulse trains may be sent to the temporal model 808 to be concatenated along the time domain and processed using similar or different CNN techniques. The temporal model 808 may calculate one or more gesture probabilities (e.g., temporal results 824-C, where the variable C represents the number of categories analyzed by the temporal model 808) to be sent to the gesture debouncer 810 for one or more gesture categories (e.g., known gestures) and / or background categories (e.g., background motion, objects not associated with known gestures). For example, the gesture debouncer 810 may be configured to recognize five possible gesture categories (e.g., tap, swipe up, swipe down, swipe right, and swipe left) and one background category (e.g., motion and objects not associated with the five possible gesture categories).
[0151] Gesture debouncer 810 may enable computing device 102 to perform unsegmented gesture detection, which may enable the device to detect gestures without first receiving an indication that a gesture is to be performed and / or without prior knowledge that a gesture is to be performed. For example, computing device 102 may detect gesture 816 without requiring user 104 to prompt the device with a “wake-up” trigger. The wake-up trigger may include a verbal, visual, or gesture prompt made by user 104 to indicate to computing device 102 that gesture 816 performance is about to occur. The wake-up trigger may be detected by any one or more sensors of computing device 102, such as antenna 214, microphone, ambient light sensor, pressure sensor, camera, etc. By not requiring a wake-up trigger event, radar system 108 continuously (e.g., in an unsegmented manner) scans and detects gesture 816 at any time.
[0152] To prevent misdetection or misrecognition of gesture 816, gesture debouncer 810 may apply one or more heuristic rules to temporal result 824-C. Misrecognition may include, for example, incorrectly relating gesture 816 to a known gesture or multiple known gestures. A first heuristic rule may include a requirement that temporal result 824-C (e.g., the result of spatiotemporal machine learning model 802) should have a value greater than an upper threshold criterion (e.g., a set value)—a maximum threshold requirement for confidence in the result—over the last three consecutive frames. A second heuristic rule may require that when multiple gestures are detected within a certain time period, temporal result 824-C has a value less than a lower threshold criterion—a minimum threshold requirement for the time elapsed between gesture executions. If user 104 performs a second gesture 816-2 quickly after performing a first gesture 816-1, gesture debouncer 810 may rely on an indication that the two gestures 816 are separate actions performed within the time period. For example, after first gesture 816-1 has been detected, time result 824-C may have a value less than a lower threshold for a period of time that elapses before second gesture 816-2 is detected. These upper and lower thresholds may be determined experimentally or customized based on user needs and implementation.
[0153] Gesture debouncer 810 may apply the one or more heuristic rules to determine a gesture result 826. This gesture result 826 may be an indication of the most likely classification of the object and / or motion (e.g., correlation of radar signal characteristics) detected by radar system 108. If gesture module 224 can classify into up to six categories, gesture debouncer 810 will indicate which of the six categories has been detected.
[0154] In a first example, assume that first radar transmit signal 402-1 is sent into proximity 106-1 of computing device 102-1 and is received by a cat walking through a room (the room is in Figure 4The first radar received signal 404-1 is reflected by the analog circuit 216 (see FIG. 1 ) of the radar system 108-1 shown. Figure 2 ) and then sent to the signal processing module 804 ( Figure 8 ). Signal processing module 804 cleans up first radar receive signal 404-1 by applying a high pass filter to remove signals associated with stationary objects (e.g., toys on the floor, not shown). First complex range-Doppler map 820-1 is sent to first frame model 806-1, which performs one or more CNN techniques and converts the map into a one-dimensional array (first frame result 822-1). First frame result 822-1 is sent to time model 808 (see Figure 8 ), where the device calculates the probability that the moving cat belongs to any one of six categories (five gesture categories and one background category). The temporal model 808 determines that the probability of each of the five gesture categories is 0.01 (range 0-1.00), and the probability of the background category is 0.95. These six values (e.g., the first temporal result 824-1) are sent to the gesture debouncer 810. The gesture debouncer 810 applies a first heuristic rule requiring that the probability of the category has a value greater than an upper threshold of 0.80. Since the probability of the background category is 0.95, the gesture debouncer 810 sends a first gesture result 826-1 indicating that the cat belongs to the background category, meaning that the cat's movement is not associated with one of the five gesture categories. Although the scope of the present teachings is not limited in this regard, it is assumed in this example that the six categories are mutually exclusive and the sum of all six probabilities is equal to 1.00.
[0155] exist Figure 4In the second example shown, a second radar transmit signal 402-2 is sent to the neighborhood 106-2 of the computing device 102-2 and is reflected by the user 104 (located four meters from the device) who is greeting a family member by waving his hand. The second radar receive signal 404-2 is received by the radar system 108-2 (using a set of techniques similar to those described in the previous example and received at the analog circuit 216), and the second frame result 822-1 is sent to the temporal model 808. The device determines that the motion of the user 104 waving his hand has the following probabilities: 0.2 for swiping up, 0.2 for swiping down, 0.2 for swiping to the left, 0.2 for swiping to the right, 0.1 for tapping, and 0.1 for the background category. The second temporal result 824-2 is sent to the gesture debouncer 810, which determines that none of the six categories has a probability value greater than the upper threshold set to 0.8. Instead of outputting second gesture result 826-2, gesture debouncer 810 may, for example, communicate (to gesture module 224, radar system 108, and / or computing device 102) that insufficient information is available regarding whether the user's swipe is a gesture or background motion. Based on this indication, gesture module 224 may instruct the device to transmit third radar transmission signal 402-3 ( Figure 4 ) to obtain additional radar signal characteristics (e.g., using a third complex range Doppler map 820-3, a third frame result 822-3, and a third time result 824-3 similar to the above examples). As described above, these techniques can use one or more additional sensors (e.g., a microphone) to collect supplemental data and / or can utilize additional information (e.g., contextual information) to discern whether the user's waving is intended to be a command to the device. One such approach includes audio that is not necessarily voice recognition. For example, even if users 104 are not distinguished, and therefore computing device 102 does not yet know which user is performing a gesture, audio and other supplemental data can be used to change or establish the probability that a particular gesture is being performed. In the following Fig.19 Examples are set forth in the description of .
[0156] In the example implementation 800, the frame model 806 and the temporal model 808 may together form a "spatial-temporal machine learning model" (indicated by 802) that utilizes CNN techniques, artificial intelligence, logic systems, residual neural networks (ResNet), dense layers, etc. Any one or more components in the spatial-temporal machine learning model may be repeated, rearranged, or ignored as desired. Fig. 9 An example structure of frame model 806 is depicted in FIG.
[0157] Fig. 9An example implementation 900 of spatiotemporal machine learning techniques utilized by the frame model 806 is shown. These techniques may utilize a neural network having several layers arranged as depicted (see Figure 7 ). Any one or more of the depicted layers may be rearranged, removed, or repeated to form alternative neural networks that may also enable detection and recognition of long-range gestures.
[0158] As depicted, the complex range-Doppler map 820 (e.g., input tensor) is first sent to an average pooling layer 902 to reduce the size of the map and the computational cost of subsequent subsequent layers. The average pooling layer 902 can be used to reduce the size of the map by averaging a set of values associated with the complex range-Doppler map 820 as input to a separable two-dimensional (2D) residual block 904. Although a separable convolution layer is depicted in the example implementation 900, the frame model 806 can alternatively or additionally utilize a standard convolution layer. The separable convolution layer is used to divide a matrix into its constituent (two) kernel portions. For example, a 3x3 complex range-Doppler map 820 may require one convolution with 9 multiplications, while the constituent kernel portions of the matrix (1x3 kernel and 3x1 kernel) may only require two convolutions with three (3) multiplications, thereby reducing computational time.
[0159] In general, residual blocks can be associated with ResNets that can skip connections (e.g., layers) within a neural network. Fig. 9 An example separable 2D residual block 904 is depicted in the dashed box of , which has two example paths. On the first path, the results from the average pooling layer 902 are input into the separable 2D convolution layer 906. The filter can slide on these inputs, performing element-by-element multiplication and summation at each position. In the example, the spatiotemporal machine learning model 802 can apply a 2x2 filter to the 3x3 input matrix of the separable 2D convolution layer 906. This 2x2 filter can slide on the values of the input matrix, resulting in a 2x2 matrix output. Alternatively, the edges of the 3x3 input matrix can be padded at the separable 2D convolution layer 906 to output a 3x3 matrix instead of a 2x2 matrix. Additionally, a step operation (e.g., skipping one or more positions when the filter slides across the input matrix) can be performed at the separable 2D convolution layer 906.
[0160] The output matrix of the separable 2D convolution layer 906 can be sent to a batch normalization layer 908, where the values of the output matrix are standardized to improve the stability and rate of the spatiotemporal machine learning model 802. For example, the values of the output matrix can be standardized by calculating the mean and standard deviation of each value. In another example, the values can be standardized by calculating the moving average of the mean and standard deviation. These standardized results can be sent to a rectifier (ReLU) 910, which is an activation function defined as the positive part of its independent variable. ReLU 910 can be used to prevent the separable 2D residual block 904 from activating all neurons 712 at the same time, thereby preventing exponential growth in computational requirements. ReLU 910 can include, for example, linear (e.g., parameterized) or nonlinear (e.g., Gaussian, S-type, analytical, logical) functions. The modified results from ReLU 910 can be sent to another separable 2D convolution layer 906, which is similar or different from the aforementioned convolution layer. Those results may be processed at another batch normalization layer 908 before being sent to a summing node 912 where the results of this first path are added with the results of the second path of the separable 2D residual block 904 .
[0161] On the second path, the result from the average pooling layer 902 bypasses the layers of the first path and is instead sent to a 2D convolution layer 914, which may include a two-dimensional standard convolution layer that does not separate the matrix into constituent kernels. The output matrix from the 2D convolution layer 914 may be sent to a summing node 912 where the results from the first and second paths are added and sent to another ReLU 910.
[0162] The frame model 806 can then implement a series of separable 2D residual blocks 904 and a maximum pooling layer 916 (collectively labeled 918). Each separable 2D residual block 904 can be used to process data using different or similar algorithms. For example, each block can use different or similar filter sizes (e.g., 1x1, 2x2, 3x3, etc.) and / or step sizes (e.g., filtering with every 1, every 2, every 3, etc.). Unlike the average pooling layer 902, one or more maximum values in a set of inputs can be determined at the maximum pooling layer 916. At the maximum pooling layer 916, for example, the maximum value of a 4x4 matrix can be calculated by sliding a 2x2 window on the matrix. Although this example uses a 2x2 window, in general, the maximum pooling layer 916 can utilize a 1x1 window, a 3x3 window, etc. Each maximum pooling layer 916 depicted in the example implementation 900 can include a window that is different or similar to the window of another maximum pooling layer 916.
[0163] At the end of the frame model 806, the data is sent to the final separable 2D convolution layer 906 and then to the flattening layer 920. The flattening layer 920 can reduce the data to a one-dimensional array (e.g., a frame summary 922). This frame summary 922 can be sent to the temporal model 808 for further processing, such as with respect to Fig.10 As described.
[0164] Fig.10 An example implementation 1000 of machine learning techniques utilized by the temporal model 808 is shown. Similar to the frame model 806, these techniques may include a neural network having several layers arranged as depicted. Any one or more of the depicted layers may be rearranged, removed, or repeated to form alternative neural networks that may also enable detection and recognition of long-range gestures.
[0165] As depicted, frame summary 922 may be sent to temporal model 808 to be correlated with the frame of radar receive signal 404 in the time domain. Frame summary 922 may first be processed at a one-dimensional (1D) residual block 1002. 1D residual block 1002 may be similar to a separable 2D residual block, except that the computation is performed in one dimension via a standard convolution (e.g., without using a separable convolution). Fig.10 An example 1D residual block 1002 is depicted in the dashed box of , which has two possible paths. On the first path, the frame summary 922 is input into a 1D convolution layer 1004. The 1D convolution layer 1004 can be similar to the separable 2D convolution layer 906, except that the calculation is performed in one dimension through a standard convolution. These results can be sent to a batch normalization layer 908 and then to a ReLU 910. The data can be processed by another 1D convolution layer 1004 that is similar or different from the previous 1D convolution layer 1004. Those results can be processed at another batch normalization layer 908 before being sent to a summing node 912, where the results of this first path are added to the results of the second path of the 1D residual block 1002.
[0166] On the second path, the frame summary 922 bypasses the layers of the first path and is sent to a 1D convolutional layer 1004, which may be similar or different than the 1D convolutional layer 1004 of the first path. The results from the first and second paths are added at a summing node 912 and sent to another ReLU 910.
[0167] The temporal model 808 can then implement a series of 1D residual blocks 1002 and maximum pooling layers 916 (collectively labeled 1006). Each 1D residual block 1002 can utilize references to related Fig. 9The data is processed by different or similar algorithms discussed in detail above. At the end of the temporal model 808, the data is sent to a dense layer 1008 and then to a softmax layer 1010. In the dense layer 1008 (e.g., a fully connected layer), each neuron can receive data from all neurons in the previous layer. The size of the data can change at the dense layer 1008 and reflect the number of categories that can be used to classify gestures. For example, if the gesture module 224 has five gesture categories and one background category, the output from the dense layer 1008 can reflect these six categories. This output can be sent to the softmax layer 1010, where a softmax function is applied to the data to assign probabilities to each category. For example, if there are six categories, each of the six categories can be assigned a probability value of 0-1. The sum of the probabilities of all six categories can be added up to 1. The gesture probability 1012 is then sent from the temporal model 808 to the gesture debouncer 810 (see the relevant Figure 8 to realize the detection and recognition of gestures.
[0168] Offline Training for Radar-Based Gesture Recognition
[0169] Gesture module 224 may be trained using an offline supervised training technique. In this case, a recording device records data generated by radar system 108. The recording device is coupled to radar system 108 to capture complex radar data. The recording device may be a separate unit connected to radar system 108. Alternatively, the recording device may be integrated within radar system 108 or computing device 102.
[0170] For offline training, radar system 108 collects positive records when the participant performs a gesture, such as swiping right with the left hand, swiping left with the right hand. In general, a positive record represents a plurality of radar data recorded by radar system 108 or a recording device during a time period when the participant performs a gesture associated with a gesture category.
[0171] Positive records can be collected using participants with various heights and handedness types (e.g., right-handed, left-handed, or ambidextrous). Moreover, positive records can also be collected by participants located at various positions relative to the radar system 108. For example, a participant can perform gestures at various angles relative to the radar system 108, including angles between approximately -45 degrees and 45 degrees. As another example, a participant can perform gestures at various distances from the radar system 108, including distances between approximately 0.3 meters and 2 meters. Additionally, positive records can be collected by participants taking various postures (e.g., sitting, standing, or lying), different recording device placements (e.g., on a table or in the hands of a participant), and various orientations of the recording device (e.g., longitudinal orientation or transverse orientation).
[0172] For offline training, radar system 108 also collects negative records while the participant performs a background task. The background task may include the participant operating a computer or computing device 102. Another background task may include the participant walking around radar system 108. In general, negative records represent complex radar data recorded by radar system 108 during a period of time when the participant performs a background task associated with a background category (or a task not associated with a gesture category).
[0173] The participant may perform background motions that resemble gestures associated with one or more gesture categories. For example, the participant may move their hand between a computer and a mouse, which may resemble a directional swipe gesture. As another example, the participant may place a cup on a table next to the recording device and then pick up the cup, which may resemble a tap gesture. By capturing these gesture-like background motions in the negative recording, the gesture module 224 may be trained to detect the difference between background tasks with gesture-like motions and intentional gestures intended to control the computing device 102.
[0174] Negative recordings can be collected in a variety of environments, including a kitchen, bedroom, or living room. In general, negative recordings capture natural behavior around radar system 108, which can include a participant reaching for computing device 102, dancing nearby, walking, cleaning a table with computing device 102 on it, or turning the steering wheel of a car while computing device 102 is in a stand. Negative recordings can also capture repetitions of hand movements similar to sliding gestures, such as moving an object from one side of radar system 108 to another. For training purposes, negative recordings are assigned background labels that distinguish these negative recordings from positive recordings. To further improve the performance of gesture module 224, negative recordings can be optionally filtered to extract samples associated with motions with speeds above a predefined threshold standard.
[0175] The positive and negative records are split or divided to form a training data set, a development data set, and a test data set. The ratio of positive records to negative records in each data set can be determined to maximize performance. In the example training program, the ratio is 1:6 or 1:8.
[0176] The recording device can fine-tune the timing of the gesture segment being recorded. To do this, the recording device detects the center of the gesture motion within the gesture segment being recorded. As an example, the recording device detects a zero Doppler intersection within a given gesture segment. The zero Doppler intersection may refer to the time situation when the gesture motion changes between a positive Doppler bin and a negative Doppler bin. In other words, the zero Doppler intersection may refer to the time situation when the Doppler determined rate of change of distance changes between a positive value and a negative value. This indicates the time when the direction of the gesture motion becomes substantially perpendicular to the radar system 108, such as during a swipe gesture. It may also indicate the time when the direction of the gesture motion reverses and the gesture motion becomes substantially stationary, such as during the execution of a tap gesture. Other indicators may be used to detect the center point of other types of gestures.
[0177] The recording device aligns the timing window based on the center of the detected gesture motion. The timing window can have a specific duration. This duration can be associated with a specific number of pulse trains, such as 12 or 30 pulse trains. Generally speaking, the number of pulse trains is sufficient to capture gestures associated with the gesture category. In some cases, additional offsets are included in the timing window. The offset can be associated with the duration of one or more pulse trains. The center of the timing window can be aligned with the center of the detected gesture motion.
[0178] The recording device adjusts the size of a given gesture segment based on its aligned timing window to generate pre-segmented data. For example, the size of the gesture segment is reduced to include samples associated with the aligned timing window. The pre-segmented data can be provided as part of a training data set, a development data set, and a test data set.
[0179] The gesture module 224 can be trained using a training data set and supervised learning. As described above, the training data set can include pre-segmented data. This training enables optimization of the internal parameters of the gesture module 224, including weights and biases.
[0180] First, the hyperparameters of the gesture module 224 are optimized using the development dataset. As described above, the development dataset may include pre-segmented data. In general, hyperparameters represent external parameters that do not change during training. A first type of hyperparameter includes parameters associated with the architecture of the gesture module 224, such as the number of layers or the number of nodes in each layer. A second type of hyperparameter includes parameters associated with the processing of the training data, such as a learning rate or the number of rounds. The hyperparameters may be selected manually, or may be automatically selected using techniques such as grid search, black box optimization techniques, gradient-based optimization, and the like.
[0181] Secondly, the gesture module 224 is evaluated using the test data set. Specifically, a two-stage evaluation process is performed. The first stage includes using the pre-segmented data in the test data set and the gesture module 224 to perform a segmentation classification task. Without using the gesture debouncer 810 to determine whether a gesture occurs, the gesture is determined based on the highest probability provided by the time model 808. Through the execution of the segmentation classification task, the accuracy, precision, and recall of the gesture module 224 can be evaluated.
[0182] The second stage includes performing an unsegmented recognition task using the gesture module 224 and the gesture de-jitter 810. Instead of using the pre-segmented data in the test data set, the unsegmented recognition task is performed using duration sequence data (or a continuous data stream). Through the execution of the unsegmented recognition task, the recognition rate and / or false alarm rate of the gesture module 224 can be evaluated. Specifically, the unsegmented recognition task can be performed using positive records to evaluate the recognition rate and using negative records to evaluate the false alarm rate. The unsegmented recognition task utilizes the gesture de-jitter 810, which implements further tuning of the threshold criteria to better achieve the desired recognition rate and the desired false alarm rate.
[0183] If the results of the segmented classification task and / or the unsegmented recognition task are not satisfactory, one or more elements of the gesture module 224 may be adjusted. These adjustments may extend to the overall architecture, training data, and / or hyperparameters of the gesture module 224. With these adjustments, the training of the gesture module 224 may be repeated. Positive and / or negative records may be reinforced to further enhance the training of the gesture module 224, as further described below.
[0184] Data augmentation techniques
[0185] In the case where a large number of interactions (e.g., 50 or more interactions) between the user 104 and the computing device 102 are not required for offline or online training, data enhancement can be used to enhance the data set of the spatiotemporal machine learning model 802 (e.g., the complex range-Doppler map 820, the frame results 822, and / or the time results 824 associated with the radar signal characteristics) to increase the number of stored radar signal characteristics. The radar enhancement techniques described in the present disclosure include the determination of random or predetermined phase rotations and / or amplitude scaling of data corresponding to one or more radar signal characteristics. By implementing these radar enhancement techniques, the computing device 102 can reduce the amount of gesture training required to accurately recognize gestures to achieve a desired confidence level. For example, the computing device 102 may only need to collect three radar signal characteristics of the user 104 performing a sliding gesture, instead of ten, to accurately recognize the command. Therefore, the user 104 can quickly enjoy using the computing device 102 without having to undergo time-consuming gesture training.
[0186] For the complex range-Doppler map 820, the absolute phase may be affected by the surface position at the range bin resolution, phase noise, errors in the sampling timing, etc. In addition, the amplitude of the complex range-Doppler map 820 may be affected by the properties of the antenna 214, the continuity between the computing devices 102 (when utilizing a computing system), the reflectivity of the signal from the scattering surface, the orientation of the scattering surface, etc. For large data sets, these absolute phases and amplitudes may be uniformly distributed. However, for small data sets (e.g., corresponding to a new user starting gesture training), these absolute phases and amplitudes may deviate, thereby reducing the accuracy of gesture detection and / or recognition.
[0187] To address these issues without requiring the user 104 to undergo time-consuming gesture training, the computing device 102 may utilize radar enhancement techniques to enhance the phase and / or amplitude of the complex range-Doppler map 820, M, based on the following relationship:
[0188] A(r,d,c)=s*M(r,d,c)*(cosθ+isinθ)
[0189] Where A is the enhanced complex range-Doppler map, r is the range bin index, d is the Doppler bin index, and c is the channel index (refer to Figure 8 ), s is a random or predetermined scaling factor selected from a normal distribution with a mean of 1, and θ is a random or predetermined rotation phase selected from a uniform distribution between -π and π. According to this equation, the complex value may be rotated by various phase values and / or scaling factors to increase the number of stored radar signal characteristics that may be used to identify a gesture. In a physical sense, the phase values may represent the angular displacement of the gesture from the computing device 102. Specifically, these phase values may represent the angular orientation of the scattering center of the user's hand (assuming that the gesture is performed by the user's hand) relative to the forward or zero-degree direction of the antenna 214 of the computing device 102. Relatedly, the amplitude value may represent the linear displacement of the gesture from the computing device 102. If the user 104 performs the gesture near the device (e.g., one foot away), the amplitude may have a larger value than when the user 104 performs the gesture far away from the device (e.g., four meters away). The random or predetermined rotation phase and scaling factor used for enhancement may be correspondingly different from the rotation phase and scaling factor of the detected and / or stored radar signal characteristics. In this manner, the enhanced data may supplement (rather than duplicate) the radar signal characteristics stored on computing device 102 .
[0190] In the example, first user 104-1 (e.g., an unregistered person) interacts with computing device 102 for the first time and begins gesture training for a swipe gesture. Computing device 102 instructs user 104 to perform a swipe gesture with their hand. Radar system 108 transmits first radar transmit signal 402-1, which is reflected by the scattering surface of the user's hand, thereby generating first radar receive signal 404-1. First antenna 214-1 receives this signal and sends it to analog circuit 216, which is then received by signal processing module 804 of gesture module 224. First complex range-Doppler map 820-1 or M of first radar receive signal 404-1 is obtained. 1 is enhanced to include two additional phase values and two additional amplitude values, thereby producing four enhanced complex range-Doppler maps (A 1 , A 2 , A 3 , A 4 ). These five graphs 1 , A 1 , A 2 , A 3 and A 4 Can be utilized by the spatiotemporal machine learning model 802 to improve detection and recognition of gestures (and background motion).
[0191] Fig.11 Experimental results 1100 are shown indicating improved performance in gesture recognition when utilizing radar enhancement techniques. In this experiment, the gesture module 224 enhances the complex range Doppler map 820 within a Keras layer. The enhanced data set 1102 includes the detected, stored, and enhanced radar signal characteristics, while the original data set 1104 includes the detected and stored radar signal characteristics. The x-axis of the experimental results 1100 represents the number of gestures performed over time in a set of gesture training. For this experiment, the number of false alarms per hour is equal to 2.0. The results indicate that the enhanced data set 1102 enables the computing device 102 to recognize known gestures more frequently at an earlier time during training compared to the original data set 1104. This means that the computing device 102 of the present disclosure can accurately recognize known gestures while requiring less interaction with the user 104 when accurately recognizing known gestures using radar enhancement techniques.
[0192] These radar enhancement techniques can be modified and / or applied to detect and distinguish users 104, and are not limited to techniques for gesture detection and recognition. Specifically, the computing device 102 can enhance a set of one or more radar signal characteristics used to distinguish users 104 and improve the detection of user presence. In the example, a first user 104-1 (e.g., a new unregistered person) is detected at two meters (2m) from the computing device 102 and at a 90-degree angle relative to the front of the device (0-degree orientation). At the user's location, the computing device 102 detects a radar signal characteristic of the first user 104-1, which can be used to distinguish it from another user. However, before the radar system 108 can determine the second radar signal characteristic, the first user 104-1 leaves the neighborhood 106 of the computing device 102. For some devices, one radar signal characteristic may not be sufficient to distinguish the presence of the first user 104-1 at a high confidence level at a future time. However, the computing device 102 of the present disclosure can enhance this first radar signal characteristic to achieve accurate distinction of the first user 104-1 at a future time. Specifically, the enhancement may include rotation phase angles θ of 0, 180, and 270 degrees, and amplitudes s corresponding to linear displacements of 0.5 m, 1 m, and 4 m.
[0193] These enhanced complex range-Doppler maps (corresponding to enhanced radar signal characteristics) may be stored with the first radar signal characteristics to enable differentiation of the first user 104-1 at a future time. When the first user 104-1 re-enters the neighboring area 106 at a future time, the user module 222 may use ten stored radar signal characteristics (nine enhanced radar signal characteristics and one detected radar signal characteristic) instead of only one radar signal characteristic to differentiate the user. The computing device 102 may additionally use information about Figure 3 The depicted communication network 302 exchanges enhanced radar signal characteristics between devices of the computing system. In this manner, a group of computing devices 102-X (forming a computing system) may improve gesture recognition by sharing enhanced data.
[0194] Experimental data using spatiotemporal machine learning models
[0195] Fig.12Experimental data 1200 of a user 104 performing a tap gesture for a computing device 102 is shown. The tap gesture may involve the user 104 pushing his hand toward the device and then pulling his hand back to its initial position. For this experiment, the user 104 performed the tap gesture at a distance of 1.5 m from the computing device 102 and an angular displacement of zero degrees. The experimental data 1200 includes real and imaginary values of a complex range Doppler map 820. The first row 1202-1, the third row 1202-3, and the fifth row 1202-5 each include real values of 30 frames (corresponding to 30 complex range Doppler maps 820) as collected by the first receiving channel 508-1, the second receiving channel 508-2, and the third receiving channel 508-3, respectively. The second row 1202-2, the fourth row 1202-4, and the sixth row 1202-6 each include 30 frames of imaginary values (corresponding to the 30 complex range Doppler maps 820) collected by the first receiving channel 508-1, the second receiving channel 508-2, and the third receiving channel 508-3, respectively. Each frame is shown as a horizontal axis (x-axis) corresponding to the distance change rate of the gesture, with zero distance change rate at the center, negative distance change rate on the left, and positive distance change rate on the right. Each frame is also shown as a vertical axis (y-axis) corresponding to the displacement of the gesture, with zero displacement (e.g., the position of the antenna 214) at the bottom and a range of 2m at the top. These displacements and distance change rates are taken relative to the position of the receiving antenna of the computing device 102. For each row 1202, the 30 frames are arranged in sequence, with time increasing from left to right.
[0196] From the experimental data 1200, it is detected in each frame that the user 104 is standing at 1.5m, as evidenced by the first circular feature that is always visible at the top of each frame. When the user 104 performs the tap gesture, their hand movement 1204 can be seen in frames 13-19. When the user 104 begins to move his hand toward the device (in frame 13), a second circular feature begins to appear. This second circular feature continues to move to the bottom of frames 14 and 15 (as the user 104 pushes his hand toward the device) until the user 104 has fully extended his arm in frame 16. In frame 17, the user 104 begins to pull his hand back toward his body and away from the computing device 102. By frame 20, the tap gesture has been completed. The gesture module 224 can determine from this data that the user 104 has performed a tap gesture, and then determine the corresponding command to be executed by the computing device 102.
[0197] When storing one or more radar signal characteristics associated with user 104 performing a tap gesture, gesture module 224 may include, for example, any one or more frames shown in experimental data 1200. In a first example, the device may select frames 13-19 (identified at 1204) of row 1202-1 to store for future reference. In a second example, the device may store frames 1-30 of row 1202-1 as radar signal characteristics of a tap gesture. In a third example, the device may store all 30 frames of each of rows 1202-1 through 1202-6 as radar signal characteristics of a tap gesture. Additional experimental data regarding the tap gesture will be provided regarding Fig.13 Give a description.
[0198] Fig.13 Experimental data 1300 is shown of a user 104 performing a tap, a right swipe, a strong left swipe, and a weak left swipe on a computing device 102 . Fig.13 The data shown in Fig.12 The experimental data 1200 of is arranged similarly, but with the following differences. Rows 1302-1, 1302-3, 1302-5 and 1302-7 correspond to absolute range Doppler maps for a tap gesture, a right swipe, a strong left swipe, and a weak left swipe, respectively. Each absolute range Doppler map can be generated by taking the average amplitude of the corresponding complex range Doppler map 820. Rows 1302-2, 1302-4, 1302-6 and 1302-8 correspond to interferometric range Doppler maps for a tap gesture, a right swipe, a strong left swipe, and a weak left swipe, respectively. Each interferometric range Doppler map can be generated by calculating the phase difference between the complex range Doppler maps 820 associated with two or more receiving channels 508.
[0199] In this experimental data 1300, the strong left swipes of rows 1302-5 and 1302-6 can be seen in frames 13-19, with clear second circular features. However, the weak left swipes of rows 1302-7 and 1302-8 lack clear second circular features, which may make the classification of this gesture challenging. In some cases, the gesture module 224 can utilize the spatiotemporal machine learning model 802, context information, online learning techniques, etc. to improve the recognition of this gesture. In other cases, this weak left swipe can be classified as "negative data" or "false gesture", which cannot be mapped to a gesture category (e.g., swipe, tap) to achieve the desired confidence level. Instead of ignoring negative data, the gesture module 224 can store this information (e.g., as background motion) to improve gesture recognition at future times.
[0200] Negative data collection
[0201] The computing device 102 of the present disclosure may store one or more radar signal characteristics to enable detection, differentiation and / or recognition of gestures and / or users. These stored radar signal characteristics are not limited to "positive data" and may also include "negative data". Positive data may include radar signal characteristics for identifying gestures (e.g., known gestures with associated commands) and / or distinguishing users to achieve a desired confidence level. Examples of positive data may include radar signal characteristics associated with a gesture category (e.g., tap, slide, flick, point) or a specific user based on radar cross section (RCS) data, for example. On the other hand, negative data may include radar signal characteristics that are not associated with one or more stored radar signal characteristics of a gesture or user 104 (not reaching the desired confidence level). Examples of negative data may include movements such as a person walking, twisting their torso, picking up an object, etc. Negative data may also include movements of animals (e.g., house cats), cleaning devices (e.g., automatic vacuum cleaners), etc. Gesture module 224 may categorize such negative data within a background category of radar signal characteristics that are not associated with known gestures (eg, gesture commands that computing device 102 is programmed or taught to recognize).
[0202] In an example, motion is detected in the vicinity 106 of the computing device 102, and the radar system 108 detects a first radar signal characteristic of the motion. If the first radar signal characteristic correlates (to a desired confidence level) with one or more stored radar signal characteristics of a tap gesture, the gesture module 224 may store this first radar signal characteristic as positive data to improve recognition of the tap gesture at a future time. If the first radar signal characteristic does not correlate with one or more stored radar signal characteristics of a known gesture, the gesture module 224 may determine that the motion is not associated with a command. Instead of discarding this data, the gesture module 224 may store the first radar signal characteristic as negative data to improve detection or recognition of gestures from background motion (e.g., movement of the user 104 or an object that may not be intended to be a gesture command). Similar techniques may be used to improve detection of user presence and differentiation of one user from another.
[0203] Fig.14 Experimental data 1400 is shown for three sets of negative data that can be stored to improve the detection of gestures from background motion. Fig.13The experimental data 1300 of FIG. 1402 is similarly arranged, but with the following differences. Rows 1402-1, 1402-3, and 1402-5 correspond to absolute range Doppler plots for user 104 moving their hands while speaking near computing device 102, user 104 twisting their torso in front of the device, and user 104 picking up and putting down an object near the device, respectively. Rows 1402-2, 1402-4, and 1402-6 correspond to interferometric range Doppler plots associated with rows 1402-1, 1402-3, and 1402-5, respectively.
[0204] Fig.15 Experimental results 1500 are shown regarding the accuracy of gesture detection and recognition in the presence of background motion. In this experiment, a user 104 performed a swipe gesture 1502 and a tap gesture 1504 over time while background motion (either natural motion of the user 104 or motion from other objects) occurred within the proximity 106 of the computing device 102. Using both positive and negative data, the gesture module 224 accurately detected that the gesture was a gesture (even if it did not always identify which known gesture it was), and in this case also identified the swipe gesture 1502 and the tap gesture 1504 at a detection and recognition rate of approximately 0.88, while false positives per hour were generated at a rate of approximately 0.10. False positives represent situations where the gesture module 224 incorrectly determined background motion to be a gesture, and here also identified background motion as a swipe or tap gesture. The false positives per hour can be affected by the above-mentioned factors as previously discussed. Figure 8 The effects of the upper and / or lower thresholds of the gesture debouncer 810 are described.
[0205] To calculate the detection and recognition rate, in general, the gesture module 224 marks motion events as "correct" or "incorrect." Correct gesture detection and recognition occurs when the gesture module 224 outputs only one accurate gesture determination associated with the command intended by the user 104 (e.g., both detection and recognition). Incorrect gesture detection occurs when the gesture module 224 does not detect a gesture (even though the gesture was performed by the user 104), detects a gesture but determines an inaccurate one (the gesture is not associated with the intended command of the user 104), or performs detection for a single gesture and determines multiple gestures. The detection and recognition rate is determined by dividing the number of correct events by the total number of events (the sum of correct events and incorrect events).
[0206] Fig.16Experimental results 1600 are shown regarding detection and recognition rates of gestures when adversarial negative data is additionally used. In this experiment, user 104 performed both gestures and adversarial motions that were similar (but not identical) to the gestures. These adversarial motions included the user 104 picking up and placing objects near computing device 102, interacting with the device's touch screen, flipping switches on and off, and moving their hands while speaking near the device. Experimental results 1600 indicate that gesture module 224 has robust performance in accurately recognizing gestures shown at a higher rate when performing adversarial motions in neighboring area 106, and having lower false alarm rates and robust results (as shown in robust results 1604), relative to results 1602 that do not use adversarial negative data.
[0207] Unsegmented gesture detection and recognition
[0208] Although the computing device 102 of the present disclosure can provide gesture training (e.g., segmented learning of gesture execution), the device can also use unsegmented learning techniques to improve the detection of gestures over time. For segmented learning, the gesture module 224 can prompt the user 104 to perform, for example, a tap gesture within a certain time period. The computing device 102 may be able to detect this tap gesture based on the "prior knowledge" that the user 104 will perform a tap gesture (rather than other gestures) within a specified time period. In contrast, unsegmented learning does not utilize prior knowledge in this way. For unsegmented learning, the gesture module 224 does not necessarily know whether and when the user 104 can perform any one or more gestures for the computing device 102. In addition, unsegmented learning can allow the computing device 102 to continuously detect gestures over time without the user 104 prompting the device (e.g., providing a wake-up trigger) before performing a gesture. Therefore, unsegmented recognition of gesture execution can be more difficult than segmented recognition.
[0209] To improve the accuracy of unsegmented recognition of gestures, the gesture module 224 can utilize one or more gesture debouncers 810 and adjust the upper and lower thresholds as needed to improve performance. Additionally, the gesture module 224 can identify gestures in time by detecting one or more zero crossings of the data in a set of two or more frames along the velocity axis (x-axis) (e.g., referring to the circular features of the experimental data 1200). For example, Fig.12The hand motion 1204 (a tap motion) performed in 1204 causes the second circular feature to move to the left (negative velocity) in frames 13-15, move back to the center (zero velocity) in frame 16, and move to the right (positive velocity) in frames 17-19. The zero crossing of this motion occurs in frame 16, and the gesture module 224 can identify this frame as the center of the motion. The computing device 102 can additionally select one or more data frames (e.g., frames 1-15 and frames 17-30) around frame 16 to form a data set for correlating this motion with a tap gesture.
[0210] Fig.17 Experimental results 1700 (confusion matrix) related to the accuracy of unsegmented gesture detection are shown. In this experiment, the user 104 performed 6 gestures over time, including background motion 1702 (e.g., movement of the user 104 or an object not associated with the gesture command), swipe left 1704, swipe right 1706, swipe up 1708, swipe down 1710, and tap 1712. The x-axis (gestures performed 1714) of this confusion matrix represents the gesture commands intended by the user 104, and the y-axis (gestures recognized 1716) represents the classification of the gesture execution by the gesture module 224. Experimental results 1700 include 36 possible results, which are quantified by the classification rate (normalized to one) over a certain period of time. The results indicate that the gesture module 224 is able to determine each gesture 1714 performed with an accuracy rate of 0.831-0.994. In this experiment, the gesture debouncer 810 used an upper threshold of 0.9.
[0211] Fig.18 Experimental results 1800 corresponding to the accuracy of unsegmented gesture detection at various linear and angular displacements from computing device 102 are shown. These results utilize Fig.17 The experimental results of 1700 are similar to the data set, and include another confusion matrix of normalized detection rates of gestures over time periods. The matrix includes 29 results on the accuracy of unsegmented gesture recognition at various angular displacements (ranging from -45 degrees to +45 degrees) and linear displacements (ranging from 0.3m to 1.5m) from the computing device 102. Fig.17 and Fig.18 The data represented in represents hundreds of motions performed by user 104, including gestures and background motion.
[0212] Additional sensors to improve fidelity of user differentiation
[0213] Fig.19An exemplary implementation 1900 of a computing device 102 is shown that uses an additional sensor (e.g., microphone 1902) to improve the fidelity of user differentiation and / or gesture recognition and user engagement. In some scenarios, radar signal characteristics associated with nearby objects (e.g., registered users or unregistered persons) or motion (e.g., gesture performance, background motion) may not provide sufficient information to distinguish users 104 and / or recognize gestures to a desired confidence level. As depicted in the example implementation 1900, the user module 222 may not be able to determine with confidence that the first user 104-1 is a registered user (e.g., a father) based solely on the radar signal characteristics. In such a scenario, the computing device 102 may direct the audio signal 1904 (e.g., sound waves emitted by the first user 104-1) to enable the radar system 108 to determine the presence of the father.
[0214] As depicted in the example implementation 1900, the user module 222 can receive the audio signal 1904 via the microphone 1902 and analyze the characteristics of these sound waves (e.g., wavelength, amplitude, time period, frequency, speed, velocity) to determine which user is present in the vicinity 106. This analysis can be performed automatically when triggered, or simultaneously or after the radar signal characteristics analysis. The audio signal 1904 can be modified by additional circuit systems and / or components before being received by the user module 222.
[0215] With or without access to private information (e.g., the content of the conversation), the user module 222 can analyze the audio signal 1904 to distinguish between users. For example, the radar system 108 can characterize the audio signal 1904 to distinguish the presence of the first user 104-1 with or without identifying the words being spoken (e.g., performing speech-to-text), as characteristics such as low pitch or fast-paced speech can be used to distinguish specific users. The radar system 108 can characterize the audio signal 1904 in terms of pitch, loudness, timbre, voice quality, rhythm, consonance, dissonance, pattern, etc. Therefore, the user 104 can comfortably discuss private information near the computing device 102 without worrying about whether the device is identifying the words, sentences, thoughts, etc. being spoken, depending on the user's settings and preferences.
[0216] The computing device 102 may store the audio detection characteristics of one or more users 104 (e.g., on a shared memory) to distinguish the presence of the users. When a registered user (e.g., a father) enters the proximity zone 106 of the radar system 108, the user module 222 may partially utilize the stored audio detection characteristics of the father to distinguish him from other users of the device. When an unregistered person enters the proximity zone 106, the user module 222 may partially utilize the stored audio detection characteristics of the registered user to determine that this is an unregistered person who has not provided an audio signal 1904 to the computing device 102, for example. The radar system 108 may then generate an unregistered user identification for this unregistered person, the unregistered user identification including the audio detection characteristics associated with one or more audio signals 1904 emitted by the unregistered person. Therefore, the radar system 108 may be able to use the audio detection characteristics stored in its unregistered user identification to distinguish this unregistered person at a later time.
[0217] Although the additional sensor of the example implementation 1900 is depicted as a microphone 1902, in general, Fig.19 The described techniques may be performed using various sensors described herein. It should be understood that it is not beyond the scope of the present teachings that, for situations or environments where privacy is not a functional concern or otherwise provides methods to eliminate privacy concerns, the additional sensor of the example implementation 1900 is a camera or video camera. In addition, additional sensors may be used to improve gesture detection and recognition. For example, the computing device 102 may additionally utilize data associated with an ambient light sensor to detect and recognize a gesture being performed by the first user 104-1. This action may be particularly useful in situations where, for example, the first user 104-1 performs an ambiguous gesture that cannot be identified using only radar signal characteristics to a desired confidence level. In another example, the computing device 102 may additionally utilize data from an ultrasonic sensor to improve recognition of an ambiguous gesture performed by the first user 104-1. Therefore, supplementary data sensed by a non-radar sensor may be used to assist in gesture recognition and other determinations, such as user presence, user differentiation, and user engagement.
[0218] In general, additional sensor inputs (e.g., audio signal 1904 from microphone 1902) are optional, and user 104 may be given privacy controls to limit the use of such additional sensors. For example, user 104 may modify their personal settings, general settings, default settings, etc. to include and / or exclude additional sensors (e.g., in addition to antenna 214 for radar). Furthermore, user module 222 may implement these personal settings after distinguishing the presence of the user. Fig. 20 Privacy controls are further described.
[0219] Adaptive privacy and other settings
[0220] Fig. 20 Example environments 2000-1 and 2000-2 are shown in which privacy settings are modified based on user presence. In example environment 2000-1, user module 222 of computing device 102 detects that first user 104-1 is present within proximity 106. Radar system 108 may implement first privacy settings 2002 for first user 104-1 in response to detecting the presence of a user or person. This first privacy setting 2002 may include information about, for example, allowed sensors (see Fig.19 ), audio reminders, calendar information, music, media, settings for home objects (e.g., lighting preferences), etc. For example, when the first privacy setting 2002 has been implemented, the first user 104-1 can receive audio reminders for calendar events.
[0221] In example environment 2000-2, computing device 102 may later detect the presence of second user 104-2 (e.g., another registered user) in addition to the continued presence of first user 104-1. Radar system 108 implements second privacy setting 2004 based on the presence of the second user to adapt the privacy of first user 104-1. Implementation may be automatic or triggered based on a command from first user 104-1. For example, second privacy setting 2004 may limit audio reminders to prevent private information from being broadcasted in the presence of others. Second privacy setting 2004 may be based on, for example, preset conditions, user input, etc.
[0222] A second privacy setting 2004 may also be implemented to protect the privacy of the first user's information and adapted based on the users in the room. For example, the presence of another registered user (e.g., a family member) may require fewer privacy restrictions than the presence of an unregistered person (e.g., a guest). Adaptive privacy settings may also be customized for each user 104. For example, the first user 104-1 may have a more stringent privacy setting (e.g., limiting audio reminders in the presence of others), while the second user 104-2 may have a less stringent privacy setting (e.g., not limiting audio reminders in the presence of others).
[0223] In addition to adaptive privacy, these techniques can also adapt other settings in a similar manner. As with the adaptive privacy described above, these adaptive settings can depend on the presence of other users such as those users (e.g., second user 104-2) near the user (e.g., first user 104-1) who is interacting with the computing device 102 or has interacted with the computing device. Adaptive settings can be applied to ongoing operations, such as when the first user 104-1 commands to play music on a stereo device. If the second user 104-2 is distinguished, or if the second user 104-2 speaks to the first user 104-1 (or vice versa), these techniques can turn down the music without explicit user interaction (e.g., a gesture to turn down the music) from the first user 104-1. Adaptive settings can also be applied to ongoing operations. In an example, the first user 104-1 (e.g., a father) can start an oven and set a timer for 20 minutes. If the second user 104-2 (e.g., a child) attempts to turn off the oven or the timer before the timer expires, the technology of the present disclosure can prevent the child from doing so. This adaptation allows the father to control the operation of the oven and the timer during the 20 minutes to prevent the baking process from being disturbed. Specifically, computing device 102 can associate the ongoing operation with user 104 who has executed a command that prevents other users from modifying the operation. In another example, a mother can execute a command to turn off the bedroom lights at 9:00pm to ensure that her child goes to bed on time. If the child executes a command to keep the lights on after bedtime, computing device 102 can prevent the child from modifying the mother's command.
[0224] Example Implementation of a Computing System
[0225] Fig.21 A technique is shown in which user differentiation may be implemented using multiple computing devices 102-1 and 102-2 forming a computing system (eg, regarding Figure 1 and Fig. 20 2. In the example environment 2100, an example residence is depicted as having a first room 304-1 and a second room 304-2. A first computing device 102-1 equipped with a first radar system 108-1 is located in the first room 304-1, and a second computing device 102-2 equipped with a second radar system 108-2 is located in the second room 304-2. The first computing device 102-1 and the second computing device 102-2 are capable of exchanging information (e.g., stored in local or shared memory) with the aid of a communication network 302, and partially form a computing system. For purposes of this illustrative example, the first proximity area 106-1 does not overlap with the second proximity area 106-2.
[0226] The first computing device 102-1 of the example environment 2100 may use the first radar system 108-1 to send a first radar transmission signal 402-1 (see above). Figure 4 ) to detect the presence of one or more users. First radar transmit signal 402-1 may be reflected by an object (e.g., first user 104-1) and modified in amplitude, phase, or frequency before being received at first computing device 102-1. First radar system 108-1 may receive first radar receive signal 404-1 (see above, Figure 4 ) (including at least one radar signal characteristic) is compared with one or more stored radar signal characteristics of registered users to determine whether the first user 104-1 is a registered user or an unregistered person. In this example, the first radar received signal 404-1 is not correlated with the one or more stored radar signal characteristics of registered users. Therefore, at 2102, the first user 104-1 is distinguished as an unregistered person.
[0227] After determining at 2102 that the first user 104-1 is an unregistered person, the first computing device 102-1 generates an unregistered user identification (e.g., a simulated identity, a pseudo identity) and assigns it to the unregistered person, as shown at 2104. The unregistered user identification may include one or more radar signal characteristics associated with the first radar received signal 404-1, which may be used to distinguish the unregistered person from other users (e.g., the second user 104-2) at a future time. The unregistered user identification may be stored on a local or shared memory, and each computing device of the computing system (e.g., the second computing device 102-2) may access the unregistered user identification even if the device has not directly detected the unregistered person. For example, even if the unregistered person has never been detected by the second computing device 102-2, the second computing device 102-2 may access the stored first radar signal characteristics associated with the unregistered user identification.
[0228] At a future time, the first user 104-1 walks into the second room 304-2 and is detected by the second computing device 102-2 at 2106. Specifically, the second radar system 108-2 of the second computing device 102-2 sends a second radar transmit signal 402-2 to detect the presence of one or more users. The second radar transmit signal 402-2 is reflected by an object (e.g., the first user 104-1) and is modified in amplitude, phase, or frequency before being received at the second computing device 102-2. The second radar system 108-2 compares this second radar receive signal 404-2 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of a registered user and an unregistered user identifier assigned to an unregistered person to determine whether the object is a registered user or an unregistered person. In this example, the at least one radar signal characteristic of the second radar receive signal 404-2 is correlated with one or more stored radar signal characteristics of an unregistered person. Therefore, based on the unregistered user identifier, the first user 104-1 is again distinguished as an unregistered person (at 2106). The second radar signal characteristic may be stored and associated with the unregistered user identification.
[0229] Although Fig.21 , but the same techniques may be applied to distinguish registered users. For example, first computing device 102-1 may transmit third radar transmit signal 402-3 to distinguish one or more additional users. Third radar transmit signal 402-3 may be reflected by another object (e.g., second user 104-2 that is different from the previously detected unregistered person) and modified before being received at first computing device 102-1. First radar system 108-1 may compare this third radar receive signal 404-3 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of registered users and unregistered user identification. In this example, the at least one radar signal characteristic of third radar receive signal 404-3 is related to one or more stored radar signal characteristics of registered users. Therefore, second user 104-2 may be distinguished as a registered user. The third radar signal characteristic may be stored to improve the distinction between registered users at a later time. First radar system 108-1 may additionally access stored settings, preferences, training history, habits, etc. of registered users to provide a customized experience.
[0230] In this example, the second user 104-2 (registered user) may later move to the second room 304-2 and enter the second adjacent area 106-2. The second radar system 108-2 of the second computing device 102-2 may transmit a fourth radar transmit signal 402-4 to distinguish the registered user. The fourth radar transmit signal 402-4 is reflected by the registered user and is modified before being received at the second computing device 102-2. The second radar system 108-2 compares this fourth radar receive signal 404-4 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of the registered user to distinguish the registered user. In this example, the second computing device 102-2 accesses one or more stored characteristics from the local memory and / or shared memory of the first computing device 102-1. Here, the at least one radar signal characteristic is correlated with the one or more stored radar signal characteristics of the registered user, and the second radar system 108-2 determines that the registered user is present within the second adjacent area 106-2 of the second computing device 102-2. About Fig. 22 The use of computing device 102 as part of a computing system is further described.
[0231] Continuity of operations across computing systems
[0232] Fig. 22 An example environment 2200 is shown in which multiple computing devices 102-1 and 102-2 across a computing system continuously perform operations. The first computing device 102-1 is depicted as being located in a bedroom 2202, which is separate from an office 2204 where the second computing device 102-2 is located. The first computing device 102-1 and the second computing device 102-2 are two or more devices (e.g., Figure 3 and Fig.21 Part of a computing system (such as those described).
[0233] At a first time, in example environment 2200-1, user 104 performs a swipe gesture to command first computing device 102-1 to read aloud the latest news headlines. User 104 may be listening to the news in bedroom 2202, and first radar system 108-1 may instruct first computing device 102-1 to continue playing the news when the presence of user 104 is detected.
[0234] At a second time, in example environment 2200-2, user 104 moves from bedroom 2202 to office 2204 and the news continues to play on second computing device 102-2. Specifically, first radar system 108-1 can detect the lack of user presence within first proximity 106-1 of bedroom 2202 and pause the news. Once user 104 moves into office 2204, second computing device 102-2 can detect the presence of this user within, for example, second proximity 106-2 and distinguish this user. Second radar system 108-2 can then automatically (e.g., without user input) continue playing the news previously paused by first radar system 108-1. In this way, user 104 can enjoy a seamless experience that is automated across multiple rooms of a residence.
[0235] The ongoing operation may follow the user 104 who performed the gesture. For example, if the user 104 of the example environment 2200-1 leaves the bedroom 2202, the first computing device 102-1 may detect the absence of this user 104 (e.g., rather than another user) and pause the news. When the user 104 is later detected and distinguished by the second computing device 102-2 in the office 2204, the news may continue to be read aloud. In addition to the user 104 of the example environment 2200-1, there may be one or more other users in the bedroom 2202 that have been detected by the first computing device 102-1. The first radar system 108-1 may determine that these other users are not the user 104 who performed the gesture command, thereby inferring that the news should follow the user 104 who performed the gesture. Alternatively, when there are other users, the news may be appended to playing in the bedroom 2202 and also follow the user 104 who performed the gesture.
[0236] Each radar system 108-1 and 108-2 may adjust ongoing operations based on the location of user 104. For example, if user 104 is lying close to first computing device 102-1 (as depicted in example environment 2200-1), first radar system 108-1 may detect this shorter distance and lower the speaker volume. Alternatively, if user 104 moves to the far side of bedroom 2202 (e.g., farther from first computing device 102-1), first radar system 108-1 may detect this greater distance and increase the speaker volume.
[0237] In another example ( Fig. 22In a first room equipped with a first computing device 102-1, user 104 performs a gesture associated with a two-part command (1) start playing news headlines aloud and (2) stop playing news headlines. In a first room equipped with a first computing device 102-1, user 104 performs a first gesture associated with the first part of the command (start playing news). User 104 then moves to a second room equipped with a second computing device 102-2 and continues to listen to the news in this room (following Fig. 22 ). At a later time, user 104 performs a second gesture associated with the second part of the command (end playing the news). In this example, first computing device 102-1 and second computing device 102-2 coordinate the two-part command across multiple rooms of the home using communication network 302. The same technology can also be applied to the first part and / or the second part of the audio input sensed by the microphone of first computing device 102-1 or second computing device 102-2.
[0238] In the additional example ( Fig. 22 In the example not depicted in FIG. 1 , user 104 may perform a gesture associated with a single command, which gesture is detected by both first computing device 102-1 and second computing device 102-2 (see related Figure 4 ). In this example, a user may provide detailed commands while moving between two rooms. User 104 begins stating their single command (schedule an appointment with their doctor) in the first room (equipped with first computing device 102-1), and first radar system 108-1 recognizes the first portion of the command. However, user 104 moves to the second room (to obtain their medical records) and continues to schedule their appointment by providing the second portion of the command to second computing device 102-2. In this example, first computing device 102-1 and second computing device 102-2 again utilize communication network 302 to coordinate a single command across multiple rooms of the home. The same techniques may also be applied to the first portion and / or second portion of the audio input sensed by the microphone of first computing device 102-1 or second computing device 102-2.
[0239] In the additional example ( Fig. 22In a first room (not depicted in FIG. 1 ), user 104 may perform a pauseable (capable of being paused and then resumed) continuous radar detection gesture that may be detected successively by first computing device 102-1 and second computing device 102-2, thereby achieving the beneficial effect of continuing intermittent continuous activities between rooms. An example of a pauseable continuous gesture may be a voice list gesture, in which a user may begin moving their hand in a rolling circular motion (hereinafter "rolling their hand" - intuitively one may think of the "keep the film rolling" gesture made by a movie director to a cameraman) in a generally vertical plane passing through themselves and the devices. In one example, in a first room (equipped with first computing device 102-1 having first radar system 108-1), a user begins rolling their hand and, while continuously rolling their hand, says: "This is my shopping list: milk, eggs, butter...," and device 102-1 will recognize the gesture using first radar system 108-1 and associate the spoken items with the user's shopping list as long as the user continues to roll their hand. If the user stops scrolling their hand, list making is suspended, even if the user continues to speak. (This suspension may occur, for example, if the user is interrupted and needs to talk to another person about a topic unrelated to their shopping list.) If the user then scrolls their hand, list making is resumed, and what they said resumes being added to the shopping list. (The user can terminate list making at any time with a "push and pull" gesture and / or an appropriate voice command.) If the user stops scrolling their hand during (unterminated) list making in the first room and then walks into the second room, the user can resume scrolling their hand in the second room and resume speaking their shopping list items, the second device 102-2 with the second radar system 108-2 will recognize the gesture, and as long as the user keeps scrolling their hand, the addition of spoken items to the shopping list will resume, and so on.
[0240] about Fig. 22 The described techniques are not limited to ongoing operations and can be applied to operations that are performed periodically over time, such as Fig.23 Further described.
[0241] Fig.23An example environment 2300 is shown in which a computing system implements operational continuity across multiple computing devices 102-1 and 102-2. In the example environment 2300-1, a user 104 is detected by a first computing device 102-1 of a computing system located in a kitchen 2302. The first radar system 108-1 determines that the user 104 is an unregistered person and assigns an unregistered user identification to the user, thereby distinguishing the user 104 from other users. In this example, the first computing device 102-1 prompts the unregistered person to begin gesture training for a first gesture. During training, radar signal characteristics associated with the manner in which the unregistered person performs the first gesture are stored and associated with the unregistered user identification. This unregistered user identification can be stored in a memory and / or accessed by any one or more devices of the computing system (e.g., the second computing device 102-2).
[0242] At a later time, as depicted in example environment 2300-2, the presence of a user is detected in restaurant 2304 by second computing device 102-2 as part of the computing system. Second radar system 108-2 distinguishes user 104 as an unregistered person based on radar signal characteristics, and accesses the training history of this user using the unregistered user identification. Second computing device 102-2 may then prompt user 104 to continue training on the second gesture. Specifically, second radar system 108-2 may determine that they have completed training on the first gesture. This sequence may continue over time using various computing devices 102 of the computing system (although Fig.23 ), until the unregistered individual has completed their hand gesture training.
[0243] However, based on the user 104's behavior or body position in the room, the user may perform gestures in different ways for some of the computing devices 102 of the computing system. When the user 104 performs a gesture in a manner that is different from how the user was taught to perform the gesture, for example, during training, the computing device 102 may determine that an ambiguous gesture has been performed. Such an ambiguous gesture may be similar to one or more gestures that may be recognized by the device, but may lack sufficient similarity to a single known gesture to allow for high confidence recognition. Therefore, each computing device 102 may utilize contextual information to improve the interpretation of ambiguous gestures, such as information about the context. Fig.24 Further described.
[0244] Example of ambiguous gesture interpretation
[0245] Fig.24A technique for radar-based ambiguous gesture determination using contextual information is shown. Example environment 2400 depicts user 104 performing ambiguous gesture 2402 detected by radar system 108 of computing device 102. It is assumed herein that user 104 intended to perform first gesture 2404, which is recognizable by the device and is associated with a first command to turn down the volume of music playing from computing device 102. However, user 104 accidentally performs ambiguous gesture 2402, which is similar to both first gesture 2404 and second gesture 2406 (associated with a second command to open a garage door), but cannot be recognized to a desired confidence level. Therefore, radar system 108 cannot determine that ambiguous gesture 2402 is first gesture 2404 and not second gesture 2406.
[0246] Specifically, the amount by which the blur gesture 2402 is correlated with the first gesture 2404 and the second gesture 2406 can be greater than the no confidence level but less than the high confidence level. For example, the correlation with each gesture can have a 40% confidence level (e.g., there is a 40% chance that the blur gesture 2402 is the first gesture 2404 and a 40% chance that it is the second gesture 2406). If the no confidence level is set to 10% and the high confidence level is set to 80%, the correlation of the blur gesture 2402 with the first gesture 2404 or the second gesture 2406 is higher than the no confidence level and lower than the high confidence level. In general, the no confidence level and the high confidence level can be modified or adapted to improve the quality of gesture detection. In the present disclosure, "desired confidence levels" will collectively refer to these no confidence levels and high confidence levels, which can be different or similar for each use case (e.g., each gesture, each user, destructiveness of gestures).
[0247] To avoid prompting user 104 to perform the gesture repeatedly until successful, radar system 108 uses contextual information 2408 to improve interpretation of ambiguous gesture 2402. In this example, radar system 108 determines that contextual information 2408 includes music playing on computing device 102 at the current time. In general, contextual information may include ongoing operations, past or planned operations, foreground or background operations, location of the device, history of common users or gestures, running applications, conditions external to computing device 102, etc. External conditions may include, for example, time of day, lighting, audio within proximity 106, etc.
[0248] Computing device 102 may determine based on contextual information 2408 that ambiguous gesture 2402 is most likely first gesture 2404. In this example, radar system 108 correlates the music being played on the device at the current time with the first command of first gesture 2404. Since the second command to open the garage door is unrelated to the music being played on the device, radar system 108 determines that first gesture 2404 is more likely to be the intended gesture of user 104. This determination may be performed using formal logic, informal logic, mathematical logic, hysteresis logic, deductive reasoning, inductive reasoning, abductive reasoning, etc. The determination may also be performed using, for example, machine learning model 700 (see Figure 7 ) and / or spatiotemporal machine learning model 802 (refer to Figure 8 ) to perform the machine learning model.
[0249] In an example, radar system 108 utilizes inductive reasoning to determine a common association by which radar system 108 can infer that ambiguous gesture 2402 is first gesture 2404. In doing so, radar system 108 can proceed according to the following logical argument:
[0250] (1) The blur gesture 2402 is the first gesture 2404 or the second gesture 2406 .
[0251] (2) The first gesture 2404 is associated with a first command to turn down the volume of music.
[0252] (3) Second gesture 2406 is associated with a second command to open the garage door.
[0253] (4) Music is currently playing (context information 2408).
[0254] (5) The first command is related to the music being played at the current time.
[0255] (6) The second command has nothing to do with the music currently being played.
[0256] (7) Ambiguous gestures are usually associated with the operation being performed at the current time.
[0257] Therefore, blur gesture 2402 is most likely the first gesture 2404 .
[0258] Upon detecting ambiguous gesture 2402, radar system 108 may determine that ambiguous gesture 2402 is not correlated to a known gesture (e.g., one or more stored radar signal characteristics) to a desired confidence level. The desired confidence level may be a quantitative or qualitative assessment of the confidence level (e.g., accuracy) required to correctly identify the gesture, or another threshold criterion or combination of threshold criteria such as symbolic, vector-based, or matrix-based.
[0259] In example environment 2400, radar signal characteristics of ambiguous gesture 2402 have a 40% chance of being associated with stored radar signal characteristics of first gesture 2404, a 40% chance of being associated with stored radar signal characteristics of second gesture 2406, and a 20% chance of being associated with stored radar signal characteristics of a third gesture. If the desired confidence level is set to 50%, radar system 108 may not be able to accurately determine which known gesture is associated with ambiguous gesture 2402 to the desired confidence level. However, radar system 108 may determine that first gesture 2404 and second gesture 2406 are more likely to be associated with ambiguous gesture 2402 than the third gesture. Specifically, radar system 108 may consider gestures that exceed a minimum confidence level (e.g., a threshold) by 35% or more. Since the third gesture has only a 20% chance of being ambiguous gesture 2402, radar system 108 may rule out this possibility. Radar system 108 may alternatively determine that first gesture 2404 and second gesture 2406 both have a 40% likelihood of exceeding a minimum confidence level.
[0260] To identify ambiguous gesture 2402 as first gesture 2404 (e.g., a known gesture) or second gesture 2406 (e.g., another known gesture), radar system 108 may utilize contextual information 2408. Specifically, radar system 108 may determine whether the first command of first gesture 2404 or the second command of second gesture 2406 is associated with contextual information 2408. The association may be quantitative or qualitative based on a strict yes / no association (e.g., binary), a varying association scale, logic or reasoning, a machine learning model (700, 802), etc. For example, radar system 108 may determine that the first command to turn down the volume of music is associated with contextual information 2408 of the music being played at the current time. On the other hand, radar system 108 may also determine that the second command to open the garage door is not associated with the music being played. Based on this determination, radar system 108 may determine that ambiguous gesture 2402 is first gesture 2404.
[0261] In some scenarios, radar system 108 may not be able to accurately identify ambiguous gesture 2402 using contextual information. If the first command of first gesture 2404 is instead associated with starting a timer, radar system 108 may determine that both the first command and the second command are not associated with the music being played at the current time. Without additional information, radar system 108 may determine that neither the first command nor the second command should be executed by computing device 102. Additionally, computing device 102 may prompt user 104 to repeat the gesture and / or provide additional input (e.g., a voice command).
[0262] Radar-enabled gesture recognition
[0263] Fig.25 Example implementations 2500-1 through 2500-3 are shown in which gesture module 224 may recognize gestures performed by user 104. To recognize gestures, gesture module 224 of radar system 108 may analyze radar receive signal 404 to determine (1) topological features, (2) temporal features, and / or (3) contextual features. Each of these features may be associated with one or more radar signal characteristics detected by computing device 102. Computing device 102 is not limited to Fig.25 The three feature categories depicted in FIG. 1 and may include other radar signal characteristics and / or categories not shown. In addition, these three feature categories are shown as example categories and may be combined and / or modified to include subcategories that implement the techniques described herein. Figures 7 to 18 The techniques discussed may additionally be included and used with Fig.25 The techniques presented in are not mutually exclusive. Fig.25 Gesture recognition discussion and Figure 6 The discussion of user differentiation is similar to that of FIG, except that it applies to gesture execution.
[0264] In example implementation 2500-1, gesture module 224 may use topology information in part to recognize gestures. Figure 6 Similar to the teachings of example implementation 600-1 of , the topological features may include RCS data associated with the height, shape, orientation, distance, material, size, and the like of the user 104. For example, the user 104 may perform the swipe gesture 2502 by forming a straight hand, but perform the pinch gesture 2504 by pinching their thumb and index finger together (e.g., making contact therebetween). The topological features associated with the swipe gesture 2502 may include a larger surface area, an orthogonal orientation, a flatter surface, and the like when compared to the topological features associated with the pinch gesture 2504. On the other hand, the topological features associated with the pinch gesture 2504 may include a smaller surface area, a more uneven (e.g., less flat) surface, a greater depth of hand position, and the like when compared to the topological features associated with the swipe gesture 2502.
[0265] In example implementation 2500-2, gesture module 224 may use, in part, temporal information to recognize gestures, as described in example implementation 600-2. Radar system 108 may recognize gestures by receiving and analyzing, for example, motion characteristics (e.g., a unique way that user 104 moves while performing a gesture). In the example, user 104 performs a swipe gesture to turn pages of a book they are reading. Radar system 108 detects one or more radar signal characteristics of the swipe gesture (e.g., as depicted in a temporal profile) and compares it to one or more stored radar signal characteristics. The temporal profile may include using Figure 5The analog circuit 216 detects the amplitude of one or more radar receive signals 404 over time. The temporal characteristic curve of the sliding gesture may have different characteristics than, for example, the swiping gesture or the pinching gesture. The swiping gesture may include two complementary movements (e.g., thereby generating two amplitude peaks), while the sliding gesture may include one movement (e.g., thereby generating one amplitude peak). Over time, the pinching gesture may be performed slower than the sliding gesture, resulting in a wider amplitude peak of the radar receive signal 404.
[0266] In example implementation 2500-3, gesture module 224 may also use context information to recognize gestures, as described above. Figure 6 Example implementations 600-3 and 600-4. Context information is not limited to Fig.25 The examples depicted in Figures 24 to 32 Various other examples are described. In this disclosure, “contextual information” refers to information that adds context (e.g., additional details) to the signals received by radar system 108. Contextual information may include user presence, user habits, location of computing device 102, commonly performed gestures and / or detected history of users, common activity in a room, ongoing operations on one or more computing devices 102, past and planned operations, foreground and background operations, etc.
[0267] Example implementation 2500-3 depicts how user presence provides additional context to recognize gestures. Computing device 102 may be configured to recognize a push-pull gesture that involves user 104 pushing their hand to a certain extent at a certain rate, and pulling their hand back to a similar extent (e.g., equal distance but opposite direction) at the same rate. In this way, when user 104 performs a push-pull gesture, the device may expect complementary push and pull motions. However, if a push-pull gesture is performed where the push and pull motions are not complementary (as depicted), gesture module 224 may determine that an ambiguous gesture has been performed.
[0268] To improve the interpretation of ambiguous gestures, user module 222 can provide additional details about the user's presence (eg, contextual information) to gesture module 224. Gesture module 224 can better interpret ambiguous gestures if it can determine which user performed the gesture.
[0269] In an example, the first user 104-1 performs a push-pull gesture, where the push motion and the pull motion are not complementary, and the gesture module 224 determines that the gesture is not associated with a push-pull gesture (e.g., the desired confidence level is not reached). The first user 104-1 may push his hand to a certain extent at a certain rate, but pull his hand back at a significantly slower rate and to a shorter extent. The gesture module 224 determines that the gesture is an ambiguous gesture, which may be a push-pull gesture or a push gesture. Instead of prompting the first user 104-1 to perform the gesture again, the gesture module 224 uses the user presence information from the user module 222 to determine that the first registered user performed the gesture. This first registered user may have performed the gesture in the past (e.g., during gesture training) and may typically perform the push-pull gesture in this modified manner. The gesture module 224 may then access one or more stored radar signal characteristics associated with the history of the first registered user performing the gesture (e.g., gesture training history). This additional information may allow computing device 102 to determine that the ambiguous gesture is a push-pull gesture commonly performed by the first registered user.
[0270] As also depicted in example implementation 2500-3, second user 104-2 may also perform a push-pull gesture in a unique manner. Second user 104-2 may push their hand to a certain extent at a certain rate, but pull their hand back at a significantly faster rate and to a longer extent. Because the push motion and the pull motion are not complementary, gesture module 224 may determine that the gesture is an ambiguous gesture, which may be a push-pull gesture or a pull gesture. Gesture module 224 may again utilize user presence information from user module 222 to determine whether a registered user performed the gesture. In this example, second user 104-2 is an unregistered person who has not previously performed gestures with computing device 102. Therefore, gesture module 224 may utilize other contextual information (such as information about Figure 26 to Figure 32 further described), prompting the unregistered person to perform the gesture again, and / or initiating gesture training.
[0271] In this example, radar system 108 may receive one or more radar receive signals 404 that include radar signal characteristics of second user 104-2 performing its version of the push-pull gesture (e.g., another ambiguous gesture). Gesture module 224 may compare these radar signal characteristics with stored radar signal characteristics to determine that the other ambiguous gesture may be a push-pull gesture (a first known gesture) or a pull gesture (a second known gesture). Specifically, the correlation of the performed version of the push-pull gesture with each of the known gestures exceeds a no confidence level (e.g., a minimum threshold), but the device determines that the correlation is also below a high confidence level.
[0272] In general, contextual information may include details determined using, for example, antenna 214, additional sensors of computing device 102, data stored on memory (e.g., user habits), local information (e.g., time, relative location), operating state, etc. In another example (not depicted), gesture module 224 may use local time as context to enable recognition of ambiguous gestures. If user 104 consistently performs a gesture to turn on the lights at 6:00am every morning, computing device 102 may note this habit to improve gesture recognition. If user 104 accidentally performs an ambiguous gesture at 6:00am (e.g., perhaps associated with turning on the lights or reading the news aloud), the device may use this contextual information to determine that user 104 most likely intended to perform a gesture to turn on the lights. In another example, if computing device 102 is located in a kitchen, gesture module 224 may determine over time that kitchen-related gestures (e.g., starting the oven) are common in the room. If user 104 performs an ambiguous gesture in the kitchen (e.g., perhaps associated with turning on a dishwasher or activating a security system in the living room), the device may use contextual information for kitchen-related gestures to determine that user 104 most likely intended to perform a gesture to turn on the dishwasher.
[0273] The gesture module 224 may additionally utilize one or more logic systems (e.g., including predicate logic, hysteresis logic, etc.) to improve gesture recognition. The logic system may be used to prioritize certain gesture recognition techniques over other gesture recognition techniques (e.g., favoring temporal features over contextual information), add weights (e.g., confidence) to certain results when relying on two or more features, etc. The gesture module 224 may also include machine learning models (e.g., 700, 802) to improve gesture recognition (e.g., interpretation of ambiguous gestures), as previously described accordingly. Figure 7 and Figure 8 As described.
[0274] Radar system 108 may use contextual information alone or in combination with topological or temporal information of radar signal characteristics to recognize gestures. In general, gesture module 224 may use the contextual information in any combination and at any time. Fig.25 For example, radar system 108 may collect topological and temporal information about a performed gesture, but lack contextual information. In another case, radar system 108 may collect topological and temporal information, but determine that the information is insufficient to recognize an ambiguous gesture. If contextual information is available, radar system 108 may utilize the contextual information to recognize an ambiguous gesture. Fig.25 Any one or more of the categories described in may take precedence over another category. Fig.26Applications of the contextual features described with respect to example environment 2500 - 3 are further described.
[0275] Contextual information associated with user habits
[0276] Fig.26 An example environment 2600 is shown in which computing devices 102-1 and 102-2 can utilize contextual information of user habits to improve gesture recognition. In the example environment 2600, a first room 304-1 (dining room 2304) contains a first computing device 102-1, and a second room 304-2 (bedroom 2202) contains a second computing device 102-2. Each computing device 102 can determine its location in a residence based on, for example, relative location relative to other devices, user input, user behavior, frequency of commands, type of commands, user presence, etc. Additionally, each computing device 102 can utilize one or more sensors to perform, for example, geo-fencing, multilateration, true range multilateration, dead reckoning, atmospheric pressure adjustment, true range inertial multilateration, angle of arrival calculation, time of flight calculation, etc. to determine the location of the device.
[0277] Based on the typical behavior or body position of the user in each room of the house, the user can perform gestures in the room in different ways. As depicted in the example environment 2600, the user 104 located in the dining room 2304 can typically perform gestures while sitting upright on a chair, while the user 104 located in the bedroom 2202 can typically perform gestures while lying horizontally on the bed. For example, when the user 104 performs a push-pull gesture in the dining room 2304, the first computing device 102-1 can detect that the push-pull gesture has been performed with the complementary push and pull motions taught during gesture training. Specifically, the user 104 can usually push and pull to a similar (but opposite) degree and at a similar rate as depicted by the complementary arrows in the example environment 2600. On the other hand, the user 104 usually performs a push-pull gesture from a lying position with non-complementary push and pull motions in the bedroom 2202. For example, the user 104 can typically push to a greater degree, a faster rate, and in a non-collinear direction relative to the pull motion, as also depicted using the non-complementary arrows in the example environment 2600.
[0278] Instead of requiring user 104 to perform the push and pull gesture perfectly (e.g., consistently, as expected, as taught during training) at each computing device 102-1 and 102-2, each device may learn the user's habits as contextual information over time to improve recognition of ambiguous gestures. In example environment 2600, first computing device 102-1 may learn over time that user 104 typically performs a first version of the push and pull gesture in restaurant 2304 with complementary push and pull motions, while second computing device 102-2 may learn over time that user 104 typically performs a second version of the push and pull gesture in bedroom 2202 with non-complementary push and pull motions. Thus, first computing device 102-1 and second computing device 102-2 may leverage this contextual information (regarding how user 104 typically performs gestures in the room) when attempting to recognize ambiguous gestures in restaurant 2304 and bedroom 2202, respectively.
[0279] In the example, user 104 is awakened by the sound of an alarm and performs a push-pull gesture to command second computing device 102-2 to turn off the alarm. However, user 104 performs the second version of the push-pull gesture in a non-complementary push and pull motion in its tired state. Second computing device 102-2 can detect one or more radar signal characteristics associated with the spatial and / or temporal characteristics of the gesture and determine that the user has performed an ambiguous gesture. Specifically, the user's performance of the push-pull gesture is similar to a first known gesture (a push-pull gesture) and a second known gesture (a waving gesture). If the device cannot recognize the performed gesture to the desired confidence level, second computing device 102-2 determines that the user has performed an ambiguous gesture that may be a first known gesture or a second known gesture. In order to identify this ambiguous gesture, second computing device 102-2 considers contextual information about the user's habits in bedroom 2202, and determines that user 104 typically performs the push-pull gesture in a tired manner with non-complementary push and pull motions. Therefore, the radar signal characteristics of the ambiguous gesture are more similar to the radar signal characteristics of the second version of the push-pull gesture typically performed by user 104 in bedroom 2202. The device determines that the ambiguous gesture is more likely to be the first known gesture (the push-pull gesture) and continues to perform the operation of turning off the alarm. Assume that in example environment 2600, two devices detect multiple radar signal characteristics (e.g., stored radar signal characteristics) associated with each version of the push-pull gesture over time and store the multiple radar signal characteristics in one or more memories.
[0280] Additionally, first computing device 102-1 may learn over time to anticipate (e.g., expect, determine that it is more common) a first version of the push-pull gesture based on the location of the first device within dining room 2304. Similarly, second computing device 102-2 may learn over time to anticipate a second version of the push-pull gesture based on the location of the second device in bedroom 2202. If first computing device 102-1 moves from dining room 2304 to bedroom 2202, first computing device 102-1 may be reconfigured (e.g., automatically upon detection of the relocation, manually by user 104) to anticipate the second version of the push-pull gesture instead of the first version. Similarly, if second computing device 102-2 moves from bedroom 2202 to dining room 2304, second computing device 102-2 may be reconfigured to anticipate the first version of the push-pull gesture. While the contextual information of example environment 2600 includes common ways for users 104 to perform gestures in one location, contextual information may also include users 104 that are typically associated with one location (e.g., room 304). Therefore, it may be useful for computing device 102 to predict the user based on the location of the device, such as with respect to Fig. 27 Further described.
[0281] Fig. 27 Prediction of user presence based on location of computing device is shown. In example environment 2700, first computing device 102-1 may learn over time that first user 104-1 (e.g., daughter) is typically present within bedroom 2202, which may allow the device to prediction of the first user's presence at 2702. If first computing device 102-1 detects the presence of a user but cannot accurately distinguish it from other users, the device may rely on the prediction of the first user's presence to distinguish user 104. For example, if an ambiguous user (e.g., a user that cannot be distinguished to a desired confidence level) is detected at a distance within bedroom 2202, first computing device 102-1 may determine that the ambiguous user is first user 104-1 (daughter) or second user 104-2 (mother). If the daughter is predominantly detected in bedroom 2202 over time (when compared to the presence of the mother), first computing device 102-1 may determine that the ambiguous user is likely to be the daughter.
[0282] Similarly, second computing device 102-2 may learn over time that second user 104-2 (the mother) is primarily present in another room (e.g., office 2204), which may allow second computing device 102-2 to anticipate the presence of the second user (here, the mother) at 2704. If second computing device 102-2 is moved to bedroom 2202, second computing device 102-2 may be reconfigured to anticipate the presence of the daughter based on the repositioning of the device. Specifically, second computing device 102-2 may access the user history detected by first computing device 102-1 to enable anticipation of the presence of the daughter. Similarly, if first computing device 102-1 is moved to office 2204, first computing device 102-1 may be reconfigured to anticipate the presence of the mother. Each computing device 102-1 and 102-2 may also learn over time that certain gestures are typically detected in each location, such as regarding Fig.28 As further described, this can improve the interpretation of ambiguous gestures.
[0283] Contextual information associated with the location of the device
[0284] Fig.28 2800. In example environment 2800, first computing device 102-1 may learn that bedroom-related gestures are more common in bedroom 2202 (bedroom-related context 2802) and kitchen-related gestures are more common in kitchen 2302 (kitchen-related context 2804). Bedroom-related gestures may include commands to control alarms, lighting, personal care, calendar events, etc., while kitchen-related gestures may include commands to control ovens, dishwashers, stoves, timers, etc. When computing device 102 determines that an ambiguous gesture has been performed, gesture module 224 may utilize contextual information about typical commands for a room (e.g., contexts 2802, 2804) to recognize the ambiguous gesture.
[0285] In a first example, first user 104-1 is awakened by an alarm clock and attempts to perform a push-pull gesture to turn it off. Because first user 104-1 is sleepy, they perform the gesture in a tired manner that is different from the push-pull gesture taught, for example, during gesture training. First computing device 102-1 detects the gesture being performed and determines that it is a push-pull gesture (to turn off the alarm clock) or a swipe gesture (to turn on the oven), but is unable to identify it to a desired confidence level. Therefore, gesture module 224 determines that an ambiguous gesture has been performed.
[0286] Instead of prompting first user 104-1 to repeat the gesture as the alarm continues to sound, gesture module 224 uses bedroom-related context 2802 (e.g., gestures commonly performed in bedroom 2202) to identify an ambiguous gesture. In this example, first computing device 102-1 determines that the ambiguous gesture is a push-pull gesture and sends a control signal to terminate the alarm. More specifically, first computing device 102-1 determines that a push-pull gesture (to turn off the alarm) is more commonly performed in bedroom 2202 than a swipe gesture (to start an oven). This determination may be based on, for example, a history of gestures performed in bedroom 2202 that is stored (e.g., recorded) to a memory.
[0287] In a second example, second user 104-2 walks into kitchen 2302 and attempts to perform a swipe gesture to start an oven. Because second user 104-2 is walking, they perform the gesture differently than, for example, a swipe gesture performed from a stationary position. Second computing device 102-2 detects the gesture being performed and determines that it is a swipe gesture (to turn on the oven) or a push-pull gesture (to turn off an alarm), but is unable to identify it to a desired confidence level. Therefore, gesture module 224 determines that an ambiguous gesture has been performed and uses kitchen-related context 2804 (e.g., gestures commonly performed in kitchen 2302) to identify the ambiguous gesture. As in the previous example, second computing device 102-2 determines that a swipe gesture (to turn on the oven) is more commonly performed in kitchen 2302 than a push-pull gesture (to turn off an alarm). The device can proceed with starting the oven.
[0288] The techniques of the example environment 2800 may include the spatiotemporal machine learning model 802 and / or the machine learning model 700 (see Figure 8 and Figure 7 ) in which the input layer 702 additionally receives context information to improve the interpretation of gestures. Figure 26 to Figure 28 The techniques are described using gesture histories and user habits detected at one or more past times, but context information may also include real-time information, such as the status of an operation being performed at the current time (e.g., an ongoing operation).
[0289] Contextual information associated with the operation being performed
[0290] Fig.29 The example environment 2900 illustrates how the state of the operation being performed at the current time can improve the recognition of an ambiguous gesture. The techniques described in the example environment 2900 can be used by the computing device 102 (see Fig.24) or a group of computing devices 102-X forming a computing system. The first computing device 102-1 is depicted as being in the first room 304-1 (kitchen 2 302), and the second computing device 102-2 is depicted as being in the second room 304-2 (restaurant 2 304). In this example, it is assumed that the first computing device 102-1 and the second computing device 102-2 are part of a computing system. Therefore, each computing device 102-1 and 102-2 can exchange, for example, context information and / or radar signal characteristics associated with a gesture being performed by the user 104.
[0291] In example environment 2900, user 104 performs a gesture to start a timer for a kitchen oven. First radar system 108-1 of first computing device 102-1 detects the gesture, determines that the gesture is associated with an operation (to start a timer), and starts a timer in kitchen 2302. User 104 then leaves kitchen 2302 and goes to restaurant 2304 to wait for their food to cook. At a later time, user 104 wants to know if the food is finished baking, so they try to perform a known gesture on second computing device 102-2 to check the status of the timer. However, user 104 performs an ambiguous gesture, which may be a known gesture (to check the status of the timer) or another known gesture (to turn off the TV), but cannot be recognized to the desired confidence level. Specifically, gesture module 224 of second computing device 102-2 compares the radar signal characteristics (e.g., time and / or topological features) of the ambiguous gesture with one or more stored radar signal characteristics to determine that the ambiguous gesture may be any of those known gestures.
[0292] Thus, second computing device 102-2 utilizes contextual information related to the state of an operation being performed by computing device 102-1 or 102-2 of the computing system at the current time. In this example, first computing device 102-1 is currently running a timer for an oven. Second computing device 102-2 detects this ongoing operation and determines that the known gesture (to check the state of the timer) is associated with the ongoing operation of first computing device 102-1 (running the timer). Additionally, gesture module 224 determines that the other known gesture (to turn off the television) is not associated with an ongoing operation of any device of the computing system (including first computing device 102-1) because the television is not turned on. Therefore, second computing device 102-2 determines that the ambiguous gesture is most likely the first known gesture (rather than the other known gesture) and reports the state of the timer as "10 minutes remaining."
[0293] However, in some scenarios, the state of the operation at the current time may include two or more operations. When such a state occurs, the computing device 102 may need to qualitatively determine the ongoing operation by, for example, prioritizing the foreground operation over the background operation, such as regarding Fig.30 Further described.
[0294] Context information associated with foreground and background operations
[0295] Fig.30 It shows how blur gestures can be recognized based on the foreground operations and background operations being performed at the current time. In the present disclosure, foreground operations 3002 will refer to operations in which the user actively participates (e.g., interacting with the screen, providing input to the screen, displaying on the screen), and background operations 3004 will refer to operations in which the user passively participates (e.g., occurring for a duration without user input). For example, foreground operations 3002 may include phone calls, video calls, scrolling websites using tactile input, typing on a display, etc. Background operations 3004 may include music being played, timers, the operating status of an appliance (e.g., the oven is on), etc.
[0296] In example environment 3000, computing device 102 detects that user 104 has performed ambiguous gesture 3006, which may be a first known gesture or a second known gesture, but cannot be recognized to a desired confidence level. Although user 104 may have intended to perform the first known gesture (e.g., a swipe gesture to turn up the volume of a phone call), user 104 performed ambiguous gesture 3006 that is associated with radar signal characteristics that are similar to both the first known gesture and the second known gesture (e.g., a swipe gesture to stop a timer). If radar system 108 cannot determine that ambiguous gesture 3006 is the first known gesture (rather than the second gesture) based on the radar signal characteristics, radar system 108 may utilize contextual information to determine the intended gesture.
[0297] The context information of example environment 3000 includes foreground operations 3002 (e.g., a phone call with Sally) and background operations 3004 (e.g., a timer with 1:05 minutes remaining) being performed by computing device 102 at the current time. Upon detecting ambiguous gesture 3006, radar system 108 may determine that a first known gesture (e.g., to increase the volume of the phone call) is associated with foreground operation 3002, and a second known gesture (e.g., to stop the timer) is associated with background operation 3004. User 104 is actively engaged in a voice call with Sally, and is passively engaged in a timer running in the background. The device may then determine, based on this contextual information, that ambiguous gesture 3006 is most likely the first known gesture, rather than the second known gesture.
[0298] Generally speaking, context information is not limited to the status of operations being performed at the current time, and may also include operations performed in the past or scheduled to be performed in the future, such as Fig.31 Further described.
[0299] Contextual information associated with past or future actions
[0300] Fig.31 1 shows how contextual information may include past operations and / or future operations of the computing device 102. The user may routinely (e.g., daily) perform gestures associated with operations to be performed (or to be caused) by the computing device 102. The example environment 3100 depicts: (1) operations that were performed during a past time period 3102; (2) operations that are being performed at the current time 3104; and (3) operations that are scheduled or predicted to be performed in a future time period 3106. The first command involves turning off the alarm 3108 every morning, and the second command includes turning on the lights 3110 immediately thereafter (and keeping them on for a certain period of time). The computing device 102 may utilize contextual information including operations performed in the past time period 3102 and / or operations to be performed in the future time period 3106 to recognize ambiguous gestures at the current time 3104. The operations for the future time period 3106 may be scheduled or predicted by the device based on, for example, the room-related context or user habits (e.g., with reference to the user's room-related context). Figure 26 to Figure 28 ).
[0301] In the first example, the context information includes operations performed in the past period 3102. The user 104 is typically awakened by the alarm at 6:00am every morning, and performs a push-pull gesture to turn off the alarm 3108. The user 104 then performs a slide gesture to turn on the lights 3110, which remain on until bedtime. These operations (depicted in the past period 3102) have been recorded by the computing device 102 to improve gesture recognition at future times. One day, in order to catch a flight, the user 104 needs to wake up early, and sets an early alarm 3112 to be awakened at 4:00am. When awakened by the alarm, the user 104 performs a push-pull gesture to turn off the alarm, and then attempts to perform a slide gesture (at the current time 3104) to turn on the lights early (e.g., early lights 3114). However, the user 104 feels tired and accidentally performs an ambiguous gesture, which may be a slide gesture or a tap gesture (to turn on the radio).
[0302] To recognize this ambiguous gesture, the gesture module 224 can refer to contextual information about past operations to determine that the user 104 intended to perform a swipe gesture, as depicted in the example environment 3100. Specifically, the gesture module 224 can determine that: (1) every morning after the user 104 performs a push-pull gesture to turn off the alarm 3108, the lights are usually turned on; (2) the user 104 recently turned off the early alarm 3112; (3) the user 104 performed an ambiguous gesture that may be a swipe gesture or a tap gesture; and (4) the swipe gesture (to turn on the lights) is associated with a recurring past operation (e.g., contextual information), and the tap gesture (to turn on the radio) is not associated with a past operation. Because the ambiguous gesture can be associated with the past operation, the gesture module 224 determines that the ambiguous gesture is most likely a swipe gesture and turns on the lights in the bedroom. In this example, the gesture module 224 correlates the past operations (to turn off the alarm 3108 and turn on the lights 3110) to improve gesture recognition.
[0303] In the second example, the context information includes an operation to be performed (e.g., scheduled) in a future period 3106. User 104 schedules an alarm 3108 at 6:00am every morning, and programs the light 3110 to automatically turn on at 6:05am. In a typical morning, user 104 is awakened by the alarm 3108, and performs a push-pull gesture to turn it off. Unlike the aforementioned example, in the absence of a gesture, the light 3110 automatically turns on at 6:05am (as scheduled). One day, in order to catch a flight, user 104 needs to wake up early, and manually sets an early alarm 3112 to be awakened at 4:00am. However, user 104 forgets to adjust the light to automatically turn on earlier (e.g., at 4:05am). The early alarm 3112 sounds at 4:00am, but the early light 3114 is still not automatically turned on at 4:05am. The user 104 attempts to perform a swipe gesture in the dark to turn on the early morning lights 3114 but accidentally performs an ambiguous gesture, which could be a swipe gesture or a tap gesture (to turn on the radio).
[0304] In this example, gesture module 224 can refer to contextual information about future operations to determine that user 104 intended to perform a swipe gesture. Specifically, gesture module 224 can determine that: (1) the lights are scheduled to automatically turn on every morning at 6:05am; (2) user 104 performed an ambiguous gesture at 4:05am that could be a swipe gesture or a tap gesture; and (3) the swipe gesture (to turn on the lights) is associated with a scheduled future operation (e.g., turn on the lights every morning at 6:05am), and the tap gesture (to turn on the radio) is not associated with any scheduled operation. Therefore, gesture module 224 determines that the ambiguous gesture is most likely a swipe gesture and turns on the lights in the bedroom.
[0305] Although Figure 26 to Figure 31 The examples of utilize context information associated with operations performed at the current time, past time, or future time, but in some cases, this information may not be sufficient to identify an ambiguous gesture. If gesture module 224 is unable to identify an ambiguous gesture based on the context information (e.g., to a certain confidence level) (or abandons doing so based on the context information), computing device 102 may determine to perform a less destructive operation, such as regarding Fig.32 Further described.
[0306] Determination of less destructive operations
[0307] Fig.32 It is shown how an ambiguous gesture can be identified based on a less destructive operation. In this document, less destructive operations and less destructive commands can generally refer to describing an operation or a command associated with an operation being less destructive. In this way, classifying an operation as a less destructive operation can be extended to a command that directs a device to perform the operation.
[0308] In example environment 3200, user 104 is awakened by an alarm and accidentally performs a blur gesture 3202 that is intended to be a snooze gesture 3204 (e.g., a waving gesture) to reset the alarm to a later time. Because user 104 performs this gesture in a tired manner, radar signal characteristics of blur gesture 3202 are similar to those of both the snooze gesture 3204 and the dismiss gesture 3206 (e.g., a push-pull gesture) to turn off the alarm. Figure 26 to Figure 31 In the example of , the gesture module 224 utilizes context information to identify the ambiguous gesture. However, in this example, the context information (e.g., an alarm going off at the current time) may be related to the operation of both the alarm stop gesture 3204 and the shutdown gesture 3206. Therefore, the context information is not sufficient to identify this ambiguous gesture 3202. Although described with respect to a single application, the possible gestures to which the ambiguous gesture 3202 may correspond may be associated with operations that affect the same application or different applications. In this way, the context information may be more or less helpful in determining the ambiguous gesture 3202 as a specific known gesture.
[0309] Instead of or in addition to prompting the user 104 to correctly perform the stop gesture 3204, the radar system 108 can determine to perform a less destructive operation. Less destructive operations can be characterized as actions that are less damaging, less persistent, less consequential, more revocable, and the like. For example, muting a phone call can be characterized as less destructive than ending a phone call. In the example environment 3200, stopping the alarm can be characterized as less destructive than turning off the alarm. Therefore, the radar system 108 can determine that the fuzzy gesture 3202 is more likely to be a stop gesture 3204, and reset the alarm to a later time. The less destructive operation can be characterized based on preset conditions, one or more logics, a history of user behavior, one or more user inputs (e.g., preferences), and the like. The less destructive operation can be determined by various techniques described herein, including the use of machine learning models (e.g., 700, 802) or the context-based and user history-based techniques described in this document.
[0310] In various aspects, determining a less destructive operation may determine whether an operation is a temporary operation or a final operation. A temporary operation may describe an operation that is revocable and does not terminate a process instance. For example, an operation that mutes a call or delays a notification may be described as a temporary operation because it only affects the characteristics of the call or notification without terminating it. In contrast, a final operation may describe an operation that is terminating or irrevocable. For example, a final operation may terminate an instance within a computing device and prevent future operations on the instance. Thus, a final operation may prevent the system from performing a temporary operation for a particular instance after a final operation has been executed. Given the finality of the final operation, a temporary operation may be determined to be a less destructive operation than a final operation.
[0311] In addition to determining a less destructive operation based on the operation itself, or as an alternative, computing device 102 may determine a less destructive command using a command received after a gesture is performed. Specifically, computing device 102 may execute a command, and user 104 may respond by performing a gesture or providing user input to the computing device. In some instances, this gesture or user input may cause computing device 102 to execute a different command to undo the original command executed on the computing device. For example, in response to a call being terminated, user 104 may redial a phone number, or receive a return call from a previous caller. Executing a command to undo an operation (e.g., redialing the call) may provide an indication that the operation is destructive (e.g., replaying a skipped song may be easier than reopening an unsaved and terminated application).
[0312] While some destructive determinations may be based on current responses to commands executed on computing device 102, computing device 102 may utilize previous executions of commands to determine less destructive operations. For example, at a previous time, computing device 102 may have received a particular response from user 104 after executing one or more commands associated with ambiguous gesture 3202 (e.g., the gesture was not recognized as one known gesture but was recognized as two known gestures, and was therefore ambiguous but associated with two of a possible many known gestures). Based on the response, computing device 102 may determine how destructive the operation is. In one example, computing device 102 may have previously executed a command to terminate a call or application, and user 104 responded to the command by relaunching the application or call. This historical actions taken by the user in response to commands received or operations performed by computing device 102 (e.g., within a certain period of time) may provide an indication of the commands intended by the user in each scenario, thereby making the command less destructive. If computing device 102 determines that an operation is likely intended by user 104, which can be determined from previous behavior, then the operation can be considered less destructive. Therefore, relying on past detection and response can enable the system to improve the determination of less destructive operations.
[0313] When user 104 performs an action or provides a command to computing device 102 to undo a previous action, an indication of the user intervention may be stored to increase the accuracy of ambiguous gesture recognition at future occurrences. For example, when user 104 takes action to undo a command executed as a result of ambiguous gesture 3202, computing device 102 may determine that ambiguous gesture 3202 was incorrectly recognized and store this determination to improve gesture recognition at future times. Such storage may include storing radar signal characteristics of ambiguous gesture 3202 in such a manner that the characteristics are disassociated from the gesture to which ambiguous gesture 3202 was incorrectly associated.
[0314] Instead of or in addition to performing a gesture or command to undo an incorrectly performed command, user 104 may repeatedly perform blur gesture 3202 to indicate that blur gesture 3202 was incorrectly recognized. This other performance of blur gesture 3202 may be determined to be similar or identical to the first performance of blur gesture 3202. Therefore, computing device 102 may determine that user 104 is attempting to correct the incorrect recognition of blur gesture 3202 by computing device 102. By determining that user 104 is re-performing blur gesture 3202 based on gesture similarity, computing device 102 may determine that the previous recognition of the blur gesture was incorrect. Once computing device 102 determines that the previous recognition of blur gesture 3202 was incorrect, it may analyze another performance of blur gesture 3202 to determine a match with a different known gesture. Therefore, blur gesture 3202 may be recognized as a gesture different from the incorrectly recognized gesture. In various aspects, the different gesture may be one of the known gestures to which the blur gesture was originally associated.
[0315] In response to determining that the original recognition of ambiguous gesture 3202 was incorrect, computing device 102 may undo or cease execution of the command associated with the incorrectly recognized ambiguous gesture 3202. Alternatively or additionally, computing device 102 may execute the command associated with a different known gesture that is correctly associated with the ambiguous gesture. As with other revisions by computing device 102, this determination may be stored to enable computing device 102 to more accurately identify future executions of gestures as the different gesture, which may include storing characteristics of the first or subsequent executions of ambiguous gesture 3202 associated with the different gesture.
[0316] Although a less destructive operation may be determined based in whole or in part on a gesture or command received following the blur gesture 3202, the preferences of user 104 may be used to determine a less destructive command. Specifically, user 104 may store data relating to which commands or operations should be characterized as less destructive. This data may include: specific user-specific decisions relating to the relative destructiveness of each command / operation; or general data that may be used to determine a less destructive operation between any set of commands / operations. User preferences may similarly determine how blur gestures 3202 are determined (e.g., whether context, destructiveness, etc. should be given a higher weight). For example, user 104 may choose to rely more on context to determine between the possible relevance of blur gestures 3202, while another user may rely more on destructiveness to avoid accidentally performing a destructive operation.
[0317] Through the described techniques, a less destructive operation can be determined, which can enable a computing system to recognize an ambiguous gesture as a known gesture in a less disruptive or minimal manner. Thus, even if a computing device cannot accurately determine an ambiguous gesture as a known gesture, the determination of a less destructive operation can increase user satisfaction with gesture control.
[0318] Continuous online learning
[0319] Fig.33 The user is shown performing an ambiguous gesture. In example environment 3300, user 104 performs an ambiguous gesture 3302 that the user intends to be a first gesture (e.g., a known gesture). Computing device 102 detects ambiguous gesture 3302 and attempts to correlate ambiguous gesture 3302 with a known gesture. Specifically, the radar system of computing device 102 may attempt to determine one or more radar signal characteristics associated with ambiguous gesture 3302. In this case, the radar system determines that first radar signal characteristic 3304 is associated with ambiguous gesture 3302. In general, when a user performs a gesture, the radar signal characteristics of the gesture may vary between different gesture executions due to slight differences in the user (or different users) in each execution of the gesture (or in the orientation or distance relative to the radar system, etc.). Therefore, an instance of a gesture performed by user 104 may have a radar signal characteristic that is different from the stored radar signal characteristic associated with a known gesture, and renders gesture module 3306 unable to determine which gesture has been performed by user 104.
[0320] In this case, ambiguous gesture 3302 is determined to have first radar signal characteristic 3304. First radar signal characteristic 3304 may be compared to one or more stored radar signal characteristics. In various aspects, the stored radar signal characteristics may be implemented in a storage medium of computing device 102, a radar system, gesture module 3306, or an external storage medium accessible by gesture module 3306. The stored radar signal characteristics may be radar signal characteristics that are associated with the first gesture, for example, based on a previous performance of the gesture, a previous calibration, etc. Due to the differences in this example of ambiguous gesture 3302, the comparison of first radar signal characteristic 3304 with one or more stored characteristics may not effectively associate ambiguous gesture 3302 with the first gesture. For example, first radar signal characteristic 3304 may be different from the one or more stored characteristics, such that two or more gestures (or no gesture) are determined to be potentially associated with ambiguous gesture 3302.
[0321] In some implementations, gesture module 3306 may determine that ambiguous gesture 3302 is likely related to the first gesture (e.g., the first gesture has the highest correlation), but the correlation cannot be determined with a desired confidence level. In other implementations, gesture module 3306 may determine that ambiguous gesture 3302 corresponds to a plurality of stored gestures that includes the first gesture, but another gesture in the plurality of stored gestures has a higher correlation than the first gesture. In response to failing to recognize the gesture with the desired confidence level, computing device 102 may not respond to ambiguous gesture 3302 (e.g., computing device 102 does not execute a command associated with the first gesture).
[0322] When computing device 102 is unable to execute the command associated with the first gesture, user 104 typically chooses to execute the command again. Fig.33 In the example shown, user 104 performs another gesture 3308, which is detected by computing device 102. In some examples, computing device 102 may display a notification prompting user 104 to repeat the gesture. In other implementations, the notification may be communicated to user 104 in other ways (e.g., using a tactile notification or an audible notification). Similar to ambiguous gesture 3302, the radar system may determine one or more radar signal characteristics of another gesture 3308. As shown, the radar system may determine that another gesture 3308 has a particular radar signal characteristic, such as second radar signal characteristic 3310. Second radar signal characteristic 3310 may be provided to gesture module 3306, where it is compared to one or more stored radar signal characteristics, similar to first radar signal characteristic 3304. It is assumed herein that second radar signal characteristic 3310 more closely corresponds to the stored radar signal characteristic when compared to first radar signal characteristic 3304, for example, because user 104 took the time to more carefully perform another gesture 3308 in response to computing device 102 being unable to recognize ambiguous gesture 3302. It is also assumed herein that comparison of second radar signal characteristic 3310 with the stored characteristic may enable gesture module 3306 to correlate another gesture 3308 with the first gesture. By correlating another gesture 3308 with a known gesture (e.g., recognizing another gesture 3308 as the first gesture), computing device 102 may execute a command associated with the known gesture. Example commands include stopping a timer, pausing / playing media executed on the device, reacting to content displayed on the device, and other commands described herein.
[0323] Given that computing device 102 (e.g., using gesture module 3306 or 224 of radar system 108) is able to determine that another gesture 3308 is related to a first gesture stored by computing device 102, ambiguous gesture 3302 is also likely intended to be the first gesture. Therefore, computing device 102 or gesture module 3306 can compare ambiguous gesture 3302 and another gesture 3308 to determine whether the two gestures are similar. Specifically, gesture module 3306 can compare first radar signal characteristic 3304 of ambiguous gesture 3302 with second radar signal characteristic 3310 of another gesture 3308 to determine whether the two gestures are similar. For more information on some of the many approaches disclosed herein, see Figures 7 to 18 Detailed description attached. If it is determined that the two gestures are similar (e.g., to a desired confidence level, which may be lower than a desired confidence level for gesture recognition), computing device 102 may determine that ambiguous gesture 3302 is a first gesture. To improve the accuracy of detecting the performance of the first gesture at a future time, first radar signal characteristics 3304 of ambiguous gesture 3302 may be associated with the first gesture and stored by computing device 102. In doing so, computing device 102 may continually increase the accuracy of gesture recognition.
[0324] In general, other factors may be combined to determine whether blur gesture 3302 is a first gesture. For example, the amount of time that has passed between blur gesture 3302 and the execution of another gesture 3308 may be determined, and the amount of time that has passed may be used to determine whether blur gesture 3302 is a first gesture. In various aspects, a shorter period of time that has passed (e.g., two seconds or less) may indicate a higher likelihood that blur gesture 3302 is a first gesture, while a longer period of time that has passed may indicate a lower likelihood that blur gesture 3302 is a first gesture. In some implementations, the time that has passed may be a period of time during which no additional gestures are detected between the blur gesture and the another gesture. In this case, if the another gesture 3308 is performed immediately after the blur gesture, without performing additional gestures in between, blur gesture 3302 is more likely to be a first gesture. Here, user 104 repeats blur gesture 3302 after it is not recognized by computing device 102.
[0325] While some implementations may simply store radar signal characteristics and correlate them with the recognized gesture, other implementations may utilize or store other data along with the radar signal characteristics themselves. For example, a weight may be determined for one or more radar signal characteristics stored by computing device 102. When first radar signal characteristic 3304 or second radar signal characteristic 3310 is stored by computing device and correlated with a first gesture, a weight may be stored or changed to indicate a confidence level of correlation of each characteristic with the first gesture.
[0326] In addition, context or other data about the gesture may be used to determine a weight for first radar signal characteristic 3304 or second radar signal characteristic 3310. For example, a shorter amount of time that has passed between ambiguous gesture 3302 and the other gesture 3308 may indicate that ambiguous gesture 3302 is likely to be the first gesture. Thus, a weight value for first radar signal characteristic 3304 or second radar signal characteristic 3310 may indicate a higher confidence in the correlation of first radar signal characteristic 3304 or second radar signal characteristic 3310 with the first gesture. When a larger amount of time has passed between ambiguous gesture 3302 and the other gesture 3308, ambiguous gesture 3302 may be less likely to be the first gesture. Thus, a weight value for first radar signal characteristic 3304 or second radar signal characteristic 3310 may indicate a lower confidence in the correlation of first radar signal characteristic 3304 or second radar signal characteristic 3310 with the first gesture. In some instances, storing the weight with first radar signal characteristic 3304 and second radar signal characteristic 3310 may improve future detection of the first gesture. Machine learning can also or instead be used to determine how to Figures 7 to 10 The various timing, context, and radar signal characteristic similarities described elsewhere in this paper are weighted.
[0327] When associating blur gesture 3302 with the first gesture, context information may be used or stored. For example, computing device 102 may determine context information about the performance of blur gesture 3302 (e.g., context information when blur gesture 3302 was performed). Context information may include, for example, the location of user 104 relative to computing device 102 or the orientation of user 104, or information in the document (e.g., information about the context of the user 104). Figures 25 to 31 ) or any other type of contextual information described in the foregoing description. The contextual information may include non-radar signal characteristics of the ambiguous gesture 3302 sensed during the performance of the ambiguous gesture. As non-limiting examples, other non-radar sensors may include ultrasonic detectors, cameras, ambient light sensors, pressure sensors, barometers, microphones, or biometric sensors, as well as other sensors described herein. The contextual information may be stored with the first radar signal characteristic 3304 or the second radar signal characteristic 3310 to enable computing device 102 to more accurately identify the first gesture in the future. For example, computing device 102 may store characteristics of the first gesture in different orientations or positions of user 104 relative to computing device 102. In this way, a more accurate comparison may be made by comparing the gesture with appropriate characteristics of the current context in which the gesture is performed. Additionally or alternatively, the contextual information may be used to adjust the weight values stored with the radar signal characteristics.
[0328] Computing device 102 may also compare ambiguous gesture 3302 to another gesture 3308 to determine whether the gestures are similar, and therefore whether ambiguous gesture 3302 may be a first gesture. For example, gesture module 3306 may compare first radar signal characteristic 3304 to second radar signal characteristic 3310 and determine that first radar signal characteristic 3304 has a correlation with second radar signal characteristic 3310 that is above a confidence threshold. While this threshold may not be sufficient to identify a gesture, it is sufficient to indicate a reasonable probability that the gestures are related (e.g., a 40, 50, or 60 percent likelihood, compared to a threshold of 80, 90, or 95 percent for identifying gestures). Upon determining that first radar signal characteristic 3304 and second radar signal characteristic 3310 are correlated, computing device 102 determines that ambiguous gesture 3302 is a first gesture.
[0329] In addition to gestures, computing device 102 may also use non-gesture commands to relate the first gesture. For example, after identifying the other gesture 3308 (e.g., relating the other gesture 3308 to the first gesture), computing device 102 may inquire whether ambiguous gesture 3302 is a command associated with the first gesture. User 104 may respond without using gestures (e.g., through touch or voice) to confirm the intended gesture that ambiguous gesture 3302 was intended to convey. In other examples, without being prompted by computing device 102, the user may provide feedback to computing device 102 through non-gesture commands to help identify ambiguous gesture 3302.
[0330] Although this example is described with respect to a single computing device 102, it should be noted that continuous online learning can be performed using multiple computing devices. For example, the blur gesture 3302 can be detected at a first computing device, and the other gesture 3308 can be detected on a second computing device. The first computing device and the second computing device can be configured to communicate information across a communication network. In this way, the blur gesture 3302 can be detected by the first computing device in a first area adjacent to the first computing device, and the other gesture 3308 can be detected by the second computing device in a second area adjacent to the second computing device. Therefore, these computing devices can take advantage of continuous online learning even if they are located in different areas (e.g., different rooms in a house).
[0331] In addition, the other gesture may not be recognized using the radar system, but still provides useful information for correlating the radar signal characteristics of the ambiguous gesture with the first gesture. Assume that computing device 102 does not recognize the ambiguous gesture as a first gesture. User 104 may perform an additional gesture that may be detected by computing device 102 using another sensor (examples listed elsewhere herein) that is not associated with the radar system (and recognized as a first gesture in another non-radar manner). Based on the indication of the additional gesture received at the other sensor, computing device 102 may determine that the additional gesture is the first gesture. Similar to the continuous online learning described above, computing device 102 may determine that ambiguous gesture 3302 is related to the first gesture based on the additional gesture being the first gesture. Therefore, computing device 102 may store the radar signal characteristics of the ambiguous gesture with the first gesture, thereby achieving improved recognition.
[0332] It should be noted that various forms of gesture recognition may be used to recognize the performed gesture. For example, there are implementations that may use unsegmented recognition to recognize ambiguous gestures. Specifically, unsegmented gesture recognition may be implemented without prior knowledge or a wake-up trigger event indicating that the user is about to perform an ambiguous gesture. However, regardless of the implementation, continuous online learning may enable computing device 102 to continuously improve the accuracy of gesture recognition, even for the most difficult gestures to recognize.
[0333] In addition, while the terms "continuous" and "continuously" are used to describe continuous online learning, it should be noted that these techniques are not required to be used at all times or forever, but rather, these techniques can operate to continuously and / or incrementally improve future gesture recognition when used. The term "online" is also used herein to describe a way in which these techniques "learn" to better recognize and / or detect gestures, etc. This "online" term is intended to express that these techniques learn how to better detect or recognize gestures as part of, accompanying, or through a process of detecting and / or recognizing gestures. In contrast, explicit training procedures (e.g., where a device trains a user to perform gestures in a specific manner) are not "online" learning. Therefore, these techniques for continuous online learning can improve gesture recognition without the need for separate training outside of normal user interaction with the device. Separate training can be used in conjunction with or before these techniques, but this is not required by these techniques.
[0334] Online learning based on user input
[0335] Fig.34An example of online learning based on user input to improve ambiguous gesture recognition is shown. In example environment 3400, user 104 is within the field of view of computing device 102. Computing device 102 may include radar system 108, which provides radar data to gesture module 224. In the example shown, user 104 performs an ambiguous gesture 3402, which is identified as having one or more specific radar signal characteristics. Gesture module 224 attempts to identify ambiguous gesture 3402 as a known gesture (e.g., through unsegmented detection in the absence of a wake-up trigger event indicating that user 104 is about to perform a gesture), which may require that the gesture be associated with a known gesture within a predetermined confidence threshold criterion.
[0336] In this example, the confidence threshold criteria are not met, and therefore ambiguous gesture 3402 cannot be identified as a specific gesture, but is instead associated with one or more known gestures (e.g., first gesture 3404 and second gesture 3406). Specifically, gesture module 224 can compare one or more radar signal characteristics of ambiguous gesture 3402 to stored characteristics associated with one or more known characteristics. If a correlation is determined between the one or more radar signal characteristics of the ambiguous gesture and the stored gestures, ambiguous gesture 3402 can be associated with the stored gestures (see for how to perform the correlation). Figures 7 to 17 , Fig.25 The ambiguous gesture may be associated with a plurality of stored gestures due to similarity to each of the plurality of stored gestures, and the gesture module 224 may not be able to identify the ambiguous gesture as a specific known gesture. The ambiguous gesture may alternatively be associated with only one stored gesture, but not identified as sufficient to meet the predetermined confidence threshold criteria.
[0337] Generally speaking, if computing device 102 cannot recognize ambiguous gesture 3402 as one of the known gestures, ambiguous gesture 3402 may not cause computing device 102 to execute a command. Fig.34 , blur gesture 3402 is associated with first gesture 3404 corresponding to play music command 3408, and with second gesture 3406 corresponding to call dad command 3410. Given that computing device 102 cannot recognize blur gesture 3402 as one of the two recognized gestures with corresponding commands, computing device 102 may not execute either command. Alternatively, computing device 102 may remain idle until a command is requested using a gesture or another form of input.
[0338] In some cases, computing device 102 may provide an indication that ambiguous gesture 3402 was not recognized, such as through a notification to user 104 using a display or speaker. However, in some instances, user 104 may execute a command or request execution of a command without being prompted (e.g., requested) by computing device 102. When computing device 102 does not execute a desired command in response to ambiguous gesture 3402, user 104 may choose to execute the command themselves or initiate the command through a different input type (e.g., non-radar input, touch input through a touch-sensitive display or keyboard, or audio input through a speech recognition system).
[0339] As shown, user 104 requests computing device 102 to “play today's top hits” using voice command 3412. Thus, computing device 102 may receive voice command 3412 from user 104 and begin playing music (e.g., using an application stored on the device or through a network connection to a media service). In general, user input may change the operating state of computing device 102 or any other connected device. For example, computing device 102 may begin playing music, maintain a timer, or perform any other operation. Any of these changes in operating state may indicate that user 104 has executed a command or requested execution of a command.
[0340] In example environment 3400, user 104 performs voice command 3412 after computing device 102 fails to respond to blur gesture 3402, and therefore, user 104 may have likely intended to cause computing device 102 to perform the same action as voice command 3412 through blur gesture 3402. Therefore, computing device 102 may determine whether the command performed through or in response to voice command 3412 is the same as or similar to the commands corresponding to gestures (e.g., first gesture 3404 and second gesture 3406) that were associated with blur gesture 3402. For example, computing device 102 may determine that voice command 3412 is the same as or similar to play music command 3408 corresponding to first gesture 3404 because both commands cause the computing device to play music.
[0341] After determining that voice command 3412 and the command associated with the associated gesture are the same, computing device 102 may store one or more radar signal characteristics of ambiguous gesture 3402 in association with first gesture 3404 to enable gesture module 224 to better identify future executions of first gesture 3404. Thus, gesture module 224 may utilize a spatiotemporal machine learning model associated with one or more convolutional neural networks to improve detection of ambiguous gesture 3402. In various aspects, computing device 102 may prompt user 104 to confirm that ambiguous gesture 3402 is the known gesture before storing the radar signal characteristics in association with the known gesture (e.g., by providing a notification on a display and accepting confirmation from user 104). In this way, computing device 102 may eliminate erroneous associations of radar signal characteristics with known gestures. In some examples, the radar signal characteristics associated with ambiguous gesture 3402 may be associated with a stored specific radar signal characteristic of a known gesture to enable gesture module 224 to increase the confidence (e.g., weight) of the association between the specific radar signal characteristic and the known gesture.
[0342] In some implementations, computing device 102 may determine a time period between performance of ambiguous gesture 3402 and performance of, or requesting performance of, a command (e.g., voice command 3412). This time period may be used to determine whether the command performed as a result of the user input is the same as a command associated with one of the known gestures to which ambiguous gesture 3402 was associated. Additionally or alternatively, this time period may be used to determine a weighting of radar signal characteristics associated with known gestures, such as Fig.33 As described.
[0343] In various aspects, computing device 102 may determine whether a different command (e.g., Fig.33 ). The different command may be different from the commands corresponding to the known gestures (e.g., first gesture 3404 and second gesture 3406) associated with blur gesture 3402 (e.g., play music command 3408 and call dad command 3410). Therefore, if it is determined that a different command is not executed or requested during a time period between execution of blur gesture 3402 and user input indicating execution or requesting execution of voice command 3412, one or more radar signal characteristics of blur gesture 3402 may be stored.
[0344] It should be noted that voice command 3412 is just one example of user input that may be used for online learning, and these techniques may instead or additionally utilize other forms of user input, such as touch commands (e.g., user 104 touching a touch-sensitive display of computing device 102, user 104 using their voice to control computing device 102 through a voice recognition system of computing device 102, or user 104 typing commands on a physical or digital keyboard). It is also important to note that user input may not be limited to being received at computing device 102. For example, user input may be received at any other device (e.g., other smart home devices) with which computing device 102 may communicate. As non-limiting examples, user input on another device may include setting a timer on a smart home appliance, activating a smart light switch or smart door lock, or adjusting a smart thermostat.
[0345] In general, the gesture module 224 can continue to improve the gesture recognition of the user's unique gestures, thereby improving the user's confidence in the gesture recognition and increasing user satisfaction. The techniques for online learning can enable the gesture module 224 to be trained without using gesture training events for segmented teaching. For example, the gesture module 224 can improve gesture recognition without the computing device 102 explicitly teaching the user 104 gestures or requesting the user 104 to perform gestures. Therefore, these techniques can eliminate the need for the user 104 to bear additional burdens to train the computing device 102.
[0346] Online Learning
[0347] Fig.35 Techniques for online learning of new gestures for a radar-enabled computing device are shown. In environment 3500, user 104 is within the field of view of computing device 102. User 104 performs gesture 3502, which is detected by radar system 108 of computing device 102, and radar signal characteristics are determined from gesture 3502 by gesture module 224 (e.g., as described above). In this example, user 104 rotates their hand to mimic the action of pouring coffee to indicate that user 104 wants computing device 102 to start the coffee machine. Upon detecting gesture 3502, computing device 102 can determine that gesture 3502 is an intentional motion performed by user 104, rather than background motion unrelated to the specific gesture intended for computing device 102.
[0348] The radar signal characteristic of gesture 3502 is compared to one or more stored radar signal characteristics associated with one or more known gestures. In environment 3500, the comparison is effective to determine a lack of correlation between gesture 3502 and the one or more known gestures. Specifically, the comparison may not be effective to correlate gesture 3502 or the associated radar signal characteristic with the one or more known gestures at a desired confidence level (e.g., at a low confidence threshold rather than a high confidence threshold). For example, the associated radar signal characteristic may not be sufficient to meet confidence threshold criteria associated with any stored radar signal characteristics.
[0349] Given the lack of correlation between gesture 3502 and the one or more known gestures, computing device 102 is unable to recognize and respond to gesture 3502. As shown, user 104 provides command 3504 (e.g., a voice command) indicating that user 104 wants computing device 102 to start the coffee machine. In some cases, computing device 102 may determine that command 3504 is not the same as or similar to one or more known commands corresponding to the one or more known gestures. Although command 3504 is shown as user 104 using their voice to ask computing device 102 to "start the coffee machine", other commands may be included, such as those entered through a touch-sensitive display or keyboard of computing device 102 or another connected device, or by a user physically executing a command (e.g., physically starting the coffee machine).
[0350] In response to determining the lack of correlation between gesture 3502 and the one or more known gestures and receiving command 3504, computing device 102 may determine that gesture 3502 is a new gesture 3506 that has not yet been learned by computing device 102 or associated with a particular command. In some cases, computing device 102 may provide some indication of such determination to user 104, such as by displaying a message asking the user whether the gesture is a new gesture (e.g., a gesture that has not yet been taught to computing device 102 or assigned to a particular command). The user may respond by indicating that the gesture is a new gesture.
[0351] Computing device 102 may determine that gesture 3502 is new gesture 3506, and then store the radar signal characteristics associated with gesture 3502 to enable computing device 102 to recognize performance of gesture 3502 in the future. Moreover, computing device 102 associates new gesture 3506 with a command to start coffee maker 3508. In this way, computing device 102 may seamlessly learn new gestures at the convenience of user 104. Moreover, by doing so, computing device 102 may not be limited to learning new gestures during gesture training.
[0352] Although specific implementations are shown in environment 3500, it should be noted that other implementations should be considered within the scope of the present disclosure. For example, after performing gesture 3502, it may not be necessary to execute command 3504. In general, command 3504 can be executed in close time to gesture 3502, such as within a predetermined time limit of two seconds, five seconds, ten seconds, or thirty seconds. Similarly, command 3504 can be executed before, after, or simultaneously with the execution of gesture 3502. As a specific example, user 104 can execute command 3504 before the execution of gesture 3502, or user 104 can execute command 3504 simultaneously with the execution of gesture 3502 (e.g., by speaking command 3504 while performing gesture 3502). In some implementations, command 3504 will be received by computing device 102 if computing device 102 does not detect an intermediate gesture or another command initiated by user 104.
[0353] In some examples, computing device 102 may learn new gestures independently for different users. For example, when computing device 102 stores the radar signal characteristics of gesture 3502, computing device 102 may associate the radar signal characteristics and new gesture 3506 with a specific user. In this way, identifying gesture 3502 in the future may include distinguishing the associated user and identifying the execution of the gesture. To achieve this determination, computing device 102 may determine another radar signal characteristic associated with the user's presence, which may be used to determine the user as a registered user. When determining a new gesture and storing its radar signal characteristics, the other radar signal characteristic associated with the user's presence may also be stored to enable the computing device to identify the user at a future time. By independently learning new gestures associated with each user, computing device 102 may associate different gestures with different commands based on the user. For example, computing device 102 determines that gesture 3502 is a new gesture for user 104, but is a known gesture or an unknown gesture for another user. In this way, computing device 102 may implement user-customizable gesture control technology.
[0354] Context-sensitive sensor configuration
[0355] Fig.36Example techniques for configuring a primary sensor for a computing device are shown. Example environment 3600 shows two different computing devices: computing device 102-1 implemented within a kitchen 3602 and computing device 102-2 implemented within an office 3604. Within kitchen 3602, user 104-1 is within a field of view of computing device 102-1. Similarly, user 104-2 is shown within a field of view of computing device 102-2. Computing device 102-1 may determine a first context 3606 associated with a condition for performing a gesture by user 104-1, and computing device 102-2 may determine a second context 3608 associated with a condition for performing a gesture by user 104-2.
[0356] User 104-1 may perform a gesture in an area within kitchen 3602. Computing device 102-1 may detect the gesture using one or more sensors configured to measure activity in the area. The manner in which each sensor is used or the accuracy or precision with which each sensor detects and recognizes the gesture may vary based on the context associated with the area in which the gesture is performed. For example, the computing device may determine the context associated with the area in which the user is performing the gesture. The context may be determined from any number of details about the area in which the gesture is performed. For example, the context may be determined based on reference to the area in which the gesture is performed. Figures 25 to 31 One or more aspects of the described context determination determine context.
[0357] In some implementations, the context may be determined based on environmental conditions in the area where the gesture is performed. For example, the computing device may determine conditions related to light, sound, interference, placement of objects, or movement in the area. Some sensors may be more capable of recognizing gestures under a particular set of conditions. For example, an optical sensor (e.g., a camera) may be more capable of recognizing gestures in a well-lit environment. Compared to an optical sensor, a radar sensor may experience less performance degradation in a poorly lit environment. Therefore, in an environment with poor lighting conditions, a camera may be less likely to be successfully used to recognize gestures.
[0358] In another example, a radar sensor may not be able to effectively recognize gestures when there is a lot of movement within the environment. For example, the radar sensor may not be able to distinguish between different movements within the environment. In a very active environment with a lot of movement, a non-radar sensor (e.g., a camera) may be more able to recognize gestures.
[0359] In another example, a radar sensor may not be able to recognize a gesture when there is a lot of interference in the environment. For example, multiple radar devices implemented in close proximity to each other may cause interference due to cross signals between the devices. When the computing device attempts to recognize a gesture, this interference may increase the noise floor in the radar receive signal and reduce the ability of the radar sensor to recognize the gesture.
[0360] Microphones can be used to supplement gesture recognition (e.g., to determine additional details about the performance of a gesture). However, in high noise environments, such audio sensors may be less effective in providing contextual details about gestures. Therefore, audio sensors may be less able to supplement gesture recognition in high noise environments.
[0361] The context may alternatively or additionally be based on the location where the computing device resides. For example, the computing device may determine that it is located in a kitchen. Based on this determination, the computing device may be able to determine which particular user is likely to perform a gesture, the type of gesture that is likely to be performed, or the environmental conditions that are likely to exist. These details may be used to determine which sensors are likely to be most capable of recognizing the gesture. For example, in a kitchen, a user may be most likely to perform gestures related to cooking (e.g., controlling an appliance, starting a timer, or searching for a recipe). The computing device may determine which sensors are most capable of detecting these particular gestures (e.g., based on past performance and recognition).
[0362] The location may provide information about environmental conditions that are likely to exist in the area where a gesture is performed. For example, a computing device located in a bedroom may be likely to be in low light conditions, a computing device located in a living room may be likely to be in a high noise or high movement environment, etc. Additionally or alternatively, the location of the computing device may provide information about which users are likely to perform gestures or which gestures are likely to be performed. For example, a particular user or a particular group of users may be more likely to be present in a particular area of a house, office, or another environment.
[0363] Instead of or in addition to using location to determine which user is performing a gesture, user detection (e.g., as described in this document) can be used to identify a specific user performing a gesture. As shown, computing device 102-1 can determine that a specific user 104-1 is performing a gesture. In doing so, first context 3606 may include information related to user 104-1, such as data related to hand size, gesture execution rate, gesture execution clarity, etc. These details can be used to determine the ability of a specific sensor to recognize a gesture performed by user 104-1. For example, a radar sensor may be less able to recognize gestures performed by a user with a smaller hand. As another example, certain sensors may be more able to accurately recognize gestures from a specific user based on previous execution and detection of gestures performed by the user. In some implementations, a user may be most likely to perform a specific gesture (e.g., based on past execution and detection of the user's gesture), and certain sensors may be more able to recognize these specific gestures.
[0364] Context may be determined based on the time of day at which the gesture is performed. This information may provide data related to any other context-related data. For example, at night, the computing device may be more likely to be in low light conditions, the gesture may be more likely to be related to an evening event (e.g., cooking), or a particular group of users may be more likely to be present. Similarly, other times of day may have other characteristics that may be useful for context determination.
[0365] The context may relate to foreground or background operations of a computing device. The operation of the device may provide useful indications of gestures that are likely to be performed or users that are likely to perform gestures. Recent operations to adjust conditions within the environment (e.g., lighting, audio, etc.) may provide indications of environmental conditions in the area where the gesture was performed. This data may be used to determine context within an area.
[0366] It should be noted that these are just some examples of ways to determine context and how the determined context can be used to determine sensor capabilities. Therefore, other examples of context determination can be used to determine sensor capabilities, such as Figures 25 to 31 The correlation between data and context or between context and sensor capabilities can be determined by any number of techniques, including any of the machine learning techniques described in this document.
[0367] Turning to the illustrated example of kitchen 3602, computing device 102-1 may determine first context 3606 for a particular area. First context 3606 may relate to the location at which the gesture is performed (e.g., kitchen 3602), the particular user performing the gesture (e.g., user 104-1), environmental conditions of the environment at which the gesture is performed, the orientation or position of computing device 102-1 relative to user 104-1 and vice versa (e.g., whether user 104-1 is within the field of view of a particular sensor), the time of day, background operations or foreground operations of the computing device (e.g., as described with respect to FIG. 36). Fig.30 Based on first context 3606, computing device 102-1 may determine the ability of one or more sensors of computing device 102-1 to provide data useful for recognizing a gesture. The computing device may include multiple sensors, such as radar sensor 3610 and a non-radar sensor (e.g., optical sensor 3612) or multiple sensors of the same type. Fig.36 A specific configuration of sensors is described, but other sensors may also be implemented and are listed elsewhere in this document.
[0368] Computing device 102-1 may determine the ability of a first sensor to recognize a gesture performed in a particular area. For example, computing device 102-1 may determine that radar sensor 3610 is capable of recognizing a gesture based on a first context. The ability of a sensor to recognize a gesture may be on a progressive scale so that different sensor capabilities may be compared to one another. The ability may be determined from a prior correlation between context and gesture recognition (e.g., from gesture recognition in an environment with a particular context). In addition to determining the ability of a first sensor to recognize a gesture performed in an area, the ability of a second sensor to recognize the gesture may also be determined.
[0369] Once the capabilities of one or more sensors are determined, the capabilities may be compared to determine which sensor is more capable of recognizing the performance of a gesture within the area. The more capable sensor may be determined as a primary sensor, and the computing device may be configured to use the primary sensor in preference to other sensors. In various aspects, the computing device may prioritize the use of the primary sensor by giving greater weight to data collected by the primary sensor. The computing device may be configured to save power by configuring the device with a primary sensor, by reducing power supplied to other non-primary sensors and, in some cases, increasing power supplied to the primary sensor.
[0370] In the particular example shown, radar sensor 3610 is configured as a primary sensor. In various aspects, radar sensor 3610 is determined to be a more capable or most capable sensor due to low light conditions that may result in an area having an insufficient amount of light to illuminate the area and allow an optical sensor to recognize a gesture. In some implementations, radar sensor 3610 may be more capable of providing data to recognize gestures based on particular preferences of user 104-1, the types of gestures most commonly performed in the area, or the orientation of radar sensor 3610 and other sensors relative to the area or to each other. In general, it should be recognized that the primary sensor may be determined based on any determined context or any correlation between a particular context and the particular capabilities of a sensor for recognizing gestures in that context.
[0371] As a result of configuring radar sensor 3610 as a primary sensor for computing device 102-1, data collected by radar sensor 3610 may be used in preference to data collected by other sensors. For example, data collected by radar sensor 3610 may have a greater weight during gesture detection or recognition than data collected by other sensors. Additionally or alternatively, a primary sensor may be configured for computing device 102-1 by changing a setting of computing device 102-1 during or after detection of a gesture, such as configuring a primary sensor different from radar sensor 3610.
[0372] As with first context 3606 determined by computing device 102-1, computing device 102-2 may determine second context 3608 related to the environment in which user 104-2 performed the gesture. Second context 3608 may be used to determine an ability of one or more sensors of computing device 102-2 to detect a gesture performed by user 104-2 within a particular area of office 3604. The capabilities of the one or more sensors may be compared to determine a most capable sensor, and the most capable sensor may be configured as a primary sensor in computing device 102-2.
[0373] As shown, optical sensor 3612 is determined to be the primary sensor. In various aspects, configuring the primary sensor for computing device 102-2 can be performed before, during, or after recognition of the gesture. In response to configuring the computing device, the user can be notified that the computing device has been configured with the primary sensor. For example, the computing device can provide a notification to the user, which can include a tone, light, text, interface notification, or vibration.
[0374] In various aspects, optical sensor 3612 may be determined to be the primary sensor because second context 3608 indicates that the area where the gesture is being performed is well-lit, user 104-2 prefers optical sensor 3612 as the primary sensor, fine resolution is required for gesture recognition based on environmental conditions or user characteristics, there is a lot of radar interference, a certain type of gesture is most likely to be performed in the area at the current time or in the future, or there are any other contexts that may favor optical sensor 3612. Computing device 102-2 may be configured such that optical sensor 3612 is used in preference to other sensors within computing device 102-2 to improve accuracy of gesture detection and / or recognition.
[0375] Through the described techniques, a computing device can be configured to optimally detect gestures based on the context in which the gesture is performed. In doing so, the computing device can adapt to each environment to accurately detect and recognize gestures, even in environments that pose great difficulties for gesture recognition.
[0376] Detecting User Engagement
[0377] Fig.37An example environment in which techniques for detecting user engagement with an interactive device may be performed is shown. Although shown as (and referred to as) a computing device 102, the interactive device may be another device that is coupled to the computing device 102 and implements a user interface to enable interaction with the user 104. In some instances, it may be beneficial for the computing device 102 to determine whether the user 104 is likely to interact with the computing device 102. For example, when an interaction between the user 104 and the computing device 102 is about to occur, the computing device 102 may display information that may be relevant to the user 104, such as time information, device status information, notifications, etc. The computing device 102 may determine the likelihood that the user 104 will interact with the device by determining the current or projected (e.g., near-future in time) engagement of the user with the computing device 102.
[0378] Computing device 102 may utilize information about user 104 or environment 3700 to determine the user's engagement with the device. Computing device 102 may utilize a radar system (e.g., radar system 108) or any other sensor system of the device to determine information about user 104 or environment 3700 that may be helpful in detecting the user's engagement. Computing device 102 may transmit radar signals and receive reflections of these radar signals reflected by user 104 or the surrounding environment. These received signals may be processed to determine characteristics of user 104 or environment 3700 that may be used to determine the user's engagement with computing device 102. For example, computing device 102 may determine a user's proximity 3702 relative to computing device 102, and the determined proximity 3702 may provide an indication of the user's engagement with the device.
[0379] In some examples, computing device 102 may determine that user 104 is more likely to engage with a device when the user is within a particular proximity of the device (e.g., less than or equal to a proximity threshold criterion). For example, computing device 102 or a sensor system (e.g., a radar system, a touch display, voice commands, etc.) that may be used to interact with user 104 may have a particular range within which it may effectively detect or identify user interaction with the device. Thus, when the user is further away from the device, user 104 may be less likely to intend to interact with the device. In this way, proximity may be used to determine a user's engagement with the device, such as determining a higher user engagement when the user is closer to the device.
[0380] However, proximity alone may not be sufficient to accurately determine user engagement with a device. In some scenarios, user 104 may happen to be proximate to computing device 102 without any intent to interact with computing device 102, such as when interacting with another device or person located near computing device 102 or when walking past computing device 102. Therefore, when proximity is used as the only factor to determine user engagement, computing device 102 may not accurately detect user engagement, thereby causing the computing device to perform suboptimally.
[0381] To increase the accuracy of the detection of user engagement, any number of factors may be used to determine user engagement, including, for example, the expected proximity of the user relative to computing device 102 or the user's physical orientation relative to computing device 102. The expected proximity of the user relative to computing device 102 may include determining a rate at which the proximity of user 104 to computing device 102 changes. In some examples, determining the expected proximity may include determining a path 3704 indicating a direction in which user 104 is moving. Path 3704 may be represented as a vector indicating a direction and, in some cases, a magnitude (e.g., rate) of the user's movement relative to the computing device. By determining the path of user 104, computing device 102 may more accurately determine whether user 104 is intending to engage with computing device 102, as opposed to simple proximity changes.
[0382] When user 104 is moving toward computing device 102, user 104 may be likely intending to interact with the device and, therefore, moving toward the device to get within the field of view of the device or to view content displayed on the device. However, other instances of user 104 approaching the device may not be related to the user's intent to interact with the device, such as when user 104 is moving to pick up an object near computing device 102. Therefore, it may be beneficial to determine user 104's path 3704, which may be used to determine where user 104 is likely moving. If path 3704 points toward computing device 102 (e.g., as shown in FIG. 1 ), user 104 may be moving toward computing device 102. Fig.37 If path 3704 is pointing toward the device (e.g., directly toward or nearly directly toward the device, as shown), user 104 may be more likely to be engaging with the device and likely to be interacting with the device. However, if path 3704 is not pointing toward the device (e.g., moving away from the device, or moving toward but not directly toward the device), then the user's proximity to the device may be more likely to be coincidental and not indicative of user engagement with the device (e.g., the user is simply moving past the computing device).
[0383] The path 3704 of the user 104 may be determined in a variety of ways, such as using subsequent position measurements or Doppler measurements. In some cases, the path 3704 may be determined by comparing the current location of the user 104 to one or more historical locations of the user 104. The path 3704 may be determined from one or more historical movements of the user 104, which may include a direction or speed. The current direction or velocity (e.g., speed when combined) may be compared to the historical movements to determine an accurate representation of the user's path 3704 toward or away from the computing device 102. These historical movements may be immediately prior to the user's current location (e.g., a previous portion of the user's path 3704), or may be recorded from previous movements, such as when the user walked past the computing device 102 the day before. This type of history may be used by building a machine learning model as described below or by heuristic rules or other means. Thus, if a user has traveled a path in the past and has repeatedly not interacted with (or has interacted with) computing device 102 while following that historical path, this information can be used to determine a low (or high) probability of an intent by user 104 to engage or interact with computing device 102.
[0384] In addition to or in lieu of proximity 3702 or estimated proximity, computing device 102 may determine a body orientation 3706 of user 104 to determine the user's engagement with computing device 102. In one example, body orientation 3706 may be based on a facial profile of user 104. A radar system or other sensor system may be used to determine where the user's face is pointed (e.g., whether the facial profile is facing toward the device or away from the device). Body orientation 3706 may be based on a user's line of sight (e.g., determining whether the user is looking toward or away from computing device 102).
[0385] Body orientation 3706 may be determined from the user's body positioning, such as the orientation of user 104's torso or head toward or away from computing device 102. A radar system or other sensor system may collect information about the user (e.g., radar reception signals) to determine the body orientation of user 104, for example, by collecting radar data to determine whether a frontal profile (e.g., a flat surface such as the abdomen, chest, torso, shoulders, etc.) of user 104 is facing computing device 102. Body orientation 3706 may be used to determine a general direction of focus of the user's body, which may be toward or away from computing device 102.
[0386] Body orientation 3706 may also or alternatively include gesture recognition of user 104. For example, computing device 102 may determine whether the user is pointing or reaching toward computing device 102, which may indicate the user's intent to engage or interact with the device. Gestures may include any number of gestures stored or not stored by computing device 102. In some instances, when computing device 102 detects or recognizes a gesture performed by user 104, user 104 may be likely to engage with the device. Computing device 102 may determine whether the gesture is a known gesture (e.g., a stored gesture). If the gesture corresponds to a known gesture, the user may be likely to engage with the device (e.g., it is not an unrelated user movement that is not a "gesture"). Computing device 102 may determine the direction of these gestures to determine whether these gestures are likely intended for computing device 102. If the gesture is directed toward computing device 102, it may be determined that user 104 has a higher degree of engagement with computing device 102.
[0387] Computing device 102 may use any combination of factors (e.g., including one or all) to estimate the user's engagement with the device or the expected (e.g., future engagement close in time). The factors may be weighted differently so that each factor can have a greater or lesser impact on the determination of the user's engagement. In some instances, computing device 102 may utilize at least two factors to estimate the user's engagement with computing device 102. Computing device 102 may utilize a machine learning model (e.g., 700, 802) based on any machine learning technique including those described herein. The machine learning may be supervised, such that user 104 interacts with computing device 102 (e.g., through touch input, gesture input, voice command, or any other input method) to confirm the user's proximity, expected proximity, or body orientation relative to the device, or an association between any of these determinations and the user's engagement.
[0388] Computing device 102 may further estimate the user's engagement based on the direction of computing device 102 relative to user 104. For example, if the field of view of a display or sensor system (e.g., a radar system) of computing device 102 is pointed toward user 104, user 104 may be likely to engage with the device. However, if the display is facing away from user 104 or the user is outside the field of view of the device's sensor system, user 104 may be less likely to engage with the device. In some instances, the direction of computing device 102 may be used in conjunction with user proximity 3702, estimated proximity, or body orientation 3706 to determine user engagement. For example, if computing device 102 determines that user 104 is moving toward computing device 102 (e.g., within a sufficiently close proximity), it may be determined that user 104 has a high level of engagement with the device. If computing device 102 is pointed in a particular direction and user 104 is opposite the device and oriented in opposite directions (e.g., such that computing device 102 and user 104 are oriented toward each other), user 104 may be likely to engage with the device. Alternatively, if user 104 is moving away from the directional orientation of the display, computing device 102 may determine that user 104 has a low level of engagement with computing device 102 .
[0389] The user's engagement level may be used to determine appropriate settings for configuring computing device 102. For example, if user 104 is determined to be engaged with the device (e.g., a high engagement level is determined), content useful to user 104 may be displayed on the device. Similarly, if user 104 is engaged and may therefore likely interact with the device, power usage of one or more sensor systems may be increased to better detect, identify, or interpret the user's interaction with computing device 102. The computing device may determine the identity of user 104 to determine appropriate settings for each engagement level. For example, when an authorized or identified user is engaged with the device, the computing device may lower the privacy settings of the device to make applications or content of the device more available to user 104 (e.g., Fig. 20 However, if it is determined that an unauthorized user is engaging with the device, the privacy setting may be changed to a more restrictive setting to prevent sensitive content from being disclosed to unauthorized or unidentified users (e.g., Fig. 20 In other cases, the privacy setting of computing device 102 may be adjusted based on user engagement without distinguishing between users 104.
[0390] When it is determined that the user is not engaged with the device (e.g., determining low engagement), the computing device can be similarly adjusted to be configured based on the user's low engagement level. For example, the sensor system of the computing device can be turned off or operated in a low power mode. Similarly, the display of the computing device 102 can be dimmed or turned off to save power for the device. When the user 104 is not engaged with the device, the device resources can be redirected to other tasks, such as maintenance tasks (e.g., updating software or firmware). In some cases, the privacy settings of the device can be increased to reduce the possibility of sensitive content being released to unauthorized users. It should be understood that although specific examples are described, other settings can be adjusted based on the user's engagement. Similarly, the computing device 102 can determine the user's engagement from other factors that can help determine the likelihood of the user interacting with the computing device. In this way, by detecting the user's engagement, the computing device can be changed based on the user's current or expected engagement with the device to perform best in any number of scenarios.
[0391] Example Method
[0392] Fig.38 An example method 3800 for a system of multiple radar-enabled computing devices is shown. The method 3800 is shown as multiple sets of operations (or actions) performed, but is not necessarily limited to the order or combination of operations shown herein, and may be performed in whole or in part with other methods described herein. In addition, any of one or more of the operations of these methods may be repeated, combined, reorganized, or linked to provide a wide range of additional methods and / or alternative methods. In the following discussion, reference may be made to Figures 1 to 37 The example environments, experimental data, or experimental results of the present invention are only referenced for the sake of example. These techniques are not limited to being performed by one or more entities operating on or associated with a computing device 102.
[0393] At 3802, a first radar transmission signal is transmitted from a first computing device of a computing system. For example, first radar transmission signal 402-1 may be transmitted from first radar system 108-1 of first computing device 102-1 to detect whether user 104 is present within first proximity 106-1 (e.g., Figure 1 and Fig. 20 ). The first computing device 102-1 may be part of a computing system including two or more computing devices 102-X, see Figure 3 , Figure 4 , Figure 21 to Figure 23 , Figure 26 to Figure 29 and Fig.36Each computing device 102 in the computing system may exchange information (e.g., radar signal characteristics, stored information, ongoing operations) with another device in the system via the communication network 302. This information may additionally be stored in one or more memories, which may be local, shared, remote, etc. The first radar transmission signal 402-1 may include a single signal, multiple signals that are similar or different, a signal pulse train, a continuous signal, etc., as described with reference to Figure 4 described.
[0394] At 3804, a first radar receive signal is received at a first computing device. For example, first radar receive signal 404-1 may be received by first radar system 108-1 of first computing device 102-1. First radar receive signal 404-1 may be associated with first radar transmit signal 402-1 that has been reflected by an object (e.g., user 104) within first proximity 106-1. This reflected signal (first radar receive signal 404-1) may represent a modification of first radar transmit signal 402-1 in time, amplitude, phase, or frequency.
[0395] At 3806, the presence of a registered user is determined based on the first radar signal characteristic of the first radar reception signal. For example, the presence of a registered user may be determined based on the first radar signal characteristic of the first radar reception signal 404-1 (see Figure 1 and Fig.21 ) presence. This first radar received signal 404-1 can be reflected by an object (e.g., a registered user) and contain one or more radar signal characteristics (e.g., associated with time, topology and / or gesture information), which can be used to distinguish this object as a registered user. Specifically, the first radar signal characteristic of the first radar received signal 404-1 can be correlated with the first stored radar signal characteristic of the registered user. The correlation can indicate the presence of the registered user within the first neighboring area 106-1. These stored radar signal characteristics can be stored in a local memory, a shared memory, or a remote memory that can be accessed by any computing device 102-X of the computing system. The first radar signal characteristic can be additionally stored in the memory to improve monitoring and differentiation of registered users at future times. As described herein, context and other information can also be used to help detect and differentiate both users.
[0396] At 3808, a second radar transmission signal is transmitted from a second computing device of the computing system. For example, second radar transmission signal 402-2 may be transmitted from second radar system 108-2 of second computing device 102-2 to detect whether user 104 is present in second proximity area 106-2, see Figure 1 and Fig.21The second computing device 102 - 2 may be part of a computing system that includes at least the first computing device 102 - 1 .
[0397] At 3810, a second radar receive signal is received at the second computing device. For example, second radar receive signal 404-2 may be received by second radar system 108-2 of second computing device 102-2. Second radar receive signal 404-2 may be associated with second radar transmit signal 402-2 that has been reflected by an object (e.g., user 104) within second proximity 106-2.
[0398] At 3812, the presence of a registered user is determined based on a correlation between a second radar signal characteristic of the second radar receive signal and one or more stored radar signal characteristics. For example, the presence of a registered user may be determined again based on the second radar receive signal 404-2. Specifically, the second radar signal characteristic of the second radar receive signal 404-2 may be correlated with a first stored radar signal characteristic or a second stored radar signal characteristic of a registered user. The correlation indicates the presence of a registered user within the second adjacent area 106-2, thereby distinguishing the user from one or more other users or potential users. The second computing device 102-2 may alternatively compare the second radar signal characteristic with the first radar signal characteristic (determined by the first computing device 102-1) to determine the presence of a registered user. In this way, information detected, determined and / or stored by the first computing device 102-1 may be used by any one or more devices of the computing system (e.g., the second computing device 102-2) to distinguish, for example, registered users.
[0399] Fig.39 An example method 3900 for radar-based ambiguous gesture determination using contextual information is shown. The method 3900 is shown as multiple sets of operations (or actions) performed, but is not necessarily limited to the order or combination of operations shown herein, and may be performed in whole or in part with other methods described herein. In addition, any of the operations in one or more of the operations of these methods may be repeated, combined, reorganized, or linked to provide a wide range of additional methods and / or alternative methods. In the following discussion, reference may be made to Figures 1 to 37 The example environments, experimental data, or experimental results of the present invention are only referenced for the sake of example. These techniques are not limited to being performed by one or more entities operating on or associated with a computing device 102.
[0400] At 3902, the computing device detects an ambiguous gesture performed by a user. For example, computing device 102 utilizes radar system 108 to detect one or more radar signal characteristics of an ambiguous gesture performed by user 104. An ambiguous gesture (e.g., ambiguous gesture 2402) may refer to a gesture that cannot be recognized by the device to a desired confidence level. Techniques for detecting and analyzing radar signal characteristics associated with gestures (or users) are described above. A gesture may be detected and determined to be a gesture with sufficient confidence, rather than a non-gesture body movement, an animal, a motionless object, etc., but its confidence is insufficient to be recognized, such as below a high confidence level and above a no confidence level. In this case, the detected gesture is an ambiguous gesture.
[0401] At 3904, the ambiguous gesture is correlated with the first gesture and the second gesture. For example, computing device 102 compares one or more radar signal characteristics of the ambiguous gesture to one or more stored radar signal characteristics. Gesture module 224 ...
Claims
1. A method, include: using a radar system at a computing device to detect an ambiguous gesture performed by a user, the ambiguous gesture being associated with a radar signal characteristic; comparing the radar signal characteristic to one or more stored radar signal characteristics, the comparing effective to correlate the ambiguous gesture to a first gesture and a second gesture, the first gesture and the second gesture being associated with a first command and a second command, respectively; determining that the first command will be less destructive than the second command; as well as In response to the determination, the computing device, an application associated with the computing device, or another device associated with the computing device is directed to execute the first command.
2. The method according to claim 1, in, Make sure the first command will be less destructive: The first command is a provisional command, and the second command is a final command.
3. The method according to claim 1, in, Determining that the first command will be less destructive Determining that the first command can be undone and the second command cannot be undone.
4. A method as claimed in any preceding claim, in, Determining that the first command is less destructive than the second command is further based on previous actions taken by the user within a time period following a previous execution of the first command or the second command.
5. A method according to any preceding claim, further comprising: include: In response to the execution of the first command, detecting, at the computing device or the other device associated with the computing device, a user input directing the computing device to execute a third command to cancel the first command, the user input not including another gesture execution; In response to detecting the user input command, determining that the ambiguous gesture is not the first gesture; as well as The determination of the ambiguous gesture is stored to improve detection of the first gesture at a future time.
6. A method according to any preceding claim, further comprising: include: In response to execution of the first command, detecting, at the computing device or the other device associated with the computing device, another gesture being performed by the user; and In response to determining that the other gesture is similar to or identical to the ambiguous gesture: determining that the user did not intend to execute the first command, the unintended execution indicating that the other gesture and the ambiguous gesture were not the first gesture; as well as The other gesture and the ambiguous gesture are determined to be the second gesture, the determination being effective to associate the other gesture with the second command.
7. The method of claim 6, further comprising, in response to associating the another gesture with the second command: directing the computing device, the application associated with the computing device, or the other device associated with the computing device to: Stop executing the first command; or executing the second command; and Another radar signal characteristic associated with the another gesture is stored to enable detection of the second gesture at the future time.
8. A method as claimed in any preceding claim, in, Determining that the first command is less destructive is based on current conditions, logic, or a history of user behavior.
9. A method as claimed in any preceding claim, in, Both the first command and the second command can affect the operation of a single application, the operation being executable by the computing device or the other device associated with the computing device.
10. The method according to claim 9, in: The single application is a phone application; The first command is to mute the telephone call; and The second command is to end the telephone call.
11. The method according to claim 9, in: The single application is a notification application; The first command is to mute, pause, or delay the notification; and The second command is to disable the notification.
12. A method as claimed in any preceding claim, in, The first command and the second command can affect the operation of different, respective applications or devices associated with the computing device.
13. A method as claimed in any preceding claim, in, Determining that the first command will be less destructive than the second command includes utilizing a machine learning model that performs unsupervised learning on commands executed by the computing device, the unsupervised learning being operable to determine the destructiveness of the command without utilizing predetermined conditions or algorithms.
14. A computing device, include: at least one antenna; a radar system configured to transmit radar transmit signals and receive radar receive signals using the at least one antenna; at least one processor; as well as A computer-readable storage medium comprising instructions, which, in response to being executed by the processor, are used to direct the computing device to perform any one of the methods of claims 1 to 13.
15. A computing system comprising a first computing device and a second computing device connected to a communication network such that: The first computing device or the second computing device is capable of performing any of the methods of claims 1 to 13; and Information can be exchanged between the first computing device and the second computing device.