Continuous online learning for radar-based gesture recognition
Through continuous online learning of radar-based gesture recognition, the computing device detects and stores fuzzy gesture characteristics, solving the problem of insufficient accuracy of gesture recognition and improving user experience and satisfaction.
Patent Information
- Application Number
- CN202280100264.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-05-06
AI Technical Summary
The existing computing devices have insufficient accuracy in gesture recognition, resulting in poor user experience, especially the difficulty in identifying misidentification problems caused by the differences and complexity of gestures performed by users.
Through a continuous online learning method of radar-based gesture recognition, the computing device detects fuzzy gestures, stores its characteristics, and improves recognition accuracy in subsequent interactions, and continuously updates gesture characteristics using radar systems and machine learning models to improve recognition capabilities.
It improves the accuracy of gesture recognition, enhances user's interaction trust and satisfaction, and allows users to use gesture control computing devices more frequently.
Smart Images

Figure CN119948429A_ABST
Abstract
Description
Background Art
[0001] Year after year, computing devices play an increasingly larger role in people's lives. However, this greater role has a correspondingly greater need for seamless and pervasive interaction between people and their devices. Gone are the days of desktop computers with only a physical keyboard through which to interact. To address this need, new ways of interacting with devices have been developed. However, many of these ways of interacting include attendant design difficulties. For example, some computing devices use gesture recognition to enable users to control their devices without requiring the user to physically contact the device or its peripherals. However, due to the complexity of gesture recognition, the computing device may fail to recognize gestures performed by the user, thereby frustrating the user, which over time may cause the user to avoid using gestures to control the device. Summary of the invention
[0002] Techniques and apparatus for continuous in-line learning of radar-based gesture recognition are described in this document. Through continuous in-line learning, a computing device can improve recognition of even the most difficult to recognize gestures by gradually storing characteristics of ambiguous gestures performed by a user. Specifically, a radar system can detect a first ambiguous gesture that the computing device fails to recognize as a known gesture and a second gesture that the computing device successfully recognizes as a known gesture. The computing device can identify similarities between the first gesture and the second gesture, and in doing so, store the characteristics of the first gesture to more accurately recognize the known gesture when it occurs in the future.
[0003] The various aspects described below include a method for continuous online learning of radar-based gesture recognition. The method may include detecting an ambiguous gesture performed by a user at a computing device. The ambiguous gesture may be associated with a first radar signal characteristic. The computing device may compare the first radar signal characteristic with one or more stored radar signal characteristics and determine that the ambiguous gesture cannot be associated with the first radar signal characteristic with a required confidence level. The method may also include detecting another gesture performed by the user, the other gesture being associated with a second radar signal characteristic. The computing device may compare the second radar signal characteristic of the other gesture with one or more stored radar signal characteristics to identify the other gesture as a first gesture. In response to this comparison, the computing device may determine that the ambiguous gesture is a first gesture, and store the first radar signal characteristic to enable the performance of the first gesture to be identified at a future time.
[0004] A device is also described, the device including a radar system capable of transmitting and receiving radar signals. The device also includes at least one computer-readable storage medium storing instructions that, when executed by at least one processor, perform continuous online learning for radar-based gesture recognition according to the method described above. A component for performing the method is also described. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Devices and techniques for continuous online learning of radar-based gesture recognition and other devices and techniques are described with reference to the following figures. The same reference numerals are used throughout the figures to refer to similar features and components: Figure 1 An example environment with a radar-enabled computing device, a user, a proximity zone, and a radar system is shown; Figure 2 Shows Figure 1 Example implementations of a radar-enabled computing device; Figure 3 An example environment is shown in which a plurality of radar-enabled computing devices are connected via a communication network to form a computing system; Figure 4 An example environment is shown in which a radar system is used by a computing device to detect, distinguish, and / or recognize a user or a gesture being performed by a user; Figure 5 An example implementation of an antenna, analog circuitry, and system processor for a radar system is shown; Figure 6 shows an example implementation in which a user module can distinguish between users; Figure 7 An example implementation of a machine learning (ML) model for distinguishing users of a computing device is shown; Figure 8 An example implementation of a gesture module is shown that utilizes a spatiotemporal machine learning model to improve detection and recognition of gestures; Fig. 9 An example implementation of deep learning techniques utilized by the frame model is shown; Fig.10 An example implementation of deep learning techniques utilized by a temporal model is shown; Fig.11 Experimental results indicating improved performance in the recognition of hand gestures when utilizing radar enhancement techniques are shown; Fig.12 Experimental data showing a user performing a tap gesture; Fig.13Experimental data showing a user performing a tap gesture, a swipe right, a strong swipe left, and a weak swipe left; Fig.14 Experimental data showing three sets of negative data that can be stored to improve detection and recognition of gestures from background motion; Fig.15 Experimental results on the accuracy of gesture recognition in the presence of background motion are shown; Fig.16 Experimental results on the accuracy of gesture detection and recognition when adversarial negative data is additionally used are shown; Fig.17 The experimental results (confusion matrix) related to the accuracy of gesture recognition are shown; Fig.18 Experimental results corresponding to the accuracy of gesture recognition at various linear and angular displacements from the antenna of the radar system are shown; Fig.19 An example implementation of a radar-enabled computing device using additional sensors to improve the fidelity of user detection and differentiation is shown; Fig. 20 An example environment is shown in which privacy settings are modified based on user presence; Fig.21 An example environment is shown in which techniques for user differentiation may be implemented using multiple computing devices forming a computing system; Fig. 22 An example environment is shown in which operations are continuously performed across multiple computing devices of a computing system; Fig.23 An example environment is shown in which a computing system implements persistence of operations across multiple computing devices; Fig.24 A technique for radar-based ambiguous gesture determination using contextual information is shown; Fig.25 An example implementation is shown in which a gesture module can recognize gestures performed by a user; Fig.26 An example environment is shown in which a computing device can utilize contextual information of a user's habits to improve gesture recognition; Fig. 27 showing an expectation of a user's presence based on a location of a computing device; Fig.28 A technique for using room-dependent context to improve recognition of ambiguous gestures is shown; Fig.29 A technique for improving recognition of ambiguous gestures using the status of the operation being performed at the current time is shown; Fig.30Techniques for distinguishing ambiguous gestures based on foreground and background operations being executed at the current time are shown; Fig.31 Techniques for using contextual information including past and / or future operations of a computing device are shown; Fig.32 Techniques for recognizing ambiguous gestures based on less destructive operations are shown; Fig.33 A user is shown performing an ambiguous gesture, the user intending the ambiguous gesture to be a first gesture (e.g., a known gesture); Fig.34 An example method for performing online learning based on user input to improve ambiguous gesture recognition is shown; Fig.35 An example technique for online learning of new gestures for a radar-enabled computing device is shown; Fig.36 An example technique for configuring a context-sensitive primary sensor to a computing device is shown; Fig.37 illustrates an environment in which techniques for detecting user interaction with a device may be implemented; Fig.38 An example method for determining the presence of a registered user is shown; Fig.39 An example method for radar-based ambiguous gesture determination using contextual information is shown; Fig.40 An example method for continuous online learning for radar-based gesture recognition is shown; Fig.41 An example method for radar-based gesture detection over long distances is shown; Fig.42 An example method for online learning based on user input is shown; Fig.43 An example method for online learning of new gestures for a radar-enabled computing device is shown; Fig.44 An example method for sensor capability determination is shown; Fig.45 An example method of recognizing ambiguous gestures based on less destructive operations is shown; and Fig.46 An example method of detecting user interaction is shown. DETAILED DESCRIPTION Overview
[0006] As computing devices are increasingly found in everyday life, users choose to rely on these devices to support a wide variety of tasks. For example, home automation using virtual assistant (VA) technology is becoming increasingly popular as a way to improve the safety, comfort, and convenience of a home. Residents can easily control, for example, lighting, climate, entertainment systems, appliances, alarms, etc. using computing devices equipped with VA technology. In order to improve the usability and user satisfaction of these devices, manufacturers aim to provide users with convenient ways to interact with their computing devices to provide efficient and accurate control of the devices when performing these functions. As the capabilities of computing devices expand to become useful in increasingly complex scenarios, additional methods of interaction are developed to enable users to communicate with their devices.
[0007] One such form of interaction that is becoming increasingly popular in computing devices is touchless gesture recognition. This is the case where the user is not required to make physical contact with the device to control the device. Instead, the user can use their entire body or part of their body to perform a gesture (e.g., a hand movement made in the air), and the computing device can recognize this gesture. The computing device then causes the execution of commands associated with the recognized gesture. Through this technology, users may be able to control their homes, for example, even when their hands are dirty or when they are located a certain distance away from the computing device. Therefore, implementing touchless gesture recognition on a computing device can improve the user's ability to interact with the device, and thereby improve user satisfaction.
[0008] While gesture recognition can enable users to more conveniently interact with their devices, computing devices may fail to accurately recognize gestures, frustrating users and making them less likely to use gestures to interact with their devices. Different users may perform gestures slightly differently, and some gestures may be more difficult to recognize than others. Some gesture recognition systems fail to recognize these gestures due to slight differences in execution between users or iterations by one user, causing users to become frustrated on their devices and avoid using these features over time. To address these challenges, the present disclosure describes continuous online learning for radar-based gesture recognition, which better enables radar-based systems to continuously improve the accuracy of gesture recognition through interaction with users.
[0009] Assume, for example, that a radar system of a computing device detects an ambiguous gesture that the radar system cannot correlate with a recognized gesture associated with a particular command. Based on the system's failure to recognize the gesture, the user may perform the gesture again. In response to this re-execution of the gesture, the radar system may recognize the gesture and execute the command associated with the re-executed gesture. Given that the user re-executed the gesture after the radar system failed to recognize the gesture, it is possible that the ambiguous gesture is the same as the recognized gesture or that the user intended the same as the recognized gesture. In this way, the computing device may correlate the characteristics of the ambiguous gesture with the recognized gesture. In doing so, the computing device may continuously update the characteristics associated with even the most difficult to detect gestures and improve gesture recognition accuracy over time.
[0010] By updating the stored characteristics for known gestures, gesture recognition can improve as the device is used, without requiring separate training outside of normal user interaction with the device. In this way, users can become more trusting of gesture recognition and rely more frequently on gestures to control their devices. Thus, techniques for continuous online learning can improve the accuracy of gesture recognition and increase user satisfaction.
[0011] It should be noted that this is only one example of continuous online learning for radar-based gesture recognition described in this document, and other examples and other techniques are described below. The present disclosure will now turn to an example operating environment, followed by an example computing device, examples of radar-based gesture detection and recognition, and a description of various techniques for utilizing and improving radar-based gesture recognition. Example Environment
[0012] Figure 1 An example environment 100 is shown in which a radar-enabled computing device 102 performs the techniques described in this document, such as detecting and distinguishing users and user interactions, detecting and recognizing gestures, and causing commands to be executed, as well as improving any of these techniques. A radar-enabled computing device 102 (computing device 102) can be used to cause tasks to be performed (e.g., turning off a light, lowering the volume of music, starting an oven, changing a TV channel) by recognizing gestures associated with commands. A swipe of a user's hand can indicate a command to change the song being played, while a push-pull gesture can indicate a command to check the status of a timer in the kitchen. Computing device 102 can allow multiple users 104, in some cases, each of which can enjoy a customized experience by being distinguished by computing device 102. In addition, multiple computing devices (e.g., computing devices 102-X, where X represents an integer value of 1, 2, 3, 4, ...) can be connected (e.g., wirelessly connected, such as by connecting to one or more wireless networks and / or by direct wireless communication) to create an interconnected network of radar systems, as described with respect to Figure 3 and Figure 4 This network of computing devices may be arranged to detect and recognize gestures performed by a user 104 (such as a first user 104-1, a second user 104-2, or a different user 104-X, where X represents an integer value such as 3, 4, 5, ...), for example, in any one or more rooms of a residence.
[0013] In particular, the techniques may include (1) detecting the presence of a user 104 within a proximity zone 106 of a computing device 102, (2) distinguishing that user from other users to enable a customized experience, and then (3) instructing the computing device 102, an application associated with the computing device 102, or another device to execute a command upon recognizing that a known gesture has been performed. Additionally, or in lieu of (3), the computing device 102 may prompt the user 104 to begin or continue gesture training based on that user's training history.
[0014] The computing device 102 may transmit radar transmit signals discretely or continuously over time (e.g., without a "wake-up trigger") to detect the presence of a user and / or the performance of a gesture within the proximity zone 106. Any one or more of these radar transmit signals may be reflected from an object in the proximity zone 106 (e.g., a user 104 making a movement or a stationary object), thereby generating one or more radar receive signals. The computing device 102 may determine radar signal characteristics (e.g., temporal information or topological information of an object or movement) based on the radar receive signals. If the determined radar signal characteristics are correlated with one or more stored radar signal characteristics, the computing device 102 may classify the object or movement. By doing so, the computing device 102 may determine that the movement may be a gesture, rather than some other moving or non-moving object. In this example, the technology compares the radar signal characteristics with one or more stored radar signal characteristics associated with a registered user or an unregistered user or a known gesture, and thereby attempts to detect and distinguish a user making a movement and detect and recognize a gesture being performed.
[0015] For example, assume that user 104-1 makes a motion with their hand. Based on one or more radar receive signals reflected from the user's hand, computing device 102 may detect that this motion is a gesture, rather than a non-gesture movement, based on radar signal characteristics of the radar receive signals. Computing device 102 correlates one or more of the radar signal characteristics with one or more stored radar signal characteristics of a known gesture (e.g., a swipe gesture associated with a command to turn on the light) simultaneously with or after detecting the gesture. If the determined radar signal characteristics correlate with a desired confidence level (e.g., a threshold criterion), the device may determine that a swipe gesture was performed and caused the light to turn on.
[0016] In another example, computing device 102 may detect a user in proximity area 106 (eg, Figure 1 The device may determine the presence of a first user within proximity zone 106 (as depicted). The device may determine at least one radar signal characteristic (e.g., height, build, movement) of this first user 104-1 and correlate it with one or more stored radar signal characteristics of “registered users.” If the determined radar signal characteristic correlates with a desired confidence level, the device may determine that first user 104-1 is a registered user and that the user is currently located within proximity zone 106.
[0017] In the present disclosure, a "registered user" generally refers to a user 104 that is associated with at least one stored radar signal characteristic and / or has an account or other registration that is accessible by the computing device 102. An account may be manually set up (e.g., by a registered user or another user of the device) or automatically set up (e.g., when interacting with the device). An account may include or be associated with one or more stored radar signal characteristics, user settings, preferences, gesture training history, user habits, etc. that may or may not be dependent on access to personally identifiable information. By virtue of an account, a registered user may have the right to modify, store, or access a large amount of information associated with the computing device (e.g., that user's settings or preferences). These rights may be granted to registered users, but not (by default) to users who have not registered the device.
[0018] For some aspects, the account registered by the registered user corresponds to their account on the cloud-based smart home service platform of a virtual assistant provider, such as Google® (which provides the "Hey Google®" or "OK Google®" voice assistant service) or Amazon® (which provides the "Alexa®" voice assistant service). In such examples, the registered user can be the primary user of an account associated with the cloud-based service platform for a particular residence, which is sometimes referred to as a primary user, a billing user, a supervisory user, an administrator user, a super user, a master user, etc., depending on the nature of the platform. Dad, Mom, another head of the household, a "home technology expert", or another designated person typically plays the role of the primary user. The registered user can alternatively be a secondary user (or tertiary user, etc.) of an account associated with the cloud-based service platform, such as a teenager or an adult who is not the primary user, who enjoys at least some of the benefits known to the cloud-based service platform, but typically has a more limited set of privileges or capabilities. However, it should be understood that other types of user registrations than those for cloud-based service platforms are within the scope of the present teachings, including but not limited to user accounts established with independent, offline or off-network device groups having their own local account establishment and registered user establishment types.
[0019] More specifically, computing device 102 uses radar system 108 (e.g. Figure 1 104) to detect the presence of a user in proximity zone 106. When an object (e.g., user 104) is detected within proximity zone 106, the radar transmit signal may be reflected from user 104 and modified (e.g., in amplitude, phase, and / or frequency) based on the appearance and / or movement of user 104. This modified radar transmit signal (e.g., radar receive signal) may be received by radar system 108 and include information for distinguishing user 104 from other users, e.g., see Figure 6 and accompanying description. The radar system 108 uses the radar received signals to determine the speed, size, shape, surface smoothness or material of the user 104 because the received signals will vary due to differences in objects, such as the object's speed (e.g., by Doppler), size (see Figure 6 ), shape (see Figure 6 Radar system 108 may also determine the distance between user 104 and computing device 102 and / or the orientation of user 104 relative to computing device 102, such as by time-of-flight analysis.
[0020] Although the proximity area 106 of the example environment 100 is depicted as a hemisphere, in general, the proximity area 106 is not limited to the morphology shown. The morphology of the proximity area 106 may also be affected by nearby obstacles (e.g., walls, large objects). In general, the computing device 102 can be located in a proximity area 106 (e.g., a bedroom) of a larger physical area (e.g., a residence, an environment) than the proximity area 106. In addition, the radar system 108 can detect and detect users and / or gestures outside the proximity area 106 depicted in the example environment 100. The boundaries of the proximity area 106 correspond to accuracy thresholds, wherein the users and / or gestures detected within this boundary are more likely to be accurately distinguished or recognized than the users or gestures detected outside this boundary, respectively. The example proximity area 106 is one centimeter to eight meters, depending on the power usage of the radar system 108, the expected confidence, and whether the radar system 108 is configured to detect users, distinguish users, detect gestures, recognize gestures, and / or detect user interactions.
[0021] For example, computing device 102 sends a first radar transmit signal into proximity zone 106 and then receives a first radar receive signal (e.g., a reflected radar transmit signal) associated with the presence of an object (e.g., first user 104-1). This first radar receive signal includes one or more radar signal characteristics (e.g., radar cross section (RCS) data, motion signatures, gesture performance, etc.) that can be used to distinguish first user 104-1 from other users 104. In particular, radar system 108 can compare the first radar receive signal with stored radar signal characteristics of registered users to determine whether first user 104-1 is a registered user who has previously interacted with the device and / or established an account with the device. Example ways of doing this include using Figure 7 and Figure 8 1 , and topology distinction 600-1, time distinction 600-2, gesture distinction 600-3, and context distinction 600-4, respectively. In this example, the radar signal characteristic of the first radar reception signal is correlated with one or more stored radar signal characteristics of registered users at a desired confidence using the machine learning model 700, which is trained using the stored radar signal characteristics and has as input the radar signal characteristic of the first radar reception signal reflected from one of the users 104.
[0022] In general, a radar transmit signal may refer to a single (discrete) signal, a burst of signals, or a continuous stream of signals transmitted over time from one or more antennas of computing device 102. A radar transmit signal may be transmitted at any time without requiring a "wake-up" triggering event (for example, Figure 8The radar receive signal may be received by the same transmit antenna, a different antenna of computing device 102, or an antenna of another device of the computing system (described below with respect to Figure 3 and Figure 4 described in more detail).
[0023] The computing device 102 in the example environment 100 may (but is not required to) forgo “personally identifying” the first user 104-1 (e.g., private or personally identifiable information of the first user) to determine that the detected object is a registered user. For example, the computing device 102 may determine that the first user 104-1 is a registered user without requiring personally identifiable information, which may include legally identifiable information (e.g., legal name). In addition, the computing device 102 may forgo identifying the first user 104-1’s personal device (e.g., a mobile phone, a device equipped with an electronic tag), collecting facial recognition information, or performing speech-to-text on a potentially private conversation to determine that the first user 104-1 is a registered user based on the user’s preferences or settings. Instead of personally identifying the first user 104-1, the computing device 102 may use radar signal characteristics that do not contain personally identifiable information and / or confidential information to “distinguish” the first user 104-1 from another user (e.g., the second user 104-2), as described below with respect to Figure 4 described.
[0024] Controls may be provided to the user 104 that allow the user 104 to make choices about whether and when the technology described herein may be able to collect information about the user (e.g., information about the user's social network, social behavior, social activities or occupation, photos taken by the user, audio recordings made by the user, preferences of the user, or the user's current location, etc.), as well as about whether to send content or communications from the server to the user 104. In addition, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the identity of the user may be processed so that personally identifiable information of the user 104 cannot be determined, or the user's geographic location may be generalized (e.g., to a city, zip code, or state level) in the case where location information is obtained so that the specific location of the user 104 cannot be determined. Thus, the user 104 may control what information is collected about the user 104, how that information is used, and what information is provided to the user 104. Gesture training
[0025] Gesture training refers to an interactive user experience through which a radar-enabled computing device helps a user learn how to make gesture commands or inputs to the device. Gesture training may involve, for example, the device performing the following steps: (i) communicating information to the user about how to make one or more specific gesture commands or inputs, (ii) suggesting / proposing that the user attempt to make that gesture, (iii) monitoring the user while the attempt is being made, and (iv) providing the user with evaluative feedback on whether that attempt was successful and whether they can try again in a different way. In an exemplary scenario, the device is a smart home display assistant (e.g., Google® NEST HUB™) with contactless gesture (e.g., air gesture) recognition based on FMCW (frequency modulated continuous wave) radar using radar signals in the 60 GHz range. For example, this smart home assistant is able to recognize gestures such as swipe left, swipe right, swipe up, swipe down, air knob turn, and push inward. The smart home assistant may provide gesture training on a swipe-left gesture by: (i) showing the user a short animation or video of a swipe-left motion while displaying or saying “This is a swipe-left”; (ii) displaying or saying “Now you try”; (iii) performing radar-based monitoring as the user attempts the gesture, and if successful, (iv) displaying or saying “Awesome, you've got it! For example, if the distance from left to right of the gesture is too short, the smart home assistant may say “Try again using alonger sweeping motion”, etc. Other methods of gesture training may be similar in general approach, but expressed in more interesting activities, such as having the user control or direct an on-screen character or other on-screen object in a game-like environment. Other kinds of interactive user experiences for gesture training may be provided without departing from the scope of the present teachings. For the purposes of user convenience, encouragement, and continued interaction with the device, it is generally desirable to avoid requiring a one-time, single-instance gesture training session in which all gesture training is completed at one time, unless the user explicitly requests such a single-instance gesture training session. Rather, it is generally desirable to suggest and provide smaller modular lessons at appropriate times over time (e.g., a first lesson on swiping left at an appropriate time on the first day, a second lesson on turning the air knob at an appropriate time on the next day, etc.).
[0026] Thus, according to one aspect, before, after, or concurrently with determining that the first user 104-1 is a registered user (or unregistered but has stored radar signal characteristics), the computing device 102 may prompt the first user 104-1 to start or continue gesture training. For example, the first user 104-1 may be halfway through training and has completed training of a first gesture (e.g., a left swipe in the previous example of the previous paragraph). The computing device 102 may have stored information about the manner in which the first user 104-1 performed the first gesture during training in a training history. When the first user 104-1 is detected, the computing device 102 may access the training history and prompt the first user 104-1 to continue training of a second gesture (e.g., an air knob turn) instead of repeating training of the first gesture. Thus, determining that the first user 104-1 is a registered user or an unregistered person with stored radar signal characteristics may allow the computing device 102 to improve the efficiency of gesture training. Differentiating the first user 104 - 1 may allow the computing device 102 to activate settings (eg, privacy settings, preferences) of registered users or unregistered persons to provide a customized experience for the first user 104 - 1 .
[0027] for Figure 1 , assuming that the second user 104-2 is sitting on a sofa in the proximity area 106 with the first user 104-1. The computing device 102 can use the radar system 108 to send a second radar transmission signal to detect the presence of another object (e.g., the second user 104-2). The radar system 108 can then compare the second radar reception signal with the stored radar signal characteristics of the registered users to determine whether the second user 104-2 is another registered user. In this example, the second radar reception signal is not found to be correlated with one or more stored radar signal characteristics of another registered user. Therefore, the second user 104-2 is distinguished as an "unregistered person."
[0028] In the present disclosure, "unregistered personnel" will generally refer to users 104 who have not registered a device and are therefore not associated with one or more accounts. Unlike registered users, unregistered personnel may not have the right to modify, store or access the information of the computing device. For example, a new visitor (e.g., a guest who is not associated with one or more accounts of the device) can be considered an unregistered person. This new visitor may not have interacted with the computing device 102 before and is therefore not associated with one or more stored radar signal characteristics, or the visitor may have done so and is associated with the stored radar signal characteristics, but has no account or other special rights. Therefore, previous visitors (e.g., nannies, housekeepers, gardeners) may be associated with one or more stored radar signal characteristics, but have no account in the device. According to one or more aspects, the computing device 102 can still store the radar signal characteristics of this user to improve their user experience. However, the device can prevent previous visitors (unregistered personnel) from exercising the rights enjoyed by registered users, such as modifying, accessing or storing information of the device.
[0029] After determining that the second user 104-2 is an unregistered person, the computing device 102 assigns an unregistered user identification (e.g., a simulated identity, a pseudo identity) to the unregistered person, and the unregistered user identification can be associated with one or more radar signal characteristics of the second radar reception signal, such as a unique random number for later identifying the unregistered user. The unregistered user identification can be stored so that the second user 104-2 can be distinguished at a future time. In particular, the unregistered user identification can be used to correlate future received radar reception signals with one or more associated radar signal characteristics of the unregistered person. The computing device 102 can also prompt the second user 104-2 to register with the computing device 102.
[0030] The computing device 102 may or may not require personally identifiable information of the second user 104-2 to determine that the other object is an unregistered person. After the second user 104-2 and the first user 104-1 have been distinguished, the computing device 102 may determine that the privacy settings of the first user 104-1 need to be adjusted (e.g., modified, restricted) to ensure that the first user's information remains private. For example, the first user 104-1 may want the device to avoid broadcasting calendar reminders (e.g., doctor's appointments) when the second user 104-2 is present. Additionally, the computing device 102 may prompt the second user 104-2 to start gesture training, which may be recorded in another training history of the second user 104-2 (e.g., associated with an unregistered user identification).
[0031] At a later time (not depicted), computing device 102 may again use radar system 108 to send a third radar transmit signal to detect whether user 104 is within proximity zone 106. If user 104 (e.g., first user 104-1, second user 104-2) is present at this time, the third radar transmit signal may be reflected from user 104, and computing device 102 may receive a third radar receive signal, which includes one or more radar signal characteristics. Radar system 108 may compare these radar signal characteristics with, for example, stored radar signal characteristics of first user 104-1 (registered user) and second user 104-2 (unregistered person associated with unregistered user identification) to determine whether first user 104-1 or second user 104-2 is present. Based on this determination, computing device 102 may customize settings and training prompts accordingly.
[0032] In one example, radar system 108 uses a third radar received signal to determine that first user 104-1 (a registered user) is again present within proximity zone 106 based on stored radar signal characteristics associated with first user 104-1. Computing device 102 may then prompt first user 104-1 to complete their gesture training and / or activate their user settings based on their training history. Alternatively, if radar system 108 determines that second user 104-2 (an unregistered person) is again present within proximity zone 106, computing device 102 may prompt second user 104-2 to continue their gesture training and / or activate a predetermined user setting. Figure 2 Computing device 102 and radar system 108 are further described. Example computing device
[0033] Figure 2 An example implementation 200 of a radar system 108 as part of a computing device 102 is shown. The computing device 102 is shown with various non-limiting example devices 202, including a home automation and control system 202-1, a smart display 202-2 associated with the home automation and control system, a desktop computer 202-3, a tablet computer 202-4, a laptop computer 202-5, a television 202-6, a computing watch 202-7, computing glasses 202-8, a gaming system 202-9, a microwave oven 202-10, a smart thermostat interface 202-11, and a car with computing capabilities 202-12. Other devices may also be used, such as security cameras, baby monitors, Wi-Fi ®A router, a drone, a trackpad, a drawing tablet, a netbook, an e-reader, other forms of home automation and control systems, a wall display, a virtual reality headset, another vehicle (e.g., an electric bicycle or airplane), and other home appliances, to name a few examples. It should be noted that the computing device 102 can be wearable, non-wearable but mobile, or relatively non-mobile (e.g., a desktop and an appliance), all without departing from the scope of the present teachings.
[0034] The computing device 102 may include one or more processors 204 and one or more computer readable media (CRM) 206, which may include memory media and storage media. Applications and / or operating systems (not shown) embodied as computer readable instructions on the CRM 206 may be executed by the processor 204 to provide some of the functionality described herein. The CRM 206 may also include radar-based applications 208 that use data generated by the radar system 108 to perform functions such as gesture-based control, human vital sign notification, collision avoidance for autonomous driving, and the like. For example, the radar system 108 may recognize a gesture performed by the user 104 indicating a command to turn off the lights in the room. This command data may be used by the radar-based application 208 to send a control signal (e.g., a trigger) to turn off the lights in the room.
[0035] The computing device 102 may also include a network interface 210 for communicating data via a wired, wireless, or optical network. For an interconnected system of multiple computing devices 102-X, each computing device 102 may communicate with another computing device 102 via the network interface 210. For example, the network interface 210 may communicate data via a local area network (LAN), a wireless local area network (WLAN), a personal area network (PAN), a wide area network (WAN), an intranet, the Internet, a peer-to-peer network, a point-to-point network, a mesh network, etc. Multiple computing devices 102-X may communicate with each other using a communication network, as described below with respect to Figure 3 The computing device 102 may also include a display.
[0036] Radar system 108 may be used as a stand-alone radar system or used with or embedded in many different computing devices or peripherals, such as in a control panel that controls home appliances and systems, in an automobile to control internal functions (e.g., volume, cruise control, or even the steering of the car), or as an accessory to a laptop computer to control computing applications on the laptop.
[0037] Radar system 108 may include communication interface 212 to transmit radar data (e.g., radar signal characteristics) to a remote device, but the communication interface may not be used when radar system 108 is integrated within computing device 102. In general, the radar data provided by communication interface 212 may be in a format that can be used to detect, distinguish, and / or recognize a user, user interaction, or gesture, such as a value of a frame of radar signal characteristics (e.g., corresponding to a complex-range Doppler map, see Figure 8 and Figure 12 to Figure 14 ) or a determination of detection or recognition by computing device 102. Communication interface 212 may also or instead communicate with a remote instance of radar-based application 208, such as commands associated with a recognized gesture or identification of a recognized gesture (e.g., indicating to radar-based application 208 on a remote computing device that a push or pull gesture has been performed).
[0038] The radar system 108 may also include at least one antenna 214 for transmitting and / or receiving radar signals. In some cases, the radar system 108 may include multiple antennas 214 implemented as antenna elements of an antenna array. The antenna array may include at least one transmitting antenna element and at least one receiving antenna element. In some cases, the antenna array may include multiple transmitting antenna elements to implement a multiple input multiple output (MIMO) radar capable of transmitting multiple different waveforms (e.g., different waveforms for each transmitting antenna element) at a given time. For implementations including three or more receiving antenna elements, the receiving antenna elements may be positioned in a one-dimensional shape (e.g., a line) or a two-dimensional shape (e.g., a triangle, a rectangle, or an L-shape). The one-dimensional shape may enable the radar system 108 to measure an angular dimension (e.g., an azimuth or an elevation), while the two-dimensional shape may enable two angular dimensions to be measured (e.g., both an azimuth and an elevation). Each antenna 214 may alternatively be configured as a transducer or a transceiver. In addition, any one or more antennas 214 may be circularly polarized, horizontally polarized, or vertically polarized.
[0039] Using an antenna array, the radar system 108 can form beams that are steered or unsteered, wide or narrow (e.g., one degree to 45 degrees, 15 degrees to 90 degrees), or shaped (e.g., shaped as a hemisphere, cube, sector, cone, or cylinder). One or more transmit antennas may have an unsteered omnidirectional radiation pattern, or may be capable of producing a wide steerable beam. Both of these technologies enable the radar system 108 to radar illuminate large volumes of space. To achieve target angular accuracy and angular resolution, the receiving antenna elements can be used to generate thousands of narrow steered beams (e.g., 2000 beams, 4000 beams, or 6000 beams) through digital beamforming. In this way, the radar system 108 can effectively monitor users and gestures in the environment.
[0040] The radar system 108 may also include at least one analog circuit 216, which includes circuitry and logic for transmitting and receiving radar signals using at least one antenna 214. The components of the analog circuit 216 may include amplifiers, mixers, switches, analog-to-digital converters, filters, etc. for conditioning radar signals. The analog circuit 216 may also include logic for performing in-phase / quadrature (I / Q) operations such as modulation or demodulation. A variety of modulations may be used to generate radar signals, including linear frequency modulation, triangular frequency modulation, stepped frequency modulation, or phase modulation. The analog circuit 216 may be configured to support continuous wave or pulse radar operations.
[0041] The analog circuit 216 can generate a radar signal (e.g., a radar transmit signal) in a spectrum (e.g., a frequency range) including frequencies between 1 gigahertz (GHz) and 400 GHz, 1 GHz and 24 GHz, 2 GHz and 6 GHz, 4 GHz and 100 GHz, or 57 GHz and 63 GHz. In some cases, the spectrum can be divided into multiple sub-spectra with similar or different bandwidths. Example bandwidths can be about 500 megahertz (MHz), 1 GHz, 2 GHz, etc. Different frequency sub-spectra can include, for example, frequencies between about 57 GHz and 59 GHz, 59 GHz and 61 GHz, or 61 GHz and 63 GHz. Although the example frequency sub-spectra described above are continuous, other frequency sub-spectra may not be continuous. In order to achieve coherence, the analog circuit 216 can use multiple frequency sub-spectra (continuous or discontinuous) with the same bandwidth to generate multiple radar signals, and the multiple radar signals are transmitted simultaneously or separated in time. In some cases, a single radar signal may be transmitted using multiple contiguous frequency sub-spectra, thereby enabling the radar signal to have a wide bandwidth.
[0042] The radar system 108 may also include one or more system processors 218 and system media 220 (e.g., one or more computer-readable storage media). For example, the system processor 218 may be implemented as a digital signal processor or a low-power processor (or both) within the analog circuit 216. The system processor 218 may execute computer-readable instructions stored within the system media 220. Example digital operations performed by the system processor 218 may include fast Fourier transforms (FFTs), filtering, modulation or demodulation, digital signal generation, digital beamforming, etc.
[0043] System media 220 may optionally include user module 222 and gesture module 224, which may be implemented using hardware, software, firmware, or a combination thereof. User module 222 and gesture module 224 may enable radar system 108 to process radar receive signals (e.g., electrical signals received at analog circuit 216) to detect the presence of user 104 and distinguish between the users and detect and recognize gestures, as well as other capabilities such as object (non-user) detection and detection of user interactions.
[0044] The user module 222 and the gesture module 224 may include one or more machine learning algorithms and / or machine learning models, such as an artificial neural network (referred to herein as a neural network), to improve user differentiation and gesture recognition, respectively. A neural network may include a set of connected nodes (e.g., neurons or perceptrons) organized into one or more layers. As an example, the user module 222 and the gesture module 224 may include a deep neural network, which includes an input layer, an output layer, and a plurality of hidden layers located between the input layer and the output layer. The nodes of the deep neural network may be partially connected or fully connected between layers.
[0045] In some cases, the deep neural network can be a recurrent deep neural network (e.g., a long short-term (LSTM) recurrent deep neural network), in which the connections between nodes form a cycle to retain information from a previous part of the input data sequence for a subsequent part of the input data sequence. In other cases, the deep neural network can be a feedforward deep neural network, in which the connections between nodes do not form a cycle. Figure 7 and Figure 8 Describe an example deep neural network. The user module 222 and the gesture module 224 may also include models capable of performing clustering (e.g., trained using unsupervised learning), anomaly detection, or regression, such as a single linear regression model, multiple linear regression models, a logistic regression model, a stepwise regression model, a multivariate adaptive regression spline, a local scatter smoothing estimation model, and the like.
[0046] In general, the machine learning architecture may be customized based on available power, available memory, or computing power. For user modules 222, the machine learning architecture may also be customized based on a certain amount of radar signal characteristics that radar system 108 is designed to recognize. For gesture modules 224, the machine learning architecture may additionally be customized based on a certain amount of gestures and / or various versions of gestures that radar system 108 is designed to recognize.
[0047] Computing device 102 may optionally (not depicted) include at least one additional sensor (other than antenna 214) to improve the fidelity of user module 222 and / or gesture module 224. In some cases, for example, user module 222 may detect the presence of user 104 with low confidence (e.g., confidence and / or accuracy metrics below a threshold). For example, such detection may occur when user 104 is far away from computing device 102 or when a large object (e.g., furniture) obscures user 104. To improve the accuracy of user detection and differentiation and / or gesture detection and recognition, computing device 102 may use one or more additional sensors (e.g., regarding Fig.19 The sensors described above) are used to verify low confidence results. These sensors can be passive, active, remote, and / or touch-based. Example sensors (some of which can sense in one or more of passive, active, remote, and touch) include microphones, ultrasonic sensors, ambient light sensors, cameras, health sensors and / or biometric sensors, barometers, inertial measurement units (IMUs) and / or accelerometers, gyroscopes, magnetic sensors (e.g., magnetometers or Hall effect), proximity sensors, pressure sensors, touch sensors, thermostats / temperature sensors, optical sensors, etc.
[0048] The user module 222 can also use the context information to distinguish between the users 104 (e.g., the first user 104-1 and the second user 104-2). This context information can also improve the interpretation of ambiguous gestures (e.g., gestures that cannot be recognized with a desired confidence level), such as Fig.24 For example, each user 104 may commonly perform their own version of a known gesture based on their personality, physical condition, ability / inability, mood, etc. This contextual information may be accessed by the gesture module 224 to improve gesture recognition.
[0049] In one example, first user 104-1 performs a gesture, and gesture module 224 determines that first user 104-1 accidentally performed an ambiguous gesture that is not associated with a known gesture (with a desired confidence level). However, if user module 222 determines that first user 104-1 (rather than second user 104-2 or another user) performed an ambiguous gesture, gesture module 224 may additionally access contextual information about stored radar signal characteristics of previous gesture performances by that user. Using this contextual information, gesture module 224 may be able to correlate the ambiguous gesture with one or more stored radar signal characteristics associated with the distinguished user to better identify it as a known gesture. In this way, user module 222 of computing device 102 may improve the fidelity of gesture recognition. Example computing system
[0050] Figure 3 An example environment 300 is shown in which a plurality of computing devices 102-1 and 102-2 are connected via a communication network 302 to form a computing system. The example environment 300 depicts a residence having a first room 304-1 (living room) and a second room 304-2 (kitchen). The first room 304-1 is equipped with a first computing device 102-1, which includes a first radar system 108-1, and the second room 304-2 is equipped with a second computing device 102-2, which includes a second radar system 108-2. In this example, the first room 304-1 is separated from the second room 304-2, but connected by a door in the home. The first computing device 102-1 in the first room 304-1 can detect a user 104 and a gesture within a first proximity zone 106-1, while the second computing device 102-2 in the second room 304-2 can detect a user 104 and a gesture within a second proximity zone 106-2.
[0051] The residence of the example environment 300 may not be limited to the arrangement and number of computing devices 102 shown. In general, an environment (e.g., a home, a building, a workplace, a car, an airplane, a public space) may include one or more computing devices 102 distributed across one or more different areas (e.g., rooms 304). For example, a room 304 may contain two or more computing devices 102 positioned close to or far away from each other (or radar systems 108 associated with a single computing device 102). Although the first proximity zone 106-1 depicted in the example environment 300 does not spatially overlap with the second proximity zone 106-2, and thus the computing devices in each area cannot sense the radar reception signals of the other area, in general, the proximity zones 106 may also be positioned to partially overlap. Although Figure 3The environment depicted in the example is a home, but in general, the environment may include any indoor and / or outdoor space, whether private or public, such as a library, office, workplace, factory, garden, restaurant, terrace, airplane, or automobile.
[0052] For environments with two or more computing devices 102, the devices may communicate with each other via one or more communication networks 302. The communication network 302 may be a LAN, WAN, a mobile or cellular communication network such as a 4G or 5G network, an extranet, an intranet, the Internet, Wi-Fi, or a wireless network. ® In some examples, computing device 102 may use a wireless communication system such as near field communication (NFC), radio frequency identification (RFID), Bluetooth ® Short distance communication.
[0053] In addition, the computing system may include one or more memories that are separate from or integrated into one or more of the constituent computing devices 102-1 and 102-2. In one example, the first computing device 102-1 and the second computing device 102-2 may include a first memory and a second memory, respectively, wherein the contents of each memory are shared between the devices using the communication network 302. In another example, the memory may be separate from the first computing device 102-1 and the second computing device 102-2 (e.g., cloud storage), but accessible to both devices. The memory may be used to store, for example, radar signal characteristics of registered users, user preferences, security settings, training history, unregistered user identification, and radar signal characteristics.
[0054] In one example, the first computing device 102-1 may use the first network interface 210-1 (see Figure 2 ) is connected to a communication network 302 to exchange information with a second computing device 102-2. Using this communication network 302, the computing devices 102 can exchange stored information about one or more users 104, which may include radar signal characteristics, training history, user settings, etc. In addition, the computing devices 102 can exchange information about ongoing operations (e.g., timers, music being played) to maintain continuity of operations and / or exchange information about operations across various rooms 304. These operations can be performed simultaneously or independently by one or more computing devices 102 based on, for example, detecting the presence of a user in the room 304. Each computing device 102 can also use the radar system 108 to associate (and store in memory) detected users and commands that are generally associated with the location of the device. About Figure 4 Radar system 108 is further described. Supports radar user detection and differentiation
[0055] Figure 4 An example environment 400 is shown in which radar systems 108 are used by computing devices 102 to detect the presence of users 104 and distinguish between the users. Example environment 400 depicts a first computing device 102-1 having a first radar system 108-1 and a second computing device 102-2 having a second radar system 108-2. The first radar system 108-1 and the second radar system 108-2 can transmit one or more radar transmit signals 402 (e.g., 402-Y, where Y represents an integer value of 1, 2, 3, ...) to detect users (and / or gestures) in a first proximity zone 106-1 and a second proximity zone 106-2, respectively. It should be noted that for simplicity, each of these zones is shown as a cone, but has a profile determined by the amplitude and quality of the radar field, where the radar receive signal can be received by the corresponding radar system 108 in that zone. Each radar transmit signal 402-Y can be referred to as a composite radar transmit signal 402-Y, which represents the composite radar transmit signal transmitted from the corresponding antenna 214 (see Figure 2) at a given time. Figure 2 ). Using radar transmit signal 402-1, first radar system 108-1 may illuminate an object (e.g., user 104) entering first proximity zone 106-1 with a wide 150° radar pulse beam (e.g., one or more radar transmit signals) operating at a frequency of 1 gigahertz to 100 gigahertz (GHz; e.g., 60 GHz). Although reference may be made to radar transmit signal 402 in the present disclosure, it will be understood that one or more radar transmit signals 402 may be transmitted and / or include one or more radar pulses over a period of time.
[0056] Upon encountering user 104, a portion of the energy associated with radar transmit signal 402-Y may be reflected back toward first radar system 108-1 and / or second radar system 108-2 in one or more radar receive signals 404-Z (where Z may represent an integer value of 1, 2, 3, ...). Each radar receive signal 404-Z may be referred to as a composite radar receive signal 404-Z that represents a superposition of multiple reflections of radar transmit signal 402-Y at one or more antennas 214 at a given time. In example environment 400, two radar receive signals 404-1 and 404-2 are depicted as being received by radar systems 108-1 and 108-2, respectively. Radar receive signals 404-1 and 404-2 may be reflected from one or more discrete dynamic scattering centers of user 104. Each radar receive signal 404 may represent a modified version of its corresponding radar transmit signal 402, where the amplitude, phase, and / or frequency are modified by one or more dynamic scattering centers. These radar received signals 404 may allow one or more of radar systems 108-1 and 108-2 to differentiate between users 104 and / or recognize gestures using, for example, radial distance, geometry (e.g., size, shape, height), orientation, surface texture, material composition, etc. For additional details on how this is performed, see at least the Figures 7 to 18 .
[0057] Although the first computing device 102-1 and the second computing device 102-2 in the example environment 400 can independently detect and distinguish the user 104 and / or recognize the gesture performance, they can also work together (e.g., have dependencies, collaboratively). This may be particularly useful if, for example, each device alone is unable to distinguish the user 104 and / or recognize the gesture performance with a desired confidence level. In this case, the two devices can exchange radar signal characteristics determined based on the radar receive signal received at each device. In one example, the first computing device 102-1 can detect an obscure user associated with a first radar signal characteristic within the first proximity area 106-1. Simultaneously or at a separate time, the second computing device 102-2 can also detect this obscure user and determine a second radar signal characteristic associated with the user's presence. If the first radar signal characteristic and the second radar signal characteristic are individually insufficient to distinguish the obscure user with a desired level of accuracy, the devices can work together and / or exchange information to achieve user differentiation.
[0058] In the first scenario, first computing device 102-1 may access the second radar signal characteristic and then compare the first radar signal characteristic and the second radar signal characteristic to one or more stored radar signal characteristics to distinguish the ambiguous user. In the second scenario, second computing device 102-2 may access the first radar signal characteristic and then compare the first radar signal characteristic and the second radar signal characteristic to one or more stored radar signal characteristics to distinguish the ambiguous user (e.g., by using Figure 7 or Figure 8 In a third scenario, first computing device 102-1 and second computing device 102-2 may work collaboratively, cooperatively, consistently, etc. to distinguish ambiguous users. This technique may also be applied to ambiguous gesture commands. Figure 5 Radar system 108 is described in greater detail.
[0059] Figure 5 An example implementation 500 is shown that includes the antenna 214, analog circuit 216, and system processor 218 of the radar system 108. In the depicted configuration, the analog circuit 216 can be coupled between the antenna 214 and the system processor 218 to implement techniques for user detection and differentiation and gesture detection and recognition. The analog circuit 216 can include a transmitter 502 equipped with a waveform generator 504 and a receiver 506 including at least one receive channel 508. The waveform generator 504 and the receive channel 508 can each be coupled between the antenna 214 and the system processor 218.
[0060] Although one antenna 214 is depicted in the example implementation 500, in general, the radar system 108 may include one or more antennas to form an antenna array. When utilizing an antenna array, the waveform generator 504 may generate similar or different waveforms for each antenna 214 to transmit into the proximity zone 106. Additionally, although one receive channel 508 is depicted in the example implementation 500, in general, the radar system 108 may include one or more receive channels. Each receive channel 508 may be configured to accept a single or multiple versions of the radar receive signal 404-Z at any given time.
[0061] During operation, transmitter 502 may communicate electrical signals to antenna 214, which may transmit one or more radar transmit signals 402-Y to detect user presence and / or gestures in proximity zone 106. In particular, waveform generator 504 may generate an electrical signal having a specified waveform (e.g., specified amplitude, phase, frequency). Waveform generator 504 may additionally communicate information about the electrical signal to system processor 218 for digital signal processing. If radar transmit signal 402-Y interacts with user 104, radar system 108 may receive radar receive signal 404-Z on receive channel 508. Radar receive signal 404-Z (or multiple versions thereof) may be sent to system processor 218 to enable user detection (using user module 222 of system media 220) and / or gesture detection (using gesture module 224). User module 222 may determine whether user 104 is within proximity zone 106 and then distinguish user 104 from other users. User 104 may be distinguished based on one or more radar receive signals 404-Z, such as regarding Figure 6 Further description.
[0062] Figure 6 Example implementations 600-1 to 600-4 are shown in which the user module 222 can distinguish between users 104. The user module 222 can use one or more radar received signals 404 in part to distinguish, for example, the first user 104-1 from the second user 104-2, with or without personally identifying the first user 104-1 or the second user 104-2. By distinguishing between the users 104, the user module 222 can enable the computing device 102 to provide each user 104 with a customized experience that retrieves, for example, training history, preferences, privacy settings, etc. In this way, the computing device 102 can improve some virtual assistant (VA)-equipped devices by meeting the privacy and / or functionality expectations of each user.
[0063] To distinguish users 104, user module 222 may analyze radar receive signal 404 to determine (1) topological distinction, (2) temporal distinction, (3) gesture distinction, and / or (4) contextual distinction of user 104. In the present disclosure, topological distinction, temporal distinction, gesture distinction, and contextual distinction may be determined based in part on one or more radar signal characteristics and, in some cases, on non-radar data from non-radar sensors. User module 222 is not limited to Figure 6 The four distinguishing categories depicted in the embodiment of the present invention may include other categories not shown. In addition, the four distinguishing categories are shown as example categories and may be combined and / or modified to include subcategories that implement the techniques described herein. Figure 6 The techniques described can also be applied as described herein (e.g., in relation to Fig.25 The gesture module 224 described in a similar manner as described in FIG.
[0064] In example implementation 600-1, user module 222 may use topological information in part to distinguish users 104. This topological information may include radar cross section (RCS) data, such as the height, stature or body type of user 104. For example, first user 104-1 (e.g., father) may be significantly larger than second user 104-2 (e.g., child). When father and child enter approaching area 106, radar system 108 may obtain radar receiving signals 404 indicating the presence of each user. These radar receiving signals 404 may partially include radar signal characteristics associated with topological information indicating the height, stature or body type of each user. In this example, the radar signal characteristics of the father may be different from the radar signal characteristics of the child. Then, user module 222 may compare the radar signal characteristics of each user with the radar signal characteristics of the stored registered users to determine whether the father and child are registered users (or unregistered personnel with relevant radar signal characteristics).
[0065] In this example, user module 222 can determine that first user 104-1 (father) is a registered user. In particular, user module 222 can correlate the father's stored radar signal characteristics (e.g., saved to a memory shared by multiple computing devices 102-X) with one or more radar reception signals 404 to determine that he is a registered user. After determining that first user 104-1 is a registered user, computing device 102 can activate the father's settings, prompt the father to continue gesture training (based on the father's training history), etc.
[0066] The user module 222 may also determine that the second user 104-2 (the child) is an unregistered person who does not have an account on the computing device 102. In particular, the user module 222 may compare the stored radar signal characteristics of the registered users to the one or more radar received signals 404 that include the radar signal characteristics of the second user 104-2. After determining that the topological information associated with the radar signal characteristics of the second user 104-2 does not correlate with one or more of the stored radar signal characteristics of the registered users (e.g., assuming a certain fidelity level), the user module 222 determines that the child is an unregistered person. The radar system 108 may assign an unregistered user identifier to the child, the unregistered user identifier including the radar signal characteristics of the child so that at a future time the topological information can be used to distinguish the child from other users (e.g., the father). This unregistered user identifier may include information such as Figures 12 to 15 The data shown, or information determined based on that data, such as the child's height range, movement data, or door, etc. The computing device 102 can also prompt the child to start gesture training and / or implement predetermined settings (e.g., standard preferences, settings programmed by the owner of the computing device 102).
[0067] In general, the stored radar signal characteristics of registered users may be collected and saved one or more times. For example, the radar system 108 may store the radar signal characteristics each time the user 104 interacts with the computing device 102 to improve user identification. The radar system 108 may also continuously store the radar signal characteristics over time to improve user and / or gesture detection. The stored radar signal characteristics of the user 104 may include topological, temporal, gesture, and / or contextual information inferred from one or more radar received signals 404 associated with that user 104. Additionally, the radar system 108 may utilize one or more models used by the user module 222 to distinguish each user based on the corresponding radar signal characteristics of each user 104. The one or more models may include machine learning (ML) models, predicate logic, hysteresis logic, etc. to improve user differentiation.
[0068] In the example implementation 600-2, the user module 222 may use temporal information in part to distinguish users 104. Unlike conventional radar detectors that may require high spatial resolution, the radar system 108 of the present disclosure may rely more on temporal resolution (rather than spatial resolution) to detect and distinguish users 104 and / or recognize gesture performance. In this way, the radar system 108 may distinguish users 104 moving into the proximity area 106 by receiving motion signatures (e.g., the unique way in which the user 104 typically moves). The motion signatures may include gait (depicted in the diagram of the example implementation 600-2), limb movements (e.g., corresponding arm movements), weight distribution, breathing characteristics, unique habits, etc. The motion signatures of a user may include limping, brisk steps, pigeon-toed walking, knock-kneed, bowed legs, etc. Using this information, the user module 222 may be able to detect the user's movements (e.g., the movement of their hands) without identifying details that may be considered private (e.g., facial features). Detecting motion signatures by radar system 108 may enable a user to maintain a higher degree of anonymity than by using, for example, devices performing facial recognition or speech-to-text technology.
[0069] Although described in the context of distinguishing a user from one or more other users or distinguishing a user as a specific registered user, these techniques can also be used in conjunction with detecting a user. Thus, the radar signal characteristics of the radar reception signal reflected from the user can be used to detect the presence of a user (e.g., any person) and whether the detected user is a specific user. Thus, the operations of detecting the presence of a user and distinguishing a user can be performed separately or as one operation.
[0070] In example implementation 600-3, user module 222 may also use, in part, gesture performance information to distinguish users 104. User module 222 may communicate with gesture module 224 when utilizing radar signal characteristics associated with gesture performance. A user may perform a gesture (e.g., a push-pull gesture) in a unique manner (or in a partially unique manner) that may help distinguish users 104 (while still sufficiently conforming to the push-pull gesture paradigm to be recognizable as a push-pull gesture, of course). For example, a push-pull gesture may include a user's hand pushing in one direction and then their hand immediately pulling in the opposite direction. While radar system 108 may expect push-pull motions to be complementary (e.g., equal extent of motion, equal rate), user 104 may perform the motion differently than desired. Each user may perform this gesture in a unique manner, which is recorded on the device for user differentiation and gesture recognition.
[0071] As depicted in example implementation 600-3, a first user 104-1 (e.g., a father) may perform a push-pull gesture differently than a second user 104-2 (e.g., a child). For example, the first user 104-1 may push their hand out to a first range (e.g., distance) at a first rate but pull their hand back to the second range at a second rate. The second range may include a shorter distance than the first range, and the second rate may be much slower than the first rate. The radar system 108 may be configured to recognize this unique or partially unique push-pull gesture based on the first user's training history (if available).
[0072] When distinguishing first user 104-1, radar system 108 may receive one or more radar receive signals 404 that include radar signal characteristics of first user 104-1 associated with their performance of the push-pull gesture. User module 222 may compare these radar signal characteristics with stored radar signal characteristics of registered users to determine if there is a correlation (see Figures 7 to 17 and Fig.25 and the accompanying description for ways to do this). If there is a correlation (e.g., assuming a certain level of fidelity), the user module 222 can determine that the first user 104-1 is a registered user (the father) based on the performance of the push-pull gesture. Similar to the teachings above regarding the example implementation 600-1, the computing device 102 can then activate the father's settings, prompt the father to continue gesture training, etc. In this example, it is assumed that the father has performed the push-pull gesture at least once in the past, and the radar signal characteristics of that performance are recorded in the father's training history to partially enable the father's presence to be distinguished from the presence of other users at a later time.
[0073] As depicted in example implementation 600-3, second user 104-2 (child) may attempt to perform a push-pull gesture. Second user 104-2 may push their hand out to a first range at a first rate but pull their hand back to a much larger third range at a third rate. In particular, radar system 108 may receive one or more radar receive signals 404, which include radar signal characteristics of second user 104-2 associated with the execution of this push-pull gesture. User module 222 may compare these radar signal characteristics with the stored radar signal characteristics of registered users to determine whether there is a correlation. Similar to the teachings of example implementation 600-1 above, user module 222 may determine that the push-pull gesture of second user 104-2 is not associated with the stored radar signal characteristics of registered users. Therefore, user module 222 may determine that the child is an unregistered person and assign an unregistered user identification to the child. However, the radar signal characteristics associated with the child's push-pull gesture may be included in the unregistered user identification to enable future distinction.
[0074] The user module 222 may also use context information in part to distinguish users 104. The context information may be determined by the user module 222 using, for example, the antenna 214, another sensor of the computing device 102, data stored on a memory (e.g., user habits), local information (e.g., time, relative position), etc. In the example implementation 600-4, the user module 222 may use the local time as a context to enable the distinction of a particular user. If a user 104 (e.g., a father) consistently sits on the sofa in the living room at 5:30 p.m. every day, the user module 222 may note this habit to improve user distinction. Whenever a user 104 is detected on the sofa at 5:30 p.m., the radar system 108 may use this context information in part to distinguish that user 104 as a father. In another example, if the computing device 102 is located in a child's room, the radar system 108 may determine over time that the child is the most common user in that room. That context information may be used to enable user distinction. Similarly, if computing device 102 is located in a shared space (e.g., backyard, entryway), radar system 108 may determine over time that unregistered persons (e.g., guests, nannies, housekeepers, gardeners, contractors, freelance helpers) are common in that area. It should be understood that while the scope of the present teachings is not necessarily limited to camera-less environments, and thus for some embodiments, cameras and facial recognition may be used to enhance contextual information, one advantageous feature provided by the camera-less embodiments described herein is that desired contextual information may indeed be derived without the use of a camera, as the presence of a camera in a home environment, particularly in more sensitive areas of the home, may evoke a sense of unease and invasion of privacy.
[0075] The context information collected by the user module 222 can be used alone to distinguish users 104 or in combination with topological information, time information, and / or gesture information. In general, the user module 222 can use any one or more of the depicted distinction categories in any combination at any time to distinguish users 104. For example, the radar system 108 can collect topological and time information about users 104 who have entered the proximity area 106, but lacks gesture and context information. In this case, the user module 222 can distinguish users 104 based on the analysis of topological and time information. In another case, the radar system 108 can collect topological and time information, but determines that the information is not enough to correctly distinguish users 104 (e.g., with a desired confidence level). If context information is available, the radar system 108 can use that context to distinguish users 104 (similar to in the example implementation 600-4). Figure 6 Any one or more of the categories depicted may take precedence over another.
[0076] The user module 222 can utilize one or more logic systems (e.g., including predicate logic, hysteresis logic, etc.) to improve user differentiation. The logic system can be used to prioritize certain user differentiation techniques over other techniques (e.g., favoring time differentiation over contextual information), add weights (e.g., confidence) to certain results when relying on two or more differentiation categories, etc. For example, the user module 222 can determine with a low confidence that the first user 104-1 may be a registered user. The logic system can determine that the low confidence is below an allowable threshold criterion (e.g., a limit) and instead prompt the radar system 108 to issue a second radar transmission signal 402-2 (or a collection of signals transmitted over a period of time) to detect the proximity area 106 again. The user module 222 can also include one or more machine learning models to improve user differentiation, such as regarding Figure 7 Further description.
[0077] Figure 7 An example implementation of a machine learning model 700 for distinguishing between users 104 and / or recognizing gestures is shown. The machine learning model 700 can perform classification, wherein the machine learning model 700 provides a numerical value for each of one or more classes that describes the degree to which the input data is believed to be classified into the corresponding class. In some instances, the numerical value provided by the machine learning model 700 can be referred to as a probability or "confidence score" that indicates a corresponding confidence associated with classifying the input into the corresponding class. In some implementations, the confidence score can be compared to one or more threshold criteria to present discrete classification predictions. In some implementations, only a certain number of classes (e.g., one) with relatively maximum confidence scores can be selected to present discrete classification predictions.
[0078] In an example implementation, the machine learning model 700 can provide probabilistic classification. For example, given a sample input, the machine learning model 700 can predict a probability distribution over a set of classes. Thus, instead of outputting the most likely class to which the sample input should belong, the machine learning model 700 can output the probability that the sample input belongs to this class for each class. In some implementations, the sum of the probability distributions of all possible classes can be one.
[0079] The machine learning model 700 can be trained using supervised learning techniques. For example, the machine learning model 700 can be trained on a training data set that includes training examples that are labeled as belonging to (or not belonging to) one or more classes. Before the user purchases the computing device 102, at least a portion of the training can be performed to initialize the machine learning model 700. This type of training is called offline training. During offline training, the training data set is not necessarily associated with the user. In some implementations, the computing device 102 enables the user to perform user gesture training. During user gesture training, the machine learning model 700 can collect a collection of new training data specific to the user and operate as a permanent learning machine by using the training data associated with the user for instant training. In this way, the machine learning model 700 can adapt to the user's unique radar feature markers and the way the user performs gestures to improve performance. This type of training is called on-line or online training.
[0080] In the depicted configuration, the machine learning model 700 is implemented as a deep neural network and includes an input layer 702, a plurality of hidden layers 704, and an output layer 706. The input layer 702 includes a plurality of inputs 708-1, 708-2 ... 708-N, where N represents a positive integer equal to the amount of the radar signal characteristic 710 associated with one or more radar receive signals 404. The plurality of hidden layers 704 may include layers 704-1, 704-2 ... 704-M, where M represents a positive integer. Each hidden layer 704 may include a plurality of neurons, such as neurons 712-1, 712-2 ... 712-Q, where Q represents a positive integer. Each neuron 712 may be connected to at least one other neuron 712 in the previous hidden layer 704 or the next hidden layer 704. The amount of neurons 712 may be similar or different between different hidden layers 704. In some cases, the hidden layer 704 may be a copy of the previous layer (e.g., layer 704-2 may be a copy of layer 704-1). The output layer 706 may include outputs 714 - 1 , 714 - 2 . . . 714 -N associated with differentiated users 716 (eg, registered users, unregistered persons) that may be detected within the proximity zone 106 .
[0081] In general, a variety of different deep neural networks can be implemented with various numbers of inputs 708, hidden layers 704, neurons 712, and outputs 714. The number of layers within the machine learning model 700 can be based on the radar signal characteristics and / or the amount of differentiation or recognition of classes (e.g., Figure 6 As an example, the machine learning model 700 may include four layers (e.g., one input layer 702, one output layer 706, and two hidden layers 704) to distinguish the first user 104-1 from the second user 104-2, as described with respect to the example environment 100 and the example implementation 600. Alternatively, the number of hidden layers may be approximately one hundred.
[0082] When utilized by the user module 222, the machine learning model 700 can improve the fidelity of user differentiation. The machine learning model 700 can collect multiple inputs 708 (e.g., radar signal characteristics 710 associated with one or more radar received signals 404) over time, the multiple inputs including topological, temporal, gesture, and / or contextual information about the user 104. For example, during a first interaction with the computing device 102, the second user 104-2 (e.g., a child) may be located away from the radar system 108, thereby generating a first set of inputs 708 for distinguishing the child as an unregistered person. The first set of inputs 708 can be included in the unregistered user identification assigned to the child. During a second interaction, the child may sit close to the radar system 108, thereby generating a second set of inputs 708 for distinguishing the child, the second set of the second inputs 708 being different from the first set and also being included in the unregistered user identification of the child. This process can continue over time, providing more inputs 708 to the machine learning model 700 to better differentiate between children at future times (e.g., with higher accuracy, at a faster rate).
[0083] When utilized by user module 222, machine learning model 700 analyzes complex radar data (e.g., phase and / or amplitude data) and generates probabilities. Some of the probabilities are associated with various gestures that radar system 108 can recognize. Another of the probabilities may be associated with a background task (e.g., background noise or gestures that radar system 108 does not recognize). Although described with respect to gestures, machine learning model 700 may be extended to indicate other events, such as whether a user is present within a given distance.
[0084] The gesture module 224 may also collect gesture execution information of the user 104 during the user gesture training as an input to the machine learning model 700 to enable the user module 222 to distinguish users based on gesture execution. If the first user 104-1 performs four push and pull gestures during the user gesture training, there may be at least four inputs 708-1, 708-2, 708-3, and 708-4 to the machine learning model 700. The user module 222 may partially utilize one or more outputs 714 of the machine learning model 700 to distinguish the first user 104-1 (e.g., father) when the first user performs a push and pull gesture at a future time.
[0085] In general, the machine learning model 700 can be integrated into the user module 222, the radar system 108, or the computing device 102, or located separately from the computing device 102 (e.g., a shared server). The gesture module 224 can also include a similar machine learning model 700, which can improve the detection and recognition of gestures performed by the user 104. For example, the gesture module 224 can detect one or more radar signal characteristics 710 associated with a gesture performed by the first user 104-1. The gesture module 224 can utilize the output 714 of the machine learning model 700 to identify the gesture as a known gesture (e.g., a push-pull gesture) associated with a command (e.g., starting the oven). The operations of the gesture module 224 can be performed simultaneously with the operations performed by the user module 222 or at a separate time. The gesture module 224 can additionally include one or more deep learning algorithms, such as a convolutional neural network (CNN), to improve the detection and recognition of gestures. About Figure 8 An example of integrating a CNN into the gesture module 224 is further described.
[0086] Although the above Figure 7 and below Figure 8 The machine learning model is described as distinguishing users and recognizing gestures, but detecting a user or a gesture can be performed as an operation of distinguishing that user or recognizing that gesture, respectively. However, in some cases, multiple or more complex operations are used, such as when an attempt to detect and recognize a gesture fails because it is sufficient to detect but not recognize the gesture (for example, where the correlation with known gestures is too low to recognize which gesture was performed, such as recognizing with a low confidence level, but the correlation is sufficient to determine that the gesture is a certain gesture rather than a non-gesture movement). Therefore, one or more radar signal characteristics of one or more radar received signals reflected from a user can be used to detect that a gesture has been performed and also to recognize that the detected gesture is a known gesture. Example spatiotemporal machine learning model
[0087] Figure 8An example implementation 800 is shown that includes a gesture module 224 that utilizes a spatiotemporal machine learning model 802 (e.g., one or more CNNs) to improve detection and recognition of gestures. This spatiotemporal machine learning model 802 can enable the computing device 102 to detect and recognize gestures at a desired confidence level over long-range extant distances (such as four meters) as well as close distances (such as a few centimeters). The gesture module 224 is depicted as having a signal processing module 804, a frame model 806, a temporal model 808, and a gesture debouncer 810.
[0088] The space-time machine learning model 802 has a multi-stage architecture, which includes a first stage (e.g., a frame model 806) and a second stage (e.g., a time model 808). In the first stage, the space-time machine learning model 802 processes complex radar data (e.g., complex range Doppler maps) across the spatial domain, which involves processing complex radar data burst by burst. In the second stage, the space-time machine learning model 802 connects the results of the frame model 806 across multiple bursts. By connecting the results, the second stage processes complex radar data across the time domain. Through the multi-stage architecture, the overall size and inference time of the space-time machine learning model 802 may be significantly less than those of other types of machine learning models. This property can enable the space-time machine learning model 802 to run on a computing device 102 with limited computing resources.
[0089] The gesture module 224 is not limited to the arrangement depicted in the example implementation 800 and may include additional or fewer components as shown. For example, the gesture module 224 may lack a gesture de-jitter 810 but include multiple signal processing modules 804, which are arranged before the frame model 806, before the time model 808, and / or after the time model 808. Additionally, any one or more of the depicted components of the gesture module 224 may be arranged separately from the gesture module 224. For example, the output of the time model 808 may be sent to the gesture de-jitter 810 that is separate from the gesture module 224. The spatio-temporal machine learning model 802 may also be separate from the gesture module 224. In one example, the spatio-temporal machine learning model 802 may be arranged within the radar system 108 but separate from the gesture module 224. In another example, the spatio-temporal machine learning model 802 may be separate from the computing device 102 (e.g., located on a remote server). This document (such as regarding Fig.25 ) describes additional details regarding the recognition of gestures.
[0090] In the example implementation 800, three radar transmit signals 812-1, 812-2, and 812-3 are transmitted using antennas 214-1, 214-2, and 214-3, respectively. The three radar transmit signals 812-1, 812-2, and 812-3 represent component signals that may be superimposed during propagation to form a composite transmit signal 402-Y. The composite transmit signal 402-Y propagates into a surrounding environment (e.g., a residence). The composite transmit signal 402-Y is reflected, such as from an environment 814 and / or a gesture 816 performed by a user 104. The environment 814 may include, for example, fixed environments, such as stationary objects (e.g., furniture), as well as non-fixed environments, such as movement of objects not associated with gestures (e.g., users 104 walking and / or interacting with their environment 814, ceiling fans, movement of livestock, etc.). Figure 8 8, composite transmit signal 402-Y is reflected from environment 814 and gesture 816 (e.g., a user's hand) to produce composite radar receive signal 404-Z. Antennas 214-1, 214-2, and 214-3 each receive a version of composite radar receive signal 404-Z, represented by radar receive signals 818-1, 818-2, and 818-3. Radar receive signals 818-1, 818-2, and 818-3 correspond to at least three respective radar signal characteristics and are sent to analog circuit 216 (see FIG. 1 ) before being sent to signal processing module 804. Figure 5 ). Analog circuitry 216 may modify (eg, digitize) radar receive signal 818 (associated with radar signal characteristics) to enable operation of signal processing module 804 .
[0091] exist Figure 8 In one example, computing device 102 transmits composite radar transmit signal 402-Y as a burst of 16 chirps at a high pulse repetition rate (PRF) of 3 kilohertz (kHz). Each burst includes a wide 150 degree radar beam of a frequency modulated continuous wave to illuminate the surrounding environment of proximity zone 106 (e.g., environment 814 and gesture 816). Each burst is transmitted periodically over time (at a rate of 30 Hz) to enable unsegmented detection of gestures. Each antenna 214 captures a superposition of reflections (corresponding to radar receive signals 404-Z) from scattering surfaces within proximity zone 106 over a long range (e.g., four meters, but other distances are also contemplated, such as approximately two meters, six meters, or eight meters). In general, radar transmit signals can be transmitted over a variable time period that is not a fixed time period (e.g., a segmented detection period).
[0092] Although three radar transmission signals 812 are depicted in the example implementation 800, in general, the computing device 102 may transmit one or more signals simultaneously from one or more antennas 214. In general, the computing device 102 may detect and recognize gestures at one or more locations within the proximity zone 106 that extend to a long range (e.g., a linear distance of one to four meters). The computing device 102 of the present disclosure does not require the user 104 to perform gestures at any particular location within the proximity zone 106 (e.g., above the interface of the device), which may allow the user 104 the freedom to comfortably perform gesture commands from a variety of locations, such as in their home, without having to approach the interface of the device.
[0093] In the example implementation 800, radar receive signals 818-1, 818-2, and 818-3 are processed by a signal processing module 804, which applies a high pass filter to remove reflections from stationary objects. This high pass filter may include, for example, one or more resistors, capacitors, inductors, operational amplifiers (op amps), etc. Alternatively or in addition, each radar receive signal 818-1, 818-2, and 818-3 may be processed with multiple stages of a Fast Fourier Transform (FFT) to generate one or more complex range Doppler maps 820-A (with respect to Fig.12 describes an example).
[0094] The complex range Doppler map 820 can be a two-dimensional representation that includes a distance dimension (e.g., a tilted distance dimension) and a Doppler dimension. The distance dimension can correspond to the displacement of the scattering surface of the gesture 816 (e.g., the surface of the user's hand) relative to the computing device 102, and the Doppler dimension can correspond to the range rate of the scattering surface relative to the computing device 102. Therefore, one or more complex range Doppler maps 820 can enable the radar system 108 to determine the relative position and movement of an object (e.g., gesture 816) within its proximity area 106. The radar system 108 can determine one or more complex range Doppler maps 820 with similar or different FFT window sizes over time. For example, the FFT window size can be set to 128×16, corresponding to the bin size of the range and Doppler data, respectively. In this example, the range resolution Δr is 0.027 meters (m), and the Doppler resolution Δf d is 0.38 meters per second (m / s), as defined by the following equation: Where c is the speed of light, B is the transmit bandwidth, which is set to 5.5 GHz, PRF is the pulse repetition frequency, which is set to 3 kHz, and f c is the center frequency, which is set to 60.75 GHz, and l is the number of chirps per burst, which is set to 16.
[0095] The signal processing module 804 then sends one or more complex range Doppler maps 820 to the frame model 806. As depicted in the example implementation 800, the signal processing module 804 sends three complex range Doppler maps 820-1, 820-2, and 820-3 to the frame model 806. Although one frame model 806 is depicted, the gesture module 224 may include one or more frame models 806 that each utilize CNN (convolutional neural network) technology. Referring to the previous example, the frame model 806 can receive the complex range Doppler maps 820-1 to 820-3, which are formatted to be 128 (the number of range bins) times 16 (the number of Doppler bins) times 6 (3 antennas). 2 values), where the "2 values" correspond to real and imaginary values as floating point representations. If proximity zone 106 is reduced to a smaller size (e.g., 1.5 m), the size of the tensor may be reduced. In this case, the number of distance bins may be clipped at the 64th bin to correspond to user 104 standing 1.7 m from computing device 102 and performing gesture 816 with their hand at a distance of 1.5 m from the device.
[0096] The frame model 806 may output a frame result 822-B, which includes a one-dimensional representation of the complex range Doppler map 820-A that has been processed for each burst. In this case, the frame results 822-1, 822-2, and 822-3 associated with different bursts may be sent to the temporal model 808 to be connected along the time domain and processed using similar or different CNN techniques. The temporal model 808 may calculate one or more gesture probabilities (e.g., temporal results 824-C, where variable C represents the amount of class analyzed by the temporal model 808) for one or more gesture classes (e.g., known gestures) and / or background classes (e.g., background motion, objects not associated with known gestures) to send to the gesture debouncer 810. For example, the gesture debouncer 810 may be configured to recognize five possible gesture classes (e.g., tap, swipe up, swipe down, swipe right, and swipe left) and one background class (e.g., motion and objects not associated with the five possible gesture classes).
[0097] Gesture debouncer 810 may enable computing device 102 to perform unsegmented gesture detection, which may enable the device to detect gestures without first receiving an indication that a gesture is to be performed and / or without prior knowledge that a gesture is to be performed. For example, computing device 102 may detect gesture 816 without requiring user 104 to prompt the device with a “wake-up” trigger. The wake-up trigger may include a verbal, visual, or gesture prompt made by user 104 to indicate to computing device 102 that performance of gesture 816 is about to occur. The wake-up trigger may be detected by any one or more sensors of computing device 102, such as antenna 214, microphone, ambient light sensor, pressure sensor, camera, etc. Because a wake-up trigger event is not required, radar system 108 continuously (e.g., in an unsegmented manner) scans and detects gesture 816 at any time.
[0098] To prevent erroneous detection or recognition of gesture 816, gesture debouncer 810 may apply one or more heuristics to temporal result 824-C. False recognition may include, for example, incorrectly associating gesture 816 with a known gesture or with multiple known gestures. A first heuristic may include a requirement that temporal result 824-C (e.g., result of spatiotemporal machine learning model 802) should have a value greater than an upper threshold criterion (e.g., a set value) within the last three consecutive frames, i.e., a maximum threshold requirement for confidence in the result. A second heuristic may require that temporal result 824-C have a value less than a lower threshold criterion when multiple gestures are detected within a certain time period, i.e., a minimum threshold requirement for the elapsed time between gesture executions. If a second gesture 816-2 is performed by user 104 quickly after execution of a first gesture 816-1, gesture debouncer 810 may trust that the two gestures 816 are indications of separate actions performed within the time period. For example, after first gesture 816-1 has been detected, time result 824-C may have a value less than a lower threshold of the elapsed time before second gesture 816-2 was detected. These upper and lower thresholds may be determined experimentally or customized based on user needs and performance.
[0099] Gesture debouncer 810 may apply the one or more heuristics to determine a gesture result 826. This gesture result 826 may be an indication of the most likely classification of the object and / or motion detected by radar system 108 (e.g., correlation of radar signal characteristics). If gesture module 224 can classify up to six classes, gesture debouncer 810 will indicate which of the six classes has been detected.
[0100] In a first example, assume that first radar transmit signal 402-1 is sent into proximity zone 106-1 of computing device 102-1 and reflects off a cat walking through the room (at Figure 4The first radar received signal 404-1 is shown in the analog circuit 216 of the radar system 108-1 (see Figure 2 ) is received and then sent to the signal processing module 804 ( Figure 8 ). Signal processing module 804 cleans up first radar receive signal 404-1 by applying a high pass filter to remove signals associated with stationary objects (e.g., toys on the floor, not shown). First complex range-Doppler map 820-1 is sent to first frame model 806-1, which performs one or more CNN techniques and transforms the map into a one-dimensional array (first frame result 822-1). First frame result 822-1 is sent to time model 808 (see Figure 8 ), where the device calculates the probability that the moving cat is any one of six classes (five gesture classes and one background class). The temporal model 808 determines that the probability of each of the five gesture classes is 0.01 (on a scale of 0 to 1.00) and the probability of the background class is 0.95. These six values (e.g., the first temporal result 824-1) are sent to the gesture de-jitter 810. The gesture de-jitter 810 applies a first heuristic, requiring that the probability of the class is greater than an upper threshold of 0.80. Since the probability of the background class is 0.95, the gesture de-jitter 810 sends a first gesture result 826-1 indicating that the cat belongs to the background class, thereby suggesting that the cat's movement is not associated with one of the five gesture classes. Although the scope of the present teachings is not limited in this regard, it is assumed in this example that the six classes are mutually exclusive and the sum of all six probabilities is equal to 1.00.
[0101] exist Figure 4In the second example shown, a second radar transmit signal 402-2 is sent into proximity zone 106-2 of computing device 102-2 and reflected from user 104 (located four meters from the device) who is greeting a family member by waving their hand. Second radar receive signal 404-2 is received by radar system 108-2 (using a similar set of techniques as described in the previous example, and at analog circuit 216), and a second frame result 822-1 is sent to temporal model 808. The device determines that the motion of user 104 waving their hand has a probability of 0.2 of swiping up, a probability of 0.2 of swiping down, a probability of 0.2 of swiping left, a probability of 0.2 of swiping right, a probability of 0.1 of tapping, and a probability of 0.1 of background class. Second temporal result 824-2 is sent to gesture debouncer 810, which determines that none of the six classes have a probability value greater than an upper threshold value set to 0.8. Instead of outputting second gesture result 826-2, gesture debouncer 810 may communicate (to gesture module 224, radar system 108, and / or computing device 102, for example) that insufficient information is available regarding whether the user's hand wave is a gesture or background motion. Based on this indication, gesture module 224 may instruct the device to transmit third radar transmission signal 402-3 ( Figure 4 ) to obtain additional radar signal characteristics (e.g., using a third complex range Doppler map 820-3, a third frame result 822-3, and a third time result 824-3, similar to the above example). As described, the technology can use one or more additional sensors (e.g., a microphone) to collect supplemental data, and / or can utilize additional information (e.g., contextual information) to discern whether the user's wave is intended as a command to the device. One such approach includes audio, which is not necessarily speech recognition. For example, even if users 104 are not distinguished and therefore computing device 102 does not yet know which user is performing the gesture, audio and other supplemental data can be used to alter or establish the probability that a particular gesture is being performed. hereinafter in Fig.19 Examples are set forth in the description of .
[0102] In the example implementation 800, the frame model 806 and the temporal model 808 may collectively form a "space-time machine learning model" (identified at 802) that utilizes CNN techniques, artificial intelligence, logical systems, residual neural networks (ResNet), dense layers, etc. Any one or more components in the space-time machine learning model may be repeated, rearranged, or ignored as desired. Fig. 9 An example structure for frame model 806 is depicted in .
[0103] Fig. 9An example implementation 900 of spatiotemporal machine learning techniques utilized by the frame model 806 is shown. These techniques may utilize a neural network having several layers arranged as depicted (see Figure 7 ). Any one or more of the depicted layers may be rearranged, removed, or repeated to form an alternative neural network that may also enable detection and recognition of gestures over long distances.
[0104] As depicted, the complex range-Doppler map 820 (e.g., input tensor) is first sent to an average pooling layer 902 to reduce the size of the map and the computational cost of subsequent subsequent layers. The average pooling layer 902 can be used to reduce the size of the map by averaging a set of values associated with the complex range-Doppler map 820 as an input to a separable two-dimensional (2D) residual block 904. Although a separable convolution layer is depicted in the example implementation 900, the frame model 806 can alternatively or additionally utilize a standard convolution layer. The separable convolution layer is used to divide the matrix into its constituent (two) kernel portions. For example, a 3×3 complex range-Doppler map 820 may require one convolution with 9 multiplications, while the constituent kernel portions of the matrix (1×3 and 3×1 kernels) may only require two convolutions with three (3) multiplications, thereby reducing computational time.
[0105] In general, residual blocks can be associated with ResNets that can skip connections (e.g., layers) within a neural network. Example separable 2D residual block 904 in Fig. 9 is depicted in the dashed box as having two example paths. On the first path, the results from the average pooling layer 902 are input to the separable 2D convolution layer 906. The filter can slide over these inputs, performing element-by-element multiplication and summation at each position. In one example, the spatio-temporal machine learning model 802 can apply a 2×2 filter to a 3×3 input matrix of the separable 2D convolution layer 906. This 2×2 filter can slide over the values of the input matrix, producing a 2×2 matrix output. Alternatively, the edges of the 3×3 input matrix can be padded at the separable 2D convolution layer 906 to output a 3×3 matrix instead of a 2×2 matrix. Additionally, striding (e.g., skipping one or more positions as the filter slides over the input matrix) can be performed at the separable 2D convolution layer 906.
[0106] The output matrix of the separable 2D convolution layer 906 can be sent to a batch normalization layer 908, in which the values of the output matrix are standardized to improve the stability and rate of the space-time machine learning model 802. For example, the values of the output matrix can be standardized by calculating the mean and standard deviation of each value. In another example, the values can be standardized by calculating the running average of the mean and standard deviation. These standardized results can be sent to a rectifier (ReLU) 910, which is an activation function defined as the positive part of its independent variable. ReLU 910 can be used to prevent the separable 2D residual block 904 from activating all neurons 712 at the same time, thereby preventing the exponential growth of computational requirements. ReLU 910 can include, for example, linear (e.g., parameter) or nonlinear (e.g., Gaussian, sigmoid, analytical, logical) functions. The modified results from ReLU 910 can be sent to another separable 2D convolution layer 906 that is similar or different from the previous convolution layer. Those results may be processed at another batch normalization layer 908 and then sent to a summing node 912 where the results of this first path are summed with the results of the second path of the separable 2D residual block 904 .
[0107] On the second path, the results from the average pooling layer 902 bypass the layers of the first path and are instead sent to a 2D convolutional layer 914, which may include a two-dimensional standard convolutional layer that does not separate the matrix into constituent kernels. The output matrix from the 2D convolutional layer 914 may be sent to a summing node 912 where the results from the first and second paths are summed and sent to another ReLU 910.
[0108] The frame model 806 can then implement a series of separable 2D residual blocks 904 and maximum pooling layers 916 (collectively labeled by 918). Each separable 2D residual block 904 can process data using different or similar algorithms. For example, each block can use different or similar filter sizes (e.g., 1×1, 2×2, 3×3, etc.) and / or strides (e.g., filtering every 1, 2, 3 values, etc.). Unlike at the average pooling layer 902, one or more maximum values from a set of inputs can be determined at the maximum pooling layer 916. At the maximum pooling layer 916, for example, the maximum value of a 4×4 matrix can be calculated by sliding a 2×2 window on the matrix. Although this example uses a 2×2 window, in general, the maximum pooling layer 916 can utilize a 1×1 window, a 3×3 window, etc. Each maximum pooling layer 916 depicted in the example implementation 900 may include a window that is different or similar to the window of another maximum pooling layer 916.
[0109] At the end of the frame model 806, the data is sent to the final separable 2D convolution layer 906 and then to the flattening layer 920. The flattening layer 920 can reduce the data to a one-dimensional array (e.g., a frame summary 922). This frame summary 922 can be sent to the temporal model 808 for further processing, such as Fig.10 described.
[0110] Fig.10 An example implementation 1000 of a machine learning technique utilized by the temporal model 808 is shown. Similar to the frame model 806, the technique may include a neural network having several layers arranged as depicted. Any one or more of the depicted layers may be rearranged, removed, or repeated to form an alternative neural network that may also enable detection and recognition of gestures over long ranges.
[0111] As depicted, frame summary 922 may be sent to temporal model 808 to correlate frames of radar receive signal 404 in the time domain. Frame summary 922 may first be processed at one-dimensional (1D) residual block 1002. 1D residual block 1002 may be similar to a separable 2D residual block, except that the calculation is performed in one dimension using a standard convolution (e.g., rather than using a separable convolution). Example 1D residual block 1002 in Fig.10 906, except that the calculation is performed with a standard convolution in one dimension. These results can be sent to a batch normalization layer 908, followed by a ReLU 910. The data can be processed by another 1D convolution layer 1004 that is similar or different from the previous 1D convolution layer 1004. Those results can be processed at another batch normalization layer 908 and then sent to a summing node 912, where the results of this first path are summed with the results of the second path of the 1D residual block 1002.
[0112] On the second path, the frame summary 922 bypasses the layers of the first path and is sent to a 1D convolutional layer 1004, which may be similar to or different from the 1D convolutional layer 1004 of the first path. The results from the first and second paths are summed at a summing node 912 and sent to another ReLU 910.
[0113] The temporal model 808 may then implement a series of 1D residual blocks 1002 and max pooling layers 916 (collectively labeled 1006). Fig. 9As discussed above, each 1D residual block 1002 can process the data using different or similar algorithms. At the end of the temporal model 808, the data is sent to a dense layer 1008 and then to a softmax layer 1010. At the dense layer 1008 (e.g., a fully connected layer), each neuron can receive data from all neurons in the previous layer. The size of the data can change at the dense layer 1008 and reflect the number of classes that can be used to classify gestures. For example, if the gesture module 224 has five gesture classes and one background class, the output from the dense layer 1008 can reflect those six classes. This output can be sent to a softmax layer 1010, where a softmax function is applied to this data to assign probabilities to each class. For example, if there are six classes, each of the six classes can be assigned a probability value from 0 to 1. The sum of all six class probabilities can add up to 1. The gesture probability 1012 is then sent from the temporal model 808 to the gesture debouncer 810 (see the description of Figure 8 ) to enable detection and recognition of gestures. Offline Training for Radar-Based Gesture Recognition
[0114] The gesture module 224 may be trained using an offline supervised training technique. In this case, the recording device records data generated by the radar system 108. The recording device is coupled to the radar system 108 to capture complex radar data. The recording device may be a separate unit connected to the radar system 108. Alternatively, the recording device may be integrated within the radar system 108 or the computing device 102.
[0115] For offline training, radar system 108 collects positive recordings when a participant performs a gesture, such as a right swipe using the left hand and a left swipe using the right hand. In general, a positive recording represents complex radar data recorded by radar system 108 or a recording device during a time period when a participant performs a gesture associated with a gesture class.
[0116] Positive records can be collected using participants with various heights and handedness types (e.g., right-handed, left-handed, or ambidextrous). Moreover, positive records can also be collected when the participant is located at various positions relative to the radar system 108. For example, the participant can perform gestures at various angles (including angles between about -45 degrees and 45 degrees) relative to the radar system 108. As another example, the participant can perform gestures at various distances (including distances between about 0.3 meters and 2 meters) from the radar system 108. Additionally, positive records can be collected with different recording device placements (e.g., on a table or in the hands of the participant) and with various orientations (e.g., longitudinal orientation or transverse orientation) of the recording device when the participant takes various postures (e.g., sit, stand, or lie down).
[0117] For offline training, radar system 108 also collects negative records while the participant performs background tasks. The background task may include the participant operating a computer or computing device 102. Another background task may include the participant walking around radar system 108. In general, negative records represent complex radar data recorded by radar system 108 during a period of time when the participant performs a background task associated with the background class (or a task not associated with the gesture class).
[0118] The participant may perform background motions that resemble gestures associated with one or more of the gesture classes. For example, the participant may move their hand between a computer and a mouse, which may resemble a directional swipe gesture. As another example, the participant may put a cup down on a table next to the recording device and then pick up the cup, which may resemble a tap gesture. By capturing these gesture-like background motions in negative recordings, the gesture module 224 may be trained to detect the difference between background tasks with gesture-like motions and intentional gestures intended to control the computing device 102.
[0119] Negative records may be collected in a variety of environments, including a kitchen, bedroom, or living room. In general, negative records capture natural behavior around radar system 108, which may include a participant reaching for computing device 102, dancing nearby, walking, cleaning a table with computing device 102 on it, or turning the steering wheel of a car with computing device 102 in a stand. Negative records may also capture repetitions of hand movements similar to a swipe gesture, such as moving an object from one side of radar system 108 to another. For training purposes, negative records are assigned background labels that distinguish them from positive records. To further improve the performance of gesture module 224, negative records may optionally be filtered to extract samples associated with motions with speeds above a predefined threshold criterion.
[0120] The positive and negative records are split or divided to form a training data set, a development data set, and a test data set. The ratio of positive records to negative records in each of the data sets can be determined to maximize performance. In the example training process, the ratio ratio is 1:6 or 1:8.
[0121] The recording device can refine the timing of gesture segments in the positive example recording. To do this, the recording device detects the center of the gesture motion within the gesture segment of the positive example recording. As an example, the recording device detects a zero Doppler crossing within a given gesture segment. A zero Doppler crossing can refer to an instance in time where the motion of the gesture changes between a positive Doppler bin and a negative Doppler bin. In other words, a zero Doppler crossing can refer to an instance in time where the Doppler-determined range rate changes between a positive value and a negative value. This indicates a time where the direction of the gesture motion becomes substantially perpendicular to the radar system 108, such as during a swipe gesture. It can also indicate a time where the direction of the gesture motion reverses and the gesture motion becomes substantially stationary, such as during the execution of a tap gesture. Other indicators can be used to detect the center point of other types of gestures.
[0122] The recording device aligns the timing window based on the detected center of the gesture motion. The timing window can have a specific duration. This duration can be associated with a specific amount of bursts (such as 12 or 30 bursts). Generally speaking, the amount of bursts is sufficient to capture gestures associated with the gesture class. In some cases, additional offsets are included in the timing window. The offset can be associated with the duration of one or more bursts. The center of the timing window can be aligned with the detected center of the gesture motion.
[0123] The recording device resizes a given gesture segment based on its aligned time window to generate pre-segmented data. For example, the size of the gesture segment is reduced to include samples associated with the aligned timing window. The pre-segmented data can be provided as part of a training data set, a development data set, and a test data set.
[0124] The gesture module 224 can be trained using a training data set and supervised learning. As described above, the training data set can include pre-segmented data. This training enables optimization of the internal parameters of the gesture module 224, including weights and biases.
[0125] First, the hyperparameters of the gesture module 224 are optimized using a development dataset. As described above, the development dataset may include pre-segmented data. In general, hyperparameters represent external parameters that do not change during training. A first type of hyperparameter includes parameters associated with the architecture of the gesture module 224, such as the number of layers or the number of nodes in each layer. A second type of hyperparameter includes parameters associated with the processing of the training data, such as a learning rate or the number of rounds. The hyperparameters may be selected manually, or may be automatically selected using techniques such as grid search, black box optimization techniques, gradient-based optimization, and the like.
[0126] Secondly, the gesture module 224 is evaluated using the test data set. In particular, a two-stage evaluation process is performed. The first stage includes performing a segmented classification task using the gesture module 224 and the pre-segmented data in the test data set. Without using the gesture debouncer 810 to determine whether a gesture occurs, the gesture is determined based on the highest probability provided by the temporal model 808. By performing the segmented classification task, the accuracy, precision, and recall of the gesture module 224 can be evaluated.
[0127] The second stage includes using the gesture module 224 and the gesture de-jitter 810 to perform an unsegmented recognition task. Instead of using pre-segmented data within a test data set, the unsegmented recognition task is performed using a duration series of data (or a continuous data stream). By performing the unsegmented recognition task, the recognition rate and / or false positive rate of the gesture module 224 can be evaluated. In particular, the unsegmented recognition task can be performed using positive example records to evaluate the recognition rate and the unsegmented recognition task can be performed using negative example records to evaluate the false positive rate. The unsegmented recognition task utilizes the gesture de-jitter 810, which enables further tuning of the threshold criteria to better achieve the desired recognition rate and the desired false positive rate.
[0128] If the results of the segmented classification task and / or the unsegmented recognition task are not satisfactory, one or more elements of the gesture module 224 may be adjusted. These adjustments may extend to the overall architecture, training data, and / or hyperparameters of the gesture module 224. With these adjustments, the training of the gesture module 224 may be repeated. The positive and / or negative records may be enhanced to further enhance the training of the gesture module 224, as further described below. Data augmentation techniques
[0129] Data augmentation can be used to enhance the data set of the spatiotemporal machine learning model 802 (e.g., complex range-Doppler maps 820, frame results 822, and / or time results 824 associated with radar signal characteristics) to increase the amount of stored radar signal characteristics without requiring a large number of interactions (e.g., 50 or more interactions) between the user 104 and the computing device 102 for offline or online training. The radar enhancement techniques described in the present disclosure include determining a random or predetermined phase rotation and / or amplitude scaling of data corresponding to one or more radar signal characteristics. By implementing these radar enhancement techniques, the computing device 102 can reduce the amount of gesture training required to accurately recognize gestures with a desired confidence level. For example, the computing device 102 may only need to collect three radar signal characteristics (instead of ten) of the user 104 performing a swipe gesture to accurately recognize the command. As a result, the user 104 can quickly enjoy using the computing device 102 without having to conduct time-consuming gesture training.
[0130] For complex range-Doppler maps 820, the absolute phase may be affected by the surface position at the range bin resolution, phase noise, sampling timing errors, etc. In addition, the amplitude of the complex range-Doppler map 820 may be affected by the properties of the antenna 214, the continuity between the computing devices 102 (when utilizing a computing system), the reflectivity of the signal from the scattering surface, the orientation of the scattering surface, etc. For large data sets, these absolute phases and amplitudes may be evenly distributed. However, for small data sets (e.g., corresponding to a new user beginning gesture training), these absolute phases and amplitudes may be biased, thereby reducing the accuracy of gesture detection and / or recognition.
[0131] To address these issues without requiring user 104 to undergo time-consuming gesture training, computing device 102 may utilize radar enhancement techniques to enhance the phase and / or amplitude of complex range-Doppler map 820 (M) based on the following relationship: Where A is the enhanced complex range-Doppler map, r is the range bin index, d is the Doppler bin index, and c is the channel index (ref. Figure 8 ), s is a random or predetermined scaling factor selected from a normal distribution with mean 1, and is from and =A random or predetermined rotation phase selected from a uniform distribution between . According to this equation, the complex value can be rotated with various phase values and / or scaling factors to increase the amount of stored radar signal characteristics that can be used to recognize gestures. In physical terms, the phase values can represent the angular displacement of the gesture from the computing device 102. In particular, these phase values can represent the angular orientation of the scattering center of the user's hand (assuming that the gesture is performed with the user's hand) relative to the forward or zero-degree orientation of the antenna 214 of the computing device 102. Relatedly, the amplitude value can represent the linear displacement of the gesture from the computing device 102. If the user 104 performs the gesture near the device (e.g., one foot away from the device), the amplitude may have a larger value than if the user 104 performs the gesture far away from the device (e.g., four meters away from the device). The random or predetermined rotation phase and scaling factor used for enhancement may be different from the rotation phase and scaling factor of the detected and / or stored radar signal characteristics, respectively. In this way, the enhancement data can supplement (rather than copy) the radar signal characteristics stored on the computing device 102.
[0132] In one example, a first user 104-1 (e.g., an unregistered person) is interacting with computing device 102 for the first time and begins gesture training on a swipe gesture. Computing device 102 instructs user 104 to perform a swipe gesture with their hand. Radar system 108 transmits a first radar transmit signal 402-1, which is reflected from the scattering surface of the user's hand, thereby generating a first radar receive signal 404-1. First antenna 214-1 receives this signal and sends it to analog circuit 216, which is then received by signal processing module 804 of gesture module 224. A first complex range-Doppler map 820-1, or M1, of first radar receive signal 404-1 is enhanced to include two additional phase values and two additional amplitude values, thereby generating four enhanced complex range-Doppler maps (A1, A2, A3, A4). These five maps M1, A1, A2, A3, and A4 can be used by spatiotemporal machine learning model 802 to improve detection and recognition of gestures (and background motion).
[0133] Fig.11Experimental results 1100 are shown indicating improved performance in the recognition of gestures when utilizing radar enhancement techniques. In this experiment, the gesture module 224 enhances the complex range Doppler map 820 within a Keras layer. The enhanced data set 1102 includes detected, stored, and enhanced radar signal characteristics, while the original data set 1104 includes detected and stored radar signal characteristics. The x-axis of the experimental results 1100 represents the number of gestures performed for the set of gesture training over time. For this experiment, the number of false positives per hour is equal to 2.0. The results indicate that the enhanced data set 1102 enables the computing device 102 to recognize known gestures more frequently at an earlier time during training compared to the original data set 1104. This means that when using radar enhancement techniques to accurately recognize known gestures, the computing device 102 of the present disclosure may be able to accurately recognize known gestures while requiring less interaction with the user 104.
[0134] These radar enhancement techniques can be modified for and / or applied to the detection and differentiation of users 104, and are not limited to techniques for gesture detection and recognition. In particular, the computing device 102 can enhance a set of one or more radar signal characteristics used to differentiate users 104 and improve detection of user presence. In one example, a first user 104-1 (e.g., a new unregistered person) is detected at a distance of two meters (2 m) from the computing device 102 and at an angle of 90 degrees relative to the forward direction (0 degree orientation) of the device. At the user's location, the computing device 102 detects a radar signal characteristic of the first user 104-1, which can be used to distinguish the first user from another user. However, before the radar system 108 can determine the second radar signal characteristic, the first user 104-1 leaves the proximity area 106 of the computing device 102. For some devices, one radar signal characteristic may not be sufficient to distinguish the presence of the first user 104-1 at a high confidence level at a future time. However, the computing device 102 of the present disclosure may enhance this first radar signal characteristic to enable accurate differentiation of the first user 104-1 at a future time. In particular, the enhancement may include rotational phase angles of 0, 180, and 270 degrees corresponding to linear displacements of 0.5 meters, 1 meter, and 4 meters. and amplitude s.
[0135] These enhanced complex range-Doppler maps (corresponding to the enhanced radar signal characteristics) may be stored with the first radar signal characteristics to enable differentiation of first user 104-1 at a future time. When first user 104-1 re-enters proximity zone 106 at a future time, user module 222 may use ten stored radar signal characteristics (nine enhanced radar signal characteristics and one detected radar signal characteristic) instead of just one radar signal characteristic to differentiate this user. Computing device 102 may additionally use information about Figure 3 The communication network 302 is used to exchange enhanced radar signal characteristics between devices in the computing system. In this way, a collection of computing devices 102-X (forming a computing system) can improve the recognition of gestures by sharing enhanced data. Experimental data using spatiotemporal machine learning models
[0136] Fig.12 Experimental data 1200 of a user 104 performing a tap gesture against a computing device 102 is shown. A tap gesture may involve the user 104 pushing their hand toward the device and then pulling their hand back to its initial starting position. For this experiment, the user 104 performed the tap gesture at a distance of 1.5 m from the computing device 102 and with an angular displacement of zero degrees. The experimental data 1200 includes real and imaginary values of a complex Doppler map 820. The first row 1202-1, the third row 1202-3, and the fifth row 1202-5 each include real values of 30 frames (corresponding to 30 complex range Doppler maps 820) as collected by the first receive channel 508-1, the second receive channel 508-2, and the third receive channel 508-3, respectively. The second row 1202-2, the fourth row 1202-4, and the sixth row 1202-6 each include imaginary values of 30 frames (corresponding to 30 complex range Doppler maps 820) collected by the first receive channel 508-1, the second receive channel 508-2, and the third receive channel 508-3, respectively. Each frame is shown with a horizontal axis (x-axis) corresponding to the range rate of the gesture, with zero range rate in the center, negative range rate on the left, and positive range rate on the right. Each frame is also shown with a vertical axis (y-axis) corresponding to the displacement of the gesture, with zero displacement (e.g., the position of the antenna 214) at the bottom and a range of 2 m at the top. These displacements and range rates are taken relative to the position of the receiving antenna of the computing device 102. For each row 1202, the 30 frames are arranged in sequence, with time increasing from left to right.
[0137] According to the experimental data 1200, the user 104 is detected standing at 1.5 m in each frame, as shown by the first circular feature that is always visible at the top of each frame. When the user 104 performs the tap gesture, their hand movement 1204 can be seen in frames 13 to 19. When the user 104 begins to move their hand toward the device (at frame 13), the second circular feature begins to appear. This second circular feature continues to move toward the bottom of frames 14 and 15 (because the user 104 pushes their hand toward the device) until the user 104 fully extends their arm at frame 16. At frame 17, the user 104 begins to pull their hand back toward their body and away from the computing device 102. By frame 20, the tap gesture has been completed. The gesture module 224 can determine from this data that the user 104 has performed the tap gesture and then determine the corresponding command to be executed by the computing device 102.
[0138] When storing one or more radar signal characteristics associated with user 104 performing a tap gesture, gesture module 224 may include, for example, any one or more of the frames shown in experimental data 1200. In a first example, the device may select frames 13 to 19 of row 1202-1 (identified at 1204) to store for future reference. In a second example, the device may store frames 1 to 30 of row 1202-1 as radar signal characteristics of a tap gesture. In a third example, the device may store all 30 frames of each of rows 1202-1 to 1202-6 as radar signal characteristics of a tap gesture. Fig.13 Additional experimental data on the tap gesture is described.
[0139] Fig.13 Experimental data 1300 is shown of a user 104 performing a tap, a right swipe, a strong left swipe, and a weak left swipe on a computing device 102. In addition to the following, Fig.13 The data shown is similar to Fig.12 1302-1, 1302-3, 1302-5, and 1302-7 correspond to absolute range Doppler maps of a tap gesture, a swipe to the right, a strong swipe to the left, and a weak swipe to the left, respectively. Each absolute range Doppler map can be generated by taking the average amplitude of the corresponding complex range Doppler map 820. Rows 1302-2, 1302-4, 1302-6, and 1302-8 correspond to interferometric range Doppler maps of a tap gesture, a swipe to the right, a strong swipe to the left, and a weak swipe to the left, respectively. Each interferometric range Doppler map can be generated by calculating the phase difference between the complex range Doppler maps 820 associated with two or more receiving channels 508.
[0140] In this experimental data 1300, the strong left swipes of rows 1302-5 and 1302-6 can be seen at frames 13 to 19 with clear second circular features. However, the weak left swipes of rows 1302-7 and 1302-8 lack clear second circular features, which makes classification of this gesture challenging. In some cases, the gesture module 224 can improve the recognition of this gesture using the spatiotemporal machine learning model 802, context information, online learning techniques, etc. In other cases, this weak left swipe may be classified as "negative data" or "wrong gesture" that cannot be mapped to a gesture class (e.g., swipe, tap) with the desired confidence level. Instead of ignoring the negative data, the gesture module 224 can store this information (e.g., as background motion) to improve gesture recognition at a future time. Negative data collection
[0141] The computing device 102 of the present disclosure may store one or more radar signal characteristics to enable detection, differentiation and / or recognition of gestures and / or users. These stored radar signal characteristics are not limited to "positive data" and may also include "negative data". Positive data may include radar signal characteristics used to identify gestures (e.g., known gestures with associated commands) and / or distinguish users at a desired confidence level. Examples of positive data may include, for example, radar signal characteristics associated with a gesture class (e.g., tap, swipe, flick, point) or a specific user based on radar cross section (RCS) data. On the other hand, negative data may include radar signal characteristics that are not associated with one or more stored radar signal characteristics of a gesture or user 104 at a desired confidence level. Examples of negative data may include movements of people walking, twisting their torsos, picking up objects, etc. Negative data may also include movements of animals (e.g., house cats), cleaning devices (e.g., automated vacuum cleaners), etc. Gesture module 224 may classify this negative data into a background class of radar signal characteristics that are not associated with known gestures (eg, gesture commands that computing device 102 is programmed or taught to recognize).
[0142] In one example, motion is detected in proximity zone 106 of computing device 102, and radar system 108 detects a first radar signal characteristic of the motion. If the first radar signal characteristic correlates (with a desired confidence level) with one or more stored radar signal characteristics of a tap gesture, gesture module 224 may store this first radar signal characteristic as positive data to improve recognition of the tap gesture at a future time. If the first radar signal characteristic does not correlate with one or more stored radar signal characteristics of a known gesture, gesture module 224 may determine that the motion is not associated with a command. Instead of discarding this data, gesture module 224 stores the first radar signal characteristic as negative data to improve detection or recognition of gestures from background motion (e.g., movements made by user 104 or an object that are not intended as gesture commands). Similar techniques may be used to improve detection of user presence and differentiation of individual users.
[0143] Fig.14 Experimental data 1400 shows three sets of negative data that can be stored to improve detection of gestures from background motion. The negative data is similar to Fig.13 1402-1, 1402-3, and 1402-5 correspond to absolute range Doppler plots for user 104 moving their hands while speaking near computing device 102, user 104 twisting their torso in front of the device, and user 104 picking up and placing an object near the device, respectively. Rows 1402-2, 1402-4, and 1402-6 correspond to interferometric range Doppler plots associated with rows 1402-1, 1402-3, and 1402-5, respectively.
[0144] Fig.15 Experimental results 1500 are shown regarding the accuracy of gesture detection and recognition in the presence of background motion. In this experiment, a user 104 performs a swipe gesture 1502 and a tap gesture 1504 over time while background motion (either natural motion of the user 104 or motion from other objects) occurs within a proximity zone 106 of the computing device 102. Using both positive and negative data, the gesture module 224 accurately detects that a gesture is a gesture (even if it cannot always identify which known gesture it is), and in this case also identifies the swipe gesture 1502 and the tap gesture 1504 at a detection and recognition rate of approximately 0.88, while false positives per hour are generated at a rate of approximately 0.10. False positives represent instances in which the gesture module 224 incorrectly determines background motion as a gesture (here also identifying the background motion as a swipe or tap gesture). The false positives per hour may be affected by the upper and / or lower thresholds of the gesture debouncer 810, as previously described with respect to Figure 8 described.
[0145] To calculate the detection and recognition rate, in general, the gesture module 224 marks motion events as "correct" or "false." Correct gesture detection and recognition occurs when the gesture module 224 outputs only one accurate gesture determination (e.g., both detection and recognition) associated with the command that the user 104 intended to perform. False gesture detection occurs when the gesture module 224 does not detect a gesture (even though the gesture was performed by the user 104), detects a gesture but determines an inaccurate gesture (a gesture that is not associated with the intended command of the user 104), or detects and determines multiple gestures for a single gesture execution. The detection and recognition rate is determined by dividing the number of correct events by the total number of events (the sum of correct events and false events).
[0146] Fig.16 Experimental results 1600 are shown regarding detection and recognition rates of gestures when adversarial negative data is additionally used. In this experiment, user 104 performed both gestures and adversarial motions that were similar (but not identical) to the gestures. These adversarial motions included motions of user 104 moving their hands while picking up and placing objects near computing device 102, interacting with the device's touch screen, turning switches on and off, and talking near the device. Experimental results 1600 indicate that, relative to results 1602 without the use of adversarial negative data, when adversarial motions are performed in proximity to region 106, the performance of gesture module 224 is robust in terms of accurately recognizing gestures as shown as a higher rate, as well as lower false positives and robust results (shown at robust results 1604). Unsegmented gesture detection and recognition
[0147] Although the computing device 102 of the present disclosure can provide gesture training (e.g., segmented learning of gesture execution), the device can also use unsegmented learning techniques to improve the detection of gestures over time. For segmented learning, the gesture module 224 can prompt the user 104 to perform, for example, a tap gesture within a certain time period. The computing device 102 may be able to detect this tap gesture based on the "prior knowledge" that the user 104 will perform a tap gesture (rather than another gesture) within a specified time period. In contrast, unsegmented learning does not utilize prior knowledge in this way. For unsegmented learning, the gesture module 224 does not have to know whether and when the user 104 can perform any one or more gestures for the computing device 102. In addition, unsegmented learning can allow the computing device 102 to continuously detect gestures over time without requiring the user 104 to prompt the device (e.g., provide a wake-up trigger) before performing a gesture. Therefore, unsegmented recognition of gesture execution may be more difficult than segmented recognition.
[0148] To improve the accuracy of the unsegmented recognition of gestures, the gesture module 224 can utilize one or more gesture debouncers 810 and adjust the upper and lower thresholds as needed to improve performance. Additionally, the gesture module 224 can recognize gestures in a timely manner by detecting one or more zero crossings of the data along the velocity axis (x-axis) over a set of two or more frames (e.g., referring to the circular features of the experimental data 1200). For example, Fig.12 The hand motion 1204 (a tap motion) performed in 1204 produces a second circular feature that moves to the left (negative velocity) at frames 13 to 15, returns to the center (zero velocity) at frame 16, and moves to the right (positive velocity) at frames 17 to 19. The zero crossing of this motion occurs at frame 16, and the gesture module 224 can identify this frame as the center of the motion. The computing device 102 can additionally select one or more data frames (e.g., frames 1 to 15 and frames 17 to 30) around frame 16 to form a set of data for correlating this motion with a tap gesture.
[0149] Fig.17 Experimental results 1700 (confusion matrix) related to the accuracy of unsegmented gesture detection are shown. In this experiment, the user 104 performed 6 gestures over time, including background motion 1702 (e.g., movement of the user 104 or an object not associated with a gesture command), swipe left 1704, swipe right 1706, swipe up 1708, swipe down 1710, and tap 1712. The x-axis (gestures performed 1714) of this confusion matrix represents the intended gesture commands of the user 104, and the y-axis (recognized gestures 1716) represents the classification of the gesture execution by the gesture module 224. Experimental results 1700 include 36 possible results quantified in terms of classification rate (normalized to 1) over a certain period of time. The results indicate that the gesture module 224 is able to determine each gesture 1714 performed with an accuracy of 0.831 to 0.994. In this experiment, the gesture debouncer 810 uses an upper threshold of 0.9.
[0150] Fig.18 Experimental results 1800 are shown corresponding to the accuracy of unsegmented gesture detection at various linear and angular displacements from the computing device 102. These results utilize Fig.17 The experimental results 1700 are a collection of similar data and include another confusion matrix of the normative detection rates of gestures over time periods. The matrix includes 29 results regarding the accuracy of unsegmented gesture recognition at various angular displacements (ranging from -45 degrees to +45 degrees) and linear displacements (ranging from 0.3 m to 1.5 m) from the computing device 102. Fig.17 and Fig.18The data represented in represents hundreds of motions performed by user 104, including gestures and background motions. Additional sensors for improved fidelity of user differentiation
[0151] Fig.19 An example implementation 1900 of a computing device 102 using an additional sensor (e.g., microphone 1902) to improve the fidelity of user differentiation and / or gesture recognition and user interaction is shown. In some cases, radar signal characteristics associated with nearby objects (e.g., registered users or unregistered persons) or motion (e.g., gesture performance, background motion) may not provide sufficient information to differentiate users 104 and / or recognize gestures with a desired confidence level. As depicted in the example implementation 1900, the user module 222 may not be able to confidently determine that the first user 104-1 is a registered user (e.g., a father) based solely on the radar signal characteristics. In this case, the computing device 102 can bootstrap the audio signal 1904 (e.g., a sound wave emitted by the first user 104-1) to enable the radar system 108 to determine that the father is present.
[0152] As depicted in the example implementation 1900, the user module 222 can receive the audio signal 1904 via the microphone 1902 and analyze the characteristics of these sound waves (e.g., wavelength, amplitude, time period, frequency, speed, velocity) to determine which user is present in the proximity zone 106. This analysis can be performed automatically when triggered, or concurrently with or after the analysis of the radar signal characteristics. The audio signal 1904 may be modified by additional circuits and / or components before being received by the user module 222.
[0153] The user module 222 may analyze the audio signal 1904 to distinguish between users, with or without access to private information (e.g., the content of a conversation). For example, the radar system 108 may characterize the audio signal 1904 to distinguish the presence of the first user 104-1, with or without recognition of the words being spoken (e.g., performing speech-to-text), because characteristics such as low pitch or fast-paced speech may be used to distinguish a particular user. The radar system 108 may characterize the audio signal 1904 in terms of pitch, loudness, tone, timbre, rhythm, consonance, dissonance, pattern, etc. Thus, the user 104 may comfortably discuss private information near the computing device 102 without worrying about whether the device is recognizing the spoken words, sentences, thoughts, etc., depending on the user's settings and preferences.
[0154] The computing device 102 may store (e.g., on a shared memory) the audio detection characteristics of one or more users 104 to enable the presence of the users to be distinguished. When a registered user (e.g., the father) enters the proximity zone 106 of the radar system 108, the user module 222 may partially utilize the stored audio detection characteristics of the father to distinguish him from other users of the device. When an unregistered person enters the proximity zone 106, the user module 222 may partially utilize the stored audio detection characteristics of the registered user to determine that this is an unregistered person who, for example, has not provided an audio signal 1904 to the computing device 102. The radar system 108 may then generate an unregistered user identification for this unregistered person, the unregistered user identification including the audio detection characteristics associated with one or more audio signals 1904 emitted by the unregistered person. Therefore, the radar system 108 may be able to use the audio detection characteristics stored in the unregistered user identification to distinguish this unregistered person at a later time.
[0155] Although the additional sensor of the example implementation 1900 is depicted as a microphone 1902, in general, Fig.19 The described techniques can be performed using various sensors described herein. It should be understood that for situations or environments where privacy is not a working concern or otherwise provides methods to eliminate privacy issues, it is not beyond the scope of this teaching that the additional sensor of the example implementation 1900 is a camera or a video camera. In addition, additional sensors can be used to improve gesture detection and recognition. For example, the computing device 102 can additionally utilize data associated with the ambient light sensor to detect and recognize gestures performed by the first user 104-1. This action may be particularly useful if, for example, the first user 104-1 performs an ambiguous gesture that cannot be recognized with a desired confidence level using only radar signal characteristics. In another example, the computing device 102 can additionally utilize data from an ultrasonic sensor to improve the recognition of ambiguous gestures performed by the first user 104-1. Therefore, supplementary data sensed by non-radar sensors can be used to assist gesture recognition and other determinations, such as user presence, user differentiation, and user interaction.
[0156] In general, additional sensor inputs (e.g., audio signal 1904 from microphone 1902) are optional, and privacy controls may be provided to user 104 to limit the use of such additional sensors. For example, user 104 may modify their personal settings, general settings, default settings, etc. to include and / or exclude additional sensors (e.g., in addition to antenna 214 for radar). Furthermore, user module 222 may implement these personal settings after distinguishing the presence of the user. Fig. 20 Further describes privacy controls. Adaptive privacy and other settings
[0157] Fig. 20 Example environments 2000-1 and 2000-2 are shown in which privacy settings are modified based on user presence. In example environment 2000-1, user module 222 of computing device 102 detects that first user 104-1 is present within proximity zone 106. Radar system 108 can implement first privacy settings 2002 for first user 104-1 in response to detecting the presence of a user or person. This first privacy setting 2002 can include information about, for example, allowed sensors (see Fig.19 ), audio reminders, calendar information, music, media, settings for home items (e.g., light preferences), etc. For example, when the first privacy setting 2002 has been implemented, the first user 104-1 can receive audio reminders for calendar events.
[0158] In example environment 2000-2, in addition to the continued presence of first user 104-1, computing device 102 may later detect the presence of second user 104-2 (e.g., another registered user). Radar system 108 implements second privacy setting 2004 based on the presence of the second user to adapt the privacy of first user 104-1. Implementation may be automatic or triggered based on a command from first user 104-1. For example, second privacy setting 2004 may limit audio alerts to prevent the publication of private information in the presence of others. Second privacy setting 2004 may be based on, for example, preset conditions, user input, etc.
[0159] The second privacy setting 2004 may also be implemented to protect the privacy of the first user's information and adapt based on the users in the room. For example, the presence of another registered user (e.g., a family member) may require fewer privacy restrictions than the presence of an unregistered person (e.g., a guest). The adaptive privacy settings may also be customized for each user 104. For example, the first user 104-1 may have a more stringent privacy setting (e.g., limiting audio reminders in the presence of others), while the second user 104-2 may have a less stringent privacy setting (e.g., not limiting audio reminders in the presence of others).
[0160] In addition to adaptive privacy, the technology can also adapt other settings in a similar manner. As with the adaptive privacy described above, these adaptive settings can depend on the presence of other users, such as those users (e.g., second user 104-2) who are close to the user (e.g., first user 104-1) who is or has interacted with the computing device 102. Adaptive settings can be applied to ongoing operations, such as when the first user 104-1 commands to play music on the stereo. If the second user 104-2 is distinguished, or if the second user 104-2 speaks to the first user 104-1 (or vice versa), the technology can turn down the music without explicit user interaction (e.g., a gesture to turn down the music) from the first user 104-1. Adaptive settings can also be applied to ongoing operations. In one example, the first user 104-1 (e.g., a father) can start an oven and set a timer for 20 minutes. If the second user 104-2 (e.g., a child) attempts to turn off the oven or timer before the timer expires, the technology of the present disclosure can prevent the child from doing so. This adaptation allows the father to control the operation of the oven and timer during these 20 minutes to prevent his baking from being interrupted. In particular, computing device 102 may associate an ongoing operation with user 104 that has executed a command that prevents another user from modifying the operation. In another example, a mother may execute a command to turn off the bedroom lights at 9:00 p.m. to ensure that her children go to bed on time. If the child executes a command to keep the lights on after their bedtime has passed, computing device 102 may prevent the child from modifying the mother's command. Example Implementation of a Computing System
[0161] Fig.21 A technique is shown in which user differentiation may be implemented using multiple computing devices 102-1 and 102-2 forming a computing system (eg, regarding Figure 1 and Fig. 20 2 . In the example environment 2100, an example home is depicted as having a first room 304-1 and a second room 304-2. A first computing device 102-1 equipped with a first radar system 108-1 is located in the first room 304-1, and a second computing device 102-2 equipped with a second radar system 108-2 is located in the second room 304-2. The first computing device 102-1 and the second computing device 102-2 are capable of exchanging information (e.g., stored in local or shared memory) with the aid of a communication network 302 and partially form a computing system. For purposes of this illustrative example, the first proximity zone 106-1 does not overlap with the second proximity zone 106-2.
[0162] The first computing device 102-1 of the example environment 2100 may use the first radar system 108-1 to send a first radar transmission signal 402-1 (see above). Figure 4 ) to detect the presence of one or more users. First radar transmit signal 402-1 may be reflected from an object (e.g., first user 104-1) and modified in amplitude, phase, or frequency before being received at first computing device 102-1. First radar system 108-1 may include first radar receive signal 404-1 (see above). Figure 4 ) (including at least one radar signal characteristic) is compared with one or more stored radar signal characteristics of registered users to determine whether first user 104-1 is a registered user or an unregistered person. In this example, first radar received signal 404-1 is not correlated with one or more stored radar signal characteristics of registered users. Therefore, at 2102, first user 104-1 is distinguished as an unregistered person.
[0163] After determining at 2102 that the first user 104-1 is an unregistered person, the first computing device 102-1 generates an unregistered user identification (e.g., a simulated identity, a pseudo identity) and assigns it to the unregistered person, as shown at 2104. The unregistered user identification may include one or more radar signal characteristics associated with the first radar received signal 404-1, which may be used to distinguish the unregistered person from other users (e.g., the second user 104-2) at a future time. The unregistered user identification may be stored on a local or shared memory, and each computing device of the computing system (e.g., the second computing device 102-2) may access the unregistered user identification, even if the device has not directly detected the unregistered person. For example, the second computing device 102-2 may access the stored first radar signal characteristics associated with the unregistered user identification, even if the unregistered person has never been detected by the second computing device 102-2.
[0164] At a future time, the first user 104-1 walks into the second room 304-2 and is detected by the second computing device 102-2 at 2106. In particular, the second radar system 108-2 of the second computing device 102-2 sends a second radar transmit signal 402-2 to detect the presence of one or more users. The second radar transmit signal 402-2 is reflected from an object (e.g., the first user 104-1) and is modified in amplitude, phase, or frequency before being received at the second computing device 102-2. The second radar system 108-2 compares this second radar receive signal 404-2 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of a registered user and an unregistered user identification assigned to an unregistered person to determine whether the object is a registered user or an unregistered person. In this example, at least one radar signal characteristic of the second radar receive signal 404-2 is correlated with one or more stored radar signal characteristics of an unregistered person. Therefore, the first user 104-1 is again distinguished as an unregistered person (at 2106) based on the unregistered user identification. The second radar signal characteristic may be stored and associated with the unregistered user identification.
[0165] Although Fig.21 , but the same techniques may be applied to distinguish registered users. For example, first computing device 102-1 may transmit third radar transmit signal 402-3 to distinguish one or more additional users. Third radar transmit signal 402-3 may be reflected from another object (e.g., second user 104-2 that is different from the previously detected unregistered person) and modified before being received at first computing device 102-1. First radar system 108-1 may compare this third radar receive signal 404-3 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of registered users and unregistered user identification. In this example, at least one radar signal characteristic of third radar receive signal 404-3 is related to one or more stored radar signal characteristics of registered users. Therefore, second user 104-2 may be distinguished as a registered user. The third radar signal characteristic may be stored to improve the distinction between registered users at a future time. First radar system 108-1 may additionally access the stored settings, preferences, training history, habits, etc. of registered users to provide a customized experience.
[0166] In this example, the second user 104-2 (registered user) may later move to the second room 304-2 and enter the second proximity zone 106-2. The second radar system 108-2 of the second computing device 102-2 may transmit a fourth radar transmit signal 402-4 to distinguish the registered user. The fourth radar transmit signal 402-4 is reflected from the registered user and is modified before being received at the second computing device 102-2. The second radar system 108-2 compares this fourth radar receive signal 404-4 (including at least one radar signal characteristic) with one or more stored radar signal characteristics of the registered user to distinguish the registered user. In this example, the second computing device 102-2 accesses one or more stored characteristics from the local memory and / or shared memory of the first computing device 102-1. Here, the at least one radar signal characteristic is correlated with the one or more stored radar signal characteristics of the registered user, and the second radar system 108-2 determines that the registered user is present within the second proximity zone 106-2 of the second computing device 102-2. About Fig. 22 The use of computing device 102 as part of a computing system is further described. Operational continuity across computing systems
[0167] Fig. 22 An example environment 2200 is shown in which multiple computing devices 102-1 and 102-2 across a computing system continuously perform operations. The first computing device 102-1 is depicted as being located in a bedroom 2202, which is separate from an office 2204 where the second computing device 102-2 is located. The first computing device 102-1 and the second computing device 102-2 are two or more devices (e.g., Figure 3 and Fig.21 Part of a computing system (such as those described above).
[0168] At a first time, in example environment 2200-1, user 104 performs a swipe gesture to command first computing device 102-1 to read aloud the latest news headlines. User 104 may be listening to the news while in bedroom 2202, and first radar system 108-1 may instruct first computing device 102-1 to continue playing the news when the presence of user 104 is detected.
[0169] At a second time, in example environment 2200-2, user 104 moves from bedroom 2202 to office 2204 and the news continues to play on second computing device 102-2. In particular, first radar system 108-1 may detect the lack of user presence within first proximity zone 106-1 of bedroom 2202 and pause the news. Once user 104 moves into office 2204, second computing device 102-2 may detect and distinguish the presence of this user within, for example, second proximity zone 106-2. Second radar system 108-2 may then automatically (e.g., without user input) continue playing the news that was previously paused by first radar system 108-1. In this way, user 104 may enjoy a seamless experience that is automated across multiple rooms in the home.
[0170] The ongoing operation can follow the user 104 who performed the gesture. For example, if the user 104 of the example environment 2200-1 leaves the bedroom 2202, the first computing device 102-1 can detect the absence of this user 104 (for example, but not another user) and pause the news. When the user 104 is later detected and distinguished by the second computing device 102-2 in the office 2204, the news can continue to be read aloud. In addition to the user 104 of the example environment 2200-1, there may be one or more other users in the bedroom 2202 that have been detected by the first computing device 102-1. The first radar system 108-1 can determine that these other users are not the user 104 who performed the gesture command, thereby deducing that the news should follow the user 104 who performed the gesture. Alternatively, in addition to playing in the bedroom 2202 when other users are present, the news can also follow the user 104 who performed the gesture.
[0171] Each radar system 108-1 and 108-2 may adjust ongoing operations based on the location of user 104. For example, if user 104 lies down close to first computing device 102-1 (as depicted in example environment 2200-1), first radar system 108-1 may detect this shorter distance and lower the speaker volume. Alternatively, if user 104 moves to the far side of bedroom 2202 (e.g., farther from first computing device 102-1), first radar system 108-1 may detect this larger distance and increase the speaker volume.
[0172] In another example ( Fig. 22In the first room equipped with the first computing device 102-1, the user 104 performs a gesture associated with the two-part command of (1) starting to play the news feed aloud and (2) stopping to play the news feed. In the first room equipped with the first computing device 102-1, the user 104 performs a first gesture associated with the first part of the command (start playing the news). The user 104 then moves to the second room equipped with the second computing device 102-2 and continues to listen to the news in this room (according to Fig. 22 ). At a later time, user 104 performs a second gesture associated with the second part of the command (end playing the news). In this example, first computing device 102-1 and second computing device 102-2 utilize communication network 302 to coordinate the two-part command across multiple rooms in the home. The same technology can also be applied to the first part and / or the second part of the audio input sensed by the microphone of first computing device 102-1 or second computing device 102-2.
[0173] In the attached example ( Fig. 22 ), user 104 may perform a gesture associated with a single command that is detected by both first computing device 102-1 and second computing device 102-2 (see Figure 4 ). In this example, a user may provide detailed commands while moving between two rooms. User 104 begins explaining their single command (to schedule an appointment with their doctor) in the first room (equipped with first computing device 102-1), and first radar system 108-1 recognizes the first portion of the command. However, user 104 moves to the second room (to get their medical records) and continues to schedule their appointment by providing the second portion of the command to second computing device 102-2. In this example, first computing device 102-1 and second computing device 102-2 again utilize communication network 302 to coordinate a single command across multiple rooms in the home. The same techniques can also be applied to the first portion and / or second portion of the audio input sensed by the microphone of first computing device 102-1 or second computing device 102-2.
[0174] In the attached example ( Fig. 22), user 104 may perform a pausable (capable of being paused and then resumed) continuous radar detection gesture that may be detected successively by both first computing device 102-1 and second computing device 102-2 to achieve the advantageous effect of continuing intermittent sustainable activities between rooms. An example of a pausable continuous gesture may be a voice list enumeration gesture, in which the user may begin to move their hand in a rolling circular motion in a generally vertical plane passing through themselves and the devices (hereinafter referred to as "rolling their hand", which may be intuitively imagined to be similar to the "keep the film rolling" gesture that a movie director makes to a cameraman). In one example, in a first room (equipped with a first computing device 102-1 having a first radar system 108-1), a user begins rolling their hand and says "This is my shopping list: milk, eggs, butter ... " while continuing to roll their hand, and device 102-1 will use first radar system 108-1 to recognize that gesture and associate those named items with that user's shopping list as long as the user keeps rolling their hand. If the user stops rolling their hand, list making is paused even if the user continues to speak. (Such a pause may occur, for example, if the user is interrupted and needs to speak to another person about a topic unrelated to their shopping list.) If the user then continues to roll their hand, list making continues, and the things they speak out will continue to be added to the shopping list. (The user can terminate list making at any time with a “push and pull” gesture and / or an appropriate voice command.) If during the user's (unterminated) list making process in the first room, the user stops rolling their hand and then walks into the second room, the user can continue rolling their hand in the second room and continue speaking their shopping list items, and the second device 102-2 with the second radar system 108-2 will recognize the gesture and continue to add the spoken items to the shopping list as long as the user keeps their hand rolling, and so on.
[0175] about Fig. 22 The described techniques are not limited to ongoing operations and can be applied to operations that are performed periodically over time, such as Fig.23 Further description.
[0176] Fig.23An example environment 2300 is shown in which a computing system implements persistence of operations across multiple computing devices 102-1 and 102-2. In the example environment 2300-1, a user 104 is detected by a first computing device 102-1 of the computing system located in a kitchen 2302. The first radar system 108-1 determines that the user 104 is an unregistered person and assigns them an unregistered user identification, thereby distinguishing the user 104 from other users. In this example, the first computing device 102-1 prompts the unregistered person to begin gesture training on a first gesture. During training, radar signal characteristics associated with the manner in which the unregistered person performs the first gesture are stored and associated with the unregistered user identification. This unregistered user identification can be stored on a memory and / or accessed by any one or more devices of the computing system (e.g., the second computing device 102-2).
[0177] At a later time, depicted in example environment 2300-2, the presence of a user is detected in restaurant 2304 by second computing device 102-2 as part of the computing system. Second radar system 108-2 distinguishes user 104 as an unregistered person based on radar signal characteristics, and uses the unregistered user identification to access the training history of this user. Second computing device 102-2 may then prompt user 104 to continue training on the second gesture. In particular, second radar system 108-2 may determine that they have completed training on the first gesture. This sequence may continue over time using various computing devices 102 of the computing system (although Fig.23 ), until the unregistered individuals have completed their hand gesture training.
[0178] However, user 104 may perform the gesture differently for some computing devices 102 of the computing system based on that user's behavior or body position in the room. When user 104 performs the gesture differently than how the gesture was taught, for example, during training, computing device 102 may determine that an ambiguous gesture has been performed. This ambiguous gesture may be similar to one or more gestures that may be recognized by the device but may lack sufficient similarity to a single known gesture to allow for high confidence recognition. Thus, each computing device 102 may utilize contextual information to improve the interpretation of ambiguous gestures, such as with respect to the context of the gesture. Fig.24 Further description. Example of ambiguous gesture interpretation
[0179] Fig.24A technique for radar-based ambiguous gesture determination using contextual information is shown. Example environment 2400 depicts user 104 performing ambiguous gesture 2402, which is detected by radar system 108 of computing device 102. It is assumed here that user 104 intended to perform first gesture 2404 that is recognizable by the device and associated with a first command to turn down the volume of music played from computing device 102. However, user 104 accidentally performed ambiguous gesture 2402, which is similar to both first gesture 2404 and second gesture 2406 (associated with a second command to open a garage door) but cannot be recognized with a desired confidence level. Therefore, radar system 108 cannot determine that ambiguous gesture 2402 is first gesture 2404 and not second gesture 2406.
[0180] In particular, the ambiguous gesture 2402 is correlated with the first gesture 2404 and the second gesture 2406 by an amount that may be greater than the no confidence level but less than the high confidence level. For example, the correlation with each gesture may have a 40% confidence level (e.g., there is a 40% chance that the ambiguous gesture 2402 is the first gesture 2404 and a 40% chance that it is the second gesture 2406). If the no confidence level is set to 10% and the high confidence level is set to 80%, then the correlation of the ambiguous gesture 2402 with the first gesture 2404 or the second gesture 2406 is greater than the no confidence level but less than the high confidence level. In general, the no confidence level and the high confidence level can be modified or adapted to improve the quality of gesture detection. In this disclosure, "desired confidence level" will refer collectively to these no confidence levels and high confidence levels, which may be different or similar for each use case (e.g., each gesture, each user, destructiveness of the gesture).
[0181] To avoid prompting user 104 to repeatedly perform the gesture until successful, radar system 108 uses contextual information 2408 to improve interpretation of ambiguous gesture 2402. In this example, radar system 108 determines that contextual information 2408 includes music playing on computing device 102 at the current time. In general, contextual information may include ongoing operations, past or planned operations, foreground or background operations, location of the device, history of common users or gestures, running applications, conditions external to computing device 102, etc. External conditions may include, for example, time of day, lighting, audio within proximity zone 106, etc.
[0182] Computing device 102 may determine based on contextual information 2408 that ambiguous gesture 2402 is most likely first gesture 2404. In this example, radar system 108 correlates the music being played on the device at the current time with the first command of first gesture 2404. Since the second command to open the garage door is unrelated to the music being played on the device, radar system 108 determines that first gesture 2404 is more likely to be the intended gesture of user 104. This determination may be performed using formal logic, informal logic, mathematical logic, hysteresis logic, deductive reasoning, inductive reasoning, abductive reasoning, etc. Machine learning models (such as machine learning model 700 (see reference 700)) may also be used. Figure 7 ) and / or spatiotemporal machine learning model 802 (reference Figure 8 )) to perform the determination.
[0183] In one example, radar system 108 utilizes inductive reasoning to determine a general association by which radar system 108 can deduce that ambiguous gesture 2402 is first gesture 2404. In doing so, radar system 108 can proceed based on the following independent variables: (1) The blur gesture 2402 is the first gesture 2404 or the second gesture 2406. (2) The first gesture 2404 is associated with a first command to turn down the music volume. (3) Second gesture 2406 is associated with a second command to open the garage door. (4) Music is playing at the current time (context information 2408). (5) The first command is related to the music being played at the current time. (6) The second command has nothing to do with the music being played at the current time. (7) Ambiguous gestures are generally associated with the operation being performed at the current time. Therefore, blur gesture 2402 is most likely first gesture 2404 .
[0184] Upon detecting ambiguous gesture 2402, radar system 108 may determine that ambiguous gesture 2402 is not associated with a known gesture (e.g., one or more stored radar signal characteristics) with a desired confidence level. The desired confidence level may be a quantitative or qualitative assessment of the confidence (e.g., accuracy) required to correctly recognize the gesture, or another symbol-like, vector-based, or matrix-based threshold criterion or combination of threshold criteria.
[0185] In example environment 2400, the radar signal characteristics of ambiguous gesture 2402 have a 40% probability of being associated with the stored radar signal characteristics of first gesture 2404, a 40% probability of being associated with the stored radar signal characteristics of second gesture 2406, and a 20% probability of being associated with the stored radar signal characteristics of the third gesture. If the desired confidence level is set to 50%, radar system 108 may not be able to accurately determine which known gesture is associated with ambiguous gesture 2402 at that desired confidence level. However, radar system 108 may determine that first gesture 2404 and second gesture 2406 are more likely to be associated with ambiguous gesture 2402 than the third gesture. In particular, radar system 108 may consider gestures that exceed a minimum confidence level (e.g., a threshold) of 35% or more. Since the third gesture has only a 20% probability of being ambiguous gesture 2402, radar system 108 may rule out this possibility. Radar system 108 may instead determine that both first gesture 2404 and second gesture 2406 have a 40% probability, which exceeds the minimum confidence level.
[0186] To identify ambiguous gesture 2402 as first gesture 2404 (e.g., a known gesture) or second gesture 2406 (e.g., another known gesture), radar system 108 may utilize contextual information 2408. In particular, radar system 108 may determine whether a first command of first gesture 2404 or a second command of second gesture 2406 is associated with contextual information 2408. The association may be quantitative or qualitative based on a strict yes / no association (e.g., binary), associations of varying scales, logic or reasoning, machine learning models (700, 802), etc. For example, radar system 108 may determine that a first command to turn down the volume of music is associated with contextual information 2408 of music being played at the current time. On the other hand, radar system 108 may also determine that a second command to open a garage door is not associated with music being played. Based on this determination, radar system 108 may determine that ambiguous gesture 2402 is first gesture 2404.
[0187] In some cases, radar system 108 may not be able to use contextual information to accurately identify ambiguous gesture 2402. If the first command of first gesture 2404 is instead associated with starting a timer, radar system 108 may determine that neither the first command nor the second command is associated with music being played at the current time. Instead of additional information, radar system 108 may determine that neither the first command nor the second command should be executed by computing device 102. Additionally, computing device 102 may prompt user 104 to repeat the gesture and / or provide additional input (e.g., a voice command). Supports radar gesture recognition
[0188] Fig.25Example implementations 2500-1 through 2500-3 are shown in which gesture module 224 may recognize gestures performed by user 104. To recognize gestures, gesture module 224 of radar system 108 may analyze radar receive signal 404 to determine (1) topological features, (2) temporal features, and / or (3) contextual features. Each of these features may be associated with one or more radar signal characteristics detected by computing device 102. Computing device 102 is not limited to Fig.25 The three feature categories depicted in the figure may include other radar signal characteristics and / or categories not shown. In addition, the three feature categories are shown as example categories and may be combined and / or modified to include subcategories that implement the techniques described herein. References may additionally be included Figures 7 to 18 The technology discussed and its Fig.25 The techniques presented are not mutually exclusive. Fig.25 Discussion and Figure 6 The user differentiation discussions in are similar, except that they apply to gesture execution.
[0189] In example implementation 2500-1, gesture module 224 may use topology information in part to recognize gestures. Figure 6 Similar to the teachings of example implementation 600-1 of , topological features may include RCS data associated with the height, build, orientation, distance, material, body type, etc. of user 104. For example, user 104 may perform swipe gesture 2502 by forming a vertical flat hand, but perform pinch gesture 2504 by pinching their thumb and index finger together (e.g., making contact therebetween). When compared to topological features associated with pinch gesture 2504, topological features associated with swipe gesture 2502 may include a larger surface area, an orthogonal orientation, a flatter surface, etc. On the other hand, when compared to topological features associated with swipe gesture 2502, topological features associated with pinch gesture 2504 may include a smaller surface area, a less uniform (e.g., less flat) surface, a greater depth of the position of the hand, etc.
[0190] In example implementation 2500-2, gesture module 224 can use time information to recognize gestures in part with reference to example implementation 600-2. Radar system 108 can recognize gestures by receiving and analyzing, for example, motion signatures (e.g., the unique way that user 104 moves when performing a gesture). In one example, user 104 performs a swipe gesture to turn pages of a book they are reading. Radar system 108 detects one or more radar signal characteristics (e.g., a time profile, as depicted) of the swipe gesture and compares them to one or more stored radar signal characteristics. The time profile can include the time profile of the swipe gesture over time, such as using Figure 5The amplitude of one or more radar receive signals 404 detected by analog circuit 216 of the present invention. The temporal profile of a swipe gesture may have different characteristics than, for example, a wave gesture or a pinch gesture. A wave gesture may include two complementary motions (e.g., producing two amplitude peaks), while a swipe gesture may include one motion (e.g., producing one amplitude peak). Over time, a pinch gesture may be performed slower than a swipe gesture, thereby producing a wider amplitude peak of radar receive signal 404.
[0191] In the example implementation 2500-3, the gesture module 224 may also refer to the above Figure 6 Example implementations 600-3 and 600-4 use context information to recognize gestures. The context information is not limited to Fig.25 The examples depicted in and reference Figures 24 to 32 Various other examples are described. In this disclosure, “contextual information” refers to information that adds context (e.g., additional details) to signals received by radar system 108. Contextual information may include user presence, user habits, location of computing device 102, commonly performed gestures and / or detected history of users, common activity in a room, ongoing operations on one or more computing devices 102, past and planned operations, foreground and background operations, etc.
[0192] Example implementation 2500-3 depicts how user presence may provide additional context to recognize gestures. Computing device 102 may be configured to recognize a push-pull gesture, which involves user 104 pushing their hand out to a certain range at a certain rate and pulling their hand back to a similar range (e.g., equal distance but opposite direction) at the same rate. In this manner, when user 104 performs a push-pull gesture, the device may expect complementary push and pull motions. However, if a push-pull gesture is performed without complementary push and pull motions (as depicted), gesture module 224 may determine that an ambiguous gesture has been performed.
[0193] To improve interpretation of ambiguous gestures, user module 222 can provide additional details about the user's presence (eg, contextual information) to gesture module 224. Gesture module 224 can better interpret ambiguous gestures if it can determine which user performed the gesture.
[0194] In one example, first user 104-1 performs a push-pull gesture with non-complementary push and pull motions, and gesture module 224 determines that the gesture is not associated with a push-pull gesture (e.g., with a desired confidence level). First user 104-1 may push their hand out to a certain range at a certain rate but pull their hand back to a shorter range at a significantly slower rate. Gesture module 224 determines that the gesture is an ambiguous gesture, which may be a push-pull gesture or a push gesture. Instead of prompting first user 104-1 to perform the gesture again, gesture module 224 utilizes user presence information from user module 222 to determine that a first registered user performed the gesture. This first registered user may have performed this gesture in the past (e.g., during gesture training) and may typically perform the push-pull gesture in this modified manner. Gesture module 224 may then access one or more stored radar signal characteristics associated with the first registered user's history of performing gestures (e.g., gesture training history). This additional information may allow computing device 102 to determine that the ambiguous gesture is a push-pull gesture as commonly performed by the first registered user.
[0195] As also depicted in example implementation 2500-3, second user 104-2 may also perform a push-pull gesture in a unique manner. Second user 104-2 may push their hand out to a certain range at a certain rate but pull their hand back to a longer range at a significantly faster rate. Due to the non-complementary push and pull motions, gesture module 224 may determine that the gesture is an ambiguous gesture, which may be a push-pull gesture or a pull gesture. Gesture module 224 may again utilize user presence information from user module 222 to determine whether a registered user performed the gesture. In this example, second user 104-2 is an unregistered person who has not previously performed gestures with computing device 102. Therefore, gesture module 224 may utilize other contextual information (such as information about Figure 26 to Figure 32 further described), prompting the unregistered person to perform the gesture again, and / or initiating gesture training.
[0196] In this example, radar system 108 may receive one or more radar receive signals 404 that include radar signal characteristics of second user 104-2 performing their version of this push-pull gesture (e.g., another ambiguous gesture). Gesture module 224 may compare these radar signal characteristics with stored radar signal characteristics to determine that the other ambiguous gesture may be a push-pull gesture (first known gesture) or a pull gesture (second known gesture). In particular, the correlation of the performed version of the push-pull gesture with each of the known gestures exceeds a no confidence level (e.g., a minimum threshold), but the device determines that the correlation is also below a high confidence level.
[0197] In general, context information may include details determined using, for example, antenna 214, additional sensors of computing device 102, data stored on memory (e.g., user habits), local information (e.g., time, relative location), operating conditions, etc. In another example (not depicted), gesture module 224 may use local time as context to enable recognition of ambiguous gestures. If user 104 consistently performs a gesture to turn on the lights at 6:00 a.m. every morning, computing device 102 may note this habit to improve gesture recognition. If user 104 accidentally performs an ambiguous gesture at 6:00 a.m. (e.g., perhaps associated with turning on the lights or reading the news aloud), the device may use this context information to determine that user 104 most likely intended to perform a gesture to turn on the lights. In another example, if computing device 102 is located in a kitchen, gesture module 224 may determine over time that kitchen-related gestures (e.g., to turn on the oven) are common in that room. If user 104 performs an ambiguous gesture in the kitchen (e.g., perhaps associated with turning on a dishwasher or turning on a security system in the living room), the device may use contextual information for kitchen-related gestures to determine that user 104 most likely intended to perform a gesture to turn on the dishwasher.
[0198] The gesture module 224 may additionally utilize one or more logic systems (e.g., including predicate logic, hysteresis logic, etc.) to improve gesture recognition. The logic system may be used to prioritize certain gesture recognition techniques over other techniques (e.g., favoring temporal features over contextual information), add weights (e.g., confidence) to certain results when relying on two or more features, etc. The gesture module 224 may also include machine learning models (e.g., 700, 802) to improve gesture recognition (e.g., interpretation of ambiguous gestures), as previously described with respect to Figure 7 and Figure 8 described.
[0199] Radar system 108 may use contextual information alone or in combination with topological or temporal information of radar signal characteristics to recognize gestures. In general, gesture module 224 may use the contextual information in any combination and at any time. Fig.25 To recognize a gesture, radar system 108 may use any one or more of the categories depicted in . For example, radar system 108 may collect topological and temporal information about a gesture being performed, but lack contextual information. In another case, radar system 108 may collect topological and temporal information, but determine that the information is insufficient to recognize an ambiguous gesture. If contextual information is available, radar system 108 may utilize that contextual information to recognize an ambiguous gesture. Fig.25 Any one or more of the categories depicted in may take precedence over another. Fig.26 The application of the contextual features described with respect to example environment 2500 - 3 is further described. Contextual information associated with user habits
[0200] Fig.26 An example environment 2600 is shown in which computing devices 102-1 and 102-2 can utilize contextual information of the user's habits to improve gesture recognition. In the example environment 2600, a first room 304-1 (dining room 2304) contains a first computing device 102-1, and a second room 304-2 (bedroom 2202) contains a second computing device 102-2. Each computing device 102 can determine its location in the home based on, for example, relative location relative to other devices, user input, user behavior, command frequency, command type, user presence, etc. Additionally, each computing device 102 can utilize one or more sensors to perform, for example, geo-fencing, multilateration, true range multilateration, dead reckoning, atmospheric pressure adjustment, true range inertial multilateration, angle of arrival calculation, time of flight calculation, etc. to determine the location of the device.
[0201] Users may perform gestures differently in each room of the home based on their typical behavior or body position in that room. As depicted in the example environment 2600, a user 104 located in the dining room 2304 may typically perform gestures while sitting upright in a chair, while a user 104 located in the bedroom 2202 may typically perform gestures while lying horizontally in bed. When the user 104 performs a push-pull gesture in the dining room 2304, for example, the first computing device 102-1 may detect that the push-pull gesture has been performed with complementary push and pull motions, as taught during gesture training. Specifically, the user 104 may commonly push and pull to similar (but opposite) ranges at similar rates, as depicted in the example environment 2600 using complementary arrows. On the other hand, the user 104 may commonly perform a push-pull gesture in the bedroom 2202 from a lying position with non-complementary push and pull motions. For example, the user 104 may typically push with a greater range, at a greater rate, and in a direction that is non-colinear with the pulling motion, as also depicted in the example environment 2600 using non-complementary arrows.
[0202] Instead of requiring user 104 to perform the push and pull gesture perfectly (e.g., as expected, consistently, as taught during training) at each computing device 102-1 and 102-2, each device may learn the user's habits as contextual information over time to improve recognition of ambiguous gestures. In example environment 2600, first computing device 102-1 may learn over time that user 104 typically performs a first version of the push and pull gesture in restaurant 2304 with complementary push and pull motions, while second computing device 102-2 may learn over time that user 104 typically performs a second version of the push and pull gesture in bedroom 2202 with non-complementary push and pull motions. Thus, first computing device 102-1 and second computing device 102-2 may leverage this contextual information (regarding how user 104 typically performs gestures in the room) when attempting to recognize ambiguous gestures in restaurant 2304 and bedroom 2202, respectively.
[0203] In one example, user 104 is awakened by the sound of an alarm and performs a push-pull gesture to command second computing device 102-2 to turn off the alarm. However, in a fatigued state of user 104, the user performs a second version of the push-pull gesture with non-complementary push and pull motions. Second computing device 102-2 may detect one or more radar signal characteristics associated with the spatial and / or temporal characteristics of the gesture and determine that the user has performed an ambiguous gesture. In particular, the user's performance of the push-pull gesture is similar to a first known gesture (a push-pull gesture) and a second known gesture (a waving gesture). If the device is unable to recognize the performed gesture with a desired confidence level, second computing device 102-2 determines that the user has performed an ambiguous gesture, which may be a first known gesture or a second known gesture. In order to recognize this ambiguous gesture, second computing device 102-2 considers contextual information about the user's habits in bedroom 2202 and determines that user 104 typically performs a push-pull gesture with non-complementary push and pull motions in a fatigued manner. Therefore, the radar signal characteristics of the ambiguous gesture more closely resemble the radar signal characteristics of the second version of the push-pull gesture that user 104 typically performs in bedroom 2202. The device determines that the ambiguous gesture is more likely to be the first known gesture (the push-pull gesture) and continues to perform the operation of turning off the alarm. Assume that in example environment 2600, both devices detect multiple radar signal characteristics associated with each version of the push-pull gesture over time (e.g., stored radar signal characteristics) and store them on one or more memories.
[0204] Additionally, first computing device 102-1 may learn over time to anticipate (e.g., expect, determine to be more common) the first version of the push-pull gesture based on the first device's location being within dining room 2304. Similarly, second computing device 102-2 may learn over time to anticipate the second version of the push-pull gesture based on the second device's location being in bedroom 2202. If first computing device 102-1 moves from dining room 2304 to bedroom 2202, first computing device 102-1 may be reconfigured (e.g., automatically reconfigured upon detecting the relocation, manually reconfigured by user 104) to anticipate the second version of the push-pull gesture instead of the first version. Similarly, if second computing device 102-2 moves from bedroom 2202 to dining room 2304, second computing device 102-2 may be reconfigured to anticipate the first version of the push-pull gesture. While the context information of the example environment 2600 includes common ways that users 104 perform gestures in a certain location, the context information may also include users 104 that are commonly associated with a certain location (e.g., room 304). Therefore, it may be useful for the computing device 102 to anticipate the user based on the location of the device, such as with respect to Fig. 27 Further description.
[0205] Fig. 27 Expectations of user presence based on the location of the computing device are shown. In example environment 2700, first computing device 102-1 may learn over time that first user 104-1 (e.g., daughter) is typically present within bedroom 2202, which may allow the device to expect the presence of the first user at 2702. If first computing device 102-1 detects the presence of a user but cannot accurately distinguish it from other users, the device may rely on expectations of the presence of the first user to distinguish user 104. For example, if an ambiguous user (e.g., a user that cannot be distinguished with a desired confidence level) is detected at a large distance within bedroom 2202, first computing device 102-1 may determine that the ambiguous user is first user 104-1 (daughter) or second user 104-2 (mother). If the daughter is primarily detected within bedroom 2202 over time (when compared to the presence of the mother), first computing device 102-1 may determine that the ambiguous user is likely to be the daughter.
[0206] Similarly, second computing device 102-2 may learn over time that second user 104-2 (the mother) is primarily present in another room (e.g., office 2204), which may allow second computing device 102-2 to anticipate the presence of the second user (here, the mother) at 2704. If second computing device 102-2 moves to bedroom 2202, second computing device 102-2 may be reconfigured to predict the presence of the daughter based on the repositioning of the device. In particular, second computing device 102-2 may access a history of users detected by first computing device 102-1 to enable anticipation of the daughter's presence. Similarly, if first computing device 102-1 moves to office 2204, first computing device 102-1 may be reconfigured to anticipate the presence of the mother. Each computing device 102-1 and 102-2 may also learn over time that certain gestures are typically detected at each location, such as with respect to gestures. Fig.28 As further described, this may improve the interpretation of ambiguous gestures. Contextual information associated with the location of the device
[0207] Fig.28 2800. In example environment 2800, first computing device 102-1 may learn that bedroom-related gestures are more common in bedroom 2202 (bedroom-related context 2802), and kitchen-related gestures are more common in kitchen 2302 (kitchen-related context 2804). Bedroom-related gestures may include commands to control alarm clocks, lighting, personal care, calendar events, etc., while kitchen-related gestures may include commands to control ovens, dishwashers, stoves, timers, etc. When computing device 102 determines that an ambiguous gesture has been performed, gesture module 224 may utilize contextual information about typical commands for a room (e.g., contexts 2802, 2804) to recognize the ambiguous gesture.
[0208] In a first example, first user 104-1 is awakened by an alarm and attempts to perform a push-pull gesture to turn it off. Because first user 104-1 is sleepy, they perform the gesture in a fatigued manner, which is different from the push-pull gesture as taught during gesture training, for example. First computing device 102-1 detects the gesture being performed and determines that it is a push-pull gesture (to turn off the alarm) or a swipe gesture (to turn on the oven), but is unable to recognize the gesture with a desired confidence level. Therefore, gesture module 224 determines that an ambiguous gesture has been performed.
[0209] Instead of prompting first user 104-1 to repeat the gesture when the alarm continues to sound, gesture module 224 uses bedroom-related context 2802 (e.g., gestures commonly performed in bedroom 2202) to recognize an ambiguous gesture. In this example, first computing device 102-1 determines that the ambiguous gesture is a push-pull gesture and sends a control signal to end the alarm. More specifically, first computing device 102-1 determines that a push-pull gesture (to turn off the alarm) is more commonly performed in bedroom 2202 than a swipe gesture (to start the oven). The determination may be based on, for example, a history of gestures performed in bedroom 2202 as stored (e.g., recorded) to a memory.
[0210] In a second example, second user 104-2 walks into kitchen 2302 and attempts to perform a swipe gesture to start an oven. Because second user 104-2 is walking, they perform the gesture in a different manner than, for example, a swipe gesture performed from a stationary position. Second computing device 102-2 detects the gesture being performed and determines that it is a swipe gesture (to start an oven) or a push-pull gesture (to turn off an alarm clock), but is unable to recognize the gesture with a desired confidence level. Therefore, gesture module 224 determines that an ambiguous gesture has been performed and uses kitchen-related context 2804 (e.g., gestures commonly performed in kitchen 2302) to recognize the ambiguous gesture. As in the previous example, second computing device 102-2 determines that a swipe gesture (to start an oven) is more commonly performed in kitchen 2302 than a push-pull gesture (to turn off an alarm clock). The device can proceed with starting the oven.
[0211] The techniques of the example environment 2800 may include the spatiotemporal machine learning model 802 and / or the machine learning model 700 (reference Figure 8 and Figure 7 ) in which the input layer 702 additionally receives context information to improve the interpretation of gestures. Figure 26 to Figure 28 The techniques are described using a history of gestures and user habits detected at one or more times in the past, but context information may also include real-time information, such as the status of an operation being performed at the current time (e.g., an ongoing operation). Contextual information associated with the operation being performed
[0212] Fig.29 2900 illustrates how the status of the operation being performed at the current time can improve the recognition of ambiguous gestures. The techniques described in the example environment 2900 can be used by the computing device 102 (see Fig.24) or a collection of computing devices 102-X forming a computing system. The first computing device 102-1 is depicted as being in the first room 304-1 (kitchen 2 302), and the second computing device 102-2 is depicted as being in the second room 304-2 (restaurant 2 304). In this example, it is assumed that the first computing device 102-1 and the second computing device 102-2 are part of the computing system. Therefore, each computing device 102-1 and 102-2 can exchange, for example, context information and / or radar signal characteristics associated with gestures performed by the user 104.
[0213] In example environment 2900, user 104 performs a gesture to start a timer for a kitchen oven. First radar system 108-1 of first computing device 102-1 detects the gesture, determines that it is associated with an operation (to start a timer), and starts a timer in kitchen 2302. User 104 then leaves kitchen 2302 and goes to restaurant 2304 to wait for their food to cook. At a later time, user 104 wants to know if the food has finished baking, so they try to perform a known gesture for second computing device 102-2 to check the status of the timer. However, user 104 performs an ambiguous gesture, which may be a known gesture (to check the status of the timer) or another known gesture (to turn off the TV), but cannot be recognized with a desired confidence level. In particular, gesture module 224 of second computing device 102-2 compares the radar signal characteristics (e.g., time and / or topological features) of the ambiguous gesture with one or more stored radar signal characteristics to determine that the ambiguous gesture may be any of those known gestures.
[0214] Thus, second computing device 102-2 utilizes contextual information about the status of an operation being performed by computing device 102-1 or 102-2 of the computing system at the current time. In this example, first computing device 102-1 is currently running a timer for an oven. Second computing device 102-2 detects this ongoing operation and determines that a known gesture (to check the status of a timer) is associated with the ongoing operation of first computing device 102-1 (running a timer). Additionally, gesture module 224 determines that another known gesture (to turn off a television) is not associated with an ongoing operation of any device of the computing system (including first computing device 102-1) because the television is not turned on. Therefore, second computing device 102-2 determines that the ambiguous gesture is most likely the first known gesture (rather than another known gesture) and reports the status of the timer as "10 minutes remaining."
[0215] However, in some cases, the status of operations at the current time may include two or more operations. When this occurs, computing device 102 may need to limit the ongoing operations by, for example, prioritizing foreground operations over background operations, such as Fig.30 Further description. Context information associated with foreground and background operations
[0216] Fig.30 It shows how blur gestures can be recognized based on the foreground and background operations being performed at the current time. In the present disclosure, foreground operations 3002 refer to operations in which the user actively participates (e.g., interacts with it, provides input to it, is displayed on the screen), and background operations 3004 refer to operations in which the user passively participates (e.g., occurs within a certain duration without user input). For example, foreground operations 3002 can include phone calls, video calls, scrolling through websites using tactile input, typing on a display, etc. Background operations 3004 can include music being played, timers, the operating status of an appliance (e.g., oven started), etc.
[0217] In example environment 3000, computing device 102 detects that user 104 has performed ambiguous gesture 3006, which may be a first known gesture or a second known gesture, but cannot be identified with a desired confidence level. Although user 104 may have intended to perform the first known gesture (e.g., a swipe gesture to turn up the volume of a phone call), user 104 performed ambiguous gesture 3006 that is associated with radar signal characteristics that are similar to both the first known gesture and the second known gesture (e.g., a swipe gesture to stop a timer). If radar system 108 cannot determine with a certain confidence level that ambiguous gesture 3006 is the first known gesture (and not the second gesture) based on the radar signal characteristics, radar system 108 can utilize contextual information to determine the intended gesture.
[0218] Contextual information for example environment 3000 includes foreground operations 3002 (e.g., a phone call to Sally) and background operations 3004 (e.g., a timer with 1:05 minutes remaining) being performed by computing device 102 at the current time. Upon detecting blur gesture 3006, radar system 108 may determine that a first known gesture (e.g., to turn up the volume on a phone call) is associated with foreground operation 3002 and a second known gesture (e.g., to stop a timer) is associated with background operation 3004. User 104 is actively engaged in the phone call to Sally and passively engaged in the timer running in the background. The device may then determine, based on this contextual information, that blur gesture 3006 is most likely the first known gesture rather than the second known gesture.
[0219] Generally speaking, context information is not limited to the status of operations being performed at the current time and may also include operations that were performed in the past or are scheduled to be performed in the future, such as information about Fig.31 Further description. Contextual information associated with past or future operations
[0220] Fig.31 1 shows how contextual information may include past and / or future operations of computing device 102. A user may routinely (e.g., daily) perform gestures associated with operations to be performed (or caused) by computing device 102. Example environment 3100 depicts: (1) operations that have been performed during a past period 3102; (2) operations that are being performed at a current time 3104; and (3) operations that are scheduled or expected to be performed in a future period 3106. A first command refers to turning off the alarm 3108 every morning, and a second command includes turning on the lights 3110 immediately afterwards (and keeping them on for a certain period of time). Computing device 102 may utilize contextual information including operations performed in past period 3102 and / or operations to be performed in future period 3106 to recognize ambiguous gestures at current time 3104. Operations in future period 3106 may be determined by the device, for example, based on room-related context or user habits (e.g., reference to a room's environment). Figure 26 to Figure 28 ) to arrange or anticipate.
[0221] In the first example, the context information includes operations performed in the past period 3102. The user 104 is typically awakened by the alarm at 6:00 am every morning and performs a push-pull gesture to turn off the alarm 3108. Then, the user 104 performs a swipe gesture to turn on the light 3110, which remains on until bedtime. These operations (depicted in the past period 3102) have been recorded by the computing device 102 to improve gesture recognition in the future. One day, the user 104 must wake up early to catch a flight and sets an early alarm 3112 to wake up at 4:00 am. When awakened by the alarm, the user 104 performs a push-pull gesture to turn off the alarm and then tries to perform a swipe gesture (at the current time 3104) to turn on the light (e.g., early light 3114) early. However, the user 104 is very tired and accidentally performs an ambiguous gesture that may be a swipe gesture or a tap gesture (to turn on the radio).
[0222] To recognize this ambiguous gesture, the gesture module 224 can refer to the contextual information about the past operations to determine that the user 104 intended to perform a swipe gesture, as depicted in the example environment 3100. In particular, the gesture module 224 can determine that: (1) the light is typically turned on every morning after the user 104 performs a push-pull gesture to turn off the alarm 3108, (2) the user 104 recently turned off the morning alarm 3112, (3) the user 104 performed an ambiguous gesture that may be a swipe gesture or a tap gesture, and (4) the swipe gesture (to turn on the light) is associated with a recurring past operation (e.g., contextual information) and the tap gesture (to turn on the radio) is not associated with a past operation. Because the ambiguous gesture may be associated with the past operation, the gesture module 224 determines that the ambiguous gesture is most likely a swipe gesture and turns on the light in the bedroom. In this example, the gesture module 224 correlates the past operations (turning off the alarm 3108 and turning on the light 3110) to improve gesture recognition.
[0223] In the second example, the context information includes an operation to be performed (e.g., scheduled) in a future period 3106. The user 104 schedules an alarm 3108 for 6:00 a.m. every morning, and presets the light 3110 to automatically turn on at 6:05 a.m. In a typical morning, the user 104 is awakened by the alarm 3108 and performs a push-pull gesture to turn it off. Unlike in the previous example, the light 3110 automatically turns on at 6:05 a.m. (as scheduled) without a gesture. One day, the user 104 must wake up early to catch a flight and manually sets the early alarm 3112 to wake up at 4:00 a.m. However, the user 104 forgets to adjust the light to automatically turn on early (e.g., at 4:05 a.m.). The early alarm 3112 rings at 4:00 a.m., but the early light 3114 does not automatically turn on at 4:05 a.m. The user 104 attempts to perform a swipe gesture in the dark to prematurely turn on the lights 3114 but accidentally performs an ambiguous gesture that may be a swipe gesture or a tap gesture (to turn on the radio).
[0224] In this example, gesture module 224 can reference contextual information about future operations to determine that user 104 intended to perform a swipe gesture. In particular, gesture module 224 can determine that: (1) the lights are scheduled to automatically turn on every morning at 6:05 AM, (2) user 104 performed an ambiguous gesture at 4:05 AM that could be a swipe gesture or a tap gesture, and (3) the swipe gesture (to turn on the lights) is associated with a scheduled future operation (e.g., turn on the lights every morning at 6:05 AM) and the tap gesture (to turn on the radio) is not associated with any scheduled operation. Therefore, gesture module 224 determines that the ambiguous gesture is most likely a swipe gesture and turns on the lights in the bedroom.
[0225] Although Figure 26 to Figure 31Examples of utilize context information associated with operations performed at the current time, past time, or future time, but in some cases this information may not be sufficient to recognize an ambiguous gesture. If gesture module 224 is unable to recognize an ambiguous gesture based on the context information (e.g., with a certain confidence level) (or abandons doing so based on the context information), computing device 102 may determine to perform a less destructive operation, such as with respect to Fig.32 Further description. Identification of less destructive operations
[0226] Fig.32 It is shown how ambiguous gestures can be recognized based on less destructive operations. In this document, less destructive operations and less destructive commands can be generally referred to to describe that an operation or a command associated with an operation is less destructive. In this way, classifying an operation as less destructive can be extended to the command that instructs a device to perform that operation.
[0227] In example environment 3200, user 104 is awakened by an alarm and accidentally performs an ambiguous gesture 3202 that is supposed to be a snooze gesture 3204 (e.g., a swipe gesture) to reset the alarm to a later time. Because user 104 performs this gesture in a fatigued manner, the radar signal characteristics of ambiguous gesture 3202 are similar to the radar signal characteristics of snooze gesture 3204 and dismiss gesture 3206 (e.g., a push-pull gesture) to turn off the alarm. Figure 26 to Figure 31 In the example of , the gesture module 224 utilizes context information to recognize the ambiguous gesture. However, in this example, the context information (e.g., the alarm rings at the current time) may be related to the operation of the snooze gesture 3204 and the dismiss gesture 3206. Therefore, the context information is not sufficient to recognize this ambiguous gesture 3202. Although described with respect to a single application, the possible gestures to which the ambiguous gesture 3202 may correspond may be associated with operations that affect the same application or different applications. In this way, the context information may be more or less helpful in determining the ambiguous gesture 3202 as a specific known gesture.
[0228] Instead of prompting the user 104 to correctly perform the snooze gesture 3204 or in addition thereto, the radar system 108 can determine a less destructive operation to be performed. A less destructive operation can be defined as an action that is less damaging, less persistent, less consequential, more reversible, and the like. For example, muting a phone call can be defined as less destructive than ending a phone call. In the example environment 3200, snoozing an alarm can be defined as less destructive than dismissing an alarm. Therefore, the radar system 108 can determine that the fuzzy gesture 3202 is more likely to be a snooze gesture 3204 and reset the alarm to a later time. A less destructive operation can be defined based on a preset condition, one or more logics, a history of user behavior, one or more user inputs (e.g., preferences), and the like. A less destructive operation can be determined by various techniques described herein, including using a machine learning model (e.g., 700, 802) or a context-based and user-based history-based technique described in this document.
[0229] In various aspects, determining a less destructive operation can determine whether an operation is a temporary operation or a final operation. A temporary operation can describe an operation that is reversible and does not terminate an instance of a process. For example, an operation that mutes a call or delays a notification can be described as a temporary operation because it only affects the characteristics of the call or notification and does not terminate it. In contrast, a final operation can describe an operation that is terminating or irreversible. For example, a final operation can terminate an instance within a computing device and prevent future operations on that instance. Therefore, with respect to a particular instance, after performing the final operation, the final operation can prevent the system from performing the temporary operation. Given the finality of the final operation, a temporary operation can be determined to be a less destructive operation than a final operation.
[0230] In addition to or as an alternative to determining a less destructive operation based on the operation itself, computing device 102 may determine a less destructive command using a command received after performing a gesture. Specifically, computing device 102 may execute a command, and user 104 may respond by performing a gesture or providing user input to the computing device. In some instances, this gesture or user input may cause computing device 102 to execute a different command to undo the original command executed on the computing device. For example, user 104 may redial a phone number, or receive a call back from a previous caller in response to the call being terminated. The command executed to undo an operation (e.g., redialing the phone) may provide an indication of the destructiveness of the operation (e.g., replaying a skipped song may be easier than reopening an unsaved and terminated application).
[0231] While some destructive determinations may be based on current responses to commands executed on computing device 102, computing device 102 may utilize previous executions of commands to determine less destructive operations. For example, at a previous time, computing device 102 may have received a particular response from user 104 after executing one or more commands associated with ambiguous gesture 3202 (e.g., a gesture that is not recognized as one known gesture but is recognized as two known gestures and is therefore ambiguous but associated with two of the possible many known gestures). Based on the response, computing device 102 may determine how destructive the operation is. In one example, computing device 102 may have previously executed a command to terminate a call or application, and user 104 responded to the command by re-initiating the application or call. This historical actions taken by the user in response to received commands or operations performed by computing device 102 (e.g., within a certain period of time) may provide an indication of the user's intended command in each case, thereby making that command less destructive. If computing device 102 determines that an operation is likely intended by user 104, then the operation may be considered less destructive, which may be determined based on previous behavior. Thus, relying on past detection and response can enable the system to improve determination of less destructive actions.
[0232] When user 104 performs an action or provides computing device 102 with a command to undo a previous action, an indication of the user intervention may be stored to improve the accuracy of future ambiguous gesture recognition. For example, when user 104 takes an action to undo a command executed due to ambiguous gesture 3202, computing device 102 may determine that ambiguous gesture 3202 was incorrectly recognized and store such determination to improve gesture recognition at a future time. Such storage may include storing a radar signal characteristic of ambiguous gesture 3202 in a manner that disassociates that characteristic from a gesture that was incorrectly associated with ambiguous gesture 3202.
[0233] Instead of performing a gesture or command to undo the incorrectly executed command, or in addition thereto, user 104 may repeat the execution of blur gesture 3202 to indicate that blur gesture 3202 was incorrectly recognized. This another execution of blur gesture 3202 may be determined to be similar or identical to the first execution of blur gesture 3202. Therefore, computing device 102 may determine that user 104 is attempting to correct computing device 102's incorrect recognition of blur gesture 3202. By determining that user 104 is re-performing blur gesture 3202 based on gesture similarity, computing device 102 may determine that the previous recognition of the blur gesture was erroneous. Once computing device 102 determines that the previous recognition of blur gesture 3202 was erroneous, it may analyze another execution of blur gesture 3202 to determine a match with a different known gesture. Therefore, blur gesture 3202 may be recognized as a gesture different from the incorrectly recognized gesture. In various aspects, the different gesture may be one of the known gestures that the blur gesture was originally associated with.
[0234] In response to determining that the original recognition of ambiguous gesture 3202 was erroneous, computing device 102 may undo or cease execution of the command associated with the incorrectly recognized ambiguous gesture 3202. Alternatively or additionally, computing device 102 may execute a command associated with a different known gesture that is correctly associated with the ambiguous gesture. As with other corrections by computing device 102, such a determination may be stored to enable computing device 102 to more accurately recognize future performances of the gesture as different gestures, which may include storing characteristics of the first or subsequent performances of ambiguous gesture 3202 in association with the different gestures.
[0235] Although less destructive operations may be determined based in whole or in part on gestures or commands received after the ambiguous gesture 3202, the preferences of the user 104 may be used to determine less destructive commands. Specifically, the user 104 may store data related to which commands or operations should be characterized as less destructive. This data may include specific user-specified decisions related to the relative destructiveness of each command / operation, or general data that may be helpful in determining less destructive operations between any set of commands / operations. User preferences may similarly determine how the ambiguous gesture 3202 is determined (e.g., whether to give higher weight to context, destructiveness, etc.). For example, a user 104 may choose to rely more on context to determine the possible relevance of the ambiguous gesture 3202, while another user may rely more on destructiveness to avoid accidental execution of a destructive operation.
[0236] Through the described techniques, a less disruptive operation can be determined, which can enable a computing system to recognize an ambiguous gesture as a known gesture in a less disruptive or minimal manner. Thus, even when a computing device fails to accurately determine an ambiguous gesture as a known gesture, determining a less disruptive operation can increase user satisfaction with gesture control. Continuous online learning
[0237] Fig.33 A user is shown performing an ambiguous gesture. In example environment 3300, user 104 performs an ambiguous gesture 3302, which the user intends to be a first gesture (e.g., a known gesture). Computing device 102 detects ambiguous gesture 3302 and attempts to correlate ambiguous gesture 3302 with a known gesture. Specifically, a radar system of computing device 102 may attempt to determine one or more radar signal characteristics associated with ambiguous gesture 3302. In this case, the radar system determines that first radar signal characteristic 3304 is associated with ambiguous gesture 3302. In general, when a user performs a gesture, the radar signal characteristics of the gesture may vary between different executions of the gesture due to slight differences in the user (or different users) each time the gesture is performed (or the orientation or distance relative to the radar system, etc.). Therefore, an instance of a gesture performed by user 104 may have a radar signal characteristic that is different from the stored radar signal characteristics associated with the known gesture, and cause gesture module 3306 to fail to determine which gesture user 104 has performed.
[0238] In this case, ambiguous gesture 3302 is determined to have first radar signal characteristic 3304. First radar signal characteristic 3304 may be compared to one or more stored radar signal characteristics. In various aspects, the stored radar signal characteristics may be implemented in a storage medium of computing device 102, a radar system, gesture module 3306, or an external storage medium accessible by gesture module 3306. The stored radar signal characteristics may be radar signal characteristics that are associated with the first gesture, for example, based on a previous performance of the gesture, a previous calibration, etc. Due to the differences in this example of ambiguous gesture 3302, the comparison of first radar signal characteristic 3304 with one or more stored characteristics may not effectively associate ambiguous gesture 3302 with the first gesture. For example, first radar signal characteristic 3304 may differ from one or more stored characteristics such that two or more gestures (or no gesture) are determined to be potentially associated with ambiguous gesture 3302.
[0239] In some implementations, gesture module 3306 can determine that ambiguous gesture 3302 is likely related to the first gesture (e.g., the first gesture has the highest correlation), but cannot determine the correlation with a required confidence level. In other implementations, gesture module 3306 can determine that ambiguous gesture 3302 corresponds to multiple stored gestures (including the first gesture), but another of the multiple stored gestures has a higher correlation than the first gesture. In response to failing to recognize the gesture with the required confidence level, computing device 102 may fail to respond to ambiguous gesture 3302 (e.g., the command associated with the first gesture is not executed by computing device 102).
[0240] When computing device 102 fails to execute the command associated with the first gesture, user 104 typically chooses to execute the command again. Fig.33 In the illustrated example, user 104 performs another gesture 3308, which is detected by computing device 102. In some examples, computing device 102 may display a notification prompting user 104 to repeat the gesture. In other implementations, the notification may be communicated to user 104 in other ways (e.g., using a tactile or auditory notification). Similar to ambiguous gesture 3302, the radar system may determine one or more radar signal characteristics of another gesture 3308. As shown, the radar system may determine that another gesture 3308 has a particular radar signal characteristic, such as second radar signal characteristic 3310. Second radar signal characteristic 3310 may be provided to gesture module 3306, where, similar to first radar signal characteristic 3304, the second radar signal characteristic is compared to one or more stored radar signal characteristics. It is assumed here that second radar signal characteristic 3310 corresponds more closely to stored radar signal characteristics when compared to first radar signal characteristic 3304, for example, because user 104 took the time to more carefully perform another gesture 3308 in response to computing device 102 failing to recognize ambiguous gesture 3302. It is also assumed here that comparison of second radar signal characteristic 3310 with the stored characteristics may enable gesture module 3306 to correlate the another gesture 3308 with the first gesture. By correlating another gesture 3308 with a known gesture (e.g., recognizing the another gesture 3308 as the first gesture), computing device 102 may execute a command associated with the known gesture. Example commands include stopping a timer, pausing / playing media being executed on the device, reacting to content being displayed on the device, and other commands described herein.
[0241] Given that computing device 102 (e.g., using gesture module 3306 or 224 of radar system 108) is able to determine that another gesture 3308 is related to a first gesture stored by computing device 102, it is likely that ambiguous gesture 3302 is also intended as a first gesture. Therefore, computing device 102 or gesture module 3306 may compare ambiguous gesture 3302 and another gesture 3308 to determine whether the two gestures are similar. Specifically, gesture module 3306 may compare first radar signal characteristic 3304 of ambiguous gesture 3302 to second radar signal characteristic 3310 of another gesture 3308 to determine whether the two gestures are similar. For more information on some of the many ways to do this disclosed herein, see the accompanying Figures 7 to 18 If it is determined that the two gestures are similar (e.g., to a required confidence level, which may be lower than the required confidence level for gesture recognition), computing device 102 may determine that ambiguous gesture 3302 is the first gesture. In order to improve the accuracy of detecting the performance of the first gesture at a future time, first radar signal characteristic 3304 of ambiguous gesture 3302 may be correlated with the first gesture and stored by computing device 102. In doing so, computing device 102 may continuously increase the accuracy of gesture recognition.
[0242] In general, other factors may be combined to determine whether blur gesture 3302 is a first gesture. For example, the amount of time that elapsed between the performance of blur gesture 3302 and another gesture 3308 may be determined, and this amount of time that elapsed may be used to determine whether blur gesture 3302 is a first gesture. In various aspects, a shorter elapsed time period (e.g., two seconds or less) may indicate a higher likelihood that blur gesture 3302 is a first gesture, while a longer elapsed time period may indicate a lower likelihood that blur gesture 3302 is a first gesture. In some implementations, the elapsed time may be a time period in which no additional gestures are detected between the blur gesture and the another gesture. In this case, if another gesture 3308 is performed immediately after the blur gesture without additional gestures being performed in between, blur gesture 3302 is more likely to be a first gesture. Here, user 104 repeats blur gesture 3302 after it is not recognized by computing device 102.
[0243] While some implementations may simply store radar signal characteristics and correlate them with recognized gestures, other implementations may utilize or store other data along with the radar signal characteristics themselves. For example, a weighting may be determined for one or more radar signal characteristics stored by computing device 102. When first radar signal characteristic 3304 or second radar signal characteristic 3310 is stored by computing device and correlated with a first gesture, a weighting may be stored or modified to indicate a confidence level of correlation of each characteristic with the first gesture.
[0244] In addition, context or other data about the gesture may be used to determine a weighting of first radar signal characteristic 3304 or second radar signal characteristic 3310. For example, a shorter amount of time elapsed between ambiguous gesture 3302 and another gesture 3308 may indicate that ambiguous gesture 3302 is likely to be a first gesture. Therefore, a weighted value of first radar signal characteristic 3304 or second radar signal characteristic 3310 may indicate a higher confidence level that first radar signal characteristic 3304 or second radar signal characteristic 3310 is associated with the first gesture. When a larger amount of time has elapsed between ambiguous gesture 3302 and another gesture 3308, ambiguous gesture 3302 may be less likely to be a first gesture. Therefore, a weighted value of first radar signal characteristic 3304 or second radar signal characteristic 3310 may indicate a lower confidence level that first radar signal characteristic 3304 or second radar signal characteristic 3310 is associated with the first gesture. In some instances, storing the weighting with first radar signal characteristic 3304 and second radar signal characteristic 3310 may improve future detection of the first gesture. Machine learning can also or instead be used to determine how to Figures 7 to 10 and described elsewhere in this document to weight various timing, context, and radar signal characteristic similarities.
[0245] When associating blur gesture 3302 with the first gesture, context information may be used or stored. For example, computing device 102 may determine context information about the performance of blur gesture 3302 (e.g., context information when blur gesture 3302 was performed). Context information may include, for example, the location of user 104 or the orientation of user 104 relative to computing device 102 or the context of the document (e.g., information about the context of the user 104). Figures 25 to 31 ) or any other type of contextual information described in the context of FIG. 3302. The contextual information may include non-radar signal characteristics of the ambiguous gesture 3302 sensed during the performance of the ambiguous gesture. As non-limiting examples, other non-radar sensors may include ultrasonic detectors, cameras, ambient light sensors, pressure sensors, barometers, microphones, or biometric sensors, among other sensors described herein. The contextual information may be stored with the first radar signal characteristic 3304 or the second radar signal characteristic 3310 to enable computing device 102 to more accurately recognize the first gesture in the future. For example, computing device 102 may store characteristics of the first gesture in different orientations or positions of user 104 relative to computing device 102. In this way, a more accurate comparison may be made by comparing the gesture to appropriate characteristics of the current context in which the gesture is performed. Additionally or alternatively, the contextual information may be used to adjust the weighting values stored with the radar signal characteristics.
[0246] Computing device 102 may also compare ambiguous gesture 3302 to another gesture 3308 to determine whether the gestures are similar, and therefore determine whether ambiguous gesture 3302 is likely to be a first gesture. For example, gesture module 3306 may compare first radar signal characteristic 3304 to second radar signal characteristic 3310 and determine that first radar signal characteristic 3304 has a correlation with second radar signal characteristic 3310 that is above a confidence threshold. While this threshold may not be sufficient to recognize a gesture, it is sufficient to indicate a reasonable probability that the gestures are related (e.g., a 40%, 50%, or 60% probability, while a threshold for recognizing gestures is 80%, 90%, or 95%). Upon determining that first radar signal characteristic 3304 and second radar signal characteristic 3310 are correlated, computing device 102 determines that ambiguous gesture 3302 is a first gesture.
[0247] In addition to gestures, computing device 102 may also use non-gesture commands to correlate with first gesture. For example, after recognizing another gesture 3308 (e.g., correlating another gesture 3308 with the first gesture), computing device 102 may query whether ambiguous gesture 3302 is a command associated with the first gesture. User 104 may respond without using a gesture (e.g., by touch or sound) to confirm the intended gesture that ambiguous gesture 3302 is intended to convey. In other cases, the user may provide feedback to computing device 102 through non-gesture commands to help identify ambiguous gesture 3302 without prompting from computing device 102.
[0248] Although the example is shown with respect to a single computing device 102, it should be noted that continuous online learning can be performed with multiple computing devices. For example, a blur gesture 3302 can be detected at a first computing device, and another gesture 3308 can be detected at a second computing device. The first computing device and the second computing device can be configured to communicate information across a communication network. In this way, blur gesture 3302 can be detected by a first computing device in a first area close to the first computing device, and another gesture 3308 can be detected by a second computing device in a second area close to the second computing device. Therefore, computing devices can take advantage of continuous online learning even when located in different areas (e.g., different rooms in a home).
[0249] Additionally, another gesture may not be recognizable using the radar system, but still provide useful information for correlating the radar signal characteristics of the ambiguous gesture with the first gesture. Assume that computing device 102 does not recognize the ambiguous gesture as a first gesture. User 104 may perform an additional gesture that is detected by computing device 102 using another sensor that is not associated with the radar system (and recognized as a first gesture in another non-radar manner) (examples listed elsewhere herein). Based on an indication of the additional gesture received at the other sensor, computing device 102 may determine that the additional gesture is a first gesture. Similar to the continuous online learning described above, computing device 102 may determine that ambiguous gesture 3302 is associated with the first gesture based on the additional gesture being the first gesture. Therefore, computing device 102 may store the radar signal characteristics of the ambiguous gesture with the first gesture, thereby enabling improved recognition.
[0250] It should be noted that various forms of gesture recognition may be used to recognize the performed gesture. For example, there are implementations in which unsegmented recognition may be used to recognize ambiguous gestures. In particular, unsegmented gesture recognition may be implemented without a priori knowledge or a wake-up trigger event indicating that the user is about to perform an ambiguous gesture. Regardless of the implementation, however, continuous online learning may enable computing device 102 to continuously improve the accuracy of gesture recognition, even for the most difficult to recognize gestures.
[0251] In addition, while the terms "continuous" and "continuously" are used to describe continuous online learning, it should be noted that the techniques are not required to be used at all times or forever, but rather the techniques can continuously and / or gradually improve future gesture recognition when used. The term "online" is also used herein to describe a way in which the techniques "learn" to better recognize and / or detect gestures, etc. This "online" term is intended to convey that the techniques learn how to better detect or recognize gestures as part of, accompanying, or throughout the process of detecting and / or recognizing gestures. In contrast, explicit training programs (e.g., in which a device trains a user to perform gestures in a specific manner) are not "online" learning. Therefore, these techniques for continuous online learning can improve gesture recognition without requiring separate training beyond normal interaction with the device. Separate training can be used in conjunction with or before these techniques, but the techniques do not require this. Online learning based on user input
[0252] Fig.34An example of online learning based on user input to improve ambiguous gesture recognition is shown. In example environment 3400, user 104 is located in the field of view of computing device 102. Computing device 102 may include radar system 108, which provides radar data to gesture module 224. In the example shown, user 104 performs an ambiguous gesture 3402, which is identified as having one or more specific radar signal characteristics. Gesture module 224 attempts to recognize ambiguous gesture 3402 as a known gesture (e.g., through unsegmented detection without a wake-up trigger event indicating that user 104 will perform the gesture), which may require that the gesture is associated with a known gesture within a predetermined confidence threshold criterion.
[0253] In this example, the confidence threshold criteria are not met, and therefore, ambiguous gesture 3402 cannot be identified as a specific gesture, but is associated with one or more known gestures (e.g., first gesture 3404 and second gesture 3406). Specifically, gesture module 224 can compare one or more radar signal characteristics of ambiguous gesture 3402 with stored characteristics associated with one or more known characteristics. If a correlation between the one or more radar signal characteristics of the ambiguous gesture and the stored gestures is determined, ambiguous gesture 3402 may be associated with the stored gestures (see for how the correlation may be performed). Figures 7 to 17 , Fig.25 and accompanying descriptions). An ambiguous gesture may be associated with multiple stored gestures due to similarity to each of the multiple stored gestures, and the gesture module 224 may not be able to recognize the ambiguous gesture as a specific known gesture. An ambiguous gesture may instead be associated with only one stored gesture, but not be recognized sufficiently to meet the predetermined confidence threshold criteria.
[0254] Generally speaking, if computing device 102 cannot recognize ambiguous gesture 3402 as one of the known gestures, ambiguous gesture 3402 may not cause computing device 102 to execute a command. Fig.34 , blur gesture 3402 is associated with first gesture 3404 corresponding to play music command 3408 and second gesture 3406 corresponding to call dad command 3410. Given that computing device 102 is unable to recognize blur gesture 3402 as one of the two recognized gestures having corresponding commands, computing device 102 may be unable to execute either command. Instead, computing device 102 may remain idle until a command is requested using a gesture or another form of input.
[0255] In some cases, computing device 102 may provide an indication of unrecognized ambiguous gesture 3402, such as by notifying user 104 using a display or speaker. However, in some instances, user 104 may execute a command or request execution of a command without being prompted (e.g., requested) by computing device 102. When computing device 102 does not execute a desired command in response to ambiguous gesture 3402, user 104 may choose to execute the command themselves or initiate the command through a different input type (e.g., non-radar input, touch input through a touch-sensitive display or keyboard, or audio input through a voice recognition system).
[0256] As shown, user 104 uses voice command 3412 to request computing device 102 to "play today's top songs". Thus, computing device 102 can receive voice command 3412 from user 104 and start playing music (e.g., using an application stored on the device or through a network connection with a media service). In general, user input can change the operating state of computing device 102 or any other connected device. For example, computing device 102 can start playing music, maintain a timer, or perform any other operation. Any of these changes in operating state can indicate that user 104 has executed a command or requested that a command be executed.
[0257] In example environment 3400, user 104 performed voice command 3412 after computing device 102 failed to respond to blur gesture 3402, and therefore, user 104 may have intended blur gesture 3402 to cause computing device 102 to perform the same action as voice command 3412. Therefore, computing device 102 may determine whether the command performed by or in response to voice command 3412 is the same as or similar to commands corresponding to gestures (e.g., first gesture 3404 and second gesture 3406) associated with blur gesture 3402. For example, computing device 102 may determine that voice command 3412 is the same as or similar to play music command 3408 corresponding to first gesture 3404 because both commands cause the computing device to play music.
[0258] After determining that voice command 3412 is the same as a command related to the associated gesture, computing device 102 may store one or more radar signal characteristics of ambiguous gesture 3402 in association with first gesture 3404 to enable gesture module 224 to better recognize future performances of first gesture 3404. Thus, gesture module 224 may utilize a spatiotemporal machine learning model associated with one or more convolutional neural networks to improve detection of ambiguous gesture 3402. In various aspects, computing device 102 may prompt user 104 to confirm that ambiguous gesture 3402 is a known gesture before storing the radar signal characteristics in association with the known gesture (e.g., by providing a notification on a display and accepting confirmation from user 104). In this way, computing device 102 may eliminate erroneous associations of radar signal characteristics with known gestures. In some examples, the radar signal characteristics associated with ambiguous gesture 3402 may be associated with specific stored radar signal characteristics of known gestures to enable gesture module 224 to increase the confidence (e.g., weight) of the association between the specific radar signal characteristics and the known gesture.
[0259] In some implementations, computing device 102 may determine a time period between performance of ambiguous gesture 3402 and user input indicating performance of a command or requesting performance (e.g., voice command 3412). This time period may be used to determine whether the command performed as a result of the user input is the same as a command associated with one of the known gestures with which ambiguous gesture 3402 is associated. Additionally or alternatively, this time period may be used to determine a weighting of radar signal characteristics associated with the known gestures, such as Fig.33 described.
[0260] In various aspects, computing device 102 may determine whether to execute or request execution of a different command (e.g., Fig.33 The different command may be different from the commands corresponding to the known gestures (e.g., first gesture 3404 and second gesture 3406) with which the blur gesture 3402 is associated (e.g., play music command 3408 and call dad command 3410). Therefore, if it is determined that the different command is not executed or requested during the time period between the execution of the blur gesture 3402 and the user input indicating the execution of the voice command 3412 or requesting the execution, one or more radar signal characteristics of the blur gesture 3402 may be stored.
[0261] It should be noted that voice commands 3412 are just one example of user input that may be used for online learning, and the techniques may instead or additionally utilize other forms of user input, such as touch commands (e.g., user 104 touching a touch-sensitive display of computing device 102, user 104 using their voice to control computing device 102 via a voice recognition system of computing device 102, or user 104 typing commands on a physical or digital keyboard). It is also important to note that user input may not be limited to being received at computing device 102. For example, user input may be received at any other device (e.g., other smart home devices) with which computing device 102 may communicate. As non-limiting examples, user input on another device may include setting a timer on a smart home appliance, actuating a smart light switch or smart door lock, or adjusting a smart thermostat.
[0262] In general, gesture module 224 may continue to improve gesture recognition for user-unique gestures, thereby improving user confidence in gesture recognition and increasing user satisfaction. Techniques for online learning may enable gesture module 224 to be trained without the need for segmented teaching using gesture training events. For example, gesture module 224 may improve gesture recognition without requiring computing device 102 to explicitly teach gestures to user 104 or request user 104 to perform gestures. Thus, the techniques may not require the additional burden of user 104 training computing device 102. Online Learning
[0263] Fig.35 Techniques for online learning of new gestures for a radar-enabled computing device are shown. In environment 3500, user 104 is positioned in a field of view of computing device 102. User 104 performs gesture 3502, which is detected by radar system 108 of computing device 102, and radar signal characteristics are determined by gesture module 224 based on gesture 3502 (e.g., as described above). In this example, user 104 rotates their hand to mimic the action of pouring coffee to indicate that user 104 wants computing device 102 to start a coffee machine. Upon detecting gesture 3502, computing device 102 can determine that gesture 3502 is an intentional action performed by user 104, rather than background motion that is not associated with a specific gesture intended for computing device 102.
[0264] The radar signal characteristic of gesture 3502 is compared to one or more stored radar signal characteristics associated with one or more known gestures. In environment 3500, the comparison is effective to determine a lack of correlation between gesture 3502 and the one or more known gestures. Specifically, the comparison may not be effective to correlate gesture 3502 or the associated radar signal characteristic with the one or more known gestures at a desired confidence level (e.g., at a low confidence threshold rather than a high confidence threshold). For example, the associated radar signal characteristic may not be sufficient to satisfy a confidence threshold criterion associated with any of the stored radar signal characteristics.
[0265] Given the lack of correlation of gesture 3502 with one or more known gestures, computing device 102 fails to recognize and respond to gesture 3502. As shown, user 104 provides command 3504 (e.g., a voice command) indicating that user 104 wants computing device 102 to start a coffee machine. In some cases, computing device 102 may determine that command 3504 is not the same as or similar to one or more known commands corresponding to one or more known gestures. Although command 3504 is shown as user 104 using their voice to ask computing device 102 to "start the coffee machine", other commands may be included, such as commands entered through a touch-sensitive display or keyboard of computing device 102 or another connected device, or by a user physically executing a command (e.g., physically starting the coffee machine).
[0266] In response to determining the lack of correlation between gesture 3502 and one or more known gestures and receiving command 3504, computing device 102 may determine that gesture 3502 is a new gesture 3506 that has not yet been learned by computing device 102 or associated with a particular command. In some cases, computing device 102 may provide some indication of such a determination to user 104, such as by displaying a message asking the user whether the gesture is a new gesture (e.g., a gesture that has not yet been taught to computing device 102 or assigned to a particular command). The user may respond by indicating that the gesture is a new gesture.
[0267] Computing device 102 may determine that gesture 3502 is new gesture 3506 and then store the radar signal characteristics associated with gesture 3502 to enable computing device 102 to recognize the performance of gesture 3502 in the future. Moreover, computing device 102 associates new gesture 3506 with a command to start coffee maker 3508. In this manner, computing device 102 may seamlessly learn new gestures at the convenience of user 104. Moreover, by doing so, computing device 102 may not be limited to learning new gestures during a gesture training period.
[0268] Although specific implementations are shown in environment 3500, it should be noted that other implementations should be considered within the scope of the present disclosure. For example, it may not be required to execute command 3504 after performing gesture 3502. In general, command 3504 can be executed close to the time of gesture 3502, such as within a predetermined time limit of two seconds, five seconds, ten seconds, or thirty seconds. Similarly, command 3504 can be executed before, after, or simultaneously with performing gesture 3502. As a specific example, user 104 can execute command 3504 before performing gesture 3502, or user 104 can execute command 3504 simultaneously with performing gesture 3502 (e.g., by speaking command 3504 while performing gesture 3502). In some implementations, command 3504 will be received by computing device 102 without computing device 102 detecting an intermediate gesture or another command initiated by user 104.
[0269] In some examples, computing device 102 can learn new gestures independently for different users. For example, when computing device 102 stores the radar signal characteristics of gesture 3502, computing device 102 can associate the radar signal characteristics and new gesture 3506 with a specific user. In this way, recognizing gesture 3502 in the future can include distinguishing the execution of the associated user and recognizing the gesture. In order to enable such a determination, computing device 102 can determine another radar signal characteristic associated with the user's presence, which can be used to determine the user as a registered user. When determining a new gesture and storing its radar signal characteristics, another radar signal characteristic associated with the user's presence can also be stored to enable the computing device to recognize the user at a future time. By independently learning new gestures associated with each user, computing device 102 can associate different gestures with different commands based on the user. For example, computing device 102 determines that gesture 3502 is a new gesture for user 104, but is a known gesture or an unknown gesture for another user. In this way, computing device 102 can enable user-customizable gesture control technology. Context-sensitive sensor configuration
[0270] Fig.36Example techniques for configuring a primary sensor to a computing device are shown. Example environment 3600 shows two different computing devices: computing device 102-1 implemented within a kitchen 3602 and computing device 102-2 implemented within an office 3604. Within kitchen 3602, user 104-1 is within the field of view of computing device 102-1. Similarly, user 104-2 is shown within the field of view of computing device 102-2. Computing device 102-1 may determine a first context 3606 associated with a condition in which a gesture is performed by user 104-1, while computing device 102-2 may determine a second context 3608 associated with a condition in which a gesture is performed by user 104-2.
[0271] User 104-1 may perform a gesture in an area within kitchen 3602. Computing device 102-1 may detect the gesture using one or more sensors configured to measure activity in the area. The manner in which each sensor is used or the accuracy or precision with which each sensor detects and recognizes the gesture may vary based on the context associated with the area in which the gesture is performed. For example, the computing device may determine the context associated with the area in which the user performs the gesture. The context may be determined based on any number of details related to the area in which the gesture is performed. For example, the context may be determined based on a reference to the area in which the gesture is performed. Figures 25 to 31 One or more aspects of context determination are described to determine context.
[0272] In some implementations, the context may be determined based on environmental conditions within the area in which the gesture is performed. For example, the computing device may determine conditions related to light, sound, interference, object placement, or movement within the area. Some sensors may be more capable of recognizing gestures under a particular set of conditions. For example, an optical sensor (e.g., a camera) may be more capable of recognizing gestures in a well-lit environment. Compared to an optical sensor, a radar sensor may experience less degradation in a poorly lit environment. Therefore, in an environment where the lighting conditions are poor, a camera may be less likely to be used to successfully recognize gestures.
[0273] In another example, when there is a lot of movement within the environment, the radar sensor may not be able to effectively recognize gestures. For example, the radar sensor may fail to distinguish different movements within the environment. In a very active environment where there is a lot of movement, a non-radar sensor (e.g., a camera) may be more able to recognize gestures.
[0274] In another example, a radar sensor may fail to recognize a gesture when there is a lot of interference in the environment. For example, multiple radar devices implemented in close proximity to each other may cause interference due to cross-signaling between the devices. When the computing device attempts to recognize a gesture, this interference may increase the noise floor in the radar receive signal and reduce the ability of the radar sensor to recognize the gesture.
[0275] Microphones can be used to supplement gesture recognition (e.g., to determine other details related to the performance of a gesture). However, in high noise environments, such audio sensors may be less effective in providing contextual details about gestures. Therefore, audio sensors may be less able to supplement gesture recognition in high noise environments.
[0276] The context may alternatively or additionally be based on the location where the computing device is located. For example, the computing device may determine that it is located in a kitchen. Based on this determination, the computing device may be able to determine which particular user is likely to perform a gesture, the type of gesture that is likely to be performed, or the environmental conditions that are likely to exist. These details may be used to determine which sensors are likely to be most capable of recognizing the gesture. For example, in a kitchen, a user may be most likely to perform gestures related to cooking (e.g., controlling an appliance, starting a timer, or searching for a recipe). The computing device may determine which sensors are most capable of detecting these particular gestures (e.g., based on past execution and recognition).
[0277] The location may provide information about environmental conditions that are likely to be present in the area where the gesture is performed. For example, a computing device located in a bedroom may be likely to be present in low light conditions, a computing device located in a living room may be likely to be in a high noise or high movement environment, etc. Additionally or alternatively, the location of the computing device may provide information about which users are likely to perform gestures or which gestures are likely to be performed. For example, a particular user or set of users may be more likely to be present in a particular area of a house, office, or another environment.
[0278] As an alternative to or in addition to using location to determine which user is performing a gesture, user detection (e.g., as described in this document) can be used to identify a specific user performing a gesture. As shown, computing device 102-1 can determine that a specific user 104-1 is performing a gesture. In doing so, first scene 3606 can include information related to user 104-1, such as data related to hand size, gesture execution rate, gesture execution clarity, etc. These details can be used to determine the ability of a specific sensor to recognize a gesture performed by user 104-1. For example, a radar sensor may be less able to recognize gestures performed by users with smaller hands. As another example, certain sensors may be more able to accurately recognize gestures from a specific user based on previous execution and detection of gestures performed by the user. In some implementations, a user may be most likely to perform a specific gesture (e.g., based on past execution and detection of the user's gesture), and certain sensors may be more able to recognize the specific gesture.
[0279] The context may be determined based on the time of day at which the gesture is performed. This information may provide data related to any of the other context-related data. For example, at night, the computing device may be more likely to be in low light conditions, the gesture may be more likely to be related to an evening event (e.g., cooking), or a particular set of users may be more likely to be present. Similarly, other times of day may have other characteristics that may be useful for context determination.
[0280] The context may be related to the foreground or background operation of the computing device. The operation of the device may provide useful indications of gestures that are likely to be performed or users that are likely to perform gestures. Recent operations to adjust conditions within the environment (e.g., lighting, audio, etc.) may provide indications of environmental conditions in the area in which the gesture was performed. This data may be used to determine the context within the area.
[0281] It should be noted that these are just some examples of ways to determine context and how the determined context can be used to determine sensor capabilities. Therefore, other examples of context determination can be used to determine sensor capabilities, such as Figures 25 to 31 The correlation between data and context or between context and sensor capabilities can be determined by any number of techniques, including any of the machine learning techniques described in this document.
[0282] Turning to the illustrated example of kitchen 3602, computing device 102-1 may determine a first context 3606 for a particular area. First context 3606 may be related to the location in which the gesture is performed (e.g., kitchen 3602), the particular user performing the gesture (e.g., user 104-1), environmental conditions of the environment in which the gesture is performed, the orientation or position of computing device 102-1 relative to user 104-1 and vice versa (e.g., if user 104-1 is within the field of view of a particular sensor), the time of day, background or foreground operations of the computing device (e.g., as described with respect to FIG. 36). Fig.30 Based on first context 3606, computing device 102-1 may determine the ability of one or more sensors of computing device 102-1 to provide data useful for recognizing gestures. The computing device may include multiple sensors, such as radar sensor 3610 and non-radar sensors (e.g., optical sensor 3612) or multiple sensors of the same type. Fig.36 A specific configuration of sensors is described, but other sensors may also be implemented, which are listed elsewhere in this document.
[0283] Computing device 102-1 may determine the ability of a first sensor to recognize a gesture performed in a particular area. For example, computing device 102-1 may determine that radar sensor 3610 is capable of recognizing a gesture based on a first context. The ability of a sensor to recognize a gesture may be progressive, so that different sensor capabilities may be compared to each other. The ability may be determined based on a prior correlation between context and gesture recognition (e.g., based on gesture recognition in an environment with a particular context). In addition to determining the ability of a first sensor to recognize a gesture performed in an area, the ability of a second sensor to recognize a gesture may also be determined.
[0284] Once the capabilities of one or more sensors are determined, the capabilities may be compared to determine which sensor is more capable of recognizing the performance of a gesture within the area. The more capable sensor may be determined as the primary sensor, and the computing device may be configured to use the primary sensor in preference to other sensors. In various aspects, the computing device may prioritize the use of the primary sensor by weighting the data collected by the primary sensor more heavily. The computing device may be configured to save power by configuring the device with a primary sensor and by reducing the power supplied to other non-primary sensors and, in some cases, increasing the power supplied to the primary sensor.
[0285] In the particular example shown, radar sensor 3610 is configured as a primary sensor. In various aspects, radar sensor 3610 is determined to be a more capable or most capable sensor due to low light conditions that may cause the amount of light in an area to be insufficient to illuminate the area and allow the optical sensor to recognize gestures. In some implementations, radar sensor 3610 may be more capable of providing data to recognize gestures based on specific preferences of user 104-1, the types of gestures most commonly performed in the area, or the orientation of radar sensor 3610 and other sensors relative to the area or to each other. In general, it should be understood that the primary sensor can be determined based on any correlation between any one of the determined contexts or a particular context and the specific capabilities of the sensor to be used to recognize gestures in that context.
[0286] Because radar sensor 3610 is configured as a primary sensor for computing device 102-1, data collected by radar sensor 3610 may be used in preference to data collected by other sensors. For example, data collected by radar sensor 3610 may be weighted more heavily during gesture detection or recognition than data collected by other sensors. Additionally or alternatively, computing device 102-1 may be configured with a primary sensor, such as being configured with a primary sensor different from radar sensor 3610, by changing a setting of computing device 102-1 during or after detection of a gesture.
[0287] As with first context 3606 determined by computing device 102-1, computing device 102-2 may determine second context 3608 related to the environment in which user 104-2 performed a gesture. Second context 3608 may be used to determine an ability of one or more sensors of computing device 102-2 to detect a gesture performed by user 104-2 within a particular area of office 3604. The capabilities of the one or more sensors may be compared to determine the most capable sensor, and the most capable sensor may be configured as a primary sensor in computing device 102-2.
[0288] As shown, optical sensor 3612 is determined to be the primary sensor. In various aspects, configuring the primary sensor for computing device 102-2 can be performed before, during, or after recognition of the gesture. In response to configuring the computing device, the user can be notified that the computing device has been configured with the primary sensor. For example, the computing device can provide a notification to the user, which can include a tone, light, text, interface notification, or vibration.
[0289] In various aspects, optical sensor 3612 may be determined to be the primary sensor because second context 3608 indicates that the area in which the gesture is performed is well lit, user 104-2 prefers optical sensor 3612 as the primary sensor, gesture recognition based on environmental conditions or user characteristics requires high resolution, there is a lot of radar interference, certain types of gestures are most likely to be performed in the area at the current or future time, or there are any other contexts that may favor optical sensor 3612. Computing device 102-2 may be configured such that optical sensor 3612 is used in preference to other sensors within computing device 102-2 to improve accuracy of gesture detection and / or recognition.
[0290] Through the described techniques, a computing device can be configured to optimally detect gestures based on the context in which the gesture is performed. In doing so, the computing device can adapt to each environment to accurately detect and recognize gestures, even in environments that pose great difficulties for gesture recognition. Detecting user interactions
[0291] Fig.37An example environment in which techniques for detecting user interactions with an interactive device may be performed is shown. Although shown as (and referred to as) a computing device 102, the interactive device may be another device that is coupled to the computing device 102 and implements a user interface to enable interaction with the user 104. In some cases, it may be beneficial for the computing device 102 to determine whether the user 104 is likely to interact with the computing device 102. For example, when an interaction between the user 104 and the computing device 102 is about to occur, the computing device 102 may display information that may be relevant to the user 104, such as time information, device status information, notifications, etc. The computing device 102 may determine the likelihood that the user 104 will interact with the device by determining a current or predicted (e.g., future, near time) interaction of the user with the computing device 102.
[0292] Computing device 102 may utilize information related to user 104 or environment 3700 to determine user interaction with the device. Computing device 102 may utilize a radar system (e.g., radar system 108) or any other sensor system of the device to determine information about user 104 or environment 3700 that may be useful for detecting user interaction. Computing device 102 may transmit radar signals and receive reflections of these radar signals reflected from user 104 or the surrounding environment. These received signals may be processed to determine characteristics of user 104 or environment 3700 that may be useful for determining user 104's interaction with computing device 102. For example, computing device 102 may determine a user's proximity 3702 relative to computing device 102, and the determined proximity 3702 may provide an indication of the user's interaction with the device.
[0293] In some examples, computing device 102 may determine that user 104 is more likely to interact with a device when the user is within a particular proximity of the device (e.g., less than or equal to a proximity threshold criterion). For example, computing device 102 or a sensor system (e.g., a radar system, a touch display, voice commands, etc.) that may be used for interaction with user 104 may have a particular distance at which it may effectively detect or recognize user interaction with the device. Thus, when the user is far away from the device, user 104 is less likely to intend to interact with the device. In this way, proximity may be used to determine user interaction with the device, such as by determining a higher user interaction when the user is closer to the device.
[0294] However, proximity alone may not be sufficient to accurately determine user interaction with a device. In some scenarios, user 104 may happen to be close to computing device 102 without ever intending to interact with computing device 102, such as when interacting with another device or person located near computing device 102 or when walking past computing device 102. Therefore, when proximity is used as the only factor in determining user interaction, computing device 102 may fail to accurately detect user interaction, resulting in suboptimal performance of the computing device.
[0295] To improve the accuracy of the detection of user interactions, any number of factors may be used to determine user interactions, including, for example, the expected proximity of the user relative to computing device 102 or the user's physical orientation relative to computing device 102. The expected proximity of the user relative to computing device 102 may include determining a rate of change of proximity of user 104 to computing device 102. In some instances, determining the expected proximity may include determining a path 3704 indicating a direction in which user 104 is moving. Path 3704 may be represented as a vector indicating the direction and, in some cases, the magnitude (e.g., rate) of movement of the user relative to the computing device. By determining the path of user 104, computing device 102 may more accurately determine whether user 104 intends or does not intend to interact with computing device 102 than a simple change in proximity.
[0296] When user 104 is moving toward computing device 102, user 104 may likely intend to interact with the device and is therefore moving toward the device to get within the field of view of the device or to view content displayed on the device. However, other instances in which user 104 is approaching the device may not be related to the user's intent to interact with the device, such as when user 104 is moving to pick up an object that is close to computing device 102. Therefore, it may be beneficial to determine user 104's path 3704, which can be used to determine where user 104 is likely to move. If path 3704 is pointing toward computing device 102 (e.g., directly toward or nearly directly toward the device, such as Fig.37 ), user 104 may be more likely to be interacting with the device and potentially interacting with the device. However, if path 3704 is not directed toward the device (e.g., away from the device or toward but not directly toward the device), the user's proximity to the device may be more likely to be coincidental and not indicative of user interaction with the device (e.g., the user is simply passing by the computing device).
[0297] The path 3704 of the user 104 may be determined in a variety of ways, such as using subsequent position measurements or Doppler measurements. In some cases, the path 3704 may be determined by comparing the current location of the user 104 to one or more historical locations of the user 104. The path 3704 may be determined based on one or more historical movements of the user 104, which may include a direction or speed. The current direction or velocity (e.g., when combined - speed) may be compared to the historical movements to determine an accurate representation of the user's path 3704 toward or away from the computing device 102. These historical movements may be immediately preceding the user's current location (e.g., a previous portion of the user's path 3704), or may be recorded from previous movements, such as when the user walked by the computing device 102 the day before. Such history may be used by constructing a machine learning model as described below or by heuristics or other means. Thus, if a user has traveled a path in the past and repeatedly did not interact with (or with) computing device 102 while following that historical path, this information can be used to determine a low (or high) probability of user 104's intent to interact or engage with computing device 102.
[0298] In addition to or in lieu of proximity 3702 or projected proximity, computing device 102 may determine a body orientation 3706 of user 104 to determine the user's interaction with computing device 102. In one example, body orientation 3706 may be based on a facial profile of user 104. A radar system or other sensor system may be used to determine where the user's face is pointing (e.g., whether the facial profile is facing toward or away from the device). Body orientation 3706 may be based on a user's line of sight (e.g., determining whether the user is looking toward or not looking at computing device 102).
[0299] Body orientation 3706 may be determined based on the user's body positioning, such as the orientation of the user's 104 torso or head toward or away from computing device 102. A radar system or other sensor system may collect information about the user (e.g., radar received signals) to determine the body orientation of user 104 by, for example, collecting radar data to determine whether a frontal profile (e.g., a flat surface such as the abdomen, chest, torso, shoulders, etc.) of user 104 is facing computing device 102. Body orientation 3706 may be used to determine a general direction of interest of the user's body, which may be toward or away from computing device 102.
[0300] Body orientation 3706 may also or alternatively include recognition of gestures of user 104. For example, computing device 102 may determine whether the user is pointing at computing device 102 or reaching toward the computing device, which may indicate the user's intention to interact or interact with the device. Gestures may include any number of gestures stored or not stored by computing device 102. In some instances, when computing device 102 detects or recognizes a gesture performed by user 104, user 104 may be likely to interact with the device. Computing device 102 may determine whether the gesture is a known gesture (e.g., a stored gesture). If the gesture corresponds to a known gesture, the user may be likely to interact with the device (e.g., the gesture is not an unrelated user movement that is not a "gesture"). Computing device 102 may determine the direction of these gestures to determine whether the gesture is likely intended for computing device 102. If the gesture is directed toward computing device 102, it may be determined that user 104 has a higher interaction with computing device 102.
[0301] Computing device 102 may use any combination of factors (e.g., including one or all of the factors) to estimate a user's interaction with the device or an expected (e.g., future, near-time) interaction. Factors may be weighted differently so that each factor can have a greater or lesser impact on determining a user's interaction. In some instances, computing device 102 may utilize at least two factors to estimate a user's interaction with computing device 102. Computing device 102 may utilize a machine learning model (e.g., 700, 802) based on any machine learning technique, including those described herein. Machine learning may be supervised, such that user 104 interacts with computing device 102 (e.g., through touch input, gesture input, voice commands, or any other input method) to confirm the user's proximity, expected proximity, or body orientation relative to the device, or an association between any of these determinations and the user's interaction.
[0302] The computing device 102 may also estimate the user's interaction based on the direction of the computing device 102 relative to the user 104. For example, if the display or field of view of the sensor system (e.g., radar system) of the computing device 102 is pointed at the user 104, the user 104 may be more likely to interact with the device. However, if the display is facing away from the user 104 or the user is outside the field of view of the sensor system of the device, the user 104 may be less likely to interact with the device. In some examples, the direction of the computing device 102 may be used in conjunction with the user's proximity 3702, estimated proximity, or body orientation 3706 to determine user interaction. For example, if the computing device 102 determines that the user 104 is moving toward the computing device 102 (e.g., within a close enough proximity), it may be determined that the user 104 has a high level of interaction with the device. If the computing device 102 is pointed in a particular direction and the user 104 is opposite the device and oriented in an opposite direction (e.g., such that the computing device 102 and the user 104 are oriented toward each other), the user 104 may be likely to interact with the device. Alternatively, if user 104 is moving in a directional orientation away from the display, computing device 102 may determine that user 104 has a low level of interaction with computing device 102 .
[0303] The user's interaction may be used to determine appropriate settings with which to configure the computing device 102. For example, if it is determined that the user 104 is interacting with the device (e.g., high interaction is determined), then content useful to the user 104 may be displayed on the device. Similarly, if the user 104 is interacting with the device and therefore may be likely to interact with the device, then the power usage of one or more sensor systems may be increased to better detect, identify, or interpret the user's interaction with the computing device 102. The computing device may determine the identity of the user 104 to determine the appropriate settings for each level of interaction. For example, when an authorized or identified user interacts with the device, the computing device may lower the privacy settings of the device to make the device's applications or content more available to the user 104 (e.g., Fig. 20 However, if it is determined that an unauthorized user is interacting with the device, the privacy setting may be changed to a more restrictive setting to prevent sensitive content from being disclosed to unauthorized or unidentified users (e.g., Fig. 20 In other cases, the privacy settings of computing device 102 may be adjusted based on user interactions without distinguishing between users 104.
[0304] When it is determined that the user is not interacting with the device (e.g., determining low interaction), the computing device can be similarly adjusted to be configured based on the user's low interaction level. For example, the sensor system of the computing device can be turned off or operated in a low power mode. Similarly, the display of the computing device 102 can be dimmed or turned off to save power of the device. When the user 104 is not interacting with the device, the device resources can be redirected to other tasks, such as maintenance tasks (e.g., updating software or firmware). In some cases, the privacy settings of the device can be improved to reduce the possibility of sensitive content being published to unauthorized users. It should be understood that although specific examples are described, other settings can be adjusted based on user interaction. Similarly, the computing device 102 can determine the user interaction based on other factors that may be useful for determining the possibility of the user interacting with the computing device. In this way, by detecting user interaction, the computing device can be changed based on the current or expected interaction of the user with the device to perform optimally in any number of situations. Example Method
[0305] Fig.38 An example method 3800 for a system of multiple radar-enabled computing devices is shown. The method 3800 is shown as a set of operations (or actions) performed, is not necessarily limited to the order or combination of operations shown herein, and may be performed in whole or in part with other methods described herein. In addition, any one or more of the operations of these methods may be repeated, combined, reorganized, or linked to provide a wide range of additional methods and / or alternative methods. In the following discussion, reference may be made to Figures 1 to 37 The example environments, experimental data, or experimental results of the present invention are provided for reference only as examples. The techniques are not limited to being performed by one or more entities operating on or associated with one computing device 102.
[0306] At 3802, a first radar transmission signal is transmitted from a first computing device of a computing system. For example, first radar transmission signal 402-1 may be transmitted from first radar system 108-1 of first computing device 102-1 to detect whether user 104 is present within first proximity zone 106-1 (e.g., Figure 1 and Fig. 20 ). The first computing device 102-1 may be part of a computing system including two or more computing devices 102-X, see Figure 3 , Figure 4 , Figure 21 to Figure 23 , Figure 26 to Figure 29 and Fig.36Each computing device 102 in the computing system may exchange information (e.g., radar signal characteristics, stored information, ongoing operations) with another device in the system via the communication network 302. This information may additionally be stored in one or more memories, which may be local, shared, remote, etc. The first radar transmission signal 402-1 may include a single signal, multiple signals that are similar or different, bursts of signals, continuous signals, etc., as shown in FIG. Figure 4 described.
[0307] At 3804, a first radar receive signal is received at a first computing device. For example, first radar receive signal 404-1 may be received by first radar system 108-1 of first computing device 102-1. First radar receive signal 404-1 may be associated with first radar transmit signal 402-1, which has been reflected from an object (e.g., user 104) within first proximity zone 106-1. This reflected signal (first radar receive signal 404-1) may represent a modification of first radar transmit signal 402-1 in time, amplitude, phase, or frequency.
[0308] At 3806, the presence of the registered user is determined based on the first radar signal characteristic of the first radar reception signal. For example, the presence of the registered user may be determined based on the first radar signal characteristic of the first radar reception signal 404-1 (see Figure 1 and Fig.21 ). This first radar receive signal 404-1 may be reflected from an object (e.g., a registered user) and include one or more radar signal characteristics (e.g., associated with time, topology, and / or gesture information) that can be used to distinguish this object as a registered user. In particular, the first radar signal characteristic of the first radar receive signal 404-1 may be correlated with the first stored radar signal characteristic of the registered user. The correlation may indicate the presence of the registered user within the first proximity area 106-1. These stored radar signal characteristics may be stored on a local, shared, or remote memory that may be accessed by any computing device 102-X of the computing system. The first radar signal characteristic may be additionally stored to the memory to improve detection and differentiation of registered users at a future time. As described herein, context and other information may also be used to assist in both detecting and distinguishing users.
[0309] At 3808, a second radar transmission signal is transmitted from a second computing device of the computing system. For example, second radar transmission signal 402-2 may be transmitted from second radar system 108-2 of second computing device 102-2 to detect whether user 104 is present within second proximity zone 106-2, see Figure 1 and Fig.21The second computing device 102 - 2 may be part of a computing system that includes at least the first computing device 102 - 1 .
[0310] At 3810, a second radar receive signal is received at a second computing device. For example, second radar receive signal 404-2 may be received by second radar system 108-2 of second computing device 102-2. Second radar receive signal 404-2 may be associated with second radar transmit signal 402-2, which has been reflected from an object (e.g., user 104) within second proximity zone 106-2.
[0311] At 3812, the presence of a registered user is determined based on a correlation of a second radar signal characteristic of the second radar receive signal with one or more stored radar signal characteristics. For example, the presence of a registered user may be determined again based on the second radar receive signal 404-2. In particular, the second radar signal characteristic of the second radar receive signal 404-2 may be correlated with a first stored radar signal characteristic or a second stored radar signal characteristic of a registered user. The correlation indicates the presence of a registered user within the second proximity area 106-2, thereby distinguishing the user from one or more other users or potential users. The second computing device 102-2 may alternatively compare the second radar signal characteristic with the first radar signal characteristic (determined by the first computing device 102-1) to determine the presence of a registered user. In this manner, information detected, determined, and / or stored by the first computing device 102-1 may be used by any one or more devices (e.g., the second computing device 102-2) in the computing system to distinguish, for example, registered users.
[0312] Fig.39 An example method 3900 for radar-based ambiguous gesture determination using contextual information is shown. The method 3900 is shown as a set of operations (or actions) performed, is not necessarily limited to the order or combination of operations shown herein, and can be performed in whole or in part with other methods described herein. In addition, any one or more of the operations of these methods can be repeated, combined, reorganized, or linked to provide a wide range of additional methods and / or alternative methods. In the following discussion, reference can be made to Figures 1 to 37 The example environments, experimental data, or experimental results of the present invention are provided for reference only as examples. The techniques are not limited to being performed by one or more entities operating on or associated with one computing device 102.
[0313] At 3902, a computing device detects an ambiguous gesture performed by a user. For example, computing device 102 utilizes radar system 108 to detect one or more radar signal characteristics of an ambiguous gesture performed by user 104. An ambiguous gesture (e.g., ambiguous gesture 2402) may refer to a gesture that cannot be recognized by the device with a desired confidence level. Techniques for detecting and analyzing radar signal characteristics associated with gestures (or users) are described above. A gesture may be detected and determined with sufficient confidence to be a gesture rather than a non-gesture body movement, an animal, a non-moving object, etc., but not with sufficient confidence to be recognized, such as below a high confidence level and above a no confidence level. In this case, the detected gesture is an ambiguous gesture.
[0314] At 3904, the ambiguous gesture is correlated with the first gesture and the second gesture. For example, computing device 102 compares one or more radar signal characteristics of the ambiguous gesture with one or more stored radar signal characteristics. Gesture module 224 correlates the ambiguous gesture with a first (known) gesture (e.g., first gesture 2404) or a second (known) gesture (e.g., second gesture 2406). The first gesture and the second gesture may correspond to a first command and a second command, respectively, to be executed by computing device 102. If gesture module 224 is unable to recognize the ambiguous gesture as the first gesture or the second gesture with a desired confidence level, computing device 102 may utilize additional information (context information 2408), such as information about the context. Figures 24 to 31 and Fig.36 In particular, the gesture debouncer 810 may use one or more heuristics to determine that the ambiguous gesture is not suffi...
Claims
1. A method comprising: detecting, at a computing device and using a radar system, an ambiguous gesture performed by a user, the ambiguous gesture being associated with a first radar signal characteristic; comparing the first radar signal characteristic to one or more stored radar signal characteristics, the comparison not effective to correlate the ambiguous gesture to a first gesture; detecting another gesture performed by the user, the other gesture being associated with a second radar signal characteristic; comparing the second radar signal characteristic of the other gesture to the one or more stored radar signal characteristics, the comparison effective to recognize the other gesture as the first gesture; and In response to the comparison: determining that the ambiguous gesture is the first gesture; and The first radar signal characteristic is stored and can be used to determine the performance of the first gesture at a future time. 2 . The method of claim 1 , wherein determining that the ambiguous gesture is the first gesture is further based on an amount of time between detecting the ambiguous gesture and detecting the other gesture.
3. The method of claim 2, wherein the amount of time is: Two seconds or less; or A time period has elapsed during which no additional gestures are detected between the ambiguous gesture and the another gesture.
4. The method according to claim 2 or 3, further comprising: determining an amount of time that has elapsed within the amount of time; as well as In response to the determining that the ambiguous gesture is the first gesture, applying a weighting value to the first radar signal characteristic and the second radar signal characteristic, the weighting value: is determined based on the amount of time that has elapsed; For a lower amount of elapsed time, a confidence level in the first radar signal characteristic and the second radar signal characteristic can be used to increase; For higher amounts of elapsed time, the confidence levels of the first radar signal characteristic and the second radar signal characteristic can be used to reduce; and is stored to improve detection of the first gesture at a future time.
5. The method according to any preceding claim, further comprising: The other gesture is determined to be similar to the ambiguous gesture by comparing the first radar signal characteristic to the second radar signal characteristic, the comparison being effective to determine that the first radar signal characteristic has a correlation with the second radar signal characteristic above a confidence threshold criterion.
6. A method as claimed in any preceding claim, wherein the comparison of the first radar signal characteristic with the one or more stored radar signal characteristics correlates the ambiguous gesture with a plurality of gestures, the correlation of the ambiguous gesture with a plurality of gestures not effectively correlating the ambiguous gesture with the first gesture with a higher confidence than at least one other gesture of the plurality of gestures.
7. The method of any preceding claim, wherein the determining that the ambiguous gesture is the first gesture is further based on receiving a prompted gesture performed by the user, the receiving being in response to providing a prompt to the user to repeat the ambiguous gesture.
8. The method of any preceding claim, wherein the determining that the ambiguous gesture is the first gesture is further based on receiving a non-gesture command.
9. A method as claimed in any preceding claim, wherein: the other gesture is detected by a second computing device of a computing system, the computing system comprising the computing device and a communications network, the communications network enabling the computing device and the second computing device to exchange information; The blur gesture is detected within a first proximity zone of the computing device; and The other gesture is detected within a second proximity zone of the second computing device.
10. The method of any preceding claim, further comprising: In response to determining that the ambiguous gesture is the first gesture, contextual information regarding performance of the ambiguous gesture is determined and stored, the stored contextual information being usable to assist in determining performance of the first gesture at a future time.
11. The method of claim 10, wherein the context information comprises a position or orientation of the user relative to the computing device during performance of the ambiguous gesture.
12. The method of claim 10 or 11, wherein the context information comprises at least one non-radar signal characteristic of the first gesture detected at another sensor.
13. The method of claim 12, wherein the another sensor comprises one or more of: an ultrasound detector, a camera, an ambient light sensor, a pressure sensor, a barometer, a microphone, or a biometric sensor.
14. A method as claimed in any preceding claim, wherein: The radar system detects the ambiguous gesture using unsegmented detection; and The unsegmented detection is performed without a priori knowledge or a wake-up trigger event indicating that the user will perform the ambiguous gesture.
15. A computing device comprising: at least one antenna; a radar system configured to transmit radar transmit signals and receive radar receive signals using the at least one antenna; at least one processor; as well as A computer-readable storage medium comprising instructions, which in response to being executed by the processor are used to instruct the computing device to perform any one of the methods of claims 1 to 14.