Training a machine learning module for radar-based gesture detection in a surrounding computing environment
By combining a radar system with an ambient computing machine learning module, the inconvenience of smart device interaction and environmental challenges have been solved, enabling low-power, contactless gesture detection, which enhances user experience and privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2022-04-08
- Publication Date
- 2026-05-12
Smart Images

Figure CN117203641B_ABST
Abstract
Description
Background Technology
[0001] As smart devices become increasingly prevalent, users are integrating them into their daily lives. For example, users might use one or more smart devices to get daily weather and traffic information, control their home's temperature, answer doorbells, turn lights on or off, and / or play background music. However, interacting with some smart devices can be cumbersome and inefficient. For instance, some smart devices may have a physical user interface that requires users to physically touch the device to navigate through one or more prompts. In this case, users must divert their attention from other primary tasks to interact with the smart device, which can be inconvenient and disruptive. Summary of the Invention
[0002] This paper describes techniques and apparatus for training machine learning modules to perform radar-based gesture detection in ambient computing environments. Compared to other smart devices that rely on physical user interfaces, smart devices with radar systems can support ambient computing by providing a gesture-based user interface with eye-free interaction and lower cognitive requirements. Radar systems can be designed to address various challenges associated with ambient computing, including power consumption, environmental variations, background noise, size, and user privacy. The radar system uses an ambient computing machine learning module to rapidly identify gestures performed by the user from a distance of up to at least two meters.
[0003] The training of the ambient computation machine learning module involves a two-stage evaluation process. The two-stage evaluation process consists of a first stage, which performs a segmented classification task using pre-segmented data. The second stage uses unsegmented or continuous time-series data and a gesture de-jitter to perform an unsegmented recognition task. The unsegmented recognition task can be significantly more challenging than the segmented classification task because it is unknown when the gesture appears in continuous time-series data. By performing both the segmented classification task and the unsegmented recognition task, the ambient computation machine learning module can be trained to filter background noise and achieve a sufficiently low false positive rate to enhance the user experience.
[0004] The aspects described below include a method for training a machine learning module to perform peripheral computations. This method includes evaluating the machine learning module using a two-stage evaluation process. The evaluation includes using the machine learning module to perform a segmented classification task using pre-segmented data to evaluate errors associated with the classification of multiple gestures. The pre-segmented data includes complex radar data with multiple gesture segments. Each gesture segment in the multiple gesture segments includes a gesture motion. The center of the gesture motion across the multiple gesture segments has the same relative timing alignment within each gesture segment. The method also includes using the machine learning module to perform an unsegmented identification task using continuous time-series data to evaluate the false alarm rate. The continuous time-series data includes other complex radar data. The method additionally includes adjusting one or more elements of the machine learning module to reduce errors and the false alarm rate.
[0005] The aspects described below also include a system comprising a radar system and a processor. The processor is configured to process complex radar data generated by the radar system according to a machine learning module trained according to any of the methods described.
[0006] The aspects described below include a computer-readable storage medium comprising computer-executable instructions that, in response to execution by a processor, cause the system to perform any of the methods described.
[0007] The aspects described below also include an intelligent device comprising a radar system and a processor. The processor is configured to process complex radar data generated by the radar system according to a machine learning module trained according to any of the methods described.
[0008] The aspects described below also include systems with means for training machine learning modules for radar systems to perform ambient computations. Attached Figure Description
[0009] The following figures illustrate an apparatus and techniques for training a machine learning module to perform radar-based gesture detection in an ambient computing environment. The same numbers are used throughout the figures to refer to similar features and components:
[0010] Figure 1-1 The diagram illustrates an example environment that enables the use of a radar system for ambient calculations.
[0011] Figure 1-2 The illustration shows an example of a swipe gesture associated with surrounding calculations;
[0012] Figure 1-3 The illustration shows an example of a tapping gesture associated with surrounding calculations;
[0013] Figure 2The illustration shows an example implementation of a radar system as part of a smart device;
[0014] Figure 3-1 The diagram illustrates the operation of a radar system;
[0015] Figure 3-2 The diagram illustrates an example Radar frame structure used for peripheral calculations;
[0016] Figure 4 The illustration shows an example antenna array and an example transceiver for a radar system used for ambient calculations.
[0017] Figure 5 The illustration shows an example scheme for ambient calculation implemented by a radar system;
[0018] Figure 6-1 The diagram illustrates an example hardware abstraction module used for peripheral computing;
[0019] Figure 6-2 The illustration shows example complex radar data generated by the hardware abstraction module for ambient calculations.
[0020] Figure 7-1 The illustration shows an example of a surrounding computation machine learning module and a gesture deshaker for surrounding computation.
[0021] Figure 7-2 The illustration shows an example graph of the probability across multiple gesture frames used for surrounding calculations;
[0022] Figure 8-1 The illustration shows a first example frame model used for ambient calculations using a radar system.
[0023] Figure 8-2 The illustration shows a first example time model for using the surrounding calculations of a radar system;
[0024] Figure 9-1 The illustration shows a second example frame model used for ambient calculations using a radar system;
[0025] Figure 9-2 The diagram illustrates an example residual block used for peripheral calculations employing a radar system.
[0026] Figure 9-3 The illustration shows a second example time model for using the surrounding calculations of a radar system;
[0027] Figure 10-1 The illustration shows an example environment that can collect data for training machine learning modules to perform radar-based gesture detection in the surrounding computing environment;
[0028] Figure 10-2The diagram illustrates an example flowchart that collects positive records to train a machine learning module to perform radar-based gesture detection in the surrounding computing environment.
[0029] Figure 10-3 The diagram illustrates an example flowchart for refining the timing of gesture segments within a confirmed record;
[0030] Figure 11 The diagram illustrates an example flowchart for enhancing positive and / or negative records;
[0031] Figure 12 The diagram illustrates an example flowchart for training a machine learning module to perform radar-based gesture detection in the surrounding computing environment.
[0032] Figure 13 The illustration shows an example of how radar systems can be used to facilitate ambient computing.
[0033] Figure 14 The illustration shows another example method for generating radar-based gesture detection events in the surrounding computing environment;
[0034] Figure 15 The diagram illustrates an example method for training a machine learning module to perform radar-based gesture detection in a surrounding computing environment; and
[0035] Figure 16 The illustration shows an example computing system that embodies, or can be implemented in, the technology that enables the use of radar systems for ambient calculations. Detailed Implementation
[0036] As smart devices become increasingly prevalent, users are integrating them into their daily lives. For example, users might use one or more smart devices to get daily weather and traffic information, control their home's temperature, answer doorbells, turn lights on or off, and / or play background music. However, interacting with some smart devices can be cumbersome and inefficient. For instance, some smart devices may have a physical user interface that requires users to physically touch the device to navigate through one or more prompts. In this case, users must divert their attention from other primary tasks to interact with the smart device, which can be inconvenient and disruptive.
[0037] To address this issue, some smart devices support ambient computing, which allows users to interact with smart devices in a non-physical and less cognitively demanding way compared to other interfaces that require physical touch and / or the user's visual attention. Leveraging ambient computing, smart devices seamlessly exist within their surroundings and provide users with access to information and services while they perform primary tasks such as cooking, cleaning, driving, talking to people, or reading books.
[0038] However, incorporating ambient computing into smart devices presents several challenges. These challenges include power consumption, environmental variations, background noise, size, and user privacy. Power consumption is a challenge because one or more sensors in a smart device that enables ambient computing must be permanently "on" to detect input from the user, which can happen at any time. For smart devices that rely on battery power, it may be desirable to use sensors that utilize relatively low power levels to ensure that the smart device can operate for a day or more.
[0039] The second challenge is the diverse environments in which smart devices can perform ambient computations. In some cases, time-based progressions (e.g., from day to night, or from summer to winter) naturally alter a given environment. These natural changes can lead to temperature fluctuations and / or alterations in lighting conditions. Thus, it is expected that smart devices will be able to perform ambient computations across such environmental changes.
[0040] The third challenge involves background noise. Compared to other devices that respond to touch-based input for user interaction, smart devices performing ambient computations may experience significantly more background noise when they operate in a permanently "on" state. For smart devices with voice user interfaces, background noise can include background conversations. For other smart devices with gesture-based user interfaces, this can include additional movements associated with daily tasks. To avoid disturbing the user, it is desirable for smart devices to filter out this background noise and reduce the probability of misidentifying background noise as user input.
[0041] The fourth challenge is size. Smart devices are expected to have a relatively small footprint. This allows them to be embedded within other objects or occupy less space on countertops or walls. The fifth challenge is user privacy. Because smart devices can be used in personal spaces (e.g., bedrooms, living rooms, or workplaces), it is expected that they can be integrated into the surrounding computing in a way that protects user privacy.
[0042] To address these challenges, techniques for training machine learning modules to perform radar-based gesture detection in an ambient computing environment are described. The radar system can be integrated into power- and space-constrained smart devices. In an example implementation, the radar system consumes 20 milliwatts or less of power and has a footprint of 4 mm × 6 mm. The radar system can also be easily housed behind materials that do not substantially affect the propagation of radio frequency signals—such as plastics, glass, or other non-metallic materials. Additionally, compared to infrared sensors or cameras, the radar system is less susceptible to changes in temperature or lighting. Furthermore, radar sensors do not produce distinguishable representations of a user's spatial structure or voice. In this way, radar sensors can provide better privacy protection compared to other image-based sensors.
[0043] To support ambient computing, the radar system uses an ambient computing machine learning module, designed to operate with limited power and computing resources. This module enables the radar system to quickly recognize gestures performed by a user at a distance of at least two meters. This allows users to flexibly interact with smart devices while performing other tasks at maximum distances away.
[0044] The training of the ambient computation machine learning module involves a two-stage evaluation process. The two-stage evaluation process consists of a first stage, which performs a segmented classification task using pre-segmented data. The second stage uses unsegmented or continuous time-series data and a gesture de-jitter to perform an unsegmented recognition task. The unsegmented recognition task can be significantly more challenging than the segmented classification task because it is unknown when the gesture appears in continuous time-series data. By performing both the segmented classification task and the unsegmented recognition task, the ambient computation machine learning module can be trained to filter background noise and achieve a sufficiently low false positive rate to enhance the user experience.
[0045] Operating environment
[0046] Figure 1-1 These are illustrations of example environments 100-1 to 100-5, in which the use of radar systems and devices including ambient calculations using radar systems can be represented. In the depicted environments 100-1 to 100-5, smart device 104 includes radar system 102 capable of performing ambient calculations. Although smart device 104 is shown as a smartphone in environments 100-1 to 100-5, smart device 104 can generally be implemented as any type of device or object, such as relative to… Figure 2 Further description.
[0047] In environments 100-1 to 100-5, the user performs different types of gestures, which are detected by radar system 102. In some cases, the user uses appendages or body parts to perform gestures. Alternatively, the user may use a stylus, a handheld object, a ring, or any type of material that can reflect radar signals to perform gestures.
[0048] In environment 100-1, the user makes a scrolling gesture by moving their hand above smart device 104 along a horizontal dimension (e.g., from the left side of smart device 104 to the right side of smart device 104). In environment 100-2, the user makes a reaching gesture, which reduces the distance between smart device 104 and the user's hand. In environment 100-3, the user makes a tapping gesture by moving their hand toward and away from smart device 104. In environment 100-4, smart device 104 is stored in a wallet, and radar system 102 provides occlusion gesture recognition by detecting gestures obscured by the wallet. In environment 100-5, the user makes a gesture to start a timer or silence an alarm.
[0049] The radar system 102 can also recognize other types of gestures or movements not shown in Figure 1. Example types of gestures include a handle-turning gesture, in which the user curls their fingers to grip an imaginary doorknob and rotates their fingers and hand clockwise or counterclockwise to mimic the action of turning the imaginary doorknob. Another example type of gesture includes a spindle-twisting gesture, which the user performs by rubbing their thumb and at least one other finger together. Gestures can be two-dimensional, such as those used with touch-sensitive displays (e.g., pinching, spreading, or tapping two fingers). Gestures can also be three-dimensional, such as many sign language gestures, for example, American Sign Language (ASL) and gestures for other sign languages worldwide. Once each of these gestures is detected, the smart device 104 can perform actions such as displaying new content, playing music, moving the cursor, activating one or more sensors, opening an application, etc. In this way, the radar system 102 provides contactless control of the smart device 104.
[0050] Some gestures can be associated with specific directions for navigating visual or auditory content presented by the smart device 104. These gestures can be performed along a horizontal plane that is generally parallel to the smart device 104 (e.g., generally parallel to the display of the smart device 104). For example, a user can perform a first swipe gesture (e.g., swipe right) to move from the left side of the smart device 104 to the right side to play the next song in the queue or skip to the next song. Alternatively, a user can perform a second swipe gesture (e.g., swipe left) to move from the right side of the smart device 104 to the left side to play the previous song in the queue or skip to the next song. To scroll through visual content in different directions, a user can perform a third swipe gesture (e.g., swipe up) to move from the bottom of the smart device 104 to the top of the smart device 104 or a fourth swipe gesture (e.g., swipe down) to move from the top of the smart device 104 to the bottom of the smart device 104. Generally speaking, gestures associated with navigation can be mapped to navigation inputs, such as changing songs, navigating card lists, and / or clearing items.
[0051] Other gestures can be associated with selections. These gestures can be performed along a vertical plane that is generally perpendicular to the smart device 104. For example, a user can use a tap gesture to select a specific option presented by the smart device 104. In some cases, a tap gesture can be equivalent to a mouse click or a tap on a touchscreen. Generally, gestures associated with selections can be mapped to "take action" intents, such as starting a timer, opening a notification card, playing a song, and / or pausing a song.
[0052] Some implementations of radar system 102 are particularly advantageous when applied in the context of smart device 104, but this presents a convergence of problems. These may include the need for spacing and layout of radar system 102, as well as low power limitations. An exemplary overall lateral dimension of smart device 104 could be, for example, approximately eight centimeters by approximately fifteen centimeters. An exemplary footprint of radar system 102 could even be further limited, such as approximately four millimeters by six millimeters including the antenna. An exemplary power consumption of radar system 102 could be on the order of several milliwatts to tens of milliwatts (e.g., between approximately two and twenty milliwatts). This limited footprint and power consumption requirement of radar system 102 allows smart device 104 to include other desired features (e.g., camera sensors, fingerprint sensors, displays, etc.) in a space-constrained package.
[0053] Using a radar system 102 that provides gesture recognition, smart devices can support ambient computing by providing shortcuts for everyday tasks. Example shortcuts include managing interruptions to alarms, timers, or smoke detectors. Other shortcuts include accelerating interactions with voice-controlled smart devices. This type of shortcut can activate voice recognition within the smart device without requiring a keyword to wake it. Sometimes, users may prefer gesture-based shortcuts over voice-activated shortcuts, especially when they are engaged in a conversation or in environments unsuitable for speaking, such as in a classroom or a quiet area of a library. When a user is driving, ambient computing using the radar system allows the user to accept or reject changes to their Global Navigation Satellite System (GNSS) route.
[0054] Ambient computing can also be applied to public spaces to control everyday items. For example, radar system 102 can recognize gestures used to control building features. These gestures could enable people to open automatic doors, select floors in elevators, and raise or lower office blinds. As another example, radar system 102 can recognize gestures used to operate faucets, flush toilets, or fountain-style water dispensers. In an example implementation, radar system 102 recognizes different types of swipe gestures, which will relate to... Figure 1-2 Further description. Optionally, the radar system can also recognize tapping gestures, which will be related to... Figure 1-3 Further description.
[0055] Figure 1-2 The illustration shows example types of swipe gestures associated with surrounding computation. Generally, a swipe gesture represents a sweeping motion across at least both sides of smart device 104. In some cases, a swipe gesture can resemble the motion of brushing crumbs off a table. A user can perform a swipe gesture using a hand with its palm facing smart device 104 (e.g., a hand positioned parallel to smart device 104). Alternatively, a user can perform a swipe gesture using a hand with its palm facing or away from the direction of motion (e.g., a hand positioned perpendicular to smart device 104). In some cases, a swipe gesture can be associated with timing requirements. For example, to be considered a swipe gesture, the user would sweep an object across two opposite points on smart device 104 within approximately 0.5 seconds.
[0056] In the depicted configuration, smart device 104 is shown with a display 106. The display 106 is considered to be located on the front side of smart device 104. Smart device 104 also includes sides 108-1 to 108-4. Radar system 102 is positioned near side 108-3. Consider a smart device 104 positioned in a portrait orientation, such that the display 106 faces the user and the side 108-3 with radar system 102 is positioned away from the ground. In this case, the first side 108-1 of smart device 104 corresponds to the left side of smart device 104, and the second side 108-2 of smart device 104 corresponds to the right side of smart device 104. Furthermore, the third side 108-3 of smart device 104 corresponds to the top side of smart device 104, and the fourth side 108-4 of smart device 104 corresponds to the bottom of smart device 104 (e.g., the side positioned near the ground).
[0057] At 110, the arrows depict the directions of a right swipe 112 (e.g., a right swipe gesture) and a left swipe 114 (e.g., a left swipe gesture) relative to the smart device 104. To perform a right swipe 112, the user moves an object (e.g., an appendage or stylus) from the first side 108-1 of the smart device 104 to the second side 108-2 of the smart device 104. To perform a left swipe 114, the user moves an object from the second side 108-2 of the smart device 104 to the first side 108-1 of the smart device 104. In this case, the right swipe 112 and the left swipe 114 traverse paths that are generally parallel to the third side 108-3 and the fourth side 108-4 and generally perpendicular to the first side 108-1 and the second side 108-2.
[0058] At 116, arrows depict the directions of an upward swipe 118 (e.g., an upward swipe gesture) and a downward swipe 120 (e.g., a downward swipe gesture) relative to the smart device 104. To perform the upward swipe 118, the user moves an object from the fourth side 108-4 to the third side 108-3. To perform the downward swipe 120, the user moves an object from the third side 108-3 to the fourth side 108-4. In this case, the upward swipe 118 and the downward swipe 120 traverse a path that is generally parallel to the first side 108-1 and the second side 108-2, and generally perpendicular to the third side 108-3 and the fourth side 108-4.
[0059] At 122, the arrow depicts the direction of an example omnidirectional swipe (e.g., an omnidirectional swipe gesture) relative to the smart device 104. Omnidirectional swipe 124 represents a swipe that is not necessarily parallel or perpendicular to a given side. In other words, omnidirectional swipe 124 represents any type of swipe motion, including directional swipes mentioned above (e.g., swipe right 112, swipe left 114, swipe up 118, and swipe down 120). Figure 1-2 In the example shown, the omnidirectional slide 124 is a diagonal slide from the point touched by sides 108-1 and 108-3 to another point touched by sides 108-2 and 108-4. Other types of diagonal movements are also possible, such as diagonal slides from the point touched by sides 108-1 and 108-4 to another point touched by sides 108-2 and 108-3.
[0060] Various swipe gestures can be defined from a device-centric perspective. In other words, a right swipe 112 typically travels from the left side to the right side of the smart device 104, regardless of the smart device 104's orientation. Consider the following example: where the smart device 104 is positioned in a landscape orientation, has a user-facing display 106, and a third side 108-3 has a radar system 102 positioned on the right side of the smart device 104. In this case, the first side 108-1 represents the top side of the smart device 104, and the second side 108-2 represents the bottom side of the smart device 104. The third side 108-3 represents the right side of the smart device 104, and the fourth side 108-4 represents the left side of the smart device 104. Thus, the user performs a right swipe 112 or a left swipe 114 by moving the object across the third side 108-3 and the fourth side 108-4. To perform an upward swipe 118 or a downward swipe 120, the user moves the object across the first side 108-1 and the second side 108-2.
[0061] At 126, a vertical distance 128 is shown between the object performing any of the swipe gestures 112, 114, 118, 120, and 124 and the front surface of the smart device 104 (e.g., the surface of the display 106) to remain relatively constant throughout the gesture. For example, the starting position 130 of the swipe gesture may be located at approximately the same vertical distance 128 from the smart device 104, as the ending position 132 of the swipe gesture. The term "approximately" may mean that the distance of the starting position 130 may be within + / - 10% or less of the distance of the ending position 132 (e.g., within + / - 5%, + / - 3%, or + / - 2% of the ending position 132). Alternatively, a swipe gesture involves movement across a path that is generally parallel to the surface of the smart device 104 (e.g., generally parallel to the surface of the display 106).
[0062] In some cases, a swipe gesture can be associated with a specific range of vertical distances 128 from the smart device 104. For example, if a gesture is performed at a vertical distance 128 that is approximately 3 to 20 centimeters from the smart device 104, the gesture can be considered a swipe gesture. The term "approximately" can mean that the distance can be within + / - 10% of a specified value or less (e.g., within + / - 5%, + / - 3%, or + / - 2% of a specified value). Although in Figure 1-2 The start position 130 and end position 132 are shown above the smart device 104, but the start position 130 and end position 132 of other swipe gestures can be positioned away from the smart device 104, especially when the user performs the swipe gesture at a horizontal distance from the smart device 104. As an example, the user can perform the swipe gesture at a distance of more than 0.3 meters from the smart device 104.
[0063] Figure 1-3 The illustration depicts an example tapping gesture associated with surrounding computation. Generally, a tapping gesture is a "bouncing" motion that first moves towards and then away from the smart device 104. This motion is generally perpendicular to the surface of the smart device 104 (e.g., generally perpendicular to the surface of the display 106). In some cases, a user may perform a tapping gesture using a hand with its palm facing the smart device 104 (e.g., a hand positioned parallel to the smart device 104).
[0064] Figure 1-3 The movement of a tapping gesture over time is depicted, with time progressing from left to right. At position 134, the user positions an object (e.g., an appendage or stylus) at a starting position 136, which is at a first distance 138 from the smart device 104. The user moves the object from the starting position 136 to an intermediate position 140. This intermediate position 140 is at a second distance 142 from the smart device 104. The second distance 142 is less than the first distance 138. At position 144, the user moves the object from the intermediate position 140 to an ending position 146, which is at a third distance 148 from the smart device 104. The third distance 148 is greater than the second distance 142. The third distance 148 may be similar to or different from the first distance 138. (The remaining text appears to be unrelated and likely refers to further details about the third distance.) Figure 2 The intelligent device 104 and the radar system 102 are further described.
[0065] Figure 2Radar system 102 is illustrated as part of smart device 104. Smart device 104 is illustrated with various non-limiting example devices, including desktop computer 104-1, tablet computer 104-2, laptop computer 104-3, television 104-4, computing watch 104-5, computing glasses 104-6, gaming system 104-7, microwave oven 104-8, and vehicle 104-9. Other devices may also be used, such as home service devices, smart speakers, smart thermostats, security cameras, baby monitors, and Wi-Fi. TM Routers, drones, trackpads, drawing tablets, netbooks, e-readers, home automation and control systems, wall displays, and other home appliances. Note that smart device 104 can be wearable, non-wearable but mobile, or relatively stationary (e.g., desktop computers and appliances). Radar system 102 can be used as a standalone radar system, or in conjunction with many different smart devices 104 or peripherals, or embedded in many different smart devices 104 or peripherals, such as embedded in control panels for home appliances and systems, embedded in a car to control internal functions (e.g., volume, cruise control, or even car driving), or used as an accessory to a laptop computer to control computing applications on the laptop computer.
[0066] The smart device 104 includes one or more computer processors 202 and at least one computer-readable medium 204, which includes a memory medium and a storage medium. An application and / or operating system (not shown), embodied as computer-readable instructions on the computer-readable medium 204, can be executed by the computer processor 202 to provide some of the functionality described herein. The computer-readable medium 204 also includes an application 206 that uses ambient computational events (e.g., gesture input) detected by the radar system 102 to perform actions associated with gesture-based touchless control. In some cases, the radar system 102 is also capable of providing radar data to support presence-based touchless control, collision avoidance for autonomous driving, health monitoring, fitness tracking, spatial mapping, human activity recognition, and so on.
[0067] The smart device 104 may also include a network interface 208 for transmitting data via wired, wireless, or optical networks. For example, the network interface 208 may transmit data via a local area network (LAN), wireless local area network (WLAN), personal area network (PAN), wired local area network (WAN), intranet, internet, peer-to-peer network, point-to-point network, mesh network, etc. The smart device 104 may also include a display 106.
[0068] Radar system 102 includes a communication interface 210 for transmitting radar data to a remote device, although this communication interface 210 is not required when radar system 102 is integrated into smart device 104. Radar data may include ambient computation events and may optionally include other types of data, such as data associated with presence detection, collision avoidance, health monitoring, fitness tracking, spatial mapping, or human activity identification. Typically, radar data provided by communication interface 210 is in a format usable by application 206.
[0069] The radar system 102 also includes at least one antenna array 212 and at least one transceiver 214 for transmitting and receiving radar signals. The antenna array 212 includes at least one transmitting antenna element and at least two receiving antenna elements. In some cases, the antenna array 212 includes multiple transmitting antenna elements and / or multiple receiving antenna elements. Using multiple transmitting antenna elements and multiple receiving antenna elements, the radar system 102 can realize a multiple-input multiple-output (MIMO) radar capable of transmitting multiple distinct waveforms at a given time (e.g., each transmitting antenna element transmits a different waveform). The antenna elements can be circularly polarized, horizontally polarized, vertically polarized, or a combination thereof.
[0070] The multiple receiving antenna elements of antenna array 212 can be positioned in a one-dimensional shape (e.g., a line) or a two-dimensional shape (e.g., a rectangular arrangement, a triangular arrangement, or an "L" shape) for implementations including three or more receiving antenna elements. A one-dimensional shape allows radar system 102 to measure one angular dimension (e.g., azimuth or elevation), while a two-dimensional shape allows radar system 102 to measure two angular dimensions (e.g., to determine the azimuth and elevation angles of an object). The spacing between elements associated with the receiving antenna elements can be less than, greater than, or equal to half the center wavelength of the radar signal.
[0071] Transceiver 214 includes circuitry and logic for transmitting and receiving radar signals via antenna array 212. Components of transceiver 214 may include amplifiers, phase shifters, mixers, switches, analog-to-digital converters, or filters for conditioning the radar signals. Transceiver 214 also includes logic for performing in-phase / quadrature (I / Q) operations such as modulation or demodulation. Various modulation methods can be used, including linear frequency modulation, triangular frequency modulation, stepped frequency modulation, or phase modulation. Alternatively, transceiver 214 may generate radar signals with a relatively constant frequency or monotone. Transceiver 214 may be configured to support continuous wave or pulse radar operation.
[0072] The transceiver 214 is used to generate a radar signal spectrum (e.g., frequency range) that can cover between 1 and 400 GHz, 4 and 100 GHz, 1 and 24 GHz, 24 GHz and 70 GHz, 2 and 4 GHz, 57 and 64 GHz, or at approximately 2.4 GHz. In some cases, the spectrum can be divided into multiple sub-spectrums with similar or different bandwidths. The bandwidth can be on the order of 500 MHz, 1 GHz, 2 GHz, 4 GHz, 6 GHz, etc. In some cases, the bandwidth for implementing ultra-wideband (UWB) radar is approximately 20% or more of the center frequency.
[0073] Different frequency sub-spectrums may include, for example, frequencies between approximately 57 and 59 GHz, 59 and 61 GHz, or 61 and 63 GHz. While the example frequency sub-spectrums described above are continuous, other frequency sub-spectrums may not be continuous. To achieve coherence, transceiver 214 can use multiple frequency sub-spectrums (continuous or discontinuous) with the same bandwidth to generate multiple radar signals, which may be transmitted simultaneously or temporally separated. In some cases, multiple continuous frequency sub-spectrums may be used to transmit a single radar signal, thereby enabling the radar signal to have a wide bandwidth.
[0074] The radar system 102 also includes one or more system processors 216 and at least one system medium 218 (e.g., one or more computer-readable storage media). In the depicted configuration, the system medium 218 optionally includes a hardware abstraction module 220. Instead of relying on techniques that directly track the position of the user's hand over time to detect gestures, the system medium 218 of the radar system 102 includes a surrounding computation machine learning module 222 and a gesture dejitter 224. The hardware abstraction module 220, the surrounding computation machine learning module 222, and the gesture dejitter 224 can be implemented using hardware, software, firmware, or a combination thereof. In this example, the system processor 216 implements the hardware abstraction module 220, the surrounding computation machine learning module 222, and the gesture dejitter 224. The hardware abstraction module 220, the surrounding computation machine learning module 222, and the gesture dejitter enable the system processor 216 to process responses from receiving antenna elements in the antenna array 212 to identify gestures performed by the user 302 in a surrounding computational context.
[0075] In an alternative implementation (not shown), the hardware abstraction module 220, the ambient computation machine learning module 222, and / or the gesture de-jitter 224 are included within the computer-readable medium 204 and implemented by the computer processor 202. This enables the radar system 102 to provide raw data to the smart device 104 via the communication interface 210, allowing the computer processor 202 to process the raw data for application 206.
[0076] Hardware abstraction module 220 transforms the raw data provided by transceiver 214 into hardware-agnostic data, which can be processed by the ambient computation machine learning module 222. Specifically, hardware abstraction module 220 conforms complex data from various types of radar signals to the expected input of ambient computation machine learning module 222. This enables ambient computation machine learning module 222 to process different types of radar signals received by radar system 102, including those utilizing different modulation schemes for frequency-modulated continuous wave radar, phase-modulated spread spectrum radar, or pulse radar. Hardware abstraction module 220 can also normalize complex data from radar signals with different center frequencies, bandwidths, transmit power levels, or pulse widths.
[0077] Additionally, the hardware abstraction module 220 accommodates complex data generated using different hardware architectures. These different hardware architectures may include different antenna arrays 212 or collections of different antenna elements positioned on different surfaces of the smart device 104. By using the hardware abstraction module 220, the surrounding computational machine learning module 222 can process complex data generated from collections of different antenna elements with different gains, collections of various numbers of different antenna elements, or collections of different antenna elements with different antenna element spacing.
[0078] By using the hardware abstraction module 220, the surrounding computational machine learning module 222 can operate in radar system 102 with different constraints affecting available radar modulation schemes, transmission parameters, or hardware architecture types. About Figure 6-1 and Figure 6-2 Further description of hardware abstraction module 220.
[0079] The ambient computation machine learning module 222 analyzes hardware-agnostic data and determines the likelihood of various gestures occurring (e.g., performed by a user). Although described relative to ambient computation, the ambient computation machine learning module 222 can also be trained to recognize gestures in non-ambient computation environments. The ambient computation machine learning module 222 is implemented using a multi-level architecture, which will relate to… Figure 7-1 and Figure 9-3 Further description.
[0080] Some types of machine learning modules are designed and trained to recognize gestures within specific or segmented time intervals, such as time intervals initiated by touch events or corresponding to time intervals when the display is active. However, to support various aspects of ambient computing, ambient computing machine learning modules are designed and trained to recognize gestures across time in a continuous and non-segmented manner.
[0081] The gesture deshake 224 determines whether a user has performed a gesture based on the likelihood (or probability) provided by the surrounding computational machine learning module 222. The gesture deshake 224 is about... Figure 7-1 and 7-2 Further description. Radar system 102 is about Figure 3-1 Further description.
[0082] Figure 3-1 The illustration shows an example operation of radar system 102. In the depicted configuration, radar system 102 is implemented as a frequency-modulated continuous wave radar. However, other types of radar architectures can be implemented, as described above. Figure 2 As described above. In environment 300, object 302 is located at a specific tilt range 304 from radar system 102 and is manipulated by a user to perform gestures. Object 302 may be a user's appendage (e.g., hand, finger, or arm), an item worn by the user, or an item held by the user (e.g., stylus).
[0083] To detect object 302, radar system 102 transmits radar signal 306. In some cases, radar system 102 may use a generally wide radiation pattern to transmit radar signal 306. For example, the main lobe of the radiation pattern may have a beamwidth of approximately 90 degrees or greater (e.g., approximately 110, 130, or 150 degrees). This wide radiation pattern allows the user more flexibility in performing gestures for ambient calculations. In an example embodiment, the center frequency of radar signal 306 may be approximately 60 GHz, and the bandwidth of radar signal 306 may be between approximately 4 and 6 GHz (e.g., approximately 4.5 or 5.5 GHz). The term "approximately" may mean that the bandwidth may be within + / - 10% or less of a specified value (e.g., within + / - 5%, + / - 3%, or + / - 2% of a specified value).
[0084] At least a portion of the radar transmitted signal 306 is reflected by the object 302. This reflected portion represents the radar received signal 308. The radar system 102 receives the radar received signal 308 and processes it to extract data for gesture recognition. As depicted, the amplitude of the radar received signal 308 is smaller than the amplitude of the radar transmitted signal 306 due to losses incurred during propagation and reflection.
[0085] The radar transmission signal 306 includes a sequence of chirps 310-1 to 310-N, where N represents a positive integer greater than one. The radar system 102 can transmit chirps 310-1 to 310-N in continuous bursts or as time-separated pulses, as per [reference needed]. Figure 3-2Furthermore, for example, the duration of each chirp 310-1 to 310-N can be on the order of tens or thousands of microseconds (e.g., between approximately 30 microseconds (μs) and 5 milliseconds (ms)). An example pulse repetition frequency (PRF) of radar system 102 can be greater than 1500 Hz, such as approximately 2000 Hz or 3000 Hz. The term "approximately" can mean that the pulse repetition frequency can be within + / - 10% or less of a specified value (e.g., within + / - 5%, + / - 3%, or + / - 2% of a specified value).
[0086] Individual frequencies of chirps 310-1 to 310-N can increase or decrease over time. In the depicted example, radar system 102 employs a dual-slope cycle (e.g., triangular frequency modulation) to linearly increase and linearly decrease the frequencies of chirps 310-1 to 310-N over time. This dual-slope cycle enables radar system 102 to measure Doppler frequency shifts caused by the motion of object 302. Typically, the transmission characteristics of chirps 310-1 to 310-N (e.g., bandwidth, center frequency, duration, and transmit power) can be tailored to achieve specific detection ranges, range resolutions, or Doppler sensitivities to detect one or more characteristics of object 302. The term "chirp" generally refers to a segment or portion of a radar signal. For pulse Doppler radar, "chirp" refers to a single pulse of the pulse radar signal. For continuous wave radar, "chirp" refers to a segment of the continuous wave radar signal.
[0087] At radar system 102, radar received signal 308 represents a delayed version of radar transmitted signal 306. The amount of delay is proportional to the tilt range 304 (e.g., distance) from antenna array 212 of radar system 102 to object 302. Specifically, this delay represents the sum of the time it takes for radar transmitted signal 306 to propagate from radar system 102 to object 302 and the time it takes for radar received signal 308 to propagate from object 302 to radar system 102. If object 302 is moving, radar received signal 308 is shifted in frequency relative to radar transmitted signal 306 due to the Doppler effect. The frequency difference between radar transmitted signal 306 and radar received signal 308 can be referred to as beat frequency 312. The value of beat frequency is based on tilt range 304 and Doppler frequency. Similar to radar transmitted signal 306, radar received signal 308 consists of one or more chirps 310-1 to 310N. Multiple chirps 310-1 to 310-N enable radar system 102 to perform multiple observations of object 302 within a predetermined time period. The radar framing structure determines the timing of chirps 310-1 to 310-N, as shown in reference... Figure 3-2 Further details are provided.
[0088] Figure 3-2The illustration shows an example radar framing structure 314 used for peripheral calculations. In the depicted configuration, radar framing structure 314 includes three different types of frames. At the highest level, radar framing structure 314 includes a sequence of gesture frames 316 (or main frames), which can be active or inactive. Generally, the active state consumes more power than the inactive state. At intermediate levels, radar framing structure 314 includes a sequence of feature frames 318, which can similarly be active or inactive. Different types of feature frames 318 include pulse pattern feature frames 320 (in... Figure 3-2 (shown in the lower left) and burst mode feature frame 322 (in Figure 3-2 (As shown in the lower right). At lower levels, the radar frame structure 314 includes a sequence of radar frames (RF) 324, which may be active or inactive.
[0089] Radar system 102 transmits and receives radar signals during active radar frames 324. In some cases, radar frames 324 are analyzed separately for basic radar operations such as search and track, clutter map generation, user location determination, etc. The radar data collected during each active radar frame 324 can be saved to a buffer after radar frame 324 is completed, or it can be provided directly to the system. Figure 2 The system processor is 216.
[0090] Radar system 102 analyzes radar data across multiple radar frames 324 (e.g., across a set of radar frames 324 associated with activity feature frames 318) to identify specific features. Examples of feature types include one or more stationary objects in the external environment, material properties of these objects (e.g., reflectivity), and physical properties of these objects (e.g., size). To perform gesture recognition during an active gesture frame 316, radar system 102 analyzes radar data associated with the multiple activity feature frames 318.
[0091] The duration of gesture frame 316 can be on the order of milliseconds or seconds (e.g., between approximately 10 milliseconds (ms) and 10 seconds (s)). After the occurrence of active gesture frames 316-1 and 316-2, the radar system 102 is inactive, as indicated by inactive gesture frames 316-3 and 316-4. The duration of inactive gesture frames 316-3 and 316-4 is characterized by a deep sleep time 326, which can be on the order of tens of milliseconds or longer (e.g., greater than 50 ms). In the example embodiment, the radar system 102 shuts down all active components within transceiver 214 (e.g., amplifiers, active filters, voltage-controlled oscillators (VCOs), voltage-controlled buffers, multiplexers, analog-to-digital converters, phase-locked loops (PLLs), or crystal oscillators) to conserve power during the deep sleep time 326.
[0092] For ambient calculations, the deep sleep time 326 can be appropriately set to achieve sufficient responsiveness and reaction time while conserving power. In other words, the deep sleep time 326 can be short enough for the radar system 102 to meet the "always-on" aspect of ambient calculations while also conserving power as much as possible. In some cases, the deep sleep time 326 can be dynamically adjusted based on the amount of activity detected by the radar system 102 or based on whether the radar system 102 determines the presence of a user. If the activity level is relatively high or the user is close enough to the radar system 102 to perform a gesture, the radar system 102 can reduce the deep sleep time 326 to increase responsiveness. Alternatively, if the activity level is relatively low or the user is far enough from the radar system 102 that a gesture cannot be performed within a specified distance interval, the radar system 102 can increase the deep sleep time 326.
[0093] In the depicted radar frame structure 314, each gesture frame 316 includes K feature frames 318, where K is a positive integer. If a gesture frame 316 is inactive, all feature frames 318 associated with that gesture frame 316 are also inactive. Conversely, an active gesture frame 316 includes J active feature frames 318 and KJ inactive feature frames 318, where J is a positive integer less than or equal to K. The number of feature frames 318 can be adjusted based on the complexity of the environment or the complexity of the gesture. For example, a gesture frame 316 may include several to one hundred feature frames 316 or more (e.g., K may be equal to 2, 10, 30, 60, or 100). The duration of each feature frame 318 can be on the order of milliseconds (e.g., between approximately 1 ms and 50 ms). In the example implementation, the duration of each feature frame 318 is between approximately 30 ms and 50 ms.
[0094] To conserve power, active feature frames 318-1 to 318-J precede inactive feature frames 318-(J+1) to 318-K. The duration of inactive feature frames 318-(J+1) to 318-K is characterized by sleep time 328. In this way, inactive feature frames 318-(J+1) to 318-K are executed consecutively, allowing the radar system 102 to remain powered down for a longer period compared to other techniques that might interleave inactive feature frames 318-(J+1) to 318-K with active feature frames 318-1 to 318-J. Generally, increasing the duration of sleep time 328 allows the radar system 102 to shut down components within transceiver 214 that require longer startup times.
[0095] Each feature frame 318 comprises L radar frames 324, where L is a positive integer that may or may not be equal to J or K. In some implementations, the number of radar frames 324 may vary across different feature frames 318 and may include several or hundreds of frames (e.g., L may be equal to 5, 15, 30, 100, or 500). The duration of the radar frames 324 may be on the order of tens or thousands of microseconds (e.g., between approximately 30 μs and 5 ms). The radar frames 324 within a particular feature frame 318 can be customized for a predetermined detection range, range resolution, or Doppler sensitivity, which helps in detecting specific features or gestures. For example, the radar frames 324 may utilize a specific type of modulation, bandwidth, frequency, transmit power, or timing. If a feature frame 318 is inactive, all radar frames 324 associated with that feature frame 318 are also inactive.
[0096] Pulse mode feature frame 320 and burst mode feature frame 322 comprise sequences of different radar frames 324. Generally, radar frames 324 within an active burst mode feature frame 320 transmit pulses that are temporally separated by a predetermined amount. This causes observation to become diffuse over time, which can make the radar system 102 more likely to recognize gestures due to the larger changes in chirp 310-1 to 310-N observed within the pulse mode feature frame 322 relative to the burst mode feature frame 322. Conversely, radar frames 324 within an active burst mode feature frame 322 transmit pulses continuously across a portion of the burst mode feature frame 322 (e.g., the pulses are not separated by a predetermined amount of time). This will result in the active burst mode feature frame 322 consuming less power than the pulse mode feature frame 320 by shutting down a larger number of components, including those with longer start-up times, as further described below.
[0097] Within each active pulse mode feature frame 320, the sequence of radar frames 324 alternates between active and inactive states. Each active radar frame 324 transmits a chirp 310 (e.g., a pulse), illustrated by a triangle. The duration of the chirp 310 is characterized by an active time 330. During the active time 330, components within transceiver 214 are energized. During a short idle time 332, which includes the remaining time within the active radar frame 324 and the duration of the subsequent inactive radar frame 324, radar system 102 conserves power by shutting down one or more active components within transceiver 214 that have an activation time during the duration of the short idle time 332.
[0098] The active burst mode feature frame 322 includes P active radar frames 324 and LP inactive radar frames 324, where P is a positive integer less than or equal to L. To conserve power, the active radar frames 324-1 to 324-P occur before the inactive radar frames 324-(P+1) to 324-L. The duration of the inactive radar frames 324-(P+1) to 324-L is characterized by a long idle time 334. By grouping the inactive radar frames 324-(P+1) to 324-L together, the radar system 102 can be powered down for a longer duration relative to the short idle time 332 occurring during the pulse mode feature frame 320. Additionally, the radar system 102 can disable additional components within the transceiver 214 that have a longer startup time than the short idle time 332 and a shorter startup time than the long idle time 334.
[0099] Each active radar frame 324 within the active burst mode feature frame 322 transmits a portion of chirp 310. In this example, active radar frames 324-1 to 324-P alternate between transmitting portions of chirp 310 with increasing frequency and portions of chirp 310 with decreasing frequency.
[0100] The radar framing structure 314 enables power savings through adjustable duty cycles within each frame type. A first duty cycle 336 is based on the number of active feature frames 318(J) relative to the total number of feature frames 318(K). A second duty cycle 338 is based on the number of active radar frames 324 relative to the total number of radar frames 324(L) (e.g., L / 2 or P). A third duty cycle 340 is based on the duration of chirp 310 relative to the duration of radar frames 324.
[0101] Consider an example radar framing structure 314 for a power state, consuming approximately 2 milliwatts (mW) of power and having a main frame update rate between approximately 1 and 4 hertz (Hz). In this example, the radar framing structure 314 includes a gesture frame 316 with a duration between approximately 250 ms and 1 second. The gesture frame 316 comprises thirty-one pulse pattern feature frames 320 (e.g., K equals 31). One of the thirty-one pulse pattern feature frames 320 is active. This results in a duty cycle 336 of approximately 3.2%. The duration of each pulse pattern feature frame 320 is between approximately 8 and 32 ms. Each pulse pattern feature frame 320 consists of eight radar frames 324 (e.g., L equals 8). Within the active pulse pattern feature frame 320, all eight radar frames 324 are active. This results in a duty cycle 338 of 100%. The duration of each radar frame 324 is between approximately 1 and 4 ms. The active time 330 within each active radar frame 324 is between approximately 32 and 128 μs. Thus, the resulting duty cycle 340 is approximately 3.2%. This example radar frame structure 314 has been found to produce good performance results, while also yielding good power efficiency results in low-power handheld smartphone applications. Moreover, this performance allows the radar system 102 to meet power consumption and size constraints associated with ambient computing while maintaining responsiveness. This power saving allows the radar system 102 to continuously transmit and receive radar signals for ambient computing for at least one hour in power-constrained devices. In some cases, the radar system 102 can operate for periods on the order of tens of hours or even days.
[0102] Although two-slope cyclic signals (e.g., triangular frequency modulated signals) in Figure 3-1 and 3-2 As explicitly shown, these techniques can be applied to other types of signals, including those related to... Figure 2 The signals mentioned. Figure 3-1 The generation of radar transmission signal 306 and ( Figure 3-1 The processing of radar received signal 308 is about Figure 4 Further description.
[0103] Figure 4The illustration shows an example antenna array 212 and an example transceiver 214 of radar system 102. In the depicted configuration, transceiver 214 includes a transmitter 402 and a receiver 404. Transmitter 402 includes at least one voltage-controlled oscillator 406 and at least one power amplifier 408. Receiver 404 includes at least two receive channels 410-1 to 410-M, where M is a positive integer greater than 1. Each receive channel 410-1 to 410-M includes at least one low-noise amplifier 412, at least one mixer 414, at least one filter 416, and at least one analog-to-digital converter 418. Antenna array 212 includes at least one transmit antenna element 420 and at least two receive antenna elements 422-1 to 422-M. Transmit antenna element 420 is coupled to transmitter 402. Receive antenna elements 422-1 to 422-M are coupled to receive channels 410-1 to 410-M, respectively.
[0104] During transmission, voltage-controlled oscillator 406 generates a frequency-modulated radar signal 424 at radio frequency. Power amplifier 408 amplifies the frequency-modulated radar signal 424 for transmission via transmitting antenna element 420. The transmitted frequency-modulated radar signal 424 is represented by radar transmit signal 306, which may include data based on... Figure 3-2 The radar achieves a frame structure 314 with multiple chirps 310-1 to 310-N. As an example, according to... Figure 3-2 The burst mode feature frame 322 generates radar transmission signal 306, and radar transmission signal 306 includes 16 chirps 310 (e.g., N equals 16).
[0105] During reception, each receiving antenna element 422-1 to 422-M receives a version of the radar received signal 308-1 to 308-M. Typically, the relative phase difference between these versions of the radar received signal 308-1 to 308-M is due to the positional differences of the receiving antenna elements 422-1 to 422-M. Within each receiving channel 410-1 to 410-M, a low-noise amplifier 412 amplifies the radar received signal 308, and a mixer 414 mixes the amplified radar received signal 308 with a frequency-modulated radar signal 424. Specifically, the mixer performs a beat operation, down-converting and demodulating the radar received signal 308 to generate a beat signal 426.
[0106] The frequency of the beat signal 426 (i.e., the beat frequency 312) represents the frequency difference between the frequency-modulated radar signal 424 and the radar received signal 308, and this frequency difference is related to... Figure 3-1The tilt range 304 is proportional. Although not shown, the beat signal 426 may include multiple frequencies representing reflections from different objects or portions of objects within the external environment. In some cases, these different objects move relative to the radar system 102 at different speeds, in different directions, or are positioned at different tilt ranges.
[0107] Filter 416 filters the beat signal 426, and analog-to-digital converter 418 digitizes the filtered beat signal 426. Receive channels 410-1 to 410-M generate digital beat signals 428-1 to 428-M, which are provided to system processor 216 for processing. Receive channels 410-1 to 410-M of transceiver 214 are coupled to system processor 216, such as... Figure 5 As shown in the image.
[0108] Figure 5 An example scheme for ambient computation implemented by radar system 102 is illustrated. In the depicted configuration, system processor 216 implements hardware abstraction module 220, ambient computation machine learning module 222, and gesture de-jitter 224. System processor 216 is connected to receive channels 410-1 to 410-M and can also interact with ( Figure 2 The computer processor 202 communicates. Although not shown, the hardware abstraction module 220, the surrounding computational machine learning module 222, and / or the gesture dejitter 224 may be alternatively implemented by the computer processor 202.
[0109] In this example, hardware abstraction module 220 receives digital beat signals 428-1 to 428-M from receive channels 410-1 to 410-M. The digital beat signals 428-1 to 428-M represent raw or unprocessed complex data. Hardware abstraction module 220 performs one or more operations to generate complex data 502-1 to 502-M based on the digital beat signals 428-1 to 428-M. Hardware abstraction module 220 transforms the complex data provided by the digital beat signals 428-1 to 428-M into a form expected by the surrounding computational machine learning module 222. In some cases, hardware abstraction module 220 normalizes the amplitudes associated with different transmit power levels or transforms the complex data into a frequency domain representation.
[0110] Complex radar data 502-1 to 502-M include amplitude and phase information (e.g., in-phase and quadrature components or real and imaginary components). In some embodiments, complex radar data 502-1 to 502-M represent range Doppler maps for each receive channel 410-1 to 410-M and for each active feature frame 318, as per [reference to...]. Figure 6-2Further described. The range-Doppler map includes implicit angle information instead of explicit one. In other embodiments, the complex radar data 502-1 to 502-M includes explicit angle information. For example, the hardware abstraction module 220 can perform digital beamforming to explicitly provide angle information, such as in the form of a four-dimensional range-Doppler-azimuth-elevation map.
[0111] Other forms of complex radar data 502-1 to 502-M are also possible. For example, complex radar data 502-1 to 502-M may include complex interferometric data for each receive channel 410-1 to 410-M. The complex interferometric data is an orthogonal representation of the range Doppler map. In yet another example, complex radar data 502-1 to 502-M includes a frequency domain representation of digital beat signals 428-1 to 428-M for the active feature frame 318. Although not shown, other embodiments of radar system 102 may provide the digital beat signals 428-1 to 428-M directly to the surrounding computational machine learning module 222. Generally, complex radar data 502-1 to 502-M includes at least Doppler information and spatial information for one or more dimensions (e.g., range, azimuth, or elevation).
[0112] Sometimes, complex radar data 502 may include a combination of any of the examples above. For instance, complex radar data 502 may include amplitude information associated with range Doppler images and complex interferometric data. Generally, the gesture recognition performance of radar system 102 can be improved if complex radar data 502-1 to 502-M include implicit or explicit information about the angular position of object 302. This implicit or explicit information may include phase information within the range Doppler image, angular information determined using beamforming techniques, and / or complex interferometric data.
[0113] The surrounding computational machine learning module 222 can perform classification, wherein the surrounding computational machine learning module 222 provides a numerical value for each of one or more categories, describing the degree to which it believes the input data should be classified into the corresponding class. In some instances, the numerical value provided by the surrounding computational machine learning module 222 may be referred to as a "confidence score," which indicates the corresponding confidence level associated with classifying the input into the appropriate class. In some implementations, the confidence score may be compared to one or more thresholds to render discrete class predictions. In some implementations, only a certain number of classes (e.g., one) with the relatively maximum confidence score may be selected to render discrete class predictions.
[0114] In an example implementation, the surrounding computational machine learning module 222 can provide probabilistic classification. For example, the surrounding computational machine learning module 222 is able to predict the probability distribution over a set of classes given a sample input. Therefore, the surrounding computational machine learning module 222 can output the probability that the sample input belongs to such a class for each class, rather than simply outputting the most probable class that the sample input should belong to. In some implementations, the sum of the probability distributions over all possible classes can be one.
[0115] Supervised learning techniques can be used to train the surrounding computational machine learning module 222. For example, a machine learning model can be trained on a training dataset that includes training examples labeled as belonging to (or not belonging to) one or more classes. Regarding... Figures 10-1 to 12 Further details regarding supervised training techniques are provided.
[0116] like Figure 5 As shown, the ambient computation machine learning module 222 analyzes complex radar data 502-1 to 502-M and generates probabilities 504. Some of the probabilities 504 are associated with various gestures that the radar system 102 can recognize. Another of the probabilities 504 may be associated with background tasks (e.g., background noise or gestures not recognized by the radar system 102). The gesture dejitter 224 analyzes the probabilities 504 to determine whether a user has performed a gesture. If the gesture dejitter 224 determines that a gesture has occurred, it notifies the computer processor 202 of an ambient computation event 506. The ambient computation event 506 includes a signal that identifies input associated with the ambient computation. In this example, the signal identifies the recognized gesture and / or passes the gesture control input to the application 206. Based on the ambient computation event 506, the computer processor 202 or the application 206 performs an action associated with the detected gesture or gesture control input. Although gestures have been described, the ambient computation event 506 can be extended to indicate other events, such as the presence of a user within a given distance. Figures 6-1 to 6-2 An example implementation of the hardware abstraction module 220 is further described.
[0117] Figure 6-1An example hardware abstraction module 220 for ambient computation is illustrated. In the depicted configuration, hardware abstraction module 220 includes a preprocessing stage 602 and a signal transformation stage 604. Preprocessing stage 602 operates on each chirp 310-1 to 310-N within the digital beat signals 428-1 to 428-M. In other words, preprocessing stage 602 performs operations on each active radar frame 324. In this example, preprocessing stage 602 includes one-dimensional (1D) Fast Fourier Transform (FFT) modules 606-1 to 606-M, which process the digital beat signals 428-1 to 428-M, respectively. Other types of modules performing similar operations are also possible, such as Fourier transform modules.
[0118] Signal transformation stage 604 operates on the sequence of chirps 310-1 to 310-M within each of the digital beat signals 428-1 to 428-M. In other words, signal transformation stage 604 performs the operation on each active feature frame 318. In this example, signal transformation stage 604 includes buffers 608-1 to 608-M and two-dimensional (2D) FFT modules 610-1 to 610-M.
[0119] During reception, one-dimensional FFT modules 606-1 to 606-M perform individual FFT operations on chirps 310-1 to 310-M within the digital beat signals 428-1 to 428-M. Assuming the radar received signals 308-1 to 308-M comprise 16 chirps 310-1 to 310-N (e.g., N equals 16), each one-dimensional FFT module 606-1 to 606-M performs 16 FFT operations to generate preprocessed complex radar data for each chirp 612-1 to 612-M. When individual operations are performed, buffers 608-1 to 608-M store the results. Once all chirps 310-1 to 310-M associated with the active feature frame 318 have been processed by the preprocessing stage 602, the information stored in the buffers 608-1 to 608-M represents the preprocessed complex radar data for each feature frame 614-1 to 614-M corresponding to the receiving channels 410-1 to 410-M.
[0120] Two-dimensional FFT modules 610-1 to 610-M process the preprocessed complex radar data of each feature frame 614-1 to 614-M to generate complex radar data 502-1 to 502-M. In this case, the complex radar data 502-1 to 502-M represent range Doppler maps, as shown in the reference... Figure 6-2 Further description.
[0121] Figure 6-2The illustration shows example complex radar data 502-1 generated by hardware abstraction module 220 for ambient calculation. Hardware abstraction module 220 is shown as processing digital beat signals 428-1 associated with receive channel 410-1. Digital beat signals 428-1 include chirps 310-1 to 310-M as time-domain signals. Chirps 310-1 to 310-M are passed to one-dimensional FFT module 606-1 in the order they are received and processed by transceiver 214.
[0122] As described above, the one-dimensional FFT module 606-1 performs an FFT operation on the first chirp 310-1 of the digital beat signal 428-1 at the first time. The buffer 608-1 stores the first portion of the preprocessed complex radar data 612-1 associated with the first chirp 310-1. The one-dimensional FFT module 606-1 continues to process subsequent chirs 310-2 to 310-N, and the buffer 608-1 continues to store the corresponding portions of the preprocessed complex radar data 612-1. This process continues until the buffer 608-1 stores the final portion of the preprocessed complex radar data 612-M associated with chirp 310-M.
[0123] At this point, buffer 608-1 stores preprocessed complex radar data associated with a specific feature frame 614-1. The preprocessed complex radar data for each feature frame 614-1 represents amplitude information (not shown) and phase information (not shown) across different chirps 310-1 to 310-N and across different range bins 616-1 to 616-A (or range intervals), where A represents a positive integer.
[0124] The two-dimensional FFT 610-1 accepts preprocessed complex radar data for each feature frame 614-1 and performs a two-dimensional FFT operation to form complex radar data 502-1 representing a range-Doppler map 620. The range-Doppler map 620 includes complex data for range segments 616-1 to 616-A and Doppler bins 618-1 to 618-B (or Doppler frequency intervals), where B represents a positive integer. In other words, each range segment 616-1 to 616-A and Doppler bin 618-1 to 618-B includes a complex number with a real and / or imaginary part representing amplitude and phase information. The number of range segments 616-1 to 616-A can be on the order of tens or hundreds, such as 32, 64, or 128 (e.g., A equals 32, 64, or 128). The number of Doppler segments can be on the order of tens or hundreds, such as 16, 32, 64, or 124 (e.g., B equals 16, 32, 64, or 124). In a first example embodiment, the number of range segments is 64 and the number of Doppler segments is 16. In a second example embodiment, the number of range segments is 128 and the number of Doppler segments is 16. The number of range segments can be reduced based on the expected tilt range of gesture 304. Complex radar data 502-1, and ( Figure 6-1 The complex radar data 502-2 to 502-M is provided to the surrounding computational machine learning module 222, such as... Figure 7-1 As shown in the image.
[0125] Figure 7-1 The illustration shows an example of a surrounding computational machine learning module 222 and a gesture de-jitter 224. The surrounding computational machine learning module 222 has a multi-level architecture, comprising a first level and a second level. In the first level, the surrounding computational machine learning module 222 processes complex radar data 502 across the spatial domain, which involves processing the complex radar data 502 based on feature frames. The first level is represented by a frame model 702. (About...) Figure 8-1 , 8-2 Example implementations of frame model 702 are further described in 9-1 and 9-2.
[0126] In the second level, the surrounding computational machine learning module 222 concatenates a summary of multiple feature frames 318. These multiple feature frames 318 can be associated with gesture frames 316. By concatenating the summaries, the second level processes complex radar data 502 across the temporal domain on a frame-by-frame basis. In some embodiments, the gesture frames 316 overlap temporally, such that consecutive gesture frames 316 can share at least one feature frame 318. In other embodiments, the gesture frames 316 are distinct and do not overlap temporally. In this case, each gesture frame 316 comprises a unique set of feature frames 318. In an example, the gesture frames 316 have the same size or duration. To explain further, the gesture frames 316 can be associated with the same number of feature frames 318. The second level is represented by a temporal model 704.
[0127] By utilizing a multi-level architecture, the overall size and inference time of the surrounding computation machine learning module 222 can be significantly smaller compared to other types of machine learning modules. This enables the surrounding computation machine learning module 222 to run on intelligent devices 104 with limited computing resources.
[0128] During operation, frame model 702 receives complex radar data 502-1 to 502-M from hardware abstraction module 220. As an example, complex radar data 502-1 to 502-M may include a set of complex numbers for feature frame 318. Assuming that complex radar data 502-1 to 502-M represents range Doppler map 620, each complex number may be associated with a specific range segment 616 (e.g., range interval or tilted range interval), Doppler segment 618 (or Doppler frequency interval), and receive channel 410.
[0129] In some implementations, the complex radar data 502 is filtered before being provided to the frame model 702. For example, the complex radar data 502 may be filtered to remove reflections associated with stationary objects (e.g., objects within one or more “center” or “slow” Doppler segments 618). Additionally or alternatively, the complex radar data 502 may be filtered based on a range threshold. For example, if the radar system 102 is designed to identify gestures at distances up to certain distances (e.g., up to approximately 1.5 meters) in a computed surrounding environment, the range threshold may be set to exclude range segments 616 associated with distances greater than the range threshold. This effectively reduces the size of the complex radar data 502 and increases the computational speed of the computed surrounding environment machine learning module 222. In a first example implementation, the filter reduces the number of range segments from 64 to 24. In a second example implementation, the filter reduces the number of range segments from 128 to 64.
[0130] Furthermore, the complex radar data 502 can be reshaped before being provided to the frame model 702. For example, the complex radar data 502 can be reshaped into an input tensor having a first dimension associated with the number of range segments 616, a second dimension associated with the number of Doppler segments 618, and a third dimension associated with the number of receive channels 410 multiplied by two. The number of receive channels 410 is multiplied by two real and imaginary values that take into account the complex radar data 502. In an example implementation, the input tensor can have a dimension of 24x16x6 or 64x16x6.
[0131] Frame model 702 analyzes complex radar data 502 and generates frame summaries 706-1 to 706-J (e.g., one frame summary 706 for each feature frame 318). Frame summaries 706-1 to 706-J are one-dimensional representations of the multidimensional complex radar data 502 for multiple feature frames 318. Temporal model 704 receives frame summaries 706-1 to 706-J associated with gesture frames 316. Temporal model 704 analyzes frame summaries 706-1 to 706-J and generates probabilities 504 associated with one or more classes 708.
[0132] Example class 708 includes at least one gesture class 710 and at least one background class 712. Gesture class 710 represents various gestures that the surrounding computational machine learning module 222 is trained to recognize. These gestures may include swiping right 112, swiping left 114, swiping up 118, swiping down 120, omnidirectional swiping 124, tapping, or some combination thereof. Background class 712 may cover background noise or any other type of motion not associated with gesture class 710, including gestures that the surrounding computational machine learning module 222 has not been trained to recognize.
[0133] In a first embodiment, the surrounding computational machine learning module 222 groups classes 708 based on three predictions. These three predictions may include longitudinal, lateral, or omnidirectional predictions. Each prediction includes two or three classes 708. For example, a longitudinal prediction includes a background class 712 and a gesture class 710 associated with a right swipe 112 and a left swipe 114. A lateral prediction includes a background class 712 and a gesture class 710 associated with an up swipe 118 and a down swipe 120. An omnidirectional prediction includes a background class 712 and a gesture class 710 associated with an omnidirectional swipe 124. In this case, classes 708 are mutually exclusive within each prediction; however, classes 708 between two predictions may not be mutually exclusive. For example, a right swipe 112 in a longitudinal prediction may correspond to a down swipe 120 in a lateral prediction. Furthermore, a left swipe 114 in a longitudinal prediction may correspond to an up swipe 118 in a lateral prediction. Additionally, the gesture class 710 associated with an omnidirectional swipe 124 may correspond to any directional swipe in the other predictions. Within each prediction, the probability of class 708 is 504, totaling 1.
[0134] In the second embodiment, the surrounding computational machine learning module 222 does not group class 708 into various predictions. Thus, class 708 are mutually exclusive, and the total probability 504 is 1. In this example, class 708 includes background class 712 and gesture class 710 associated with swiping right 112, swiping left 114, swiping up 118, swiping down 120, and tapping.
[0135] Gesture dejitter 224 detects ambient computational events 506 by evaluating probability 504. Gesture dejitter 224 enables radar system 102 to identify gestures from a continuous data stream while maintaining a false alarm rate below a false alarm rate threshold. In some cases, gesture dejitter 224 may utilize a first threshold 714 and / or a second threshold 716. The value of the first threshold 714 is determined to ensure that radar system 102 can quickly and accurately identify different gestures performed by different users. The value of the first threshold 714 can be determined experimentally to balance the responsiveness and false alarm rate of radar system 102. Generally, gesture dejitter 224 can determine that a gesture has been performed if the probability 504 associated with a corresponding gesture class 710 is higher than the first threshold 714. In some embodiments, the probability 504 of a gesture must be higher than the first threshold 714 for multiple consecutive feature frames 318—such as two, three, or four consecutive feature frames 318.
[0136] The gesture dejitter 224 can also use a second threshold 716 to keep the false alarm rate below a false alarm rate threshold. Specifically, after a gesture is detected, the gesture dejitter 224 prevents another gesture from being detected until the probability 504 associated with gesture class 710 is less than the second threshold 716. In some implementations, for multiple consecutive feature frames 318, such as two, three, or four consecutive feature frames 318, the probability 504 of gesture class 710 must be less than the second threshold 716. As an example, the second threshold 716 can be set to approximately 0.3%. Regarding... Figure 7-2 A sample operation of gesture deshake 224 is further described.
[0137] Figure 7-2 An example diagram 718 illustrates probabilities 504 across multiple gesture frames 316. For simplicity, three probabilities 504-1, 504-2, and 504-3 are shown in diagram 718. These probabilities 504-1 to 504-3 are associated with different gesture classes 710. Although only three probabilities 504-1 to 504-3 are shown, the operations described below can be applied to other implementations with other numbers of gesture classes 710 and probabilities 504. For simplicity, the probability 504 associated with the background class 712 is not explicitly shown in diagram 718.
[0138] As shown in Table 718, for gesture frames 316-1 and 316-2, probabilities 504-1 to 504-3 are lower than the first threshold 714 and the second threshold 716. For gesture frame 316-3, probability 504-2 is greater than the first threshold 714. Furthermore, probabilities 504-1 and 504-3 are between the first threshold 714 and the second threshold 716 for gesture frame 316-3. For gesture frame 316-4, probabilities 504-1 and 504-2 are higher than the first threshold 714, and probability 504-3 is between the first threshold 714 and the second threshold 716. For gesture frame 316-5, probability 504-1 is greater than the first threshold 714, probability 504-2 is lower than the second threshold 716, and probability 504-3 is between the first threshold 714 and the second threshold 716.
[0139] During operation, gesture deshake 224 can detect a surrounding computed event 506 in response to one of probabilities 504-1 to 504-3 being greater than a first threshold 714. Specifically, gesture deshake 224 identifies the highest probability among probabilities 504. If the highest probability is associated with one of the gesture classes 710 but not the background class 712, gesture deshake 224 detects a surrounding computed event 506 associated with the gesture class 710 having the highest probability. In the case of gesture frame 316-3, gesture deshake 224 can detect a surrounding computed event 506 associated with the gesture class 710 corresponding to a probability 504-2 greater than the first threshold 714. If more than one probability 504 is greater than the first threshold 714, such as in gesture frame 316-4, gesture deshake 224 can detect a surrounding computed event 506 associated with the gesture class 710 corresponding to the highest probability 504, which in this example is probability 504-2.
[0140] To reduce false alarms, the gesture deshake 224 can detect a surrounding computational event 506 in response to a probability 504 greater than a first threshold 714 for multiple consecutive gesture frames 316—such as two consecutive gesture frames 316. In this case, the gesture deshake 224 does not detect a surrounding computational event 506 at gesture frame 316-3 because the probability 504-2 is lower than the first threshold 714 for the previous gesture frame 316-2. However, the gesture deshake 224 detects a surrounding computational event 506 at gesture frame 316-4 because the probability 504-2 is greater than the first threshold 714 for consecutive gesture frames 316-3 and 316-4. Using this logic, the gesture deshake 224 can also detect another surrounding computational event 506 as occurring during gesture frame 316-5 based on a probability 504-1 greater than the first threshold 714 for consecutive gesture frames 316-4 and 316-5.
[0141] After a user performs a gesture, the user may make other movements that could increase the probability 504 of gesture class 710 to the expected level. To reduce the likelihood that these other movements might cause the gesture dejitter 224 to incorrectly detect subsequent surrounding computation events 506, the gesture dejitter 224 may apply additional logic referencing a second threshold 716. Specifically, the gesture dejitter 224 may prevent subsequent surrounding computation events 506 from being detected until the probability 504 associated with gesture class 710 is less than the second threshold 716 of one or more gesture frames 316.
[0142] Using this logic, the gesture deshake 224 can detect the surrounding computed event 506 at gesture frame 316-4 because probabilities 504-1 to 504-3 are less than the second threshold 716 in one or more gesture frames prior to gesture frame 316-4 (e.g., at gesture frames 316-1 and 316-2). However, because the gesture deshake 224 detects the surrounding computed event 506 at gesture frame 316-4, it does not detect another surrounding computed event 506 at gesture frame 316-5, even though probability 504-1 is greater than the first threshold 714. This is because after detecting the surrounding computed event 506 at gesture frame 316-4, probabilities 504-1 to 504-3 have no opportunity to decrease to below the second threshold 716 for one or more gesture frames 316.
[0143] about Figures 8-1 to 8-3 A first example implementation of the ambient computation machine learning module 222 is described. This ambient computation machine learning module 222 is designed to recognize... Figure 1-2 Directional sliding and omnidirectional sliding 124. About Figures 9-1 to 9-3 A second example implementation of the ambient computation machine learning module 222 is described. This ambient computation machine learning module 222 is designed to recognize directional swipe and tap gestures. Additionally, with Figures 8-1 to 8-3 Compared to the surrounding computational machine learning module 222, Figures 9-1 to 9-3 The surrounding computational machine learning module 222 enables the recognition of gestures at greater distances.
[0144] Figure 8-1 and 8-2 The illustration shows an example frame model 702 used for surrounding computation. Typically, frame model 702 includes convolutional, pooling, and activation layers utilizing residual blocks. Figure 8-1 In the illustrated configuration, frame model 702 includes an average pooling layer 802, a split 804, and a first residual block 806-1. The average pooling layer 802 accepts an input tensor 800 including complex radar data 502. As an example, the input tensor 800 may have dimensions of 24x16x6. The average pooling layer 802 performs downsampling, which reduces the size of the input tensor 800. By reducing the size of the input tensor 800, the average pooling layer 802 can reduce the computational cost of the surrounding machine learning module 222. The split 804 splits the input tensor along the range dimension.
[0145] The first residual block 806-1 performs calculations similar to interferometry. In an example implementation, the first residual block 806-1 can be implemented as a 1x1 residual block. The first residual block 806-1 includes a main path 808 and a bypass path 810. The main path 808 includes a first block 812-1, which includes a first convolutional layer 814-1, a first batch normalization layer 816-1, and a first rectifier layer 818-1 (e.g., a rectified linear unit (ReLU)). The first convolutional layer 814-1 can be implemented as a 1x1 convolutional layer. Typically, the batch normalization layer is a generalization technique that helps reduce overfitting of the surrounding computational machine learning module 222 to the training data. In this example, the bypass path 810 does not include another layer.
[0146] The main path 808 also includes a second convolutional layer 814-2 and a second batch normalization layer 816-2. The second convolutional layer 814-2 can be similar to the first convolutional layer 814-1 (e.g., it can be a 1x1 convolutional layer). The main path 808 additionally includes a first summarization layer 820-1, which uses summarization to combine the outputs from the main path 808 and the bypass path 810. After the first residual block 806-1, the frame model 702 includes a second rectifier layer 818-2, a linker layer 822, the second residual block 806-2, and a third rectifier layer 818-3. The second residual block 806-2 can have the same structure as the first residual block 806-1 described above. Regarding... Figure 8-2 The structure of frame model 702 is further described.
[0147] exist Figure 8-2 In the illustrated configuration, frame model 702 also includes a residual block 824. Residual block 824 differs from the first residual block 806-1 and the second residual block 806-2. For example, residual block 824 can be implemented as a 3x3 residual block. Along the main path 808, residual block 824 includes a second block 812-2, which has the same structure as the first block 812-1. The main path 808 also includes a wraparound fill layer 826, a depthwise convolutional layer 828, a third batch normalization layer 816-3, a fourth rectifier layer 818-4, the third block 812-3, and a max-pooling layer 830. The depthwise convolutional layer 828 can be implemented as a 3x3 depthwise convolutional layer. The third block 812-3 has the same structure as the first block 812-1 and the second block 812-2. The second block 812-2, the surrounding fill layer 826, the depthwise convolutional layer 828, the third batch normalization layer 816-3, the fourth rectifier layer 818-4, the third block 812-3, and the max pooling layer 830 represent the first block 832-1.
[0148] Along the bypass path 810, the residual block 822 includes a third convolutional layer 814-3 and a rectifier layer 818-4. The third convolutional layer 814-3 can be implemented as a 1x1 convolutional layer. The residual block 822 also includes a second summarization layer 820-2, which uses summarization to combine the outputs of the main path 808 and the bypass path 810.
[0149] Following residual block 824, frame model 702 includes a second block 832-2, which has a structure similar to the first block 832-1. Frame model 702 also includes a planarization layer 834, a first dense layer 836-1, a fifth rectifier layer 818-5, a second dense layer 836-2, and a sixth rectifier layer 818-6. Frame model 702 outputs a frame summary 706 associated with the current feature frame 318 (e.g., one of frame summaries 706-1 to 706-J). In the example implementation, frame summary 706 has a single dimension with 32 values. Frame summary 706 can be stored in memory, such as within system medium 218 or computer-readable medium 204. Multiple frame summaries 706 are stored in memory over time. Time model 704 processes multiple frame summaries 706, as per [the previous sentence, likely a date or time period]. Figure 8-3 Further description.
[0150] Figure 8-3 The illustration shows an example time model 704 used for peripheral calculations. Time model 704 accesses previous frame summaries 706-1 to 706-(J-1) from memory and links these summaries to the current frame summary 706-J, as indicated by link 838. This set of linked frame summaries 840 (e.g., linked frame summaries 706-1 to 706-J) can be associated with the current gesture frame 316. In the example implementation, the number of frame summaries 706 is 12 (e.g., J equals 12).
[0151] The temporal model 704 includes a Long Short-Term Memory (LSTM) layer 842 and three branches 844-1, 844-2, and 844-3. Branches 844-1 to 844-3 include corresponding dense layers 836-3, 836-4, and 836-5, and corresponding softmax layers 846-1, 846-2, and 846-3. The softmax layers 846-1 to 846-3 can be used to compress the sets of actual values associated with possible classes 708 into a set of actual values in the range of 0 to 1, which sum to one. For a class 708 associated with a particular prediction 848, each branch 844 generates a probability 504. For example, the first branch 844-1 generates a probability 504 associated with the longitudinal prediction 848-1. The second branch 844-2 generates a probability 504 associated with the lateral prediction 848-2. The third branch 844-3 generates a probability 504 associated with the omnidirectional prediction 848-3. Typically, the time model 704 can be implemented with any number of branches 844, including one branch, two branches, or eight branches.
[0152] Figure 8-1 and Figure 8-2 The convolutional layer 814 described in the paper can use cyclic padding in the Doppler dimension to compensate for Doppler aliasing. Figure 8-1 and Figure 8-2 The 814 convolutional layers can also use zero padding in the range dimension. Regarding... Figures 9-1 to 9-3 Another implementation of the surrounding computational machine learning module 222 is further described.
[0153] Figure 9-1 The illustration shows another example frame model 702 used for surrounding calculations. (Compared to...) Figure 8-1 and 8-2 Compared to the 702 frame model, Figure 9-1 The frame model 702 employs separable residual blocks 902. Specifically, Figure 9-1 The frame model 702 uses a series of separable residual blocks 902 to process the input tensor, and the max pooling layer 830 generates the frame summary 706.
[0154] like Figure 9-1 As shown, frame model 702 includes an average pooling layer 802 and a separable residual block 902-1 that operates across multiple dimensions. Figure 9-1 The average pooling layer 802 can be compared with Figure 8-1The average pooling layer 802 operates in a similar manner. For example, the average pooling layer 802 accepts an input tensor 800 that includes complex radar data 502. As an example, the input tensor 800 may have dimensions of 64x16x6. The average pooling layer 802 performs downsampling, which reduces the size of the input tensor 800. By reducing the size of the input tensor, the average pooling layer 802 can reduce the computational cost of the surrounding machine learning module 222. The separable residual block 902-1 may include layers that operate using 1x1 filters. Figure 9-2 Further description of the separable residual block 902-1.
[0155] Figure 9-2 The illustration shows an example separable residual block 902 used for peripheral computation. Separable residual block 902 includes a main path 808 and a bypass path 810. Along the main path 808, separable residual block 902 includes a first convolutional layer 908-1, a first batch normalization layer 816-1, a first rectifier layer 818-1, a second convolutional layer 908-2, and a second batch normalization layer 816-2. Convolutional layers 908-1 and 908-2 are implemented as separable two-dimensional convolutional layers 906 (or more generally, separable multi-dimensional convolutional layers). By using separable convolutional layers 906 instead of standard convolutional layers in separable residual block 902, the computational cost of separable residual block 902 can be significantly reduced at the cost of a relatively small decrease in accuracy performance.
[0156] The separable residual block 902 also includes a third convolutional layer 908-3 along the bypass path 810, which can be implemented as a standard two-dimensional convolutional layer. The summing layer 820 of the separable residual block 902 uses summarization to combine the outputs of the main path 808 and the bypass path 810. The separable residual block 902 also includes a second rectifier layer 818-2.
[0157] return Figure 9-1 Frame model 702 includes a series of blocks 904, implemented using a second separable residual block 902-2 and a first max-pooling layer 830-1. The second separable residual block 902-2 may have a similar structure to the first separable residual block 902-1 and uses a layer operating with a 3x3 filter. The first max-pooling layer 830-1 may perform operations using a 2x2 filter. In this example, frame model 702 includes three cascaded blocks 904-1, 904-2, and 904-3. Frame model 702 also includes a third separable residual block 902-3, which may have a layer operating with a 3x3 filter.
[0158] Additionally, frame model 702 includes a first separable two-dimensional convolutional layer 906-1, which can operate using a 2x4 filter. The first separable two-dimensional convolutional layer 906-1 compresses complex radar data 502 across multiple receive channels 410. Frame model 702 also includes a flattening layer 846.
[0159] The output of frame model 702 is frame summary 706. In the example implementation, frame summary 706 has a single dimension with 36 values. Frame summary 706 can be stored in memory, such as in system medium 218 or computer-readable medium 204. Over time, multiple frame summaries 706 are stored in memory. Time model 704 processes multiple frame summaries 706, as per [the example description]. Figure 9-3 Further description.
[0160] Figure 9-3 The illustration shows an example time model 704 used for peripheral calculations. Time model 704 accesses previous frame summaries 706-1 to 706-(J-1) from memory and links these summaries to the current frame summary 706-J, as indicated by link 838. This set of linked frame summaries 840 (e.g., linked frame summaries 706-1 to 706-J) can be associated with the current gesture frame 316. In the example implementation, the number of frame summaries 706 is 30 (e.g., J equals 30).
[0161] The temporal model 704 includes a series of blocks 914, which include residual blocks 912 and a second max-pooling layer 830-2. The residual block 912 can be implemented as a one-dimensional residual block, and the second max-pooling layer 830-2 can be implemented as a one-dimensional max-pooling layer. In this example, the temporal model 704 includes blocks 914-1, 914-2, and 914-3. (Regarding...) Figure 9-2 Further description of residual block 912.
[0162] like Figure 9-2 As seen in the diagram, residual block 912 can have a similar structure to separable residual block 902. However, there are some differences between separable residual block 902 and residual block 912. For example, separable residual block 902 uses multidimensional convolutional layers, and some of these convolutional layers are separable convolutional layers. Conversely, residual block 912 uses one-dimensional convolutional layers, and these convolutional layers are standard convolutional layers instead of separable convolutional layers. In the context of residual block 912, the first, second, and third convolutional layers 908-3 are implemented as one-dimensional convolutional layers 910. In an alternative implementation, blocks 914-1 to 914-3 are replaced with long short-term memory (LSM) layers. LSM layers can improve performance at the cost of increased computational cost.
[0163] return Figure 9-3The temporal model 704 also includes a first dense layer 828-1 and a softmax layer 846-1. The temporal model 704 outputs a probability 504 associated with class 708.
[0164] Training machine learning modules for radar-based gesture detection in ambient computing environments
[0165] Figure 10-1 The illustrated examples are environments 1000-1 to 1000-4, where data can be collected to train a machine learning module for radar-based gesture detection in the surrounding computing environment. In the depicted environments 1000-1 to 1000-4, recording device 1002 includes radar system 102. Recording device 1002 and / or radar system 102 are capable of recording data. In some cases, recording device 1002 is implemented as smart device 104.
[0166] In environments 1000-1 and 1000-2, radar system 102 collects positive records 1004 when participant 1006 performs a gesture. In environment 1000-1, participant 1006 performs a rightward swipe 112 using their left hand. In environment 1000-2, participant 1006 performs a leftward swipe 114 using their right hand. Typically, positive records 1004 represent complex radar data 502 recorded by radar system 102 or recording device 1002 during the time period in which participant 1006 performs a gesture associated with gesture class 710.
[0167] Positive records 1004 can be collected using participants 1006 with varying heights and hand preferences (e.g., right-handed, left-handed, or ambidextrous). Furthermore, positive records 1004 can be collected using participants 1006 located at various positions relative to the radar system 102. For example, participant 1006 can perform gestures at various angles relative to the radar system 102, including angles between approximately -45 degrees and 45 degrees. As another example, participant 1006 can perform gestures at various distances from the radar system 102, including distances between approximately 0.3 and 2 meters. Additionally, positive records 1004 can be collected using participants 1006 in various postures (e.g., sitting, standing, or lying down), in different positions of the recording device 1002 (e.g., on a table or in the hands of participant 1006), and in various orientations of the recording device 1002 (e.g., longitudinal orientation, longitudinal orientation of the side 108-3 to the right of the participant, or lateral orientation of the side 108-3 to the left of the participant).
[0168] In environments 1000-3 and 1000-4, radar system 102 collects negative records 1008 while participant 1006 performs a background task. In environment 1000-3, participant 1006 operates a computer. In environment 1000-4, participant 1006 moves a cup around recording device 1002. Typically, negative record 1008 represents complex radar data 502 recorded by radar system 102 during the time period when participant 1006 performs a background task associated with background class 712 (or a task not associated with gesture class 710).
[0169] In environments 1000-3 and 1000-4, participant 1006 may perform background movements similar to gestures associated with one or more gesture classes 710. For example, participant 1006 in environment 1000-3 may move their hand between the computer and the mouse, which could be similar to a directional swipe gesture. As another example, participant 1006 in environment 1000-4 may place a cup face down on a table next to recording device 1002 and pick it up again, which could be similar to a tapping gesture. By capturing these gesture-like background movements in negative recording 1008, radar system 102 can be trained to distinguish between background tasks with gesture-like movements and intentional gestures intended to control smart device 104.
[0170] Negative records 1008 can be collected in various environments, including kitchens, bedrooms, or living rooms. Typically, negative records 1008 capture natural behaviors around recording device 1002, which may include a participant 1006 reaching for recording device 1002 while it is in the holder's possession, dancing nearby, walking, cleaning a table with recording device 1002 on a table, or turning a car's steering wheel. Negative records 1008 may also capture repetitive hand movements similar to swiping gestures, such as moving an object from one side of recording device 1002 to the other. For training purposes, negative records 1008 are assigned background labels that distinguish them from positive records 1004. To further improve the performance of the surrounding computation machine learning module 222, negative records 1008 may optionally be filtered to extract samples associated with movements having a velocity above a predefined threshold.
[0171] Positive record 1004 and negative record 1008 are split or partitioned to form training, development, and testing datasets. The ratio of positive record 1004 to positive record 1008 in each dataset can be determined to maximize performance. In the example training process, the ratio is 1:6 or 1:8. About Figure 12 The training and evaluation of the surrounding computational machine learning module 222 are further described. Regarding... Figure 10-2 Further description confirms the capture of record 1004.
[0172] Figure 10-2 The diagram illustrates an example flowchart 1010 for collecting positive records 1004 to train a machine learning module to perform radar-based gesture detection in the surrounding computing environment. At 1012, the recording device 1002 displays an animation illustrating a gesture (e.g., one of the gestures associated with gesture class 710). For example, the recording device 1002 may show the participant 1006 an animation of a specific swipe or tap gesture.
[0173] At 1014, recording device 1002 prompts participant 1006 to perform the illustrated gesture. In some cases, recording device 1002 displays a notification or plays an audible tone to prompt participant 1006. Participant 1006 performs the gesture after receiving the prompt. Additionally, radar system 102 records complex radar data 502 to generate a positive record 1004.
[0174] At 1016, recording device 1002 receives a notification from the proctor monitoring the data collection process. This notification indicates the completion of a gesture segment. At 1008, recording device 1002 marks a portion of positive recording 1004 that occurs between the time the participant is prompted at 1004 and the time the notification is received at 1006 as a gesture segment. At 1020, recording device 1002 assigns a gesture label to the gesture segment. The gesture label indicates the gesture class 710 associated with the animation displayed at 1012.
[0175] At 1022, recording device 1002 (or another device) preprocesses positive record 1004 to remove gesture fragments associated with invalid gestures. These gesture fragments can be removed if their duration is longer or shorter than expected. This may happen if participant 1006 performs gestures too slowly. At 1024, recording device 1002 (or another device) splits positive record 1004 into training, development, and test datasets.
[0176] The affirmative record 1004 may include a delay between when the recording device 1002 prompts the participant 1006 at 1014 and when the participant 1006 begins to perform a gesture. Additionally, the affirmative record 1004 may include a delay between when the participant 1006 completes to perform the gesture and when the recording device 1002 receives the notification at 1016. To refine the timing of the gesture segments within the affirmative record 1004, additional operations may be performed, such as regarding... Figure 10-3 Further description.
[0177] Figure 10-3The illustration shows an example flowchart 1026 for refining the timing of gesture segments within the affirmative recording 1004. At 1028, the recording device 1002 detects the center of the gesture movement within the gesture segment of the affirmative recording 1004. As an example, the recording device 1002 detects zero Doppler crossover within a given gesture segment. Zero Doppler crossover can refer to a time instance where the movement of the gesture changes between positive and negative Doppler segments 618. In other words, zero Doppler crossover can refer to a time instance where the Doppler-determined rate of change of distance changes between positive and negative values. This indicates the time when the direction of the gesture movement becomes substantially perpendicular to the radar system 102, such as during a swipe gesture. It can also indicate the time when the direction of the gesture movement is reversed and the gesture movement becomes substantially stationary, such as at the middle position 140 of a tapping gesture. Figure 1-3 As shown. Other indicators can be used to detect the center point of other types of gestures.
[0178] At 1030, the recording device 1002 aligns a timing window based on the center of the detected gesture motion. The timing window may have a specific duration. This duration may be associated with a specific number of feature frames 318—such as 12 or 30 feature frames 318. Typically, the number of feature frames 318 is sufficient to capture the gesture associated with gesture class 710. In some cases, an additional offset is included within the timing window. The offset may be associated with the duration of one or more feature frames 318. The center of the timing window may be aligned with the center of the detected gesture motion.
[0179] At 1032, recording device 1002 adjusts the size of the gesture segment based on a timing window to generate pre-segmented data. For example, the size of the gesture segment is reduced to include samples associated with the alignment timing window. The pre-segmented data can be provided as part of a training dataset, a development dataset, and a test dataset. Positive records 1004 and / or negative records 1008 can be augmented to further enhance the training of the surrounding computational machine learning module 222, as per [reference to...]. Figure 11 Further description.
[0180] Figure 11 The illustration shows an example flowchart 1100 for enhancing positive records 1004 and / or negative records 1008. Using data augmentation, the training of the surrounding computational machine learning module 222 can be more generalized. Specifically, it can mitigate the effects of potential biases within positive records 1004 and negative records 1008 specific to recording device 1002 or the radar system 102 associated with recording device 1002. In this way, data augmentation enables the surrounding computational machine learning module 222 to ignore certain kinds of noise inherent in the recorded data.
[0181] Example biases can exist within the amplitude information of complex radar data 502. For instance, the amplitude information can depend on variations in the manufacturing process of the antenna array 212 of the radar system 102. Moreover, the amplitude information can be biased by the signal reflectivity of scattering surfaces and the orientation of these surfaces.
[0182] Another example of bias can exist within the phase information of complex radar data 502. Two types of phase information are associated with complex radar data 502. The first type of phase information is absolute phase. The absolute phase of complex radar data 502 can depend on surface location, phase noise, and errors in sampling timing. The second type of phase information includes the relative phase across the different receive channels 410 of complex radar data 502. The relative phase can correspond to the angle of the scattering surface around radar system 102. Typically, it is expected that the surrounding computational machine learning module 222 will be trained to evaluate relative phase rather than absolute phase. However, this can be challenging because the absolute phase may be biased, and the surrounding computational machine learning module 222 can identify and rely on this bias to make correct predictions.
[0183] To improve the positive record 1004 or the negative record 1008 in a way that reduces the impact of these biases on the training of the surrounding computational machine learning module 222, additional training data is generated by augmenting the recorded data. At 1102, for example, magnitude scaling is used to augment the recorded data. Specifically, the magnitudes of the positive record 1004 and / or the negative record 1008 are scaled using a scaling factor selected from a normal distribution. In the example, the normal distribution has a mean of 1 and a standard deviation of 0.025.
[0184] At 1104, random phase rotation is additionally or alternatively used to enhance the recorded data. Specifically, random phase rotation is applied to complex radar data 502 within positive record 1004 and / or negative record 1008. The random phase values can be selected from a uniform distribution between -180 degrees and 180 degrees. Typically, data augmentation allows the recorded data to be artificially altered in a cost-effective manner without requiring the collection of additional data.
[0185] Figure 12 The diagram illustrates an example flowchart 1200 for training a machine learning module to perform radar-based gesture detection in an ambient computing environment. At 1202, the ambient computing machine learning module 222 is trained using a training dataset and supervised learning. As described above, the training dataset can include pre-segmented data generated at 1032. This training enables the optimization of the internal parameters of the ambient computing machine learning module 222, including weights and biases.
[0186] At 1204, the hyperparameters of the surrounding computational machine learning module 222 are optimized using the development dataset. As mentioned above, the development dataset may include pre-segmented data generated at 1032. Typically, hyperparameters represent extrinsic parameters that remain unchanged during training at 1202. First-type hyperparameters include parameters associated with the architecture of the surrounding computational machine learning module 222, such as the number of layers or the number of nodes in each layer. Second-type hyperparameters include parameters associated with processing the training data, such as the learning rate or the number of epochs. Hyperparameters can be manually selected or automatically selected using techniques such as grid search, black-box optimization techniques, gradient-based optimization, etc.
[0187] At 1206, the ambient computation machine learning module 222 is evaluated using a test dataset. Specifically, a two-stage evaluation process is performed. The first stage, described at 1208, involves performing a segmented classification task using the ambient computation machine learning module 222 and pre-segmented data within the test dataset. Instead of using a gesture de-jitter 224 to determine the ambient computation event 506, the ambient computation event 506 is determined based on the highest probability provided by the temporal model 704. By performing the segmented classification task, the accuracy, precision, and recall of the ambient computation machine learning module 222 can be evaluated.
[0188] The second stage, described at 1210, includes performing an unsegmented identification task using the surrounding computational machine learning module and gesture de-jitter 224. Instead of using pre-segmented data from the test dataset, continuous time-series data (or a continuous data stream) is used to perform the unsegmented identification task. By performing the unsegmented identification task, the detection rate and / or false positive rate of the surrounding computational machine learning module 222 can be evaluated. Specifically, the unsegmented identification task can be performed using positive records 1004 to evaluate the detection rate, and negative records 1006 can be used to perform the unsegmented identification task to evaluate the false positive rate. The unsegmented identification task utilizes the gesture de-jitter 224, which allows for further tuning of the first threshold 714 and the second threshold 716 to achieve the desired detection rate and the desired false positive rate.
[0189] If the results of the segmented classification task and / or the unsegmented recognition task are unsatisfactory, one or more elements of the surrounding computational machine learning module 222 can be adjusted. These elements may include the overall architecture of the surrounding computational machine learning module 222, adjustments to the training data, and / or adjustments to the hyperparameters. Using these adjustments, the training of the surrounding computational machine learning module 222 can be repeated at 1202.
[0190] Example Method
[0191] Figures 13 to 15Example methods 1300, 1400, and 1500 are depicted for various aspects of using a radar system to perform ambient calculations. Methods 1300, 1400, and 1500 are shown as sets of operations (or actions) performed, but are not necessarily limited to the order or combination of operations shown herein. Furthermore, any one or more operations may be repeated, combined, reorganized, or linked to provide a wide variety of additional and / or alternative methods. In the sections discussed below, reference may be made to environments 100-1 through 100-5 of Figure 1, and... Figure 2 , Figure 4 , Figure 5 The entities detailed in 7-1 are referenced here for illustrative purposes only. The technology is not limited to being performed by one or more entities operating on a single device.
[0192] exist Figure 13 At position 1302, a radar transmission signal comprising multiple frames is transmitted. Each frame in the multiple frames includes multiple chimes. For example, radar system 102 transmits radar transmission signal 306, such as... Figure 3-1 As shown. The radar transmitted signal 306 is associated with multiple feature frames 318, such as... Figure 3-2 As shown. Each feature frame 318 includes multiple chirps 310 depicted within the active radar frame 324. Multiple feature frames 318 may correspond to the same gesture frame 316.
[0193] At 1304, a radar received signal is received, including a version of the radar transmitted signal reflected by the user. For example, radar system 102 receives radar received signal 308, which represents a version of the radar transmitted signal 306 reflected by the user (or more generally, object 302), such as... Figure 3-1 As shown.
[0194] At 1306, complex radar data for each frame in multiple frames is generated based on the radar received signal. For example, the hardware abstraction module 220 of radar system 102 generates complex radar data 502 based on the digital beat frequency signal 428 associated with the radar received signal 308, such as... Figure 6-1 As shown. The hardware abstraction module 220 generates complex radar data 502 for each of the multiple feature frames 318. The complex radar data 502 can represent a range Doppler image 620, such as... Figure 6-2 As shown.
[0195] At point 1308, complex radar data is provided to the machine learning module. For example, hardware abstraction module 220 provides complex radar data 502 to the surrounding computational machine learning module 222, such as... Figure 7-1 As shown.
[0196] At 1310, a frame summary for each of the multiple frames is generated by the first level of the machine learning module based on complex radar data. For example, the frame model 702 of the surrounding computational machine learning module generates a frame summary 706 for each feature frame 318 based on complex radar data 502, as shown below. Figure 7-1 , 8-2 As shown in 9-1.
[0197] At position 1312, multiple frame summaries are linked by the second-level machine learning module to form a set of links for the frame summaries. For example, temporal model 704 links frame summaries 706-1 to 706-J to form a set of links for frame summaries 840, as shown below. Figure 8-3 and 9-3 As shown.
[0198] At 1314, the probabilities associated with multiple gestures are generated by the second level of the machine learning module and based on the set of connections in the frame summary. For example, the temporal model 704 generates probability 504 based on the set of connections in the frame summary 840. Probability 504 is associated with multiple gestures or multiple gesture classes 710. Example gestures include directional swipe, omnidirectional swipe, and tap. One of the probabilities 504 can also be associated with a background task or background class 712.
[0199] At 1316, it is determined that the user has performed one of the multiple gestures based on the probability associated with the gestures. For example, gesture deshake 224 determines that the user has performed one of the multiple gestures based on probability 504 (e.g., detecting an ambient computational event 506). In response to determining that the user has performed a gesture, the smart device 104 may perform an action associated with the determined gesture.
[0200] exist Figure 14 At position 1402, radar system 102 receives radar signals reflected by the user. For example, radar system 102 receives radar signals 308 reflected by the user (or more generally, object 302), such as... Figure 3-1 As shown.
[0201] At position 1404, complex radar data is generated based on the received radar signal. For example, the hardware abstraction module 220 of radar system 102 generates complex radar data 502 based on the digital beat frequency signal 428 associated with the radar received signal 308, such as... Figure 6-1 As shown. The hardware abstraction module 220 generates complex radar data 502 for each of the multiple feature frames 318. The complex radar data 502 can represent a range Doppler image 620, such as... Figure 6-2 As shown.
[0202] At position 1406, a machine learning module is used to process complex radar data. Supervised learning has been used to train the machine learning module to generate probabilities associated with multiple gestures. For example, the surrounding computation machine learning module 222 is used to process complex radar data 502, such as... Figure 5 As shown, supervised learning has been used to train the surrounding computational machine learning module 222 to generate probabilities 504 associated with multiple gestures (e.g., multiple gesture classes 710). Figure 7-1 As shown.
[0203] At position 1408, the gesture with the highest probability among multiple gestures is selected. For example, gesture deshake 224 selects the gesture with the highest probability among multiple gestures with probability 504. Consider targeting... Figure 7-2 The gesture frames 316-1 to 316-5 in the example show an example probability of 504. In this case, the gesture deshake device 224 selects the third probability 504-3 as the highest probability for gesture frame 316-1. For gesture frames 316-3 and 316-4, the gesture deshake device 224 selects the second probability 504-2 as the highest probability. For gesture frame 316-5, the gesture deshake device 224 selects the first probability 504-1 as the highest probability.
[0204] At point 1410, the highest probability is determined to be greater than the first threshold. For example, gesture deshake 224 determines that the highest probability is greater than the first threshold. Consider targeting... Figure 7-2 The gesture frames 316-1 to 316-5 in the example show an probability of 504. In this case, the gesture de-jitter 224 determines that the highest probability (e.g., probabilities 504-3 and 504-2) within gesture frames 316-1 and 316-2 is below a first threshold 714. However, for gesture frames 316-3 and 316-4, the gesture de-jitter 224 determines that the highest probability (e.g., probability 504-2) is greater than the first threshold 714.
[0205] A first threshold 714 can be predetermined to achieve the target responsivity, target detection rate, and / or target false alarm rate of the radar system 102. Generally, increasing the first threshold 714 reduces the false alarm rate of the radar system 102, but may reduce responsivity and detection rate. Similarly, decreasing the first threshold 714 may increase the responsivity and / or detection rate of the radar system 102 at the cost of increasing the false alarm rate. In this way, the first threshold 714 can be selected in a manner that optimizes the responsivity, detection rate, and false alarm rate of the radar system 102.
[0206] At 1412, in response to determining that the highest probability is greater than a first threshold, it is determined that the user has performed one of multiple gestures. For example, gesture deshake 224 determines that the user has performed one of multiple gestures in response to the selection of a gesture and determining that the highest probability 504 is greater than the first threshold 714 (e.g., detecting surrounding computational events 506). In response to determining that the user has performed a gesture, smart device 104 can perform an action associated with the determined gesture.
[0207] Sometimes, the gesture deshake 224 has additional logic for determining whether a user is performing a gesture. This logic may include determining, for more than one consecutive gesture frame 316, that the highest probability is greater than a first threshold 714. Optionally, the gesture deshake 224 may also require that probability 504 is less than a second threshold 716 for one or more consecutive gesture frames preceding the current gesture frame for which the highest probability is selected at 1408.
[0208] exist Figure 15 At point 1502, a two-stage evaluation process is used to evaluate the machine learning module. For example, using information about... Figure 12 The two-stage evaluation process described is used to evaluate the surrounding computational machine learning module 222.
[0209] At point 1504, pre-segmented data and a machine learning module are used to perform a segmented classification task to evaluate the errors associated with the classification of multiple gestures. The pre-segmented data consists of complex radar data with multiple gesture segments. Each of the multiple gesture segments includes a gesture motion. The center of the gesture motion across the multiple gesture segments has the same relative timing alignment within each gesture segment.
[0210] For example, a segmentation task as described at point 1208 is performed using pre-segmented data from the test dataset. The pre-segmented data includes complex radar data 502 with multiple gesture fragments. Figure 10-3 Flowchart 1026 aligns with the center of the gesture movement within each gesture segment. Errors can indicate errors in the gestures performed by the user.
[0211] At point 1506, continuous time-series data, a machine learning module, and a gesture de-jitter are used to perform an unsegmented identification task to evaluate the false positive rate. For example, continuous time-series data is used to perform the unsegmented identification task, such as... Figure 12 As described at point 1210 in the document. Continuous time series data is not pre-segmented.
[0212] At 1508, adjust one or more elements of the machine learning module to reduce errors and false positive rates. For example, the overall architecture of the surrounding computational machine learning module 222, adjustments to the training data, and / or adjustments to hyperparameters can be made to reduce errors and / or false positive rates.
[0213] Example computing system
[0214] Figure 16 The illustration shows various components of an example computing system 1600, which can be implemented as described above. Figure 2 Any type of client, server, and / or computing device described herein is used to implement aspects of training machine learning modules to perform environmental computing.
[0215] The computing system 1600 includes a communication device 1602 that enables wired and / or wireless communication of device data 1604 (e.g., received data, data being received, data scheduled for broadcast, or data packets of data). The communication device 1602 or the computing system 1600 may include one or more radar systems 102. Device data 1604 or other device content may include device configuration settings, media content stored on the device, and / or information associated with the user of the device. Media content stored on the computing system 1600 may include any type of audio, video, and image data. The computing system 1600 includes one or more data inputs 1606 that can receive any type of data, media content, and / or input, such as human speech, user-selectable input (explicit or implicit), messages, music, television media content, recorded video content, and any other type of audio, video, and / or image data received from any content and / or data source.
[0216] The computing system 1600 also includes a communication interface 1608, which can be implemented as a serial and / or parallel interface, a wireless interface, any type of network interface, one or more of a modem, and any other type of communication interface. The communication interface 1608 provides a connection and / or communication link between the computing system 1600 and a communication network through which other electronic, computing, and communication devices exchange data with the computing system 1600.
[0217] The computing system 1600 includes one or more processors 1610 (e.g., any of a microprocessor, controller, etc.) that process various computer-executable instructions to control the operation of the computing system 1600. Alternatively or additionally, the computing system 1600 may be implemented using any or a combination of hardware, firmware, or fixed logic circuitry systems implemented together with processing and control circuitry typically identified at 1612. Although not shown, the computing system 1600 may include a system bus or data transfer system that couples various components within the device. The system bus may include any or a combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of various bus architectures.
[0218] The computing system 1600 also includes a computer-readable medium 1614, such as one or more memory devices that enable persistent and / or non-transitory data storage (i.e., as opposed to signal transmission only), examples of which include random access memory (RAM), non-volatile memory (e.g., any one or more of read-only memory (ROM), flash memory, EPROM, EEPROM, etc.), and disk storage devices. The disk storage device can be implemented as any type of magnetic or optical storage device, such as a hard disk drive, a recordable and / or rewritable optical disc (CD), any type of digital universal disc (DVD), etc. The computing system 1600 may also include a mass storage medium device (storage medium) 1616.
[0219] Computer-readable medium 1614 provides a data storage mechanism for storage device data 1604, as well as various device applications 1618 and any other types of information and / or data related to the operation of computing system 1600. For example, operating system 1620 may be maintained as a computer application and executed on processor 1610 using computer-readable medium 1614. Device application 1618 may include device managers, such as any form of control application, software application, signal processing and control modules, device-specific native code, hardware abstraction layer for a device, and so on.
[0220] Device application 1618 also includes any system components, engines, or managers that enable surrounding computing. In this example, device application 1618 includes... Figure 2 Applications 206, surrounding computation machine learning module 222, and gesture de-shake 224.
[0221] in conclusion
[0222] Although techniques and apparatus for training machine learning modules to perform radar-based gesture detection in an ambient computing environment have been described in feature- and / or method-specific language, it should be understood that the subject matter of the appended claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as exemplary embodiments for training machine learning modules to perform radar-based gesture detection in an ambient computing environment.
[0223] The following are some examples.
[0224] Example 1: A method for training a machine learning module, the method comprising:
[0225] The machine learning module is evaluated using a two-level evaluation process, which includes:
[0226] The machine learning module is used to perform a classification task using segments of pre-segmented data to evaluate the error associated with the classification of multiple gestures. The pre-segmented data includes complex radar data with multiple gesture segments, each of which includes a gesture motion, and the center of the gesture motion across the multiple gesture segments has the same relative timing alignment within each gesture segment.
[0227] The machine learning module is used to perform an unsegmented identification task using continuous time-series data to assess the false alarm rate, the continuous time-series data including other complex radar data; and
[0228] Adjust one or more elements of the machine learning module to reduce the errors and the false positive rate.
[0229] Example 2: The method according to claim 1, wherein the continuous time series data is not segmented in time.
[0230] Example 3: The method according to claim 1 or 2, wherein the continuous time series data includes negative records associated with at least one user’s natural movement or performance of repetitive movements similar to the plurality of gestures.
[0231] Example 4: The method according to any of the preceding claims further comprises:
[0232] The intrinsic parameters of the machine learning module are trained using second pre-segmented data before evaluation; and
[0233] The hyperparameters of the machine learning module are optimized using third pre-segmented data before evaluation.
[0234] Example 5: The method according to any of the preceding claims further comprises:
[0235] Before evaluating the machine learning module, random offsets are applied to the pre-segmented data and the continuous time series data.
[0236] Example 6: The method of claim 5, wherein applying the random offset comprises at least one of the following:
[0237] Apply phase rotation to the complex radar data; or
[0238] Amplitude scaling is applied to the complex radar data.
[0239] Example 7: The method according to any of the preceding claims further comprises:
[0240] Before performing the segmented classification task, the pre-segmented data is generated, and the generation of the pre-segmented data includes:
[0241] Detect the center of the gesture motion within each recorded gesture segment;
[0242] Alignment timing window based on the center of the detected gesture movement; and
[0243] The size of the gesture segment is adjusted based on the timed window to generate the pre-segmented data.
[0244] Example 8: The method of claim 7, wherein detecting the center of the gesture motion comprises detecting zero Doppler crossovers within each gesture segment of the affirmative record.
[0245] Example 9: The method according to any of the preceding claims, wherein the pre-segmented data comprises:
[0246] Positive records associated with at least one user performing the plurality of gestures; and
[0247] Negative records associated with the at least one user’s natural movement or repetitive movements similar to the multiple gestures.
[0248] Example 10: The method of claim 9, wherein the positive record and the negative record comprise the complex radar data recorded by the radar system.
[0249] Example 11: The method of claim 10, wherein the complex radar data represents a complex range Doppler map associated with multiple receiving channels.
[0250] Example 12: The method according to any one of claim 10 or 11, wherein the affirmative record is associated with each of the plurality of gestures performed multiple times by the at least one user at different distances from the radar system.
[0251] Example 13: The method according to any one of claims 10 to 12, wherein the affirmative record is associated with each of the plurality of gestures performed multiple times by the at least one user at different angles relative to the radar system.
[0252] Example 14: The method according to any of the preceding claims, wherein the error represents a situation in which the machine learning module incorrectly classifies a gesture performed by the user or incorrectly classifies a background motion performed by the user as a gesture.
[0253] Example 15: The method according to any of the preceding claims further comprises:
[0254] Perform another unsegmented identification task to evaluate the detection rate; and
[0255] Adjust one or more of the elements to increase the detection rate.
[0256] Example 16: A system including a radar system and a processor, the processor being configured to process complex radar data generated by the radar system according to a machine learning module trained according to any one of the methods of claims 1 to 15.
[0257] Example 18: A computer-readable storage medium including instructions that, in response to execution of a processor, cause a system to execute any one of the methods described in Examples 1 to 15.
[0258] Example 19: A smart device includes a radar system and a processor configured to process complex radar data generated by the radar system according to a machine learning module trained according to any one of the methods of claims 1 to 15.
[0259] Example 20: The smart device according to Example 19, wherein the smart device comprises:
[0260] Smartphone;
[0261] Smartwatch;
[0262] Smart speaker;
[0263] Intelligent thermostat;
[0264] Security camera;
[0265] Game system; or
[0266] Home appliances.
[0267] Example 21: The smart device according to Example 19, wherein the radar system is configured to consume less than 20 milliwatts of power.
[0268] Example 22: The smart device according to Example 19, wherein the radar system is configured to operate using a frequency associated with a millimeter wavelength.
[0269] Example 23: The smart device according to Example 19, wherein the radar system is configured to transmit and receive radar signals for a period of at least one hour.
Claims
1. A method for training a machine learning module, the method comprising: The machine learning module is evaluated using a two-stage evaluation process, the evaluation including: The machine learning module is used to perform a segmented classification task using pre-segmented data to evaluate the error associated with the classification of multiple gestures. The pre-segmented data includes complex radar data with multiple gesture segments, each of which includes a gesture motion, and the center of the gesture motion across the multiple gesture segments has the same relative timing alignment within each gesture segment. The machine learning module is used to perform an unsegmented identification task using continuous time-series data to assess the false alarm rate, the continuous time-series data including other complex radar data; and Adjust one or more elements of the machine learning module to reduce the errors and the false positive rate, wherein the one or more elements include the overall architecture of the machine learning module, training data, and / or hyperparameters.
2. The method according to claim 1, wherein, The continuous time series data is not segmented in time.
3. The method according to claim 1, wherein, The continuous time series data includes negative records associated with at least one user’s natural movement or repetitive movements similar to the multiple gestures.
4. The method according to claim 1, further comprising: The internal parameters of the machine learning module are trained using data from a second pre-segmented dataset before evaluation. as well as The hyperparameters of the machine learning module are optimized using third pre-segmented data before evaluation.
5. The method of claim 1, further comprising: Before evaluating the machine learning module, random offsets are applied to the pre-segmented data and the continuous time series data.
6. The method according to claim 5, wherein, Applying the random offset includes at least one of the following: Applying phase rotation to the complex radar data; and Amplitude scaling is applied to the complex radar data.
7. The method of claim 1, further comprising: Before performing the segmented classification task, the pre-segmented data is generated, and the generation of the pre-segmented data includes: Detect the center of the gesture motion within each recorded gesture segment; Alignment timing window based on the center of the detected gesture movement; and The size of the gesture segment is adjusted based on the timed window to generate the pre-segmented data.
8. The method according to claim 7, wherein, Detecting the center of the gesture motion includes detecting zero Doppler crossovers within each gesture segment of the affirmative record.
9. The method according to claim 1, wherein, The pre-segmented data includes: Positive records associated with at least one user performing the plurality of gestures; and Negative records associated with the at least one user’s natural movement or repetitive movements similar to the multiple gestures.
10. The method according to claim 9, wherein, The positive and negative records include the complex radar data recorded by the radar system.
11. The method according to claim 10, wherein, The complex radar data represents a complex range Doppler map associated with multiple receiving channels.
12. The method according to claim 10, wherein, The affirmative record is associated with each of the plurality of gestures performed multiple times by the at least one user at different distances from the radar system.
13. The method according to claim 10, wherein, The affirmative record is associated with each of the plurality of gestures performed multiple times by the at least one user at different angles relative to the radar system.
14. The method according to claim 1, wherein, The error refers to a situation where the machine learning module incorrectly classifies a gesture performed by the user or incorrectly classifies a background motion performed by the user as a gesture.
15. The method according to any one of claims 1 to 14, further comprising: Perform another unsegmented identification task to evaluate the detection rate; as well as Adjust one or more of the elements to increase the detection rate.
16. A system comprising a radar system and a processor, the processor being configured to process complex radar data generated by the radar system according to a machine learning module trained according to any one of claims 1 to 15.