Gesture control system using biopotential and vector

By using bioelectric potential and position sensors on the wrist to sense user gestures and positions, the interaction problem of traditional user interfaces in environments without displays or where input devices are difficult to manipulate is solved, enabling safe and intuitive machine interaction.

CN115335796BActive Publication Date: 2026-07-31PISON TECHNOLOGY INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PISON TECHNOLOGY INC
Filing Date
2021-01-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the existing technology, graphical user interfaces for machine-to-user interaction are difficult to use in certain environments, especially when there is no monitor or the user cannot easily manipulate the input device, and traditional input devices are not intuitive, feasible, or secure.

Method used

The gesture control system employs bioelectric potential and position sensors. It senses the user's gestures and location by wearing bioelectric potential and position sensors on the wrist, combines them with a processor to determine the geographic location, and provides feedback by interacting with the machine through a response device.

Benefits of technology

It enables safe, intuitive, and convenient machine interaction in challenging environments, allowing users to operate without requiring additional hand movements, thus enhancing the intuitiveness and safety of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115335796B_ABST
    Figure CN115335796B_ABST
Patent Text Reader

Abstract

A device includes a sensor configured to be worn by a person and adapted to sense information indicating a direction indicative of the person's intention. The device includes a processor and memory for instructions executable by the processor. When instructions are executed, the processor determines the direction based on the information sensed by the sensor and determines a geographic location based on the direction and the person's position. The sensor may include bioelectrical signals, based on which the person's intended gesture can be determined.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-citation of related applications

[0002] This application is a national phase entry of International Application No. PCT / US2021 / 014723, filed January 22, 2021, which claims priority to U.S. Patent Application No. 11,157,086, filed June 2, 2020 and now published October 26, 2021, which is a continuation-in-part of U.S. Patent Application Serial No. 16 / 774,825, filed January 28, 2020.

[0003] This application relates to pending U.S. Patent Application Serial No. 16 / 104,273, filed August 17, 2018, which is a continuation of U.S. Patent Application Serial No. 15 / 826,131, now published September 11, 2018, and is a non-provisional application of Provisional Patent Application 62 / 566,674, filed October 7, 2017, and Provisional Patent Application 62 / 429,334, filed December 2, 2016, all of which are incorporated herein by reference. This application also relates to U.S. Patent Application Serial No. 16 / 055,123, filed August 5, 2018; U.S. Patent Application Serial No. 16 / 246,964, filed January 14, 2019; and PCT Patent Application Serial No. PCT / US19 / 061421, filed November 4, 2019. This international application is an international application that designates U.S. Patent Application Serial No. 16 / 196,462, filed November 20, 2018; and U.S. Patent Application Serial No. 16 / 737,252, filed January 8, 2020.

[0004] All of the foregoing is incorporated herein by reference. Technical Field

[0005] This application relates to a system for bioelectric potential-based gesture control. Background Technology

[0006] Most machines have a "user interface" through which people interact with them. People provide input through one or more devices, from which the machine interprets the person's intent. In response to this input, the machine provides feedback to the person, for example, through the machine's actions or through one or more devices that present information to the person.

[0007] When the machine is a computer system with a monitor or includes a computer system with a monitor, a common example of a user interface is the "graphical user interface" (GUI). Using a GUI, a person manipulates a user interface device that provides input to applications running the computer. In turn, the computer provides visual output on the monitor. The computer may provide other outputs, such as audio. Through the user interface device, a person can control the position of a graphical object, often referred to as a cursor, within the monitor and can instruct on actions to be performed. The actions to be performed are often based in part on the position of the graphical object within the information presented on the monitor. Various user interface devices have been created for computers. Similarly, many kinds of graphical user interfaces and other types of user interface paradigms have been used for computers. Summary of the Invention

[0008] This invention provides a simplified overview of some concepts, which will be further described in the detailed embodiments described below. This invention does not designate any features as critical or essential, nor does it limit the scope of the claimed subject matter.

[0009] The graphical user interface paradigm used for interacting with machines or computers may be difficult to use in certain environments. In some environments, the computing device may have a small display or no display at all. Furthermore, in some environments, users may not be able to easily manipulate any input device because they are grasping or holding another object. For various reasons, in some environments, the typical combination of input devices and displays does not provide a sufficiently intuitive, feasible, convenient, or safe form of machine interaction.

[0010] Furthermore, most graphical user interfaces (GUIs) for computers are based on the interaction between the user and graphical objects presented on the monitor, or objects known to the computer and presented on the monitor by the computer. Input from the user is often associated with a location within the display space that the computer knows and controls.

[0011] As described herein, a human-computer interface is provided in which one or more users provide one or more inputs indicating a real-world geographic location, and determine the geographic location based on these inputs. Each of the one or more users has a corresponding user interface device from which it receives corresponding input. A response device receives and processes the input from at least one user interface device. There may be multiple response devices, each receiving input from one or more user interface devices. The response device, or another computing device in communication with the response device, may determine the geographic location based on the received one or more inputs. Feedback regarding the determined geographic location may be provided to at least one user.

[0012] Some exemplary scenarios in which a geographic location can be determined include, but are not limited to, the following scenarios.

[0013] A person might point in a specific direction, such as north or south, and may need information about objects in that general direction. A person might point to a street corner or building, possibly wanting information about that corner or building, or simply to identify it. A person might point in a direction and may want to set a point on a map. A person might point in a direction to identify objects, such as spotting a deer or other animal while hunting, or communicating during a game or other activity.

[0014] Multiple people may point to roughly the same location or the same object. Multiple users may want to cooperate to select or identify the same object. Multiple users may want to cooperate to select or identify multiple objects. One or more objects may be moving. For example, multiple firefighters may point to the same window or other location on a building. Multiple people on a street may point to the same vehicle on the street. For example, several police officers may point to the same vehicle. Multiple people in a location may point to roughly the same location. Multiple people in a city park may point to an object. For example, security personnel in a location may point to the same person, an open door, or an object. Multiple people in a search and rescue mission may identify objects or people or animals to be rescued, or hazards or obstacles.

[0015] The selected geographic location can be used to control communications. For example, a person can point to another person and initiate communication directly with that person. A person can select a group of people, for example, by pointing at them individually or by using a gesture to select them as a group, to communicate with that group. A person can use a gesture (e.g., pointing to the sky) to initiate communication with a predetermined group of individuals. Similarly, communication with or control of an object, or both, can be performed by pointing to a single object, selecting a group of objects within a set of objects, or selecting a predetermined group of objects.

[0016] A geographic location can be defined as an absolute location or a relative location. Examples of absolute location include, but are not limited to, coordinates in a map system, GPS coordinates, or longitude and latitude at various resolutions, or any combination thereof. Examples of relative location include, but are not limited to, reference points and directions, or reference points and directions and distances, multiple reference points and their corresponding directions, or multiple reference points and their corresponding directions and distances, or any combination thereof.

[0017] Geographic location can be determined based on input from sensor data from sensors associated with at least one person. Additional sensor data from sensors associated with one or more other people, or one or more other objects, or a combination thereof, can be used to determine geographic location. Examples of such sensors include at least any one or more of the following: a position sensor (described below), a geomagnetic sensor, a compass, a magnetometer, an accelerometer, a gyroscope, an altimeter, an eye-tracking device, an encoder providing information indicating changes in position, a camera providing image data or motion video, a reference to a known landmark, or a sound sensor (e.g., one or more microphones), or other sensors that can provide indications of direction from a person, or any combination of two or more of these. Sensor data can be interpreted as, for example, any of a person's multiple raw hand gestures, an expected action, an arm angle or other indication of relative vertical position, direction or orientation, distance or height, or any combination thereof. Geographic location can be determined based on such sensor data received from sensors associated with one or more people.

[0018] In some implementations, contextual data can be used to assist in determining geographic location. Such contextual data may include information about one or more objects (e.g., geographic location), speech, sound, or text data, image data from one or more cameras, map data or user input about a map, or any combination of these or other contextual information.

[0019] In some implementations, additional information, such as contextual data, may be used to provide input for selecting, controlling, initiating, or performing additional actions related to or based on a geographic location selected as an object. For example, voice, gestures, or other data may be used as modifiers. For instance, a firefighter pointing to a building may communicate via voice, such as, “Two people are near the second-floor window!” For instance, voice, gestures, or other data may be used to annotate a selected object, with different gestures used to provide different labels for that object, such as “safe” or “dangerous.” For instance, voice, gestures, or other data may be used to activate commands about an object, activating different actions in different ways. This activation may be similar to a “click,” but can be obtained from any other data received by the user.

[0020] Various types of processing techniques can be used to determine geographic location based on sensor data and optional contextual data. In some implementations, various forms of triangulation can be applied using directions from multiple reference points. In some implementations, the input can allow the user to specify distances associated with the determined location and direction. In some implementations, techniques from computational geometry can be applied, such as variations of the Simultaneous Localization and Mapping (SLAM) algorithm for calculating the agent's location, where the geographic location to be determined is considered as the location of the agent to be calculated. In some implementations, various forms of odometry can be used, where changes in location over time are determined based on sensor data. In some implementations, visual odometry can be performed on images from a camera. Any combination of these implementations can also be used.

[0021] The geographic location can be used for a variety of purposes, including but not limited to the following, and may include combinations of two or more of the following: The geographic location can be used as input to an application, for example, to identify or select an object. Various actions can be performed by the application in conjunction with the selected object, non-limiting examples of which include invoking a camera to track the object or retrieving information about the object from a database. The geographic location can be data passed to another user or another machine. When combined with data based on signals received from one or more user interface devices to select an action to be performed, the geographic location can be considered contextual information. The geographic location can be provided to the user as feedback, for example, by displaying a point on a map.

[0022] In some embodiments described herein, a human-machine interface (e.g., an interface for interacting with a computer or computer-controlled device) may be implemented via a combination of one or more user interface devices worn by a person and one or more response devices associated with one or more machines or a person. The machine may include, for example, any type of computing device or computer-controlled device. The person performs, attempts, or intends to perform an action sensed by sensors in the user interface device; the sensors output signals. The response device interprets and uses, or acts on, data received based on the output signals from the sensors in the user interface device, to enable the person to interact with the machine. Feedback is provided to the person. This feedback may be provided by the machine, the response device, or the user interface device, or other feedback devices, or a combination of two or more of these devices. As described in more detail below, one action that may be performed by the response device and for which feedback may be provided includes identifying a geographic location.

[0023] In some embodiments, the user interface device is worn by a person on a body part and includes at least a bioelectrical potential sensor and a position sensor. The body part on which the user interface device is worn is an appendage that the person can typically move relative to their torso. Furthermore, at this appendage, bioelectrical potentials associated with the activity of muscles controlling the body part can be sensed on the skin surface, whereby the muscle activity enables the person to induce or intend to induce certain postures of the body part, as explained herein.

[0024] In an example of a wrist-worn user interface device, the device is configured to be worn by a person such that at least one bioelectric sensor is placed on top of the wrist, as described herein, to sense bioelectric potentials at the top of the wrist. The bioelectric potentials sensed at the top of the wrist are related to the muscle activity that controls the hand and fingers. As described herein, any hand posture is a result of the activation or relaxation of these muscles. A bioelectric potential signal is generated in response to a person's desire to form any hand posture; in some cases, hand postures may not be formed, for example, when the person has a neuromuscular disease. The bioelectric potential sensor has an output that provides a bioelectric potential signal indicating the sensed bioelectric potentials.

[0025] In an example of a wrist-worn user interface device, as described herein, a position sensor senses the position of the wrist relative to a reference point and has an output that provides a position signal indicating the sensed position. Different types of position sensors can use different reference points.

[0026] A response device is a machine, typically a computing device, that receives data from a user interface device based on bioelectrical potential signals and positional signals. The response device performs actions based on the received data. The received data may be output signals from sensors, or data representing information obtained from processing these signals, such as data obtained from filtering or transforming these signals, data indicating features extracted from these signals, gesture or position information obtained from the foregoing, data representing a raw gesture detected from these signals, or other data obtained from processing these signals, or any combination of two or more of these data.

[0027] The response device responds to received data to detect the original gesture in the received data and prompts an action to be performed according to the original gesture. The original gesture may include a first posture of the hand or a first position of the wrist at a first moment, or both; and a second posture of the hand or a second position of the wrist at a second moment after the first moment, or both. The original gesture may include a "hold" gesture, wherein the second posture is the same as the first posture.

[0028] In some cases, the action to be performed is based on a series of detected primitive gestures. In other cases, the action is based on a specific detected primitive gesture. In some cases, the action is based on the co-occurrence of two or more detected primitive gestures. In still others, the action is based on the occurrence of one or more primitive gestures in a specific context.

[0029] Typically, the action to be performed is an operation executed by a computing device. This operation is usually performed by an application running on the computing device. The computing device may run several applications. When interacting with multiple applications through a user interface device, the computing device is instructed to provide input to the selected application, and this input causes the selected application to perform the selected action. Raw gestures detected in the data received from the user interface can be used to select an application, or to select the input to be provided to the application, or both.

[0030] For example, the response device can direct input to a selected application, which can then perform an action in response to a detected hand posture after a period of time during which no wrist movement is detected. Conversely, when wrist movement is detected, the response device can ignore data based on signals from a bioelectric sensor. This behavior ensures that wrist movement does not generate signals that interfere with the detection of hand posture.

[0031] For example, the responding device can select an application based on the detected wrist position. Within the selected application, while the wrist remains in the detected position, the application can select and perform an action based at least on the detected hand posture. For example, this behavior can enable switching between different applications. For instance, a person can raise their wrist above their head and lift their index finger to activate an application and have it perform an action. Subsequently, the person can lower their wrist and lift their index finger to activate another application and have that other application perform a different action.

[0032] For example, the response device can direct input to a selected application, which can then choose an action based on the detected wrist position. When the wrist remains in the detected position, the selected action can be performed in response to the detected hand posture. For instance, this behavior could allow selecting a menu item or choosing a value within a range. For example, an audio playback system could allow selecting a song from a playlist or adjusting the volume.

[0033] For example, the selected application could perform an action in response to a detected hand gesture, which occurs after a period of time when no wrist movement is detected, followed by the detection of a held gesture and the detection of wrist movement in one direction. This behavior could, for example, enable selecting objects by circling around them, selecting the next or previous item in a list, adjusting scrolling or sliding elements, and other actions. For instance, a person could raise their arm, then lift and grasp their fingers, and then make a circular motion to select an object.

[0034] For example, the response device can use detected hand gestures to initiate and terminate operating modes that primarily rely on wrist position for execution. This behavior, for instance, could enable the initiation of operating modes based on wrist movements to control movement on objects such as unmanned aerial vehicles. For example, in response to raising and holding fingers, a remote control mode for a drone could be activated, where wrist position controls the drone's movement.

[0035] It can also operate based on any other information from other sensors on the user interface device. It can also perform operations based on any contextual information available to the response device.

[0036] Such behaviors in responsive devices enable a wide range of user experiences for machine interaction in various environments. When hand gestures for machine interaction can be performed while a person grasps or holds another object (e.g., raising a finger), the user interface device allows the person to interact with the machine while performing other activities. When hand gestures for machine interaction can be performed discreetly, or with minimal movement, such as when a person has their hand in their pocket, the user interface device allows the person to interact with the machine conveniently, safely, or discreetly in challenging environments. When the hand gestures or wrist positions for machine interaction are intuitive (e.g., pointing and swiping), the user interface device allows the person to interact with the machine intuitively.

[0037] Therefore, in one aspect, a device includes a sensor configured to be worn by a person and adapted to sense information indicating the direction of the person's intention. The device includes a processor and a memory for storing instructions executable by the processor. When executing instructions, the processor determines the direction based on the information sensed by the sensor, and determines the geographic location based on the direction and the person's position.

[0038] On the other hand, the response device is configured to communicate with a user interface device. The response device includes a processing system comprising a processing unit and a computer memory. The computer memory has computer program instructions stored thereon, which, when processed by the processing unit, configure the response device. The response device is configured to receive sensor data from the user interface device and determine its geographic location.

[0039] In some embodiments, the device includes a bioelectric potential sensor and a position sensor. The bioelectric potential sensor is configured to be worn by a person and respond to an intention expressed by the person. The bioelectric potential sensor is configured to sense bioelectric potentials present on the person. These bioelectric potentials are associated with muscles controlling a body part or an appendage near the body part. The position sensor is configured to sense the person's position. A processor storing instructions executable by a processor detects a raw gesture representing the person's expressed intention based on the sensed bioelectric potentials and the sensed position. In response to the raw gesture, the processor determines the geographic location.

[0040] In some implementations, the response device is configured to receive data based on bioelectrical potential signals and data based on location signals. The response device is configured to detect a raw gesture indicating the intention expressed by a person. In response to the raw gesture, the processor determines the geographic location.

[0041] Any of the foregoing may include one or more of the following features: The sensor includes a bioelectric sensor configured to be worn by a person and configured to sense bioelectric potentials present on the person. The sensor includes a position sensor configured to be worn by a person and configured to sense the person's position. The sensor includes an eye-tracking device. The sensor includes a geomagnetic sensor. The sensor includes an altimeter. The sensor includes an encoder that provides information indicating changes in position.

[0042] Any of the foregoing may include one or more of the following features: A body part is the wrist. An appendage is the hand. An appendage includes one or more fingers of the hand. A posture is the posture of the hand. A hand posture includes the posture of one or more fingers. A hand posture includes the movement of one or more fingers. A hand posture includes the posture of a single finger. A hand posture includes the movement of a single finger. A single finger is the thumb. A single finger is the index finger. A hand posture includes the position of the hand relative to the wrist. A primitive gesture includes finger sliding. A primitive gesture includes finger lifting. A primitive gesture includes finger grasping. The position of the wrist includes the position of the top of the wrist. The position of the wrist includes the direction of the top of the wrist. The position of the wrist includes the movement of the top of the wrist.

[0043] Any of the foregoing may include one or more of the following features: The position sensor includes an accelerometer that generates a position signal, the position signal including signals indicating the direction and magnitude of acceleration of a body part. The position sensor includes a gyroscope that generates a position signal, the position signal including signals indicating the orientation and angular velocity of a body part. The position sensor includes a geomagnetic sensor that generates a position signal, the position signal including signals indicating direction.

[0044] Any of the foregoing may include one or more of the following features: Determining a geographic location based on a first geographic location of a person, a first direction based on at least sensor data, and a second direction derived from a second geographic location. The second geographic location is that of a second person. The second direction is the direction indicated by the second person. The second geographic location is a location derived from signals from a device carried by this person. The second direction is the direction derived from signals from a device carried by this person.

[0045] Any of the foregoing may include one or more of the following features: Geographic location is determined based on data indicating the person's geographic location, direction, and distance along that direction. The distance along that direction is based on at least one of a bioelectrical signal or a positional signal from the person. The determined geographic location specifies a point on a map. In response to a detected initial gesture, the processor modifies the position of the specified point on the map.

[0046] Any of the foregoing may include one or more of the following features: Determining geographic location based on the selected object. Selecting an object with a geographic location in response to a person's geographic location and in response to a detected initial gesture. The geographic location of the selected object is the determined geographic location. The geographic location of the selected object is based on information about the geographic location of the selected object accessed from a database. The geographic location of the selected object is calculated based on the object's time-of-flight data.

[0047] Any of the foregoing may include one or more of the following features: The position of the person includes the position of the top of the wrist. The position of the person includes the orientation of the top of the wrist. The position of the person includes the movement of the top of the wrist.

[0048] Any of the foregoing may include one or more of the following features. The processor determines a geographic location by receiving data indicating a first geographic location of a person, determining a first direction based on sensor data received from a user interface device worn by the person, determining a second direction from a second geographic location, and determining a geographic location based on the first geographic location and the first direction, as well as the second geographic location and the second direction. The second geographic location includes the geographic location of a second person. The second direction includes the direction indicated by the second person. The second geographic location includes a location derived from signals from a device carried by the second person. The second direction includes a direction derived from signals from a device carried by the second person. The second geographic location includes the location of a second sensor. The second direction includes a direction based on signals from a second sensor.

[0049] Any of the foregoing may include one or more of the following features. The processor determines the geographic location by receiving data indicating the person's geographic location and orientation, determining the distance along a first direction based on sensor data, and determining the geographic location based on the person's geographic location, orientation, and the determined distance. The determined geographic location specifies a point in coordinates associated with a map. In response to a detected initial gesture, the processor modifies the position of the specified point.

[0050] Any of the foregoing may include one or more of the following features. To determine a geographic location, the processor is configured to, in response to the geographic location of a person and in response to sensor data, select objects having geographic locations, and provide the geographic location of the selected objects as the determined geographic location. The geographic location of the selected objects includes accessing information about the geographic location of the selected objects from a database. The geographic location of the selected objects includes calculating the geographic location of the selected objects based on the objects' time-of-flight data.

[0051] Any of the foregoing can be embodied as a computer system, a component of such a computer system, a process executed by such a computer system, or a component of such a computer system, or an article of manufacture including a computer memory storing computer program code therein, wherein, when processed by a processing system of one or more computers, the computer memory configures the processing system of one or more computers to provide such a computer system or a component of such a computer system.

[0052] The following detailed description refers to the accompanying drawings, which form part of this application, and illustrate specific exemplary embodiments by way of illustration. Other embodiments can be implemented without departing from the scope of this disclosure. Attached Figure Description

[0053] Figure 1A This is a block diagram of the user interface device;

[0054] Figures 1B to 1C This is a perspective view of an exemplary user interface device;

[0055] Figures 2A to 2D This is a block diagram of an example implementation of the response device;

[0056] Figure 3 This is a schematic diagram of the signal processing stage from the user interface device;

[0057] Figure 4A An exemplary hand gesture is shown;

[0058] Figure 4B An exemplary location of the top of the wrist is shown;

[0059] Figure 4C This is a state diagram showing the processing of data received from the user interface device;

[0060] Figure 5 This is a data flow diagram illustrating the interaction between the user interface device and the response device to provide input to the application on the response device;

[0061] Figure 6 This is a diagram illustrating how the application selects different actions based on wrist position;

[0062] Figure 7 A state diagram showing how the application performs different operations based on hand gestures or wrist positions, or both.

[0063] Figure 8 This demonstrates how to map wrist position to different operations;

[0064] Figure 9A This is a data flow diagram illustrating the determination of geographical location;

[0065] Figure 9B and Figure 9C An exemplary scenario for determining a geographical location is shown;

[0066] Figure 9D It is a data flow diagram of an application that selects operations to be performed on selected objects within a 3D scene;

[0067] Figure 10 This is an example graphical user interface;

[0068] Figure 11 This is an example graphical user interface;

[0069] Figure 12 This is a block diagram of a general-purpose computer. Detailed Implementation

[0070] As described herein, a human-machine interface (e.g., an interface for interacting with a computer or computer-controlled device) is implemented by a combination of one or more user interface devices worn by a person and one or more response devices associated with one or more machines or a person. The machine may include, for example, any type of computing device or computer-controlled device.

[0071] A person performs, attempts, or intends to perform an action, which is detected by sensors in a user interface device; the sensors output signals. A response device interprets the received data based on the output signals from the sensors in the user interface device to enable the person to interact with the machine. Feedback is provided to this person. This feedback can be provided by a machine, a response device, a user interface device, or other feedback devices, or a combination of two or more of these devices. As described in more detail below, one action that can be performed by a response device and for which feedback can be provided includes identifying a geographic location.

[0072] User interface devices, response devices, and machines have various possible configurations and interrelationships, which will be discussed below. Figures 2A to 2D Some of these are described in more detail. Typically, the user interface device, response device, and machine can be separate devices. Two or more devices can be integrated into the same device, for example: a single device containing both a user interface device and a response device (e.g., a smartwatch controlling another machine); a single device containing both a response device and a machine (e.g., a user interface device interacting with a smartphone, tablet, or computer); and a single device containing a user interface device, a response device, and a machine (e.g., a smartwatch). These devices can be partially integrated into the same device, for example, when functions described herein as residing in a response device are located in a user interface device (as combined below). Figure 3 (For more detailed explanation). Multiple instances of user interface devices, response devices, machines, or combinations thereof may exist. For example, in some cases, multiple user interface devices may be worn by a single person, with the respective user interface devices worn on corresponding parts of the body, such as different wrists or ankles. In other cases, multiple user interface devices may be worn by multiple different persons; for example, a first person wears a first user interface device, a second person wears a second user interface device, and both user interfaces are connected to one or more response devices. Connections between devices can be implemented in various ways, including wired and wireless communications, computer networks, computer bus architectures, and combinations thereof.

[0073] Figure 1A A schematic circuit block diagram of a user interface device 100 worn by a person is shown. The body part on which the user interface device can be worn may be an appendage that a person can typically move relative to their torso, such as the arm above the wrist. Furthermore, at this appendage, bioelectrical potentials associated with the activity or anticipated activity of muscles controlling the body part can be sensed on the skin surface, wherein the muscle activity enables the person to induce or anticipate certain postures of the body part.

[0074] While such user interface devices can be worn on many different parts of a person's body, this paper will use an example of a wearable user interface device that places at least one bioelectric sensor on the top of the wrist to describe the user interface device. The phrase "top of the wrist" as used herein refers to the posterior surface of the arm, located at or near the distal end of the arm, adjacent to the joint that connects the hand to the arm (i.e., the wrist) (hereinafter combined with...). Figures 4A to 4C (To be described in more detail).

[0075] In some embodiments, when the user interface device is worn on the top of the wrist, the user interface device 100 outputs data 120, including at least a signal 102 indicating the position of the wrist relative to a reference point and a signal 104 indicating bioelectrical potentials related to muscle activity of the hand controlling the hand and the fingers connected to the wrist. To generate these signals, the user interface device 100 has at least two sensors: a bioelectrical potential sensor 106 and a position sensor 108. When the user interface is worn on a body part other than the wrist, the user interface device outputs a signal indicating the position of that body part relative to a reference point and a bioelectrical potential sensed at the skin surface of that body part.

[0076] The phrase “position” refers to any data that at least partially describes a change in position or orientation, heading, or any of these relative to a reference point, such as velocity, angular velocity, acceleration, or angular acceleration, or any combination of two or more of these. Within a frame of reference (e.g., a coordinate system associated with a person's body), the reference point can be absolute (e.g., the origin in the frame of reference) or relative (e.g., a previously known position). Different types of position sensors may use different reference points. Position can be represented by coordinates, one-dimensional, two-dimensional, or three-dimensional, whether Cartesian, radial, or spherical, or as a vector. Position can also be represented by relative terms, such as high, low, left, right, far, near, and combinations thereof. As used herein, “motion” of a body part, such as the wrist, is intended to represent the change in position or orientation of the body part relative to a reference point, or both, over time. Typically, a change in the position or orientation of the tip of the wrist is due to arm movement and is referred to herein as wrist movement or arm movement. Although other movements of the hand or fingers may occur in or near the wrist (or other body parts), such as bending of the hand or extension or contraction of one or more fingers (or similar body parts), these movements are referred to as “movements” in this article.

[0077] "Biopotential" is an electrical signal that propagates within the body, can be non-invasively sensed on the skin surface, and indicates neural signals transmitted to muscles, the electrical activity of muscles in response to those neural signals, and other electrical activity within tissues. When sensed at the top of the wrist, this biopotential is associated with the activity of the muscles controlling the hand and fingers, including movement of a single finger or group of fingers relative to the hand or relative to each other, and movement of the hand relative to the arm, and any combination of these movements. The phrase "associated with muscle activity" is intended to include any one or more of neural signals, muscle tissue signals, signals associated with anticipated muscle movement (whether or not), or signals associated with unintended muscle movement (whether or not), or any combination of two or more of these signals. Muscle movement may or may not occur, depending on the individual's condition. The individual may have a disease, such as amyotrophic lateral sclerosis (ALS), in which muscles do not move as expected or do not move at all, but neural signals may still be generated for anticipated muscle movement. The individual may be constrained in other ways, thus preventing muscle movement, but generating neural signals and possible muscle tissue signals. In some instances of this text, the text refers to movement of body parts, such as finger movement. These examples should be understood as brief references and include not only actual movements of body parts but also intentional but unrealized movements of those body parts, in which case bioelectrical signals indicate intentional muscle movements used to control those body parts.

[0078] User interface device 100 can also output feedback to the user via one or more feedback devices 110, examples of which are described in detail below. User interface device 100 can receive feedback signals 112 associated with feedback devices 110. User interface device 100 can generate (e.g., via controller 130) feedback signals 112 for feedback devices 100. User interface device 100 may not be the only source of user feedback, as such feedback can be provided from response devices, machines, or other feedback devices. In some embodiments, user interface device 100 may not have any feedback devices. For example, a flash of light may indicate that the user interface device is communicating with a response device, or vibration may indicate that the response device has completed an action.

[0079] User interface device 100 may also have one or more input devices 114, such as buttons, through which a person may intentionally or unintentionally provide input directly to user interface device 100.

[0080] User interface device 100 typically communicates with response device via a wireless connection. User interface device may be battery powered. When disconnected from response device, user interface device may communicate with response device via wireless transceiver (not shown) (sending data 120 and receiving feedback signal 112).

[0081] Various components of the user interface device (e.g., 106, 108, 110, 114, and 122) interact with controller 130. The controller is connected to these components to control their operation. The controller may include an analog-to-digital converter that samples, downsamples, or upsamples the output of sensors 106, 108, and 122 to provide a sequence of digital samples at a desired sampling rate based on the sensor's output signal. Controller 230 may include any other circuitry that processes or regulates the data 102, 104, and 124 prior to sampling, such as any amplifier or bandpass, notch, high-pass, or low-pass filters.

[0082] Biopotential sensor 106 senses biopotential at a body location (e.g., the top of the wrist) of the wearable user interface device. Biopotential sensor 106 provides one or more output signals, namely biopotential signals 104, indicating the sensed biopotential.

[0083] The biopotential sensor 106 includes one or more pairs of electrodes. Figure 1C As described above, each pair of electrodes is referred to herein as a "channel". When the user interface device 100 is worn on the wrist, the electrodes are arranged to contact the skin at the top of the wrist, and the electrodes within each pair are oriented such that the first electrode in that pair is closer to the finger than the second electrode. In some embodiments, as described in the related applications listed above or in the above... Figure 1C The types of electrodes and circuits described herein can be used to provide each channel of the biopotential sensor 106.

[0084] A differential signal can be generated from the output of each pair of electrodes. As described in U.S. Patent 10,070,799, a differential signal can be generated from the output of each pair of electrodes by using an operational amplifier. When using a reference electrode (e.g., Figure 1C In embodiments where additional noise suppression is applied to the electrodes 180, the differential signal may also include a signal from a reference electrode. For example, given a signal (e1) from the first electrode, a signal (e2) from the second electrode, and a signal (r) from the reference electrode, the reference signal (r) may be subtracted from the signals (e1, e2) of the first and second electrodes before generating the differential signal for the electrode pair. This can be performed in the analog domain, such that the differential signal for the electrode pair can be expressed as: (e1-r) - (e2-r). The differential signal may be the output of a channel. The outputs from one or more channels constitute the biopotential signal 104.

[0085] Bioelectric potential signal 104 is used to detect hand posture or the posture of another body part near the wearable user interface device. "Hand posture" includes the position or movement of one or more fingers relative to the hand, or the position or movement of the hand relative to the wrist, or any combination thereof. Therefore, in a relatively simple case, a system is designed and implemented to detect at least a first posture and distinguish it from the absence of any posture. If a second posture is introduced, the system is designed to distinguish between the first and second postures, as well as to distinguish between the first and second postures and the absence of either one. Therefore, the bioelectric potential sensor 106 is designed based on the application such that the bioelectric potential sensed by the bioelectric potential sensor 106 carries sufficient information relative to noise and relative to other specific gestures to be detected, enabling the detection of a person's specific posture.

[0086] In applications where the user interface device is worn on the top of the wrist, bioelectric signals can be processed to detect hand posture. This processing may include distinguishing between movement of any finger and inactivity of any finger. It may include differentiating one movement of a finger from different movements of the same finger or the same or different movements of different fingers. It may include distinguishing one hand posture from another hand posture. The bioelectric sensor may be single-channel, comprising a single pair of electrodes. It may have only two channels, each with a pair of electrodes. It may have only three channels. It may have four or more channels. Information from other sensors can be used to process the output signal from the bioelectric sensor to help distinguish different hand postures.

[0087] Within a single channel, a larger spacing between electrodes improves the channel's ability to output differential signals, providing more meaningful information. For two or more channels, a larger spacing between the two channels improves signal separation and reduces crosstalk. The electrode spacing ("electrode spacing") and the channel spacing ("channel spacing") are related to and selected based on the characteristics of the biopotential propagating within the body part on which the user interface device is intended to be worn.

[0088] For some applications, the placement of each electrode pair can be optimized for capturing neural signals. Capturing neural signals is helpful for applications where the person is known to have a neuromuscular disease such as amyotrophic lateral sclerosis (ALS) or may be limited in actual muscle movement. For example, electrode pairs can be placed at the top of the wrist to emphasize the capture of neural signals associated with the median, ulnar, or radial nerves, or any combination of these nerves. However, in many implementations, placing two differential channels at a sufficient interval at the top of the wrist provides sufficient information without the need for careful placement of specific nerves.

[0089] In some applications, acquiring signals from the electrodes may include a form of dynamic impedance matching, as described in U.S. Patent 10,070,799. An operational amplifier providing differential signals from the two electrodes in a channel can be connected to a variable resistor adapted to match the impedance of the skin. Without dynamic impedance matching, the sensor can still naturally adapt to the wearer's skin impedance after several minutes of wear. However, implementing dynamic impedance matching in hardware allows the sensor to adapt to the person more quickly.

[0090] As described in U.S. Patent 10,070,799, bioelectrical signals can include signal portions from various sources, including nerves, neuromuscular systems, muscles, and external sources (e.g., radio waves). Several processes can be performed to reduce the influence of unwanted signals, for example, by using notch filters and bandpass filters. Typically, external noise tends to be in the range of 60 Hz and integer multiples thereof; therefore, notch filters can be used for these different frequencies. Muscle contraction electrical signals have a frequency range of 20 Hz to 300 Hz, with the main energy in the range of 50 Hz to 150 Hz. Nerve signals typically have a frequency range of about 200 Hz to about 1000 Hz. A bandpass filter can be applied to attempt to focus on the nerve signal. Another bandpass filter can be applied to attempt to focus on the muscle signal. Bandpass filters with frequencies from 20 Hz to 1000 Hz can be used to focus on both nerve and muscle signals. As described in more detail below, such filtering can be performed by the bioelectrical sensor 106 or another component on the user interface device, or it can be part of signal processing performed by another part of the system.

[0091] As described above, position sensor 108 senses the position of a body part wearing the user interface device relative to a reference point, such as the top of the wrist. When sensing at the top of the wrist, the sensed position represents the position of the wrist relative to the reference point. The reference point can vary depending on the type of position sensor. Examples include, but are not limited to, absolute points (e.g., any reference point relative to the body, head, or mouth or other specific body part) or relative points, such as previously sensed positions. Position sensor 108 provides an output signal indicating the sensed position, such as position signal 102. This will be discussed in conjunction with the following... Figure 4A A more detailed explanation of the concept of body part location with a wrist-worn user interface device.

[0092] For example, position sensor 108 may include one or more accelerometers, gyroscopes, geomagnetic sensors, or other forms of magnetometers, or combinations of two or more of these. For example, an inertial measurement unit (IMU) may be used, which includes at least an accelerometer and a gyroscope, but may also include a geomagnetic sensor. An example of a commercially available product that can be used as position sensor 108 is a 9-axis sensor (e.g., the Bosch BNO055 absolute orientation sensor) that combines three types of devices (accelerometer, gyroscope, and geomagnetic sensor). Such a sensor can be used to provide sensor data for determining geographic location.

[0093] An accelerometer includes a device that measures its own acceleration along one or more axes and provides an output signal in response to stress or other forces caused by acceleration within the device. Taking into account the accelerometer's initial position and acceleration data on three axes, a processing device can determine the position and orientation of the wrist in three-dimensional space relative to its initial position. This sensor can be used to provide sensor data for determining geographic location.

[0094] A gyroscope comprises a device that measures direction and angular velocity along one or more dimensions. Taking into account the three-dimensional signals from the gyroscope, the three-dimensional orientation of the wrist relative to an initial orientation can be determined. This sensor can be used to provide sensor data for determining geographic location.

[0095] A geomagnetic sensor includes a device that measures the Earth's magnetic field to provide a signal indicating the measurement and direction of the magnetic field. Taking into account the signal from the geomagnetic sensor, a processing device can determine location, such as orientation or heading relative to the Earth's magnetic poles. Such a sensor can be used to provide sensor data for determining geographic location.

[0096] The position sensor 108 may include one or more sound detection devices, such as a single microphone or multiple microphones. The sound detection device may be a microphone array, including multiple microphones spaced apart on the user interface device 100. The sound detection device may be a microphone and resonant device on the user interface device 100. The sound detection device outputs an audio signal. The audio signal can be processed to determine the position of the wrist relative to a sound source (e.g., the mouth). Such an exemplary embodiment can be found in a co-pending patent application entitled “A Wearable Device Including a Sound Detection Device Providing Location Information for a BodyPart”, filed January 8, 2020, by Dexter Ang and David Cipolletta, serial number 16 / 737,242, which is incorporated herein by reference. Sound source localization techniques can be used to calculate such a position. For example, the sound source may be the mouth of a person wearing the user interface device. The sound detection device may be used in conjunction with any one or more other types of position sensors. The sound detection device can be used to provide sensor data to determine geographic location.

[0097] The user interface device may also include one or more additional sensors (e.g., 122) of other types that output additional sensor data 124. Examples of sensor types that can be used as additional sensors include, but are not limited to, any device capable of sensing, detecting, or measuring physical parameters, phenomena, or events, such as a Global Positioning System (GPS) sensor, humidity sensor, temperature sensor, visible light sensor, infrared sensor, image capturing device (e.g., camera), audio sensor, proximity sensor, ultrasonic sensor, radio frequency signal detector, skin impedance sensor, pressure sensor, electromagnetic interference sensor, touch capacitive sensor, and combinations thereof.

[0098] Geographic location can be determined based on input from sensor data from sensors associated with at least one person. Additional sensor data from sensors associated with one or more other people, or one or more other objects, or a combination thereof, can be used to determine geographic location. Examples of such sensors include at least any one or more of the following: a position sensor (described below), a geomagnetic sensor, a compass, a magnetometer, an accelerometer, a gyroscope, an altimeter, an eye-tracking device, an encoder providing information indicating changes in position, a camera providing image data or motion video, a reference to a known landmark, or an audio sensor (e.g., one or more microphones), or other sensors that can provide directional indications from a person, or any combination of two or more of these. Sensor data can be interpreted as, for example, any of a person's multiple raw hand gestures, an expected action, an arm angle or other indication of relative vertical position, direction or orientation, distance or height, or any combination thereof. Geographic location can be determined based on such sensor data received from sensors associated with one or more people.

[0099] In some implementations, the user interface device may include sensors for sensing data to determine geographic location, eliminating the need for bioelectrical sensors or position sensors. A person may wear multiple user interface devices, each with different sensors providing different data.

[0100] Turning now to feedback device 110, any device (or combination of devices) that provides feedback to a person wearing user interface 100 can be used. Feedback is any stimulus that is perceptible to the person and is generated by an action performed by the response device. In order to... Figure 1A For the purposes of this description, the user interface device 100 is shown to include a feedback device 110. However, in some embodiments, the feedback device may be located outside the user interface device 100, in place of any feedback device 110 in the user interface device 100 or in addition to any feedback device 110 in the user interface device 100.

[0101] For example, feedback device 110 may include a tactile feedback source, also known as a tactile device. A tactile device has physical contact with a person, typically through non-invasive contact with the skin, whether direct or indirect through some material. The tactile device generates stimuli that a person can perceive, typically through the skin (tactile stimulation) or through the sense of movement and position of the limbs (kinesthetic stimulation). Examples of tactile feedback may include, but are not limited to: vibration, motion, force, pressure, acceleration, heat, texture, shape, or other forms of tactile or kinesthetic sensation. User interface device 100 may receive data from an external source, such as feedback signal 112, for controlling the state of the tactile device. User interface device 100 may generate or store data for controlling the state of the tactile device via controller 130.

[0102] As another example, feedback device 110 may include a light source, such as one or more light-emitting diodes (LEDs). For example, from the wearer's perspective, a single LED may be placed on a visible surface of user interface device 100. Of course, multiple LEDs can be used. LEDs can produce monochromatic or polychromatic light. LEDs can produce light of a single intensity or multiple intensities. User interface device 100 may receive data from an external source, such as feedback signal 112, for controlling the state of the light source. User interface device 100 may generate or store data for controlling the state of the light source via controller 130.

[0103] As another example, feedback device 110 may include a sound source, such as one or more speakers. For example, a small speaker may be placed on a surface facing a person wearing user interface device 100. Multiple speakers may, of course, be used. The speaker may be selected based on the desired sound signal to be generated, such as a single tone or frequency, up to the full range of audible sounds, such as music. Sound may be generated at a single set volume or at different volumes. The volume may be adjusted by a person, user interface device, or other means (e.g., via feedback signal 112). User interface device 100 may receive data from an external source, such as as indicated by feedback signal 112, for controlling sound playback on one or more speakers. User interface device 100 may generate or store data for controlling sound playback on one or more speakers via controller 130. Such feedback signal may include audio data or control data, or both.

[0104] As another example, feedback device 110 may include a display, such as an N-pixel × M-pixel array for displaying images. For example, a small display may be placed on a surface facing a person wearing the user interface device 100. The display may be selected based on the desired resolution of the image to be produced on the user interface device, which may depend on the application and may support simple text or graphics up to full-motion video. Controls may be used to adjust the contrast, brightness, color balance, and other characteristics of the display, which may be adjusted by a person, the user interface device, or other devices (e.g., via feedback signal 112). User interface device 100 may receive data from an external source, such as as indicated by feedback signal 112, for controlling the display of content. User interface device 100 may generate or store data for controlling the display of content. Feedback signal 112 may include any combination of image data, text data, graphic data, video data, or control data.

[0105] User interface device 100 may also have one or more input devices 114 through which a person can provide input to user interface device 100. While such input devices are sensors in some sense (e.g., sensor 122 mentioned above), input devices 114 are intended to represent devices that a person actively manipulates to provide specific input directly to the user interface device, such as buttons or touchscreens. Such input devices typically do not generate signals that are processed as part of the human-machine interface, but are used to control the user interface device, such as powering the device, initiating network connections, and controlling settings.

[0106] exist Figures 1B to 1C In the exemplary configuration of the wrist-worn user interface device shown, the housing 150 includes Figure 1A The circuit shown and its corresponding mechanical components. Figure 1B It is a top-down perspective view; Figure 1C This is a bottom perspective view. The fastening device 152 allows a wristband or other device (not shown) to be attached to the housing 150 to allow the housing to be worn. Button 154 is an example of the input device 114. Figure 1A ); Light-emitting diode (LED) 156 is an example of feedback device 110.

[0107] like Figure 1C As shown, electrode 160 is typically made of stainless steel or other conductive material. In some embodiments, the electrode may be a flat disc. In some embodiments, the electrode may have a curved surface, for example, a conical shape. The diameter of a circular electrode is typically greater than 2 mm and less than 10 mm. In some embodiments, the electrode may have a diameter of approximately 6 mm.

[0108] Within the channel (e.g., a pair of electrodes 160 and 162), the electrodes have an electrode spacing measured between their centers. The electrode spacing 164 can be limited at the lower end by the size of the electrodes and the desired spacing between them, for example, approximately 15 mm (15 mm) in the case of an electrode diameter of 6 mm, and with approximately 9 mm of material provided between the near edges of the electrodes, and at the higher end by the size of the device, for example, approximately 25 mm (25 mm) in the case of an electrode diameter of 6 mm, and the device itself is approximately 30 mm. Typically, the maximum possible spacing is desired because reducing the electrode spacing results in a lower signal amplitude. In some embodiments, the electrode spacing can be approximately 25 mm (25 mm).

[0109] For two or more channels, the channel pair has a channel spacing measured between a first line drawn by electrodes (e.g., 160 and 162) passing through the first channel and a second line drawn by electrodes (e.g., 166 and 168) passing through the second channel. The channel spacing 170 can range from approximately 10 mm to 20 mm. This range is limited by high-end anatomy and therefore depends on the body part where the device is worn. At the top of the wrist, the high end of the channel spacing range is approximately 20 mm. Reducing the channel distance generally decreases source separation between channels. In some embodiments, the channel spacing can be approximately 15 mm. In some embodiments, the biopotential sensor may include one or more reference electrodes 180, which can be used for noise reduction, as described in more detail below.

[0110] Turn now Figures 2A to 2D Several response devices will be described. Response device 200 is a machine, typically a computing device, which may be associated with one or more machines or one or more other computing devices, or both. The response device responds to user interface device 100 to allow a person to interact with the machine, which may be the response device itself or a machine or computing device associated with the response device.

[0111] like Figure 2A As shown, the response device 200 receives data 202 based on the bioelectric potential signal 104 and the position signal 102 generated by the user interface device 100. As described in more detail below, the received data 202 may be the output signals 102 and 104 from the sensors, or it may be data representing information derived from processing these signals, such as data obtained from filtering or transforming these signals, data indicating features extracted from these signals, data representing the posture or position of a body part derived from these signals, or the original gesture detected from these signals, or other data obtained from processing these signals, or any combination thereof.

[0112] The response device 200 performs an action based on the received data 202. The response device 200 may also perform an action, for example, through the behavior of the response device or its associated machine, or by sending a feedback signal 112 to the user interface device. Figure 1A Feedback 204 is generated to a person via user interface device 100, or via other feedback device 206 controlled by response device 200. As described in more detail below, the actions to be performed and the feedback provided to the user depend on the application, computing device, machine, or other systems with which the person interacts using the user interface device and response device. This will now be discussed in conjunction with the appendix... Figures 2B to 2D Here are some examples to describe this kind of interaction.

[0113] exist Figure 2B In this context, the response device 200 is a computing device with which a person interacts. For example, the response device can be a computer (e.g., a general-purpose desktop or laptop computer), a smartphone, a tablet, a game console, a set-top box, a smart TV, etc. Some computing devices (e.g., smartwatches) may include both a user interface device and a response device within a single physical device. Figure 2B In the response device 200, the computing device typically has an operating system 208 that supports the running of one or more applications 210. The user interface device 100 is used by a person to interact with the operating system or applications. Feedback 204 from the operating system or applications can be provided via a display, speakers, or other feedback devices 206 accessible to the computing device, typically in the form of images and sounds.

[0114] In such Figure 2B In the illustrated configuration, the user interface device enables a person to interact with various applications running on the response device, such as media playback applications, camera, map applications, weather applications, search engines, augmented reality applications, virtual reality applications, video games, or other applications.

[0115] In such Figure 2B In the illustrated implementation, various different applications may exist, which the user can interact with via a user interface device. In such an environment, selection within these applications, and especially actions performed in response to the user interface device, are possible.

[0116] exist Figure 2CIn this system, response device 200 is a computing device running a remote control application 212 that interacts with machine 220, which is served or controlled by the computing device. A person interacts with machine 220 via the response device using a user interface device; in response to data received from the user interface device, the response device sends input to the remote control application 212. The behavior of machine 220 is the primary feedback of the system, but the computing device or user interface device may also provide other feedback. For example, the remote control application 212 may provide additional feedback 204 to feedback device 206 (e.g., the display of the computing device) or the user interface device.

[0117] For example, another machine 220 could be a drone, also known as an unmanned autonomous vehicle, whose application 212 includes a remote controller. Application 212 generates control commands, which are transmitted to the drone by a computing device under the control of application 212. When a person is using a user interface device, the application can generate these control commands in response to received data 202, based on signals generated by the user interface device. For example, the application can activate a remote control session in response to detecting an initial gesture associated with the function. During the remote control session, the position of the wrist is interpreted by the application as providing directional control of the drone. The application can terminate the remote control session in response to detecting an initial gesture associated with the function.

[0118] As another example, the other machine 220 may be another computer or a collection of computers. As another example, the other machine 220 may be an articulated robotic arm, which may or may not be considered a robotic person, for example, for manufacturing, telemedicine, or other applications. As another example, the other machine 220 may be a mobile robotic person capable of traversing land or water, having a motor and a propulsion system driven by that motor.

[0119] exist Figure 2D In this context, the response device is a computing device 200 with a communication application 230. The communication application 230 can send a message 232 to a user via a feedback device 206. When this user is using a user interface device, the communication application 230 can send the message 232 to the user interface device 100. The computing device 200 or another computing device (not shown) can then receive data 202 from the user interface device. The received data 201 can respond to the message 232, and... Figure 2D In the example shown, the user can be directed to a communication application. Feedback 204 can also be provided for various purposes, such as indicating that the computing device 200 has received data 202. Such a communication application can take many forms.

[0120] exist Figure 2DIn the example shown, communication application 230 is used to collect labeled data for training machine learning system 250. A response device sends messages to instruct a person to operate a user interface device, for example, by performing a specified action. In turn, the response device receives data from the user interface device. The response device can label this data as corresponding to the action instructing the person to perform and provide the labeled data to machine learning system 250.

[0121] As another example, to monitor the health status of a user interface device wearer, a response device may send a message requesting a response from the wearer. The response device then receives data from that person's user interface device. Based on the received data, the response device may perform additional actions, such as notifying a care provider about the person's status via a monitoring application 260. An exemplary implementation of such a system is described in co-pending PCT patent application PCT / US19 / 061421, filed November 14, 2019, entitled "Communication Using Biopotential Sensing and Tactile Actions," which is incorporated herein by reference.

[0122] As another example, an incoming message in a messaging application (e.g., a voice call or text messaging application) can cause a response device to send a notification message to a user interface device. The response device then receives data from that person's user interface device. In response to the data received from the user interface device, the response device can take an action in the messaging application, such as answering a call or replying to a text message.

[0123] In any such embodiment of the user interface device 100 and the response device 200, signals 102 and 104 generated at the user interface device are converted into inputs used by an application on the response device. This conversion can be performed by one or more processing devices that may reside on the user interface device, the response device, or both. In some embodiments, this conversion can be performed by one or more processing devices residing on a computing device separate from the user interface device and the response device.

[0124] In any of the foregoing configurations, one action that can be performed by a response device and that can provide feedback to it includes identifying a geographic location. This geographic location can be used in a variety of ways, examples of which include, but are not limited to, the following, and may include combinations of two or more of the following: The geographic location can be used as input to the application. The geographic location can be data passed to another user or another machine. When combined with data based on signals received from one or more user interface devices to select the action to be performed, the geographic location can be considered contextual information. The geographic location can be provided to the user as feedback.

[0125] Converting signals 102 and 104 from sensors 106 and 108 into inputs for an application can involve multiple processing stages. Now refer to... Figure 3 These processing stages will now be described in more detail.

[0126] The first processing stage is performed by sensors (300) in the user interface device. Some sensors output digital signals. Some sensors output analog signals. In some embodiments, the biopotential sensors for the channel may include dynamic impedance matching circuitry and amplifiers to output analog differential signals. The output of this first stage is the raw sensor signal.

[0127] The next processing stage 302 may involve conditioning the signal output from the sensor. After any conditioning, the signal output from the sensor is sampled and synchronized. For sensors that output analog signals, the analog signals may be subject to signal conditioning operations performed by analog circuitry, such as filtering and amplification. The analog signal is then converted into a digital signal using an analog-to-digital converter (ADC). For sensors that output digital signals, or for digital signals output from an ADC, the digital signals may undergo signal conditioning operations performed by digital circuitry, such as filtering and amplification. Regardless of whether the conditioned signal is analog or digital, signals from various sensors can be sampled at the same sampling rate, thereby providing an output stream of digital sensor signals at that sampling rate, such that the data from various sensors are synchronized and have the same sampling rate. This sampling can be performed by an ADC or a digital-to-analog converter (DAC) and may include upsampling with or without interpolation or downsampling with or without filtering. The digital signal output stream from this second stage 302 at that sampling rate is referred to herein as the “normalized digital sensor signal”.

[0128] The normalized digital sensor signal can be subjected to some digital signal processing (304) without any filtering or other conditioning applied before sampling, and after sampling can be applied. For example, the bioelectric potential signal can be filtered by applying a notch filter to eliminate or reduce the influence of electrical interference. For example, the bioelectric potential signal can be filtered by applying a bandpass filter to eliminate or reduce the influence of other electrical activities affecting the body, thus focusing only on signals that may be related to muscle movement. If the bioelectric potential signal is to be processed in a context that emphasizes neural signals, it can be filtered by using a bandpass filter to eliminate or reduce the influence of signals associated with muscle activation. In this digital signal processing stage 304, any further filtering or processing can also be applied to the signal derived from the output of any other sensor. If used, the output of this third stage 304 is referred to herein as a “filtered” normalized digital sensor signal, which remains the output stream of the digital signal at the sampling rate.

[0129] In the next stage 306, the optional filtered, normalized digital sensor signal is processed to extract the signal's "features." Figure 3 In this context, this processing is represented by feature extraction layer 306. A feature is any property of a signal that can be quantified by applying operations to a signal or a transformation of a signal (e.g., a frequency domain representation of the signal) or a combination of two or more signals or their transformations. Typically, considering the signal as a stream of digital signals output at the sampling rate, features are computed against a subset of the sample stream by applying a window function (called a window). The window function can be specified by the time and interval in the digital signal stream. The result of applying the window function is that values ​​outside the interval are zero. The simplest window function is a rectangular window, which preserves the original samples within the interval. Window functions can be used to apply mathematical functions instead of identity functions to the samples within the interval. For a sequence of windows in the digital sample stream, another parameter is the amount of overlap from one window to the next, i.e., the offset from the beginning of one window to the beginning of the next immediately following window. The offset can be equal to, less than, or greater than the interval specified for the window. The output of this fourth stage comprises a stream of feature sets, outputting one feature set for each window of the (optionally filtered) normalized digital sensor signal. Several examples of features are described in more detail below.

[0130] exist Figure 3 In the next processing stage 308, referred to as the "gesture detection" layer, the next processing stage involves processing normalized digital sensor signals, filtered normalized digital sensor signals, or a stream of feature sets extracted therefrom to identify the pose of a body part, the position of a body part, or the original gesture, or a combination or subset thereof, collectively referred to as "gesture detection." The responding device then prompts the execution of an action. Note that although... Figure 3A feature extraction layer 306 is shown between the gesture detection layer 304 and the digital signal processing layer 304. However, in some embodiments, the gesture detection layer 308 may use data from the feature extraction layer, the digital signal processing layer, and the normalized digital sensor signal output. In response to the detection of an initial gesture, the response device directs the input to the application, as shown in 310, and the application performs an operation in response.

[0131] like Figure 3 As shown, in some implementations, a computational model, referred to as a machine learning model, can be used to detect poses, positions, or primitive gestures. A machine learning model is a model trained using labeled data by machine learning processing, examples of which are described in more detail below. Typically, data from one or more user interface devices is collected in a manner that also labels the data to represent known poses, positions, or primitive gestures; for example, normalized digital sensor signals, filtered normalized digital sensor signals, or a stream of feature sets extracted from these signals. A machine learning (ML) training module 320 receives the labeled data and generates parameters for a trained machine learning (ML) model 322. The ML training module 320 can process this data outside of real-time interaction between the user interface device and the response device. Typically, the trained ML model 322 is provided to the gesture detection layer 308 by providing the parameters of the trained ML model to the gesture detection layer, which applies these parameters in its model implementation. Thus, the parameters of the trained model are provided to components within the user interface device or the response device, such as an operating system, device driver, or application responsible for performing gesture detection.

[0132] The interactions between the various processing layers have now been described; some exemplary implementations of these different layers will now be described.

[0133] At certain points during these processing phases, information is transmitted between the user interface and the response device. This processing can be considered as the signal transmission layer (STL). Figure 3 (Not shown in the diagram), but according to embodiments, the signal transmitted from the user interface device to the response device on the signal transmission layer can be the output of any of layers 302, 304, 306, or 308. In other words, the user interface device can transmit normalized digital sensor signals, filtered normalized digital sensor signals, feature set streams derived from these signals, or gesture or position information derived from these signals, or raw gestures detected based on these data, or any combination thereof. The suitable design of any particular embodiment depends on how much processing is performed within the user interface device before data is transmitted to the response device and on the information used on the response device to perform any further processing, which depends on the design constraints of the embodiment.

[0134] In this document, this signal transmission layer is referred to as a device interface. The device interface includes specifications for the type and format of information transmitted between the user interface device and the response device, the protocol for transmitting such information, and the communication system for transmitting the information. In some embodiments, the user interface device implements a streaming data protocol to transmit digital sample streams based on the output of bioelectrical and position sensors. The response device can send messages and feedback to the user interface device via the device interface. Therefore, the device interface can be bidirectional. Alternatively, a separate device feedback interface can be present. This feedback interface can have its own specifications for the type and format of information transmitted between the response device and the user interface device, the protocol for transmitting such information, and the communication system for transmitting the information. In some embodiments, the feedback interface is provided by asynchronous messaging. In some embodiments, the communication system is a wireless Bluetooth connection between the user interface device and the response device.

[0135] At certain points during these processing phases, information may also be passed within the response device from a first application running on the response device (which may be the operating system or device driver of the user interface device) to a second application running on the response device or a machine associated with the response device. In some embodiments, the first application includes a gesture detection layer, and the second application is designed to respond to the output of the first application. In some embodiments, the second application may have a set of expected inputs that can influence the behavior of the application, and data received from the user interface device is mapped to these inputs by the first application. In some embodiments, the first application does not perform gesture detection, but instead passes data such as digital samples or feature set streams to the second application, and the second application performs gesture detection.

[0136] In this document, an application programming interface (API) refers to the information passed within a response device from a first application to a second application (or more generally, to an application running on the response device), and the protocols (including data formats and message protocols) used for such communication. Because there may be outputs (e.g., messages and feedback) generated by the second application that need to be sent from the response device to the user interface device, this application API can be bidirectional. Alternatively, there can be a separate application feedback interface with its own data format and message protocol.

[0137] In some implementations, for example, a device driver (first application) for a user interface device that runs on the responding device as part of an operating system may include a device interface for handling interactions with the user interface device and an application interface for handling interactions with a second application.

[0138] In some implementations, the first and second applications are combined into a single application, such that the single application implements a device interface to handle interactions with the user interface device, or otherwise directly processes information from the user interface without the intervention of the other application.

[0139] Multiple application programming interfaces (APIs) can provide multi-layered processing. For example, a device driver within the operating system (a first application) can process data received from a user interface device into information about detected features, or detected postures, or location information, or detected raw gestures; another application (a second application) can further process such information, for example, mapping such information to corresponding events to simulate another input device, and providing such events to yet another third application.

[0140] In this context, user interface device 100 ( Figure 1A ) generates position signal 102 and biopotential signal 104, and responds to device 200 ( Figures 2A to 2D The device receives data 202 based on these signals via a device interface. Within the response device, information based on the received data 202 can be passed to one or more applications via an application programming interface (API). The content of the data transmitted via the device interface and the API depends on the location where different processing stages occur.

[0141] When the user interface device is battery-powered and wirelessly connected to transmit data to the response device, in order to reduce the amount of processing performed on the user interface device, the user interface device can be configured to implement a device interface to transmit normalized digital sensor signals. In this embodiment, further data processing occurs on the response device.

[0142] In many computing devices that can be used to implement a response device, the operating system typically uses a device driver to manage data transfer from an input device (e.g., a user interface device) to an application. Some operating systems allow data from an input device (e.g., a user interface device) to bypass the operating system and be transferred directly to the application's memory. The device driver implements a device interface that communicates with the user interface device and an application interface that communicates with the application; the application implements the application interface.

[0143] In embodiments where the device interface provides normalized digital sensor signals to the responding device, components on the responding device process these signals to generate input for the application. In some embodiments, the device driver can pass the received data to the application, in which case the application implements a feature extraction and gesture detection layer. In some embodiments, the device driver can implement a feature extraction layer, and the application can implement a gesture detection layer. In some embodiments, the device driver can implement both a feature extraction layer and a gesture detection layer. In some embodiments, a first application implements a portion of the gesture detection layer and passes the processed information to a second application.

[0144] The pairing of this user interface device with this response device supports many different contexts in which a person can interact with the machine. Examples of such usage contexts include, but are not limited to, any one or more of the following.

[0145] The combination of a response device and a user interface device may have a very small display or no display at all. The person may not be able to use the display. The person may be focused on an object and unable to look at both the object and the display simultaneously. The person may be holding or grasping another object. The person may not be able to provide touch input to a touchscreen or similar device. The person may need to free their hand to hold an object instead of the user interface device. The person may not be able to provide auditory input. The person may not be able to provide voice input. The person may not be able to speak. It may be unsafe or undesirable for the person to speak or make noise in their current environment. The person may not be able to receive auditory feedback. It may be unsafe or undesirable for the person to receive auditory feedback in their current environment. The person may be physically restricted or impaired, or may have limited mobility. The person may be in a situation requiring limited perceptible activity, including movement or motion. In this context and other contexts, the person may attempt to communicate or interact with other people in similar contexts.

[0146] To support this type of human-computer interaction in such a context, the response device performs an operation in response to received data, based on positional and bioelectrical signals generated on the user interface device (e.g., the result of finger and hand movements and wrist and arm movements), user input to the user interface device, data received from other sensors on the user interface device, contextual data from any source, and combinations thereof. In turn, the person receives feedback through one or more of the user interface device, the response device, a machine, or other feedback devices, or combinations thereof.

[0147] The operation to be performed may also be based on information from any other sensors on the user interface device, or on one or more other contextual information available in the user interface device, response device, or machine. Examples of other contextual information that can be used to perform the operation include, but are not limited to, the following and combinations thereof: Information from another user interface device worn by another person may be used. Information from sensors of another device may be used, such as GPS information from a device associated with the person. Information accessed from a local or remote database may be used.

[0148] More specifically, based on positional and bioelectrical signals, the primitive gesture occurring at a given moment is detected. In some cases, the operation to be performed is selected based on the occurrence of the primitive gesture. In others, the operation is selected based on the occurrence of multiple primitive gestures in the sequence. In still others, the operation is selected based on the co-occurrence of two or more primitive gestures. In yet another case, the operation is selected based on one or more contexts surrounding this occurrence. Combinations of these underlying principles for the operation to be performed are also possible.

[0149] To illustrate the various primitive gestures that can be detected based on bioelectrical and positional signals, examples based on user interface devices worn on a person's wrist are described below.

[0150] As follows Figure 4A In more detail, a primitive gesture can be characterized as the following sequence: a first posture of the hand or a first position of the wrist, or both, at a first moment; and a second posture of the hand or a second position of the wrist, or both, at a second moment after the first moment. In some primitive gestures (e.g., hold), the first and second positions are identical. A primitive gesture can also be considered as detecting such a transition, or as not detecting a transition.

[0151] The original gesture can be characterized by a set of features extracted from position signals or a set of features extracted from bioelectrical signals, or both, thereby distinguishing the original gesture from other original gestures. This set of features can be extracted from a single window of signal samples or from a sequence of consecutive windows of signal samples. In some implementations, this set of features characterizing the original gesture corresponds to one or more hand postures or one or more wrist positions, or a combination thereof.

[0152] The original gesture can be represented by one or more states, as in the following combination. Figure 4CMore specifically: the initial state of the signal in the window before the original gesture occurs (first set of features); the start of the signal change in the window indicating the occurrence at the beginning of the original gesture (second set of features); the state or change of the signal in one or more windows immediately following the start (third set of features); and the signal change in the window at the end of the original gesture (fourth set of features). The original gesture may have a moment when it is considered to have been detected. The original gesture may exist for a period of time, and therefore may have a start time and an end time.

[0153] Due to the biopotential sensor located at the top of the wrist ( Figure 1A The output of the bioelectric sensor (106) is related to the muscle activity controlling the hand and fingers, thus the bioelectric signal is an important source of information about hand posture. However, if the position of the wrist (i.e., location or orientation) changes due to arm movement, this arm movement can cause inconsistencies in the degree and position of contact between the electrodes of the bioelectric sensor and the skin, leading to inconsistent bioelectric signal generation. Therefore, arm movement reduces the ability to detect hand posture from the bioelectric signal. In applications where this inconsistency occurs, data based on the bioelectric signal can be ignored when significant arm movement is detected. In this document, this operation of ignoring the bioelectric signal is referred to as "motion silencing". Therefore, in some implementations, in order to detect the original gesture based on hand posture, the initial state of the signal in the first window, just before the second window where the posture can be detected, can be that the wrist position has not changed significantly. When the wrist position has not changed significantly, the bioelectric signal is analyzed to detect the hand posture.

[0154] While many possible hand postures can be detected, this paper describes several postures associated with raising one or more fingers. Finger movements can include extension or contraction, or both. Differences in the bioelectrical signals generated for different primitive finger gestures can be used to distinguish between them. By processing the bioelectrical signals to indicate postures associated with raising one or more fingers, very subtle finger movements can be detected, which can be performed when holding an object or when the hand is concealed (e.g., in a person's pocket).

[0155] refer to Figure 4AThe hand 450 is shown in the intermediate position 452. The user interface device 454 is shown worn near the top of the wrist 456. As shown in 460, when the fingers are lifted from this intermediate position 452, the first joint 458 closest to the hand rotates toward the back of the hand and away from the palm, which is the opposite of retracting the fingers toward the palm. This movement is also different from a flick or tap when the fingers extend or straighten without rotating the first joint toward the back of the hand. Any finger or multiple fingers can be lifted sequentially or together. Lifting a finger can also be understood as the action of lifting the finger off a surface, such as a table or a grasped object, while the palm is placed on the surface. The fingers are lifted off the surface while the hand remains in place. The “start” or second set of characteristics characterizing this movement is a sudden increase in amplitude in the bioelectrical signal channels. Differences between bioelectrical signal channels allow the distinction between the lifting of one finger (e.g., the index finger) and the lifting of other fingers.

[0156] The initial gesture, referred to herein as “lifting,” occurs briefly while the finger is raised (460), defined by a fixed or adjustable threshold, and then returns to a resting state (e.g., 452), where the bioelectrical signal no longer indicates finger lifting. Therefore, the state or change of the signal immediately following the start (third set of features) is a sudden drop in the bioelectrical signal associated with the raised finger. Lifting can be considered detected when a window of slight increase in the bioelectrical signal associated with the raised finger posture is detected, followed by multiple windows of significant decrease in the bioelectrical signal associated with the raised finger. When detected, lifting can be considered to have ended, and therefore, detecting the fourth set of features to consider lifting to have ended is not involved.

[0157] The original gesture, referred to herein as "hold," occurs when the finger is raised (460) and held in that position for a period of time (460). The time period for raising and holding can be defined by a fixed or adjustable threshold. Thus, the signal state or change in one or more windows immediately following the start (third set of features) is a sustained increase in the amplitude of the bioelectrical signal associated with the raised finger, without further substantial wrist movement. Holding is considered to have been detected when two or more windows of increased amplitude are detected in the bioelectrical signal associated with the raised finger. Holding is considered to have ended when the bioelectrical signal associated with the raised finger decreases significantly, which indicates the fourth set of features defining holding.

[0158] When the fingers are raised (460) and in a "hold" gesture (460), the original gesture, referred to herein as "finger gliding," occurs, followed by a movement of the arm such that the wrist position changes with a sweep in one direction, for example, to the right or left as shown at 462, or up or down as shown at 464. Therefore, the signal state or change in one or more windows immediately following the start (third set of features) is a change in wrist position by an amount exceeding a threshold in a given direction. A finger gliding can be considered detected when wrist movement is detected after a raised and held gesture is detected, without involving the detection of a fourth set of features to consider the finger gliding to have ended. A similar original gesture to a finger gliding is a combination of a raised and held posture followed by a clockwise or counterclockwise rotation of the wrist tip as shown at 466.

[0159] Over a period of time, as the finger slowly lifts from an initial position 452 to a lifted position 454, a primitive gesture, referred to herein as a “simulated lift,” occurs. Compared to no movement, the amplitude of the bioelectric signal typically increases for a simulated lift, initially lower than the amplitude indicated by the bioelectric signal at the time of lift, but gradually increasing over time. Therefore, the signal state or change in one or more windows immediately following the start (third set of features) is a continuous increase in the bioelectric signal associated with the lifted finger. A simulated lift is considered detected when two or more consecutive windows of change (increase or decrease) in the bioelectric signal associated with the lifted finger are detected. The simulated lift is considered to end when the bioelectric signal associated with the lifted finger returns to its initial state, signifying the definition of the fourth set of features for a simulated lift. Therefore, a simulated lift can be used as input to slowly increase or decrease the value while the simulated lift continues for a period of time.

[0160] Due to position sensor ( Figure 1A When the position signal (108) is located at the top of the wrist, its output is related to the wrist's position; therefore, the position signal can serve as a primary source of information about the position, orientation, and any changes in the wrist's orientation. The position signal can be used to distinguish different positions, orientations, and movements of the wrist's orientation. The wrist's position and positional changes can be interpreted as various primitive hand gestures.

[0161] Here are two illustrative examples of position-based primitive gestures: relative position and rotation. Relative position can be based on position along one or more axes. In the examples below, relative position is defined by referring to a grid of two or more elements mapped to in one, two, or three dimensions. In the example below, the grid is described as having three vertical elements and three horizontal elements. Similarly, rotation can be based on orientation. In the example below, rotation is represented as the wrist relative to the elbow in two or more 180° increments. o The rotation angle in the subdivision. In the examples below, the subdivision is described as positive or negative N.o Rotate (from +90) o to -90 o Various other primitive gestures, primarily location-related, can be defined based on location signals, either alone or in combination with other signals.

[0162] It can detect the orientation or rotation of the wrist and hand relative to the elbow, for example, based on "rolling" detected by a gyroscope sensor located at the wrist. For example, as... Figure 4A As shown, the wrist can be in the "middle" position 452, or rotated 90 degrees clockwise. o Alternatively, position 470 degrees with palm facing up, or rotate 90 degrees counterclockwise. o Or, in the "palm down" position 472, or in the process of rotating from one orientation to another. Using three different relative rotation positions (orientations) is merely illustrative; any number of different relative rotation positions can be used depending on the granularity and reliability of the rotation detection. The amount of rotation can be expressed as N. o Represented by rotation.

[0163] It can also detect the hand's orientation relative to the wrist (the amount of forearm flexion at the wrist). For example, the hand's orientation can be "up," "down," or "middle" relative to the wrist. Using three different relative orientations is merely illustrative; any number of different relative orientations can be used depending on the granularity and reliability of the orientation detection. Orientation can be determined by, for example, a distance N from the center. o It can be represented by offset, or by direction classification, or by orientation. It can also detect the left or right rotation of the hand relative to the wrist.

[0164] The relative vertical position can be detected based on the start, end, or duration of the arm's vertical movement, for example, by raising or lowering the arm. For example, as... Figure 4B As shown, the wrist can be "high" or "upward," for example, at 480 above the shoulder; "low" or "downward," for example, at 482 below the hip; or "medium" or "intermediate," for example, at 484 between the hip and shoulder, or during movement from one position to another. Using the shoulder and hip as thresholds is merely illustrative; other threshold heights or reference points can be used. For example, the height of the wrist relative to the elbow can be determined instead of the shoulder. Similarly, using three different relative heights is merely illustrative; any number of different relative vertical positions can be used depending on the granularity and reliability of the position detection. This direction can be represented by a radial direction relative to anchor points such as the elbow, waist, or hip. This information can be provided as an angle, which can represent the arm angle when the user interface device is worn on the arm.

[0165] Similarly, the relative horizontal position can be detected based on the start, end, or duration of a horizontal arm movement, for example, by sweeping the arm to the left or right. For example, the wrist can be in the wearer's "left," "middle," or "right," or during a movement from one such position to another. Using three different relative horizontal positions is merely illustrative; any number of different relative horizontal positions can be used depending on the granularity and reliability of the position detection. The direction can be represented by a radial direction relative to an anchor point (e.g., the elbow or spine). In some implementations, the horizontal direction can be provided, for example, by a geomagnetic sensor. In some implementations, the horizontal direction can be derived from accelerometer or gyroscope data from an initial reference position.

[0166] The application may include a graphical user interface (GUI) that allows users to change the definition of relative positions. For example, a grid with multiple grid elements may be defined based on thresholds used to define relative vertical and horizontal positions. The GUI can present a graphical representation of this grid to the user. The user can manipulate the number and size of the grid elements defining the relative horizontal or vertical positions. For example, as... Figure 6 and Figure 8 As shown (described in more detail below), when displayed on a touchscreen or other conventional graphical user interface that allows selection of displayed objects, the user can select and move the lines defining the grid. In response to the user's selection, the application on the device updates the size of the grid elements. During this process, the updated grid elements can be interactively displayed. When complete, the updated grid elements can be used to process wrist position information into relative vertical and horizontal positions.

[0167] As described above, other detectable primitive gestures are based on one or more postures or positions detected at a particular moment, in chronological order, or simultaneously at a particular moment. A particularly useful primitive gesture is one designed to initiate the processing of signals from a user interface device, or from another input device or sensor, or both. This primitive gesture is referred to herein as a “wake word.” In some embodiments, the user interface device may be configured to not send data from a sensor other than the one providing the signal, before processing signals from the user interface device, from which the primitive gesture defining the wake word is detected. The wake word can be used to “unlock” the user interface device and initiate processing of data received from it. Similarly, the same wake word or another primitive gesture (“sleep word”) can be used to “lock” the user interface device and stop processing of data received from it. Such a wake word or sleep word may control the processing performed by the response device, or may be dedicated to an application running on the response device, or both.

[0168] A wake-up word or sleep word can be any sequence of primitive gestures detectable from actively monitored sensors. Ideally, the sequence of primitive gestures defining a wake-up word should not arise naturally from ordinary movements of the wrist or arm, or from movements of the hand or fingers; rather, such a sequence of primitive gestures should most likely be intentionally generated by a person. For example, with a specific N... o Rotating your wrist in a sequence (as described below) can be used as a wake word.

[0169] In some embodiments, the original gesture sequence defining the wake word or sleep word may be processed by a low-level driver of the operating system of the response device as a signal to activate or deactivate the response device's processing of data received from the user interface device or from another sensor or input device. In some embodiments, the original gesture sequence defining the wake word or sleep word may be processed by an application as a signal to activate or deactivate the application's input processing based on the response device's processing of data received from the user interface device. In some embodiments, the original gesture sequence defining the wake word or sleep word may be processed on the user interface device to activate or deactivate data transmission from the user interface device to the response device. In such embodiments, the response device may be configured to continuously process data received from the user interface device. In some embodiments, the wake word or sleep word may activate processing of another form of input device or sensor data, for example, activating speech recognition applied to audio signals received from a microphone on the user interface device, the response device, or other devices.

[0170] In some embodiments, where the wake word or sleep word used to unlock and lock the user interface device is handled by a low-level device driver of the operating system of the responding device, the responding device may receive signals from the user interface device via an input port. The device driver of this input port can process the received signals to determine whether an unlock command or a lock command has been received. In response to receiving an unlock command, subsequent data can be passed to the operating system of the responding device for further processing by the operating system or other applications managed by the operating system. Similarly, in response to receiving a lock command, the device driver can stop passing data from the user interface device to the operating system of the responding device.

[0171] For reference Figure 4CA state diagram is used to illustrate the operation of detecting the original gesture. Typically, the data to be processed is in the form of a sequence of digital samples, with the amplitude and sampling rate of each sample. The system defines a series of sample "windows." A sample window can be specified by a. a specific moment in the signal and b. a period of time or multiple samples surrounding that moment. For each window, features are calculated based on the samples within that window using digital signal processing techniques. For each window, the sampled data within that window and any features extracted from the sampled data are processed to determine whether a feature associated with the original gesture has been detected in that window. Any given window or set of windows can be labeled as representing or not representing the original gesture. Information about features detected in previously processed windows can be used to determine whether the original gesture has been detected in the current window. More details on how windows are processed are provided below. Processing the received data into windows of samples and corresponding features allows for a state- and sequence-based approach to process the received data to detect the occurrence of original gestures, combinations of original gestures, and sequences of original gestures. An exemplary implementation of this state-based processing is as follows... Figure 4C As shown.

[0172] Before receiving any wake word from the user interface device, the process begins in initial state 400. In response to data other than the wake word, as shown in 402, the process remains in initial state 400. In response to the detection of a wake word, the process transitions to state 404, as shown in 406. A wake word received in this state can be used to transition 440 back to initial state 400.

[0173] In state 404, the processing primarily monitors information from the position sensor on the wrist. As long as the information from the position sensor indicates substantial movement, the processing remains in state 404, as shown in 408, effectively achieving motion muting. In this state, information about the wrist's position and orientation can be updated based on the information from the position sensor. If there is no substantial wrist movement, information from the bioelectric sensor can be analyzed to determine if there is any indication of hand or finger movement, such as finger lifting. Upon detection of finger lifting, a transition to lift state 410 occurs, as shown in 412. Similar states can be provided for the detection of other features related to other movements.

[0174] In the raised state 410, by analyzing subsequently received data, the process determines whether the detected movement indicates a raised position, simulates a raised position, or holds, and which of these indicates. If the detected movement indicates a raised position, an operation can be performed in response to the raised position, and a transition 414 back to state 404 can occur. For example, the operation performed can be based on the current position and orientation of the wrist when the raised position is detected.

[0175] If the detected motion indicates a simulated lift, a transition 416 occurs to simulated lift state 418. In this state, an operation in response to the simulated lift can occur while the simulated lift continues to be detected, as shown in 420. The operation performed can be based on the current position and orientation of the wrist when the simulated lift is detected. When the simulated lift ends, a transition 422 occurs back to state 404.

[0176] If the detected motion indication persists (the bioelectric sensor detects a sustained level of activation in the finger muscles), a transition 424 to hold state 426 occurs. In this state, an operation in response to the hold can occur while the hold continues to be detected, and no wrist movement is detected, as shown in 428. When the hold is no longer detected, a transition 432 back to state 404 can occur. If wrist movement is detected while the hold continues to be detected, a transition 430 to state 434 occurs, and a finger slide is indicated as detected in state 434, and the corresponding operation can be performed. After performing the operation related to the finger slide, a transition 436 back to state 404 can occur. The operation performed in state 428 or 434 can be based on the current position and orientation of the wrist when the hold or finger slide is detected.

[0177] Now for reference Figure 5 The data flow diagram will now be described, showing the response device ( Figures 2A to 2D 200 in the middle) responds to based on the user interface device ( Figure 1A The response device receives data (202) from the sensor (100) to provide input to the application. After activating the response device to receive and process information from the user interface device, the response device can perform operations based on information 502 and information 504, and combinations thereof, where information 502 is derived from wrist position detected by the position sensor, and information 504 is derived from bioelectric potentials related to muscle movement detected by the bioelectric potential sensor. Information 502 and 504 can be any digital sample, feature, posture, position, or detected raw gesture and state information, as well as any other available contextual information, generated by any processing layer of data from the user interface device. In some embodiments, the response device may include a selection module 706 that can select an application (508-1, 508-2, ..., 508-N). Within the application (508-1, 508-2, ..., 508-N), the application performs actions based on information 510, which may include information 502 or 504, and any additional information derived from this information by the response device or the application.

[0178] The application can select actions based on one or more detected wrist positions or detected hand gestures, or combinations thereof. In some implementations, the response device can associate actions with wrist positions, and when the wrist is in a given position, those actions can be activated in response to detected hand gestures. As described above, settings defining thresholds for determining relative wrist positions can be manipulated in the graphical user interface. Such settings can also include associating one or more actions with a position having a gesture or other primitive gesture that might occur when the wrist is in that position.

[0179] Several additional examples will now be described to further illustrate this type of behavior of the response device.

[0180] First Example

[0181] In one exemplary embodiment, different applications are assigned to different wrist positions relative to the body. For example, a first application might be assigned to an "up" position, a second application to a "down" position, and a third application to a "middle" position. The response device directs input to the first application when it determines the user interface device is in the up position. The response device directs input to the second application when it determines the user interface device is in the down position. The response device directs input to the third application when it determines the user interface device is in the middle position.

[0182] For example, input directed to the application may be based on one or more primitive gestures detected when the user interface device is in a given position. For instance, when the user interface device is in an upward position, the user can perform a lift with their index finger. Input based on this primitive gesture is then passed to the first application. As another example, when the user interface device is in a downward position, the user can perform N... o The wrist rotates. Then, input based on this initial gesture is passed to the second application. As another example, when the user interface device is in the middle position, the user can perform N... o The wrist rotates, then the index finger lifts. Input based on this original gesture sequence is passed to a third application.

[0183] As a specific illustrative use case, a person might carry a mobile computing device (response device), such as a mobile phone, while walking, and simultaneously wear a wristband (user interface device) with an inertial measurement unit and bioelectrical potential sensors. The mobile computing device can be in the person's pocket, held in their hand, worn on their wrist, or located elsewhere on their body. The mobile computing device and the user interface device can be combined in the same device. The mobile computing device may have a display, but this display may or may not be visible to the person. The mobile computing device typically has an operating system, such as iOS or Android, and can run one or more applications. For example, the first application could be a weather application, the second a music playback application, and the third a map application. The person might wear a headset to communicate with the mobile computing device, or the mobile computing device might have a speaker to play audio under the guidance and control of the mobile computing device.

[0184] Using this example setup, the person can raise their hand above their head and then lift their index finger. In this case, the mobile computing device will lift its index finger to point at the weather app, which in turn allows the mobile computing device to play audio through headphones, for example, using a voice-generated version or a recording of the current weather forecast.

[0185] This user can position their arm, wrist, and hand in a neutral position and raise their index finger. In this situation, the mobile computing device will raise its index finger and point it at the audio playback application, which will then cause the audio playback application to start playing music. The current music selection for playback can be determined by various default settings. Once playback has started, the audio playback application can accept further input, which will provide various playback controls (see the second example below).

[0186] The person can lower their hand to a low position below the waist or hips and raise their index finger. In this case, the mobile computing device will raise its index finger and point it at the map application, which in turn allows the mobile computing device to play audio through headphones, for example, using a voice-generated version or a recording of the user's current location on the map.

[0187] In this exemplary embodiment, the operating system of the mobile computing device or the control application running on the mobile computing device may include Figure 5The selection module 506 may include mapping the detected wrist position to an application. For example, the selection module may maintain a data structure such as a table that stores data representing the relationship between the detected wrist position and the corresponding application. In response to receiving data indicating the detected wrist position from the user interface device, the operating system or control application may, based on the data received from the user interface device, select the application corresponding to that wrist position as the target for any action associated with the detected original gesture.

[0188] Second example

[0189] As another exemplary implementation, consider an application that offers a variety of selectable functions, such as an audio playback application. Examples of audio playback operations include playing, stopping, pausing, skipping forward, skipping backward, selecting a playlist, selecting an item from the playlist, etc.

[0190] In this exemplary embodiment, different operations can be assigned to different elements in an M×N grid, with a minimum grid size of 1×2 or 2×1, i.e., at least two grid elements. The grid can be rectangular, and the grid elements within it can be rectangular, uniform, and adjacent. Grid elements can be non-rectangular. Grid elements can be non-uniform. Grid elements can be non-adjacent. Grid elements can have different shapes and sizes. Grid elements can be arranged in two or three dimensions. The shape and size of the grid elements can be user-defined and user-modifiable, or can be defined by the computer system and can be fixed, or a combination of both. Each element in the grid can be assigned to one or more operations to be performed. The application can include multiple grids, wherein, in response to an initial gesture occurring while one grid is active, the application activates another grid to process a further initial gesture.

[0191] Considering the currently active application and its set of operations, the detected wrist position can optionally be combined with one or more detected poses or primitive gestures and passed to the currently active application. The currently active application then performs one or more operations based at least on the detected wrist position. The application uses the detected position to select a mesh element of the currently active mesh and then performs an operation associated with that mesh element. The performed operation may depend on the detected pose or primitive gesture.

[0192] The grid representing the operation set and its mapping to the detection location may not be visible to the user. However, although invisible to the user, one or more grids may be conceptually or virtually known or understood by the user, allowing the user to select locations that they know will be interpreted as associated with the corresponding elements of the grid. In some applications, the grid may be visible on the display.

[0193] As a specific illustrative use case, consider an audio playback application that provides various operations, such as start, stop, and pause playback; skip forward; skip backward; increase volume; and decrease volume. (Reference) Figure 6 This illustrates an exemplary grid representing a conceptual spatial mapping of operations to locations. In this example, start, stop, and pause are mapped to the central area 600. Jump forward is mapped to the right area 602; jump backward is mapped to the left area 604. Volume up is mapped to the top area 606; volume down is mapped to the bottom area 608. This mapping of operations to locations can be explained to the user in some way, allowing the user to conceptually understand where the different areas are located and the operations associated with them. On a device with a display, a visualization of this mapping can be presented on the display.

[0194] When the audio application is an active application used to receive input, the responding device can provide input to the audio application based on a detected initial gesture. In one exemplary embodiment, in response to the detection of the initial gesture, the input can activate a spatial (virtual or "real") menu, such as... Figure 6 As shown. When the index finger is raised, wrist movement (i.e., finger sliding) allows for additional operations. For example, sliding the finger to the right or left allows for selecting a jump forward or backward operation. Sliding the finger up or down allows for selecting a volume up or volume down operation. Rotating the wrist by rotating the forearm around the elbow can be used to provide additional input. When the user stops holding the index finger in the raised position, the processing of wrist movements can end.

[0195] For example, a user can place their wrist in a mid-position relative to their body, raise their index finger, hold it in the raised position, move their wrist to the right, and then release the raised finger. In response to this combination of wrist position and gesture, the responding device directs wrist position and gesture information to the audio application. The audio application then activates the grid 800 and processes the wrist position and gesture information as a command from the user to jump to the next song on the playlist. The audio application begins playing the next song in the playlist. After the user stops raising their index finger, the audio application stops processing input about the grid and continues playing the selected song. As another example, a simulated raise can be used to gradually increase or decrease the volume of the sound played on the device.

[0196] The application can maintain status information based on information received from the user interface device. (See now for reference.) Figure 7The application can be in a first listening state 700, in which a first primitive gesture, such as raising and holding the index finger, can activate monitoring of further input associated with multiple grid elements of the first grid. The first primitive gesture invokes a transition 702 to a second state 704. In this second state, wrist movement detected after the first primitive gesture can be used to determine a spatial position, which is then mapped to one or more grid elements, referred to as selected grid elements. Specific states (e.g., 706, 708) can be assigned to each grid element. In some embodiments, when the wrist is held in the spatial position corresponding to the selected grid element, for example, in state 708, an operation associated with the selected grid element can be performed, as shown in 710, for example, increasing the volume. The application can remain in this state until the user moves the wrist position so that it is no longer on that grid element, or until the grid or application is no longer activated.

[0197] In some implementations, when the wrist is in a spatial position corresponding to the selected grid element, for example, in state 706, the operation associated with the selected grid element can be processed after another second initial gesture occurs, as shown in transition 712 to state 714. Therefore, by the user placing their wrist in a mid-vertical position (corresponding to the audio application in state 700), performing an index finger lift (activating the command grid and transitioning to state 704), sliding the wrist upwards (causing a transition to state 706), and then performing another action, such as N... o Rotation (resulting in a transition to state 714) places the application in state 714. After processing an operation in state 714, a transition to another state can occur, such as a transition to state 704 or 700.

[0198] By switching between states, such as between states 706 and 708, wrist movement when the index finger is raised allows the application to switch its selection between different grid elements. For example, after the user initially moves their wrist to the right and the application initially selects a grid element to the right, the user can move their wrist to the left, causing the application to select a grid element to the left. A third primitive gesture, such as ending the index finger raising and holding, can terminate the processing of wrist movement information associated with the spatial element, resulting in a transition back to the initial state 700.

[0199] Third Example

[0200] As another illustrative use case, consider a controller for a micro, unmanned aerial vehicle (i.e., a drone) that has a controller that responds to commands received wirelessly via radio frequencies to control the movement of the aircraft. For such an aircraft, various control operations can be provided to control its movement, such as guiding it to move up or down, hover, rotate left or right, move forward or backward, or some combination thereof. This type of control is often referred to as “low-level” control. A wake word can be used to initiate this low-level control via a user interface device.

[0201] refer to Figure 8 The diagram illustrates an exemplary grid representing a spatial mapping from wrist position to operation. In this example, the absence of wrist movement is mapped to a hover operation, as shown in the central region 800. Rightward wrist movement mapped to the right center region 802 is mapped to moving the drone to the right. Leftward wrist movement mapped to the left center region 804 is mapped to moving the drone to the left. Upward wrist movement mapped to the top center region 806 is mapped to moving the drone upward or forward. Downward wrist movement mapped to the bottom center region 808 is mapped to moving the drone downward or backward. The distinction between upward / forward or downward / backward can be based on detected finger movement. For example, a hold combined with an upward or downward wrist position can signal upward or downward movement of the drone; however, the absence of a hold results in forward or backward movement of the drone. The top left and top right corners, as well as the bottom left and bottom right corners, can indicate combined wrist movements or can be mapped to another command. Clockwise or counterclockwise wrist movement N o Rotation ( Figure 8 (Not shown in the image) is also a wrist movement, which can be mapped to the corresponding rotation of the drone in the corresponding direction and amount.

[0202] Examples of using geolocation

[0203] In some implementations, the response device may use information based on data received from one or more user interface devices to determine the geographic location. In some implementations, additional sensor information or additional contextual information, or both, received from one or more other user interface devices, one or more other sensors, or other information sources, or any combination thereof, may be used to assist in the determination.

[0204] Some exemplary scenarios in which a geographic location can be determined include, but are not limited to, the following.

[0205] A person might point in a specific direction, such as north or south, and may need information about objects in that general direction. A person might point to a street corner or building, possibly wanting information about that corner or building, or simply to identify it. A person might point in a direction and may want to set a point on a map. For example, a person might point to their feet to set a point on a map or to transmit their location to another device. A person might point in a direction to identify objects, such as spotting a deer or other animal while hunting, or communicating during a game or other activity.

[0206] Multiple people may point to roughly the same location or the same object. Multiple users may want to cooperate to select or identify the same object. Multiple users may want to cooperate to select or identify multiple objects. One or more objects may be moving. For example, multiple firefighters may point to the same window or other locations on a building. Multiple people on a street may point to the same vehicle on the street. For example, several police officers may point to the same vehicle. Multiple people in a location may point to roughly the same location. Multiple people in a city park may point to an object. For example, security personnel may point to the same person, an open door, or an object. Multiple people in search and rescue missions may identify objects or people or animals to be rescued, or hazards or obstacles.

[0207] The selected geographic location can be used to control communications. For example, a person can point to another person and initiate communication directly with that person. A person can select a group of people, for example, by pointing at them individually or by using a gesture to select them as a group, to communicate with that group. A person can use a gesture (e.g., pointing to the sky) to initiate communication with a predetermined group of individuals. Similarly, communication with or control of an object, or both, can be performed by pointing to a single object, selecting a group of objects within a set of objects, or selecting a predetermined group of objects.

[0208] A geographic location can be defined as an absolute location or a relative location. Examples of absolute location include, but are not limited to, coordinates in a map system, GPS coordinates, or longitude and latitude at various resolutions, or any combination thereof. Examples of relative location include, but are not limited to, reference points and directions, or reference points and directions and distances, multiple reference points and their corresponding directions, or multiple reference points and their corresponding directions and distances, or any combination thereof.

[0209] Geographic location can be determined based on input from sensor data from sensors associated with at least one person. Additional sensor data from sensors associated with one or more other people, or one or more other objects, or a combination thereof, can be used to determine geographic location. Examples of such sensors include at least any one or more of the following: a position sensor for locating body parts (described below), a geomagnetic sensor, a compass, a magnetometer, an accelerometer, a gyroscope, an altimeter, an eye-tracking device, an encoder providing information indicating changes in position, a camera providing image data or motion video, a reference to a known landmark, or other sensors that can provide indications of direction from a person, or any combination of one or more of these. Sensor data can be interpreted, for example, any of a person's multiple raw hand gestures, an expected action, an arm angle or other indication of relative vertical position, direction or orientation, distance or height, or any combination thereof. Geographic location can be determined based on such sensor data received from sensors associated with one or more people.

[0210] In some implementations, contextual data can be used to assist in determining geographic location. Such contextual data may include information about one or more objects (e.g., geographic location), speech, sound, or text data, image data from one or more cameras, map data or user input about a map, or any combination of these or other contextual information.

[0211] In some implementations, additional information, such as contextual data, may be used to provide input for selecting, controlling, initiating, or performing additional actions related to or based on a geographic location selected as an object. For example, voice, gestures, or other data may be used as modifiers. For instance, a firefighter pointing to a building may communicate via voice, such as, “Two people are near the second-floor window!” For instance, voice, gestures, or other data may be used to annotate a selected object, with different gestures used to provide different labels for that object, such as “safe” or “dangerous.” For instance, voice, gestures, or other data may be used to activate commands about an object, activating different actions in different ways. This activation may be similar to a “click,” but can be obtained from any other data received by the user.

[0212] Figure 9A This is an illustrative data flow diagram of a system for determining geographical location. Figure 9AIn this embodiment, response device 952 receives received data 950-1, 950-2, ..., hereinafter referred to as "950" from one or more user interface devices. In some embodiments, data 950-2 received from a second user interface device (if available) can be transmitted via its own response device 954, which in turn provides information to response device 952. In some embodiments, response devices 952 and 954 can provide information to yet another device to determine a geographical location.

[0213] To aid in determining geographic location based on data received from the user interface device, the received data 950 includes at least sensor data from one or more sensors on the user interface device. Examples of such sensors include at least any one or more of the following: a position sensor for locating body parts (described below), a geomagnetic sensor, a compass, a magnetometer, an accelerometer, a gyroscope, an altimeter, an eye-tracking device, an encoder providing information indicating changes in location, image data from a camera, a reference to known landmarks, or a sound sensor (e.g., one or more microphones), or other sensors that can provide directional indications from a person, or any combination of two or more of these. In some embodiments, the response device 952 may use the data 950 received from the user interface device to determine hand posture and wrist movement, from which information such as pointing, pointing direction, and pointing with movement can be determined. Other primitive gestures may also be detected to indicate selection of an object or an action to be performed relative to the selected object.

[0214] The response device 952 may also receive other context data 956. The other context data 956 may be any one or more of the following types of data.

[0215] Context data can include information about one or more objects, such as their geographic location. For example, a database containing information about objects can be accessed.

[0216] Context data 956 may include speech or other audio data from one or more microphones. Context data 956 may include text or other data derived from speech or other audio signals identified from such audio data.

[0217] Context data 956 may include image or video data from one or more cameras. Video data may provide depth information provided by time-of-flight processing. Similarly, other time-of-flight sensors may provide such depth information. Context data 956 may include results of object recognition or surface detection based on this image data.

[0218] Contextual data 956 may include user input about the map or other map-related information. This user input may include, for example, location, region, or curve relative to the map. Points of interest data associated with the map of the area surrounding the user may also serve as contextual data 956.

[0219] Now turning Figure 9B and Figure 9C This shows an example use case where a geographic location can be determined. Figure 9B The diagram shows area 980, with three individuals 982, 984, and 986 depicted at different geographical locations within area 980. This diagram is not drawn to scale but is intended to show a certain perspective or depth from which each individual stands on the ground within area 980. Each individual is pointing towards object 988 on the ground. It should be understood that the individual providing this information can be located anywhere, not necessarily on the ground, as these techniques can be applied to individuals on the ground, in vehicles, underwater (e.g., diving), or in the air (e.g., in air vehicles, such as airplanes or even skydiving). Each individual is shown pointing towards object 988, thus providing a different direction from their corresponding position within area 980. The angle of the individual's arm may also be provided. Based on the information available from each individual (i.e., their position and direction) and the optional arm angle, the geographical location of object 988 can be determined. The position and direction of each individual can be used to determine a two-dimensional location on the ground corresponding to area 980. The height of the object can be further determined using the arm angle from each individual.

[0220] exist Figure 9C The geographical location of object 988 on an exemplary display is shown in a two-dimensional manner. For example, one of the persons may have a device with a display that provides a graphical indication 990 of the determined geographical location of object 988. The display may optionally include the geographical location of the person, for example, P1. The display may optionally include the geographical locations of other persons, for example, P2 and P3.

[0221] return Figure 9A Taking into account the received data 950 and any context data 956, the responding device can determine the geographic location 960. There are many methods for calculating geographic location. Below are a few examples.

[0222] Various types of processing techniques can be used to determine geographic location from sensor data and optional contextual data. In some implementations, various forms of triangulation can be applied using directions from multiple reference points. In some implementations, the input can allow the user to specify distances associated with the determined location and orientation. In some implementations, techniques from computational geometry can be applied, such as variations of the Simultaneous Localization and Mapping (SLAM) algorithm for calculating the agent's location, where the geographic location to be determined is considered as the location of the agent to be calculated. In some implementations, various forms of odometry can be used, where changes in location over time are determined based on sensor data. In some implementations, visual odometry can be performed on images from a camera. Any combination of these implementations can also be used.

[0223] Sensor data used to determine geographic location may also have timestamps or other time stamps. In implementations using sensor data associated with multiple users or devices, time information can be used to eliminate or weight sensor data. For example, newer data may be emphasized or used, while less recent data may be de-emphasized or discarded.

[0224] Determining a geographic location can also include data indicating certainty, validity, credibility, or the defined geographic area. For example, such data can be based on the quantity of inputs used to determine the geographic location. For instance, if the data is obtained from a large number of sources, such as a large number of individuals pointing to the same general area, the data might mean that the object is more certain to be the expected object or that the object is more meaningful, or both. As another example, the data can be based on the resolution of the available data used to calculate the geographic location.

[0225] For example, considering a first person's first geographic location and first direction, a first vector is defined; considering a second person's second geographic location and second direction, a second vector is defined; the intersection of the first and second vectors provides the location. The first and second geographic locations, as well as the first and second directions, can be based on the time when an initial gesture is detected from the corresponding user interface worn by the person—for example, a finger lift occurring when the person points in a direction. In response to the initial gesture, data available before and after this time can be used as one or more of the geographic location or direction. Therefore, two people can point to an object, and their information can be used to determine the object's geographic location.

[0226] As another example, considering a person's geographical location, orientation and distance can be derived from data received by the user interface device. Orientation can be derived from signals from a geomagnetic sensor or from left-right wrist movements, or both. Distance can be based on up-down wrist movements. The vector to the determined geographical location can be displayed to the user, for example, on a map on a display or in a heads-up display, to provide feedback on the determined geographical location. A response device can update the location in response to further input from the user.

[0227] As another example, considering a person's geographic location, orientation and distance can be derived from data received by the user interface device and data received by another device carried by the person. For example, a person wearing the user interface device may have a mobile device (e.g., a telephone) with its own geomagnetic sensor or other sensors, or sensors from which a first orientation can be derived. A person wearing the user interface device may have a heads-up display including eye tracking. A person wearing the user interface device may wear a second user interface device on a different arm than the first. A first orientation can be calculated such that it is intended to represent, for example, the direction of the person's gaze or the direction in which the person's head, chest, or fingers are pointing. The user interface device also provides one or more signals from which a second orientation can be derived, for example, an orientation based on a geomagnetic sensor indicating the direction in which a finger is pointing. The vectors representing the first and second orientations, projected onto a common plane, and the difference between the origins of the first and second orientations can be used to determine the intersection point, thereby providing the distance from one or both origins to the geographic location. The vector to the determined geographic location can be displayed to the user, for example, on a map on a display or in a heads-up display, to provide feedback to the user about the determined geographic location. The response device can update the position in response to further input from the user.

[0228] As another example, considering a person's geographic location, orientation, and the selected object, the time-of-flight data of the selected object can be used to determine distance information. Based on the person's geographic location, their orientation, and distance to the selected object, a vector to the determined geographic location can be displayed to the user, for example, on a map on a display or in a heads-up display, to provide feedback to the user about the determined geographic location. A response device can update the location in response to further input from the user.

[0229] In some implementations, contextual data can be used to refine the calculated geographic location. For example, if the user says "bank," a location corresponding to "bank" can be selected. As another example, a region selected on a map can be used to ensure that the determined geographic location is within that region.

[0230] Geographic location can be used in a variety of ways, including but not limited to the following, and may include combinations of two or more of the following: Geographic location can be used as input to an application, for example, to identify or select an object. Various actions can be performed by the application in conjunction with the selected object, non-limiting examples of which include invoking a camera to track the object or retrieving information about the object from a database. Geographic location can be data passed to another user or another machine. When combined with data based on signals received from one or more user interface devices to select an action to be performed, geographic location can be considered contextual information. Geographic location can be provided to the user as feedback, for example, by displaying a point on a map.

[0231] Considering the geographic location 960, in addition to providing feedback to the user on a map display, this information can be used for various operations. For example, in some implementations, this point of interest data can be provided as feedback data along with any defined geographic location 960. Such geographic location 960, along with other information derived from data received from the user interface device, can be provided to the application for processing. One or more objects near the defined location can be detected, tracked, or selected. Operations can be performed in conjunction with the detected or selected objects, such as invoking a camera to track the objects or obtaining information about the objects from a database.

[0232] In one exemplary implementation, consider an application where a person might view objects in three-dimensional space. The viewer might wear a virtual reality display or other immersive display. Alternatively, the viewer might be viewing a traditional two-dimensional display. This user interface device allows the user to intuitively select and perform actions regarding the displayed objects, even when grasping other objects.

[0233] Now for reference Figure 9D In the case of displaying a virtual scene, the computer has a 3D model of the scene being rendered, represented by scene data 903. This scene data 903 can be processed using a scene data processor 900 to generate data representing a 3D surface 902. Taking into account the position 904 of the user interface device or user in the scene and a vector 906 emanating from that position, the intersection of this vector and the 3D surface can be determined. This intersection point corresponds to an object in the 3D scene. An "object selector" 908 can process the 3D surface, vector, and position information to select an object 910.

[0234] Input 911, based on signals from the user interface device, can be used to calculate vector 906 using vector calculation module 918. In some applications, vectors can be provided by the application, such as the viewpoint direction of an avatar in a video game. Similarly, input 911, based on signals from the user interface device, can be used to calculate position 904 using position calculation module 920.

[0235] Taking into account the selected object 910 and the input 911 based on the signal from the user interface device, an operation 914 can be performed on the selected object. The operation selection module 916 receives the input 911 and the selected object 910, and determines the selected operation 914 to be performed.

[0236] For each object, the application can have a set of possible inputs and corresponding actions. The input received by the user interface device can be one or more possible inputs associated with the selected object. In response to these inputs, the application can select the operation to perform on the selected object.

[0237] As another exemplary implementation, consider an application in which a person views real-world objects through a heads-up or augmented reality display. This user interface device allows the user to intuitively select and perform actions regarding real-world objects, even when grasping other objects. This can be combined with the above. Figure 9B A similar approach described herein enables the selection of real-world objects and the performance of operations on those selected objects within an application. However, in this case, there is no man-made scene data 903 based on which the surfaces of the objects have already been defined in three dimensions. Instead, the responding device can receive video information representing the real-world location from headphones worn by the user, such as a camera. In this case, the video information is processed to extract surface information to provide scene data 903.

[0238] User interface example is as follows: Figure 10 As shown. For example, this interface can be provided in a heads-up display. In this image, the user can see the wrist-worn user interface device 1000 and a line 1002 shown as emanating from the wrist-worn device to the selected object 1004. A two-dimensional top view 1006 can be displayed. Other objects that can be selected are indicated at 1008 and 1010. Any of these objects can be selected based on the position of the wrist, and then another basic gesture (e.g., raising the index finger) can specify the operation to be performed on the selected object, such as retrieving information about that object. In this example, objects 1004, 1008, and 1010 could be a drone or other objects that can be controlled after selection using the above-described technology.

[0239] In another exemplary embodiment, consider an application in which a person views real-world objects from a two-dimensional third-person perspective, such as... Figure 11As shown. For example, the display may present a two-dimensional aeronautical map 1100, in which the user's location or the location of a selected object is also marked, for example, by an arrow 1102. The use of arrows allows for the reflection of geomagnetically sensed directions in the display. The user's geographic location may be input by the user or determined using Global Positioning System (GPS) data or other data suitable for determining such a geographic location, which may be provided by a mobile phone or other mobile device that the user may carry. GPS or other geographic location information may be provided by a responding device, such as a tablet computer that the user is using.

[0240] In this user interface, the responding device can respond to wrist movements, such as raising and lowering the arm, to change the distance between a location 1102 along a line on the map and another geographic location 1104 far from the user. The responding device can use geolocation sensors on the user interface device or direction or heading information from another source to change the direction of the line. For example, the user can move their wrist left and right to change the orientation information, and can move their wrist up and down to change the distance.

[0241] After manipulating a point on the map by heading and distance from position 1102, an object at that point can be selected using another primitive gesture, and an operation can be performed on the selected object. For example, lifting can be used to initiate an operation on the selected object.

[0242] As another example, taking into account the initial heading and distance, the responding device can respond to a finger swipe by defining an arc based on the initial position 1104 at the start of the swipe and the ending position 1106 at the end of the swipe. In response to another primitive gesture, the responding device can allow selection of one or more objects within a region defined by points 1102, 1104, and 1106.

[0243] In the examples above, multiple user interface devices may be used in some applications. In some implementations, the multiple user interface devices house different parts of the user interface device, such as bioelectrical sensors and position sensors, and can be worn on the same body part. In some implementations, the multiple user interface devices are worn on different body parts. For example, a first user interface device may be worn on the top of the wrist of the right arm, while a second user interface device may be worn on the top of the wrist of the left arm. When using multiple user interface devices, these devices can generally be identical, for example, having the same shape factor, sensors, and feedback mechanisms, but they can also be different. For example, a user interface device worn on the top of the wrist can be configured differently from a user interface device worn on the ankle. In implementations using multiple user interface devices, the selection of an application or the selection of an operation to be performed by the application, or both, can also be based on which user interface device is providing data based on bioelectrical or positional signals.

[0244] Several examples of human interaction with input and response devices provided by user interface devices have now been described, and further details of exemplary implementations of signal processing and raw gesture detection will now be provided.

[0245] To detect various primitive gestures, the conditioning signals from the sensors are further processed to detect patterns in those signals corresponding to the primitive gestures. In some embodiments, signals received from the bioelectric sensor can be processed to classify a portion of the bioelectric signal occurring at a given moment as data indicating the hand posture at that moment. Signals received from the position sensor can be processed to classify a portion of the position signal at a given moment as data indicating the position of the wrist tip at that moment. In some embodiments, signals from the bioelectric sensor and the position sensor are processed together to obtain data indicating the hand posture and wrist tip position at a given moment. Data regarding hand posture and wrist position can be processed to allow for the detection of primitive gestures. In some embodiments, features are extracted from the bioelectric and position signals, and these features can be directly processed into data indicating the primitive gesture.

[0246] The signal from the sensor can be in the form of a digital sample sequence, with the amplitude and sampling rate of each sample. To perform various signal processing operations on the signal, the system defines a series of sample "windows". A sample window can be specified by a. a specific moment in the signal and b. a period of time or multiple samples surrounding that moment. For each window, features are computed using digital signal processing techniques. Any given window or set of windows can be labeled as representing the original gesture or not.

[0247] For example, a window could be 200 milliseconds long, with adjacent windows overlapping by 50%. This means that the samples forming the last 100 milliseconds of the first window are also the samples forming the first 100 milliseconds of the second window immediately following the first window in the window sequence. Different window structures can be used for data collection and labeling, rather than for real-time signal processing and classification. For example, for real-time classification, a 200-millisecond sliding window with no overlap could be used. The size of the window and the overlap of adjacent windows can be adjusted for specific user interface designs and applications.

[0248] Considering the sample windows of the conditioned sensor signals, each window is processed to extract features, which are then used to define and distinguish different patterns associated with the type of information that is to be extracted from the sensor signals, such as wrist position or movement or hand posture or other features.

[0249] The processing module within the response device performing feature extraction or gesture detection can implement a device interface to receive sampled sensor data from a user interface device and an application programming interface to provide features or other data to other applications within the response device. These features can be output as a data stream, indicating the features associated with each window.

[0250] Features that can be obtained from a signal include, for example, features obtained from the time-domain representation of the processed signal and features obtained from the frequency-domain representation of the processed signal. Features that can be obtained from the time-domain representation of a signal include, but are not limited to, Hudgin's time-domain features.

[0251] Tables I through V list exemplary features that can be calculated for each window. Not all listed features are required; there may also be some unlisted features available:

[0252] Table I - Basic Feature Set of a Single Channel

[0253]

[0254] Table II - Extended Feature Set for Single Channels (Supplement to Table I)

[0255]

[0256] Table III: Exemplary features derived by comparing two channels:

[0257]

[0258] Table IV: Features extracted from IMU channels:

[0259]

[0260] Table V: Characteristic combinations of signals from all gyroscope channels or from all accelerometer channels:

[0261]

[0262] With enough data, the signal itself can be fed into deep learning models, which will learn its own features.

[0263] Model training

[0264] One technique for processing signals to detect patterns is to use a parametric model, training the model's parameters based on signals corresponding to known patterns. These signals are "labeled," meaning there is data associated with the signal indicating a known pattern it contains. A computer system processes this labeled data through a process called "training" to generate a model. The resulting model can then detect whether a specific pattern appears in real-time signals from sensors.

[0265] One of the challenges in training such models is obtaining a large amount of labeled data, called the training set. The training set ideally includes samples with both positive and negative labels. Positively labeled samples are those labeled as representing known features of a signal. Negatively labeled samples are those known not to represent known features of a signal. Having such negatively labeled samples in the training set allows for training a model that can better distinguish between intentional and everyday activities.

[0266] There are multiple methods to obtain positively labeled samples for the training set. In one implementation, multiple user interface devices can be equipped to receive signals from a computer. In response to these signals, the user interface devices generate prompts to a person wearing the device to perform a specified action. The user then performs the specified action. The user interface devices send sensor signals to the computer, which labels the signal based on the specified action.

[0267] More specifically, each sample window in the received signal is based on a selected motion marker. Considering a set of samples comprising several windows, the windows are labeled as corresponding to or not corresponding to a hand gesture or wrist position or original hand gesture, or any combination thereof. Considering a set of samples and a label, the window marking component identification begins. Sample windows that appeared before the window with the beginning are not labeled. For training purposes, the sample set in windows that were not labeled before the beginning was detected can be discarded.

[0268] Another exemplary method for collecting training data is to use an application running on a computing device that has a touchscreen or other device that can detect the actual hand and finger positions.

[0269] Model training is typically dedicated to a specific version of the sensor and the signal processing performed on the signals from those sensors. Multiple users with devices having the same configuration (i.e., the same sensor, sensor configuration, low-level signal processing, and feature extraction hardware) can provide data for training a model with that configuration. However, given the multiple labeled data windows that include their respective features, and optionally additional sensor data corresponding to these windows, conventional training algorithms can be applied to models that have been adapted to the specific type of signal being processed.

[0270] Training can be performed to create a training model for a single user using labeled data from a single user, or to create a training model for multiple users using labeled data from multiple users. Model training for a single user can be performed on a computing device accessible to the user interface device of that single user. For multiple users, the labeled data can be transferred to and stored on a server computer connected to each user's labeled data source via a computer network.

[0271] After training and testing the model, the trained model can be deployed. User devices (whether response devices or user interface devices) download the trained model from a server computer or from the computing device that trained the model.

[0272] By training such a model, several features that influence classification can emerge. For example, the energy ratio between the first and second channels of the first sensor has been found to be highly predictive for distinguishing thumb and index finger extensions. In some cases, this ratio can also be used to filter out other "false positives," such as wrist extension, little finger extension, etc. There is a dependency between sensor configuration and signal quality on one hand and the ability to distinguish between certain postures, positions, original gestures, or combinations thereof on the other hand (e.g., between thumb and index finger extensions in terms of energy ratio).

[0273] Several exemplary implementations have been described. Figure 12 An example of a general-purpose computing device is shown, which can be used to implement response devices or other computer systems in conjunction with such response devices. This is merely one example of a computer and is not intended to limit the scope or functionality of such a computer. Figure 12 As shown, the above system can be implemented in one or more computer programs that execute on one or more such computers.

[0274] Figure 12This is a block diagram of a general-purpose computer that uses a processing system to process computer program code. Computer programs on a general-purpose computer typically include an operating system and application programs. The operating system is a computer program that runs on the computer and manages the application programs' and the operating system's access to various computer resources. These resources typically include memory, storage devices, communication interfaces, input devices, and output devices.

[0275] Examples of such general-purpose computers include, but are not limited to, large computer systems (e.g., server computers, database computers, desktop computers, laptop and notebook computers) and mobile or handheld computing devices, such as tablet computers, handheld computers, smartphones, media players, personal data assistants, audio or video recorders, or wearable computing devices.

[0276] refer to Figure 12 An exemplary computer 1200 includes a processing system comprising at least one processing unit 1202 and a memory 1204. The computer may have multiple processing units 1202 and multiple means for implementing the memory 1204. The processing unit 1202 may include one or more processing cores (not shown) that operate independently of each other. Additional cooperative processing units (e.g., a graphics processing unit 1220) may also be present in the computer. The memory 1204 may include volatile means (e.g., dynamic random access memory (DRAM) or other random access storage means), non-volatile means (e.g., read-only memory, flash memory, etc.), or some combination of both, and optionally include any memory available in the processing means. Other memory, such as dedicated memory or registers, may also reside in the processing unit. Such memory is configured in Figure 12 The image is shown by dashed line 1204. Computer 1200 may include additional memory (removable or non-removable), including but not limited to magnetically or optically recorded disks or tapes. This additional memory... Figure 12 The image is shown in the form of a removable memory 1208 and a non-removable memory 1210. Figure 12 The various components in the system are typically interconnected via interconnection mechanisms (e.g., one or more buses 1230).

[0277] Computer storage media are any media that can store and retrieve data by a computer at an addressable physical storage location. Computer storage media include volatile and non-volatile storage devices, as well as removable and non-removable storage devices. Memory 1204, removable memory 1208, and non-removable memory 1210 are examples of computer storage media. Some examples of computer storage media are RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical or magneto-optical recording storage devices, magnetic tape, magnetic tape, disk storage, or other magnetic storage devices. Computer storage media and communication media are mutually exclusive media categories.

[0278] Computer 1200 may also include a communication connection 1212 that allows the computer to communicate with other devices via a communication medium. The communication medium typically transmits computer program code, data structures, program modules, or other data via wired or wireless means by propagating modulated data signals (e.g., carrier waves or other transmission mechanisms) over a material. The term "modulated data signal" refers to a signal having one or more characteristics set or altered in a manner that encodes information in the signal, thereby changing the configuration or state of a signal receiving device. By way of example, and not limitation, the communication medium includes wired media, such as wired networks or direct-line connections, and wireless media includes any non-wired communication medium that allows signals to propagate, such as acoustic, electromagnetic, electrical, optical, infrared, radio frequency, and other signals. The communication connection 1212 is a device, such as a network interface or a radio transmitter, that engages with the communication medium to send and receive data from signals propagated through the communication medium.

[0279] A communication connection may include one or more radio transmitters for telephone communication over a cellular telephone network or wireless communication interfaces for wireless connection to a computer network. For example, a computer may have cellular connections, Wi-Fi connections, Bluetooth connections, and other connections. This connection supports communication with other devices, such as voice or data communication.

[0280] Computer 1200 may have various input devices 1214, such as various pointer (whether single-pointer or multi-pointer) devices, such as a mouse, writing tablet and pen, touchpad and other touch-based input devices, stylus, image input devices (e.g., still and moving cameras), and audio input devices (e.g., microphones). Computer may also have various output devices 1216, such as a monitor, speakers, printer, etc. These devices are well known in the art and do not need to be discussed in detail here.

[0281] Various memories 1210, communication connections 1212, output devices 1216, and input devices 1214 may be integrated within the computer casing or connected via various input / output interface devices on the computer. In this case, reference numerals 1210, 1212, 1214, and 1216 may indicate the interface connected to the device or the device itself, as appropriate.

[0282] A computer's operating system typically includes computer programs, often called drivers, for managing access to various memory devices 1210, communication connections 1212, output devices 1216, and input devices 1214. This access may include managing the input and output of these devices. In the case of a communication connection, the operating system may also include one or more computer programs for implementing communication protocols for transferring information between the computer and the devices via the communication connection 1212.

[0283] Each component (also referred to as a "module" or "engine," etc.) of a computer system running on one or more computers can be implemented as computer program code processed by the processing system of one or more computers. Computer program code includes computer-executable instructions or computer-interpreted instructions, such as program modules, which are processed by the computer's processing system. These instructions define routines, programs, objects, components, data structures, etc., and when processed by the processing system, instruct the processing system to perform operations on data or configure the processor or computer to implement various components or data structures in computer memory. Data structures are defined in the computer program, specifying how data is organized in computer memory, such as in memory devices or storage devices, so that the computer's processing system can access, manipulate, and store the data.

[0284] It should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific embodiments described above. The specific embodiments described above are disclosed as examples only.

Claims

1. A system comprising: A biopotential sensor, configured to be worn by a person on their wrist and to sense the biopotential at the wrist; A wrist position sensor, the wrist position sensor being configured to be worn on the wrist and providing wrist position data indicating the position of the wrist; as well as processor; The system is configured as follows: The direction vector of the person's intention is determined based at least on the wrist position data from the wrist position sensor, wherein the wrist position sensor includes an accelerometer that provides the acceleration of the wrist, a gyroscope that provides the angular velocity of the wrist, and a geomagnetic sensor that provides the heading of the wrist, and the wrist position data is three-dimensional attitude information that fuses the acceleration, the angular velocity, and the heading. Based at least on the biopotential sensed by the biopotential sensor, a gesture indicating an intention to make a selection is detected; Based on the direction vector and the detected gesture, a target is selected, the selected target being indicated by the intersection of the direction vector determined when the gesture indicating the intention to make the selection is detected and a 3D surface, wherein the 3D surface corresponds to the geometric outer boundary of an object in a three-dimensional scene.

2. The system according to claim 1, wherein, The target is a geographic location.

3. The system according to claim 1, wherein, The target is the object.

4. The system according to claim 3, wherein, The system is further configured as follows: After selecting the object based on the direction vector and the detected gesture, the object is controlled; The steps for controlling the object include: One or more commands are determined based at least on the wrist position data from the wrist position sensor and the biopotential sensed by the biopotential sensor; Send one or more commands to the object.

5. The system according to claim 4, wherein, The system is further configured as follows: Initiate a mode capable of executing steps that control the object; The mode that initiates steps capable of controlling the object includes: A wake word gesture is detected based on one or more of (i) wrist position data from the wrist position sensor and (ii) biopotentials sensed by the biopotential sensor, the wake word gesture indicating an intention to initiate the mode capable of performing steps to control the object.

6. The system according to claim 1, wherein, The step of determining the direction vector of the person's intention is performed based on the following two factors: The wrist position data from the wrist position sensor; as well as Data indicating the location of the person in question.

7. The system according to claim 6, wherein, The data indicating the location of the person is Global Positioning System (GPS) data.

8. The system according to claim 6, wherein, The data indicating the location of the person is based on triangulation using directions from multiple reference points.

9. The system according to claim 6, wherein, The step of selecting the target is based at least on: The direction vector is based on (i) wrist position data from the wrist position sensor and (ii) data indicating the position of the person; The detected gesture, which is based at least on the biopotential sensed by the biopotential sensor; as well as Context data, which includes information related to the target.

10. The system according to claim 9, wherein, The context data includes the location of the target.

11. The system according to claim 9, wherein, The context data includes image data or video data.

12. A system comprising: A biopotential sensor, configured to be worn by a person on their wrist and to sense the biopotential at the wrist; A wrist position sensor, the wrist position sensor being configured to be worn on the wrist and providing wrist position data indicating the position of the wrist; as well as The system is configured to provide output data to one or more processors, the output data being based on biopotentials sensed by the biopotential sensor and wrist position data sensed by the wrist position sensor, the one or more processors being configured to: The direction vector of the person's intention is determined based at least on the wrist position data from the wrist position sensor, wherein the wrist position sensor includes an accelerometer that provides the acceleration of the wrist, a gyroscope that provides the angular velocity of the wrist, and a geomagnetic sensor that provides the heading of the wrist, and the wrist position data is three-dimensional attitude information that fuses the acceleration, the angular velocity, and the heading. Based at least on the biopotential sensed by the biopotential sensor, a gesture indicating an intention to make a selection is detected; Based on the direction vector and the detected gesture, a target is selected, the selected target being indicated by the intersection of the direction vector determined when the gesture indicating the intention to make the selection is detected and a 3D surface, wherein the 3D surface corresponds to the geometric outer boundary of an object in a three-dimensional scene.

13. The system according to claim 12, wherein, The target is a geographic location.

14. The system according to claim 12, wherein, The target is the object.

15. The system according to claim 14, wherein, The one or more processors are further configured to: After selecting the object based on the direction vector and the detected gesture, the object is controlled; The steps for controlling the object include: One or more commands are determined based at least on the wrist position data from the wrist position sensor and the biopotential sensed by the biopotential sensor; Send one or more commands to the object.

16. The system according to claim 15, wherein, The one or more processors are further configured to: Initiate a mode capable of executing steps that control the object; The mode that initiates steps capable of controlling the object includes: A wake word gesture is detected based on one or more of (i) wrist position data from the wrist position sensor and (ii) biopotentials sensed by the biopotential sensor, the wake word gesture indicating an intention to initiate the mode capable of performing steps to control the object.

17. The system according to claim 12, wherein, The step of determining the direction vector of the person's intention is performed based on the following two factors: The wrist position data from the wrist position sensor; as well as Data indicating the location of the person in question.

18. The system according to claim 17, wherein, The data indicating the location of the person is Global Positioning System (GPS) data.

19. The system according to claim 17, wherein, The data indicating the location of the person is based on triangulation using directions from multiple reference points.

20. The system according to claim 17, wherein, The step of selecting the target is based at least on: The direction vector is based on (i) wrist position data from the wrist position sensor and (ii) data indicating the position of the person; The detected gesture, which is based at least on the biopotential sensed by the biopotential sensor; as well as Context data, which includes information related to the target.

21. The system according to claim 20, wherein, The context data includes the location of the target.

22. The system according to claim 20, wherein, The context data includes image data or video data.