Pose tracking system

By combining event cameras and frame-based cameras, leveraging the high frame rate of event cameras and the color analysis of frame-based cameras, the accuracy and efficiency issues of pose recognition in traditional systems are solved, achieving fast and effective pose recognition suitable for computer-generated human-computer interaction in real-world environments.

CN114365187BActive Publication Date: 2026-04-24APPLE INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APPLE INC
Filing Date
2020-09-01
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional frame-based camera systems lack accuracy and efficiency in hand movement and pose recognition. Event camera data processing is complex and difficult to recognize poses quickly and effectively.

Method used

By combining event cameras and frame-based cameras, user poses are identified through subset processing and block grouping techniques of event camera data. The high frame rate of event cameras and color analysis of frame-based cameras are used to reduce noise, identify regions of interest, and track entity paths.

Benefits of technology

It improves the accuracy and efficiency of posture recognition, enabling rapid and effective identification of user postures, and is suitable for computer-generated human-computer interaction in real-world environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114365187B_ABST
    Figure CN114365187B_ABST
Patent Text Reader

Abstract

Various implementations disclosed herein include devices, systems, and methods to identify poses based on event camera data and based on frame-based camera data (e.g., for a CGR environment). In some implementations, at an electronic device with a processor, event camera data corresponding to light (e.g., IR light) reflected from a physical environment and received at an event camera is obtained. In some implementations, frame-based camera data corresponding to light (e.g., visible light) reflected from the physical environment and received at a frame-based camera is obtained. In some implementations, a subset of the event camera data is identified based on the frame-based camera data, and a pose (e.g., of a person in the physical environment) is identified based on the subset of the event camera data. In some implementations, a path (e.g., of a hand) is identified by tracking groupings of event camera event blocks in the subset of the event camera data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in general to systems, methods, and apparatus for object tracking or pose recognition using camera data. Background Technology

[0002] Hand movements and other poses can involve relatively rapid motion. Therefore, pose recognition systems using images from traditional frame-based cameras may lack accuracy and efficiency or be otherwise limited by the relatively slow frame rates used by such cameras. Event cameras, on the other hand, capture events pixel-by-pixel and can provide the ability to capture data about user movement significantly faster than many frame-based cameras. However, pose recognition systems using event camera data may face challenges in attempting to identify a user's pose because, in a given physical environment, there may be numerous events unrelated to the pose being tracked. The need to analyze and interpret the potentially large amounts of event data from the physical environment can significantly reduce the ability of such event camera-based pose recognition systems to identify poses quickly, efficiently, and accurately. Summary of the Invention

[0003] The various embodiments disclosed herein include apparatuses, systems, and methods for performing event-camera-based pose recognition using a subset of event camera data. In some embodiments, pose is identified based on a subset of event camera data. In some embodiments, this subset of event camera data is identified by a frame-based camera having an FOV that overlaps with the event camera's field of view (FOV). In some embodiments, the event camera FOV and the frame-based camera FOV are temporally or spatially related. In one embodiment, at an electronic device with a processor, the event camera data is generated from light (e.g., infrared (IR) light) reflected from the physical environment and received at the event camera. In some embodiments, the frame-based camera data is generated from light (e.g., visible light) reflected from the physical environment and received at the frame-based camera. In some embodiments, the frame-based camera data is used to identify regions of interest (e.g., bounding boxes) for analysis by the event camera. For example, the frame-based camera data can be used to remove background based on considerations in event camera data analysis. In some implementations, identifying a subset of event camera data involves selecting only event camera data corresponding to a portion of the user's physical environment (e.g., their hand). In some implementations, the event camera is tuned to IR light to reduce noise.

[0004] The various specific embodiments disclosed herein include devices, systems, and methods for identifying paths (e.g., hands) by tracking groupings of blocks of event camera events. Each block of event camera events can be an area having a predetermined number of events occurring at a given time. Grouping of blocks can be performed based on grouping radius, distance, or time, for example, including blocks within a given 3D distance of another block occurring at an instant. Tracking a group of blocks when the grouping of blocks recurs at different points in time may be easier and more accurate than attempting to track individual events moving over time, such as related events associated with the tip of the thumb at different points in time.

[0005] The various embodiments disclosed herein include apparatuses, systems, and methods for acquiring event camera data corresponding to light reflected from a physical environment (e.g., IR or a first wavelength range). In some embodiments, event blocks associated with multiple times are identified based on blocking criteria. For example, each block may be a region having a predetermined number of events occurring at a given time or over a time period of predetermined length. In some embodiments, an entity (e.g., a hand) is identified at each of the multiple times. In some embodiments, a path is determined by tracking the position of the entity at the multiple times. In some embodiments, the entity at each of the multiple times comprises a subset of the event block associated with the corresponding time. In some embodiments, a subset of the block is identified based on grouping criteria. In some embodiments, tracking the position of the entity at each of the multiple times provides a path as the entity moves over time, such as the path of a hand as it moves over time. In some embodiments, frame-based camera data corresponding to light reflected from a physical environment is received, and event camera data is identified based on the frame-based camera data.

[0006] In some embodiments, the electronic device is an electronic device in which event camera data is processed. In some embodiments, the electronic device is the same electronic device that includes the event camera (e.g., a laptop computer). In some embodiments, the electronic device is a different electronic device that receives event data from an electronic device that has an event camera (e.g., a server that receives event data from a laptop computer). In some embodiments, a single electronic device including a processor has an event camera, an IR light source, and a frame-based camera (e.g., a laptop computer). In some embodiments, the event camera, the IR light source, and the frame-based camera are located on more than one electronic device.

[0007] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Attached Figure Description

[0008] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.

[0009] Figure 1 It is a block diagram based on some specific implementations of exemplary systems.

[0010] Figure 2 It is a block diagram of an exemplary controller based on some specific implementations.

[0011] Figure 3 It is a block diagram of an exemplary electronic device based on some specific implementations.

[0012] Figures 4 to 6 This is a block diagram of an exemplary configuration of an electronic device, based on some specific implementations, including an event camera to track the path of an entity (e.g., a hand).

[0013] Figures 7A to 7B This is a diagram showing examples of events detected by an event camera at a given time according to some specific implementations, and a grouping of blocks of the detected events.

[0014] Figure 8 This is a flowchart illustrating examples of pose recognition methods based on some specific implementations.

[0015] Figure 9 These are block diagrams and exemplary circuit diagrams of pixel sensors for exemplary event cameras, based on some specific implementations.

[0016] Figure 10 This is a flowchart illustrating an example of a method for identifying the path of an entity based on some specific implementations.

[0017] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation

[0018] Numerous details have been described to provide a thorough understanding of the exemplary embodiments shown in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will recognize that other effective aspects or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein. Although Figures 1 to 3 Exemplary specific implementations relating to electronic devices are described, including but not limited to watches and other wearable electronic devices, mobile devices, laptop computers, desktop computers, HMDs, gaming devices, home automation devices, accessory devices, and other devices that include or use image capture devices.

[0019] Figure 1 This is a block diagram of an exemplary operating environment 100 according to some specific implementations. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for brevity and to avoid obscuring further relevant aspects of the exemplary implementations disclosed herein. For this purpose, as a non-limiting example, operating environment 100 includes a controller 110 and an electronic device (e.g., a laptop computer) 120, one or both of which may be located within a physical environment 105. A physical environment refers to the physical world that people can sense and / or interact with without the assistance of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with physical environments through senses such as sight, touch, hearing, taste, and smell.

[0020] In some implementations, controller 110 is configured to manage and coordinate the user's computer-generated reality (CGR) environment. In some implementations, controller 110 includes a suitable combination of software, firmware, or hardware. See below for reference. Figure 2 The controller 110 is described in more detail. In some implementations, the controller 110 is a computing device located locally or remotely relative to the physical environment 105.

[0021] In one example, controller 110 is a local server located within physical environment 105. In another example, controller 110 is a remote server (e.g., a cloud server, central server, etc.) located outside physical environment 105. In some implementations, controller 110 is communicatively coupled to a corresponding electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.).

[0022] In some implementations, controller 110 and the corresponding electronic device (e.g., 120) are configured to present the CGR environment to the user together.

[0023] In some implementations, electronic device 120 is configured to present a CGR environment to a user. In some implementations, electronic device 120 includes a suitable combination of software, firmware, or hardware. See below for reference. Figure 3 The electronic device 120 is described in more detail. In some specific implementations, the functions corresponding to the controller 110 are provided by or combined with the electronic device 120, for example, in the case of the electronic device being used as a stand-alone unit.

[0024] According to some specific implementations, when a user appears within physical environment 105, electronic device 120 presents a CGR environment to the user. A CGR environment refers to a fully or partially simulated environment sensed and / or interacted with by a person via an electronic system. In a CGR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the CGR environment are adjusted in a manner consistent with at least one physical law. For example, a CGR system may detect head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of characteristics of virtual objects in the CGR environment may be done in response to a representation of physical motion (e.g., a voice command).

[0025] Humans can use any of their senses to sense and / or interact with CGR objects, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, audio objects can enable audio transparency, which selectively introduces ambient sound from the physical environment, with or without computer-generated audio. In some CGR environments, a person can sense and / or interact only with audio objects.

[0026] Examples of CGR include virtual reality and mixed reality. A virtual reality (VR) environment is a simulated environment designed to provide one or more senses entirely based on computer-generated sensory input. A VR environment includes virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of a person's presence within the computer-generated environment and / or through the simulation of a subgroup of physical movements of a person within the computer-generated environment.

[0027] Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments are simulated environments designed to incorporate sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between a purely physical environment as one end and a virtual reality environment as the other end, but not including either end.

[0028] In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present the MR environment can track position and / or orientation relative to the physical environment, enabling virtual objects to interact with real objects (i.e., physical objects or representations of them from the physical environment). For example, the system can cause movement so that virtual trees appear stationary relative to the physical ground.

[0029] Examples of mixed reality include augmented reality and augmented virtual. An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment using the system. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via images or videos of the physical environment and perceive the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto a physical environment, such as as a hologram or on a physical surface, so that a person can use the system to perceive the virtual objects superimposed on the physical environment.

[0030] Augmented reality environments also refer to simulated environments where the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, a system can transform images from one or more sensors to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. Alternatively, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portion is a representative but not realistic version of the original captured image. Furthermore, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.

[0031] Augmented virtual (AV) environments are simulated environments in which a virtual or computer-generated environment is combined with one or more sensory inputs from a physical environment. Sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could adopt shadows that correspond to the sun's position within the physical environment.

[0032] Many different types of electronic systems enable people to sense and / or interact with a variety of CGR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.

[0033] Figure 2This is a block diagram of an example controller 110 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, or similar type interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0034] In some embodiments, the one or more communication buses 204 include circuitry for communication between interconnecting system components and control system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image capture devices or other sensors, one or more displays, etc.

[0035] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (CGRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures or subsets thereof, including optional operating system 230, computer-generated reality (CGR) module 240, and attitude recognition unit 250.

[0036] The operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks.

[0037] In some implementations, the CGR module 240 is configured to create, edit, present, or experience a CGR environment. The CGR module 240 is configured to present virtual content that will be used as part of a CGR environment to one or more users. For example, users can view and otherwise experience a CGR-based user interface that allows users to select, place, move, and otherwise present the CGR environment based on the location of the virtual content, such as through gestures, voice commands, or input devices.

[0038] In some embodiments, the pose recognition unit 250 is configured to use event camera data to determine the path of an entity or for pose recognition. In some embodiments, the pose recognition unit 250 uses frame-based camera data. In some embodiments, the pose recognition unit 250 is used as a functional I / O device or as part of a CGR environment for one or more users. Although these modules and units are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of these modules and units may reside in a separate computing device.

[0039] also, Figure 2 This is used more as a functional description of various features present in a specific implementation, and differs from the structural diagrams of the specific implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, or firmware chosen for that particular implementation.

[0040] Figure 3This is a block diagram of an example of an electronic device 120 according to some specific embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C, or similar interfaces), one or more programming (e.g., I / O) interfaces 310, one or more displays 312, one or more internal or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0041] In some implementations, one or more communication buses 304 include circuitry for interconnecting and communicating between control system components. In some implementations, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, magnetometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0042] In some embodiments, one or more displays 312 are configured to present a CGR environment to a user. In some embodiments, one or more displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), or similar display types. In some embodiments, one or more displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, an electronic device may include a single display. Alternatively, an electronic device may include a display for each of the user's eyes.

[0043] Memory 320 includes high-speed random access memory, such as DRAM, CGRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optionally an operating system 330, a computer-generated reality (CGR) module 340, and a gesture recognition unit 350.

[0044] Operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks.

[0045] In some implementations, the CGR module 340 is configured to create, edit, present, or experience a CGR environment. The CGR module 340 is configured to present virtual content that will be used as part of a CGR environment to one or more users. For example, users can view and otherwise experience a CGR-based user interface that allows users to select, place, move, and otherwise present the CGR environment based on the location of the virtual content, such as through gestures, voice commands, or input devices.

[0046] In some embodiments, the pose recognition unit 350 is configured to use event camera data to determine the path of an entity or for pose recognition. In some embodiments, the pose recognition unit 350 uses frame-based camera data. In some embodiments, the pose recognition unit 350 is used as a functional I / O device or as part of a CGR environment for one or more users. Although these modules and units are shown residing on a single device (e.g., electronics 120), it should be understood that in other embodiments, any combination of these modules and units may reside in a separate computing device.

[0047] also, Figure 3 This is used more as a functional description of various features present in a specific implementation, and differs from the structural diagrams of the specific implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 3Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, or firmware chosen for that particular implementation.

[0048] Figures 4 to 6 This is a block diagram illustrating an exemplary configuration of an electronic device, according to some specific implementations, including an event camera to track the path of an entity (e.g., a hand). In some specific implementations, the entity's path is used for pose recognition.

[0049] like Figure 4 As shown, the electronic device 420 in physical environment 405 includes a processor 424 and a tracking array 422. In some specific embodiments, the tracking array 422 includes an event camera 422b and a frame-based camera 422c. Figure 4 As shown, the tracking array 422 includes a light source (e.g., an IR LED), an event camera 422b, and a frame-based camera 422c (e.g., RGB-D). In some embodiments, the tracking array 422 is operatively mounted on an electronics device 420. In some embodiments, the tracking array 422 is mounted on the electronics device 420, located below, to the side, or above the display 426.

[0050] In some implementations, the event camera 422b is always on due to its low power consumption. In some implementations, the event camera 422b operates at a very high frame rate when monitoring the physical environment 405. In some implementations, the event camera 422b detects events at a frame rate faster than 1000 Hz. Therefore, the electronic device 420 can smoothly track a person 410 or the hand 412 of person 410 using the event camera 422b in the physical environment 405.

[0051] like Figure 4As shown, the event camera data corresponds to light reflected from the physical environment 405 and received at the event camera 422b. In some embodiments, the event camera data corresponds to light reflected and received at the event camera 422b in a first wavelength band. In some embodiments, the event camera 422b is tuned (e.g., using a filtered light source or a filter at the light source or event camera) to the first wavelength band. In some embodiments, the event camera 422b uses IR light. In some embodiments, the event camera 422b uses IR light to reduce noise in the event camera data (e.g., interference with visible light). In some embodiments, the IR light can be focused within a preset operating range (e.g., distance, height, or width). In some embodiments, the IR light is provided by a light source 422a. In some embodiments, the light source 422a can be an infrared (IR) light source. In some embodiments, the light source 422a can be a near-infrared (NIR) light source. In some specific implementations, the light source 422a is an IR light source, including an IR LED, a diffuser optics, or a focusing optics.

[0052] In some embodiments, the field of view (FOV) of the event camera 422b and the emitting cone of the light source 422a are spatially correlated or overlapped. In some embodiments, the light source 422a is configured as an IR light source, which can be varied in size, intensity, or maneuverability. For example, because the intensity of reflected IR light is based on distance (e.g., distance...). 2 In some specific implementations, the intensity of the IR source 422a is adjusted to control or optimize the number of events or reduce noise.

[0053] like Figure 4 As shown, frame-based camera data corresponds to light (e.g., visible light) reflected from physical environment 405 and received at frame-based camera 422c. In some embodiments, frame-based camera data corresponds to light (e.g., visible light) reflected and received at frame-based camera 422c in a second wavelength band. In some embodiments, the second wavelength band used by frame-based camera 422c differs from the first wavelength band used by event camera 422b. In some embodiments, the field of view (FOV) of physical environment 405 for event camera 422b and the FOV of physical environment 405 for frame-based camera 422c are spatially correlated. In various embodiments, the FOV of event camera 422b and the FOV of frame-based camera 422c overlap. In some embodiments, frame-based camera data and event camera data are temporally and spatially correlated.

[0054] like Figure 5As shown, a subset of event camera data from event camera 422b in physical environment 405 can be determined. In some embodiments, this subset of event camera data is based on frame-based camera data. In some embodiments, frame-based camera 422c identifies a bounding box 510 in physical environment 405, within which event camera 422b analyzes event camera data. For example, frame-based camera 422c identifies and then reduces or removes background in physical environment 405 based on considerations in the analysis by event camera 422b. In some embodiments, frame-based camera 422c selects foreground portions of physical environment 405 for considerations in the event camera analysis. For example, frame-based camera 422c identifies person 410 as foreground in physical environment 405. Alternatively, for example, frame-based camera 422c identifies hand 412 as foreground and person 410 as background in physical environment 405. In some implementations, the frame-based camera 422c uses color analysis of the frame-based camera data to identify a subset of the event camera data. For example, the color of a hand can be identified in the physical environment to determine the region of interest for the event camera 422b. In some implementations, the frame-based camera 422c uses a low frame rate. In some implementations, the frame-based camera 422c uses a frame rate of 2 to 10 frames per second.

[0055] like Figure 6 As shown, in some implementations, event blocks associated with multiple times within a subset of event data (e.g., bounding box 510) are identified based on block formation criteria. In some implementations, each block in the event block is a region in the physical environment 405, which includes events associated with multiple times T1, T2, ..., T... N Each time point in the event block represents a predetermined number of events occurring at a given time (e.g., a small time interval). In some implementations, the number of events determines the size or resolution of the event block. In some implementations, a larger number of events per block reduces the resolution of the event camera 422b analysis. In some implementations, the number of event blocks at each time point across multiple time periods is the same. In some implementations, the number of event blocks at each time point across multiple time periods is different. For example, there might be 1000 event blocks associated with time T1 and 1400 event blocks associated with a second time T2, etc. In some implementations, the size of the event block is based on the object being tracked (e.g., a finger, hand, body, etc.).

[0056] In some implementations, an event block at each of the multiple times is used to determine the entity. In some implementations, the entity at each of the multiple times (e.g., a particle, a hand) comprises a subset of the event block associated with the corresponding time. Figure 6 As shown, at multiple times T1, T2, ..., TN At each time point in the process, at least one entity 600 in the physical environment is identified (e.g., a hand 412 representing a person 410). In some implementations, a subset of event blocks that identify entity 600 at each time point in multiple time periods is based on a grouping radius, a distance factor between multiple such event blocks, a smoothing factor, a time factor, etc.

[0057] In some implementations, path 610 is determined by tracking the position of entity 600 at multiple time points. In some implementations, event camera 422b detects an initial subset of event blocks and an additional subset of event blocks that move through the physical environment 405 to the terminating event block to determine path 610. Figure 6 As shown, tracking the position (e.g., center) of entity 600 as it moves over time provides path 610. In some implementations, entity 600 is created (e.g., by moving), tracked, and then discarded (e.g., after a period of inactivity) to form path 610. In some implementations, event camera 422b identifies a cluster of intensity changes that causes an event, starting at a first time, and tracks the cluster of intensity changes that propagate through the physical environment 405 at multiple intermediate times until no additional cluster of intensity changes appears at a second time.

[0058] In some embodiments, the path 610 of an entity is drawn and displayed as path 610' on the display 426 of the electronic device 420. In some embodiments, the posture of a person 410 (e.g., a predetermined movement of a body part of the person 410) is identified based on path 610. Therefore, the posture of a person 410 (e.g., hand 412) is identified based on event camera data. In some embodiments, the posture of a person 410 is identified using paths of two or more entities tracked in the event camera data. In some embodiments, moving fingers, moving body parts, or moving limbs (e.g., arm only) are tracked to determine the posture of a person 410. In some embodiments, moving body parts of two or more individuals are tracked to determine the posture of a person. In some embodiments, path 610' is part of a CGR environment. In some embodiments, path 610' determines the posture as an operator command of the electronic device or CGR environment.

[0059] like Figure 7A As shown, event camera 422b detects multiple events 730 within the bounding box. In some specific implementations, the events 730 detected by event camera 422b include positive events (e.g., generated by the leading edge of a hand or finger) and negative events (e.g., generated by the trailing edge of a hand or finger). For example, event camera 422b detects multiple events 730 within bounding box 510.

[0060] like Figure 7BAs shown, in some specific implementations, a subset of block 732 of event 730 is grouped to identify entity 700. Entity 700 can be tracked at multiple times to determine a path. Figure 7B In the given time T X A portion of path 710 of entity 700 is shown. In some implementations, such as those using machine learning (ML), the path is used to determine the pose. In some implementations, a first neural network is trained to recognize the pose when the input path is given.

[0061] In some implementations, the entity is a hand, and a first configuration of the hand determines a first hand pose at the start of the path, and a second configuration of the hand determines a second hand pose at the end of the path, and the pose includes the first hand pose, the path, and the second hand pose. In some implementations, the second hand pose differs from the first hand pose. In some implementations, an ML or second neural network can be trained to recognize or output the hand pose when a given configuration of the hand is entered or input.

[0062] In some embodiments, the electronic device 420 is an electronic device in which event camera data is processed. In some embodiments, the electronic device 420 is the same electronic device (e.g., a laptop computer) that includes an event camera 422b. In some embodiments, the electronic device 420 is a different electronic device (e.g., a server that receives event data from a laptop computer) that receives event data from an electronic device having an event camera 422b. In some embodiments, a single electronic device including a processor has an event camera 422b, a light source 422a, and a frame-based camera 422c. In some embodiments, the event camera 422b, the light source 422a, and the frame-based camera 422c are located on more than one electronic device.

[0063] Figure 8 This is a flowchart illustrating an exemplary method for gesture recognition according to some specific implementations. In some specific implementations, gesture recognition is performed by an electronic device (e.g., Figures 1 to 3 Method 800 is executed by a controller 110 or electronic device 120. Method 800 may be executed at a mobile device, HMD, desktop computer, laptop computer, server device, or by multiple devices communicating with each other. In some embodiments, method 800 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, method 800 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0064] At block 810, method 800 obtains event camera data corresponding to light reflected from the physical environment and received at the event camera. In some embodiments, method 800 obtains event camera data corresponding to light in a first wavelength band (e.g., IR light) reflected from the physical environment and received at the event camera. In some embodiments, the event camera detects IR light to reduce noise in the event camera data. In some embodiments, event camera data is generated based on changes in light intensity detected at the pixel sensor of the event camera. In some embodiments, an event detected by the event camera is triggered by a change in light intensity exceeding a comparator threshold.

[0065] At block 820, method 800 obtains frame-based camera data corresponding to light reflected from the physical environment and received at a frame-based camera. In some embodiments, method 800 obtains frame-based camera data corresponding to light (e.g., visible light) in a second wavelength band reflected from the physical environment and received at a frame-based camera. In some embodiments, the second wavelength band is different from the first wavelength band. In some embodiments, the FOV of the event camera and the FOV of the frame-based camera are temporally or spatially related.

[0066] At box 830, method 800 identifies a subset of event camera data based on frame-based camera data. In some implementations, at box 830, the frame-based camera identifies a region of interest in the physical environment within which the event camera can analyze event data. For example, at box 830, the frame-based camera removes the background from the physical environment based on considerations in the event camera analysis. For example, at box 830, the frame-based camera selects a foreground portion of the physical environment (e.g., a person's hand) for considerations in the event camera analysis. In some implementations, a subset of event camera data is identified through color analysis of the frame-based camera data.

[0067] At box 840, method 800 identifies a person's pose (e.g., a body part or a predetermined movement of an object held by a person's body part) based on a subset of event camera data. In some implementations, at box 840, method 800 identifies the pose by determining a path by tracking clusters of events in a subset of event camera data over time. In some implementations, at box 840, the path is determined by grouping blocks of event camera events (e.g., hands) that are tracked at multiple different time points, where each block has a predetermined number of events at each of the multiple time points.

[0068] In some implementations, at box 840, a machine learning (ML) or ML model can be used to analyze entity paths to identify or detect poses. In some implementations, a first neural network is trained to identify moving poses given the entity's path as input. In some implementations, movement includes the entity's pose (e.g., hand position, such as a fist, a hand with one or more fingers extended, a flat hand, or hand orientation such as a vertical or horizontal flat hand). In some implementations, the pose includes the entity's pose at the start or end of the entity path. In some implementations, the pose includes one or more poses of the entity along the entity path. In various implementations, the ML model can be, but is not limited to, a deep neural network (DNN), an encoder / decoder neural network, a convolutional neural network (CNN), or a generative adversarial neural network (GANN). In some implementations, an event camera used for pose detection is designated for automated processing (e.g., machine viewing instead of human viewing) to address privacy concerns. In some implementations, images from a frame-based camera can be combined with event camera data for ML. In some implementations, frame-based camera data includes image data that is first analyzed to detect entities such as body parts, and then that body part is analyzed for movement (e.g., pose recognition). In some implementations, a hand is detected first, and then a pose is detected during hand movement.

[0069] In some implementations, the event camera data, or a subset of the event camera data, corresponding to reflected IR light from the physical environment is the result of the IR light source in the physical environment. In some implementations, the IR light can be focused within a preset operating range (e.g., distance, height, or width).

[0070] Figure 9 These are block diagrams and exemplary circuit diagrams of pixel sensors for exemplary event cameras, based on some specific implementations. Figure 9 As shown, the pixel sensor 915 can be positioned relative to an electronic device (e.g., by arranging the pixel sensor 915 into a 2D matrix 910 of rows and columns). Figure 1 The electronic device 120 is set on the event camera at a known location. Figure 9 In the example, each pixel sensor in pixel sensor 915 is associated with an address identifier defined by a row of values ​​and a column of values.

[0071] Figure 9 An example circuit diagram of a circuit 920 suitable for implementing the pixel sensor 915 is also shown. Figure 9In the example, circuit 920 includes a photodiode 921, a resistor 923, a capacitor 925, a capacitor 927, a switch 929, a comparator 931, and an event compiler 932. In operation, a voltage is generated across photodiode 921 that is proportional to the intensity of light incident on the pixel sensor. Capacitor 925 is parallel to photodiode 921, and therefore the voltage across capacitor 925 is the same as the voltage across photodiode 921.

[0072] In circuit 920, switch 929 is positioned between capacitors 925 and 927. Therefore, when switch 929 is in the closed position, the voltage across capacitor 927 is the same as the voltage across capacitor 925 and photodiode 921. When switch 929 is in the open position, the voltage across capacitor 927 is fixed at the previous voltage across capacitor 927 when switch 929 was last in the closed position. Comparator 931 receives and compares the voltages of capacitors 925 and 927 on its input side. If the difference between the voltage on capacitor 925 and the voltage on capacitor 927 exceeds a threshold amount (“comparator threshold”), an electrical response (e.g., voltage) indicating the intensity of light incident on the pixel sensor is present on the output side of comparator 931. Otherwise, there is no electrical response on the output side of comparator 931.

[0073] When an electrical response is present on the output side of comparator 931, switch 929 turns to the closed position, and event compiler 932 receives the electrical response. Upon receiving the electrical response, event compiler 932 generates a pixel event and adds the pixel event with information indicating the electrical response (e.g., the value or polarity of the electrical response). In one embodiment, event compiler 932 also fills the pixel event with one or more of the following: timestamp information corresponding to the time point of pixel event generation and an address identifier corresponding to the specific pixel sensor that generated the pixel event.

[0074] Event cameras typically include multiple pixel sensors, such as pixel sensor 915, each of which outputs a pixel event in response to detecting a change in light intensity exceeding a comparison threshold. When aggregated, the pixel events output by the multiple pixel sensors form a pixel event stream output by the event camera.

[0075] Figure 10 This is a flowchart illustrating an exemplary method for identifying the path of an entity according to some specific implementations. In some specific implementations, this is achieved by an electronic device (e.g., Figures 1 to 3Method 1000 is executed by a controller 110 or electronic device 120. Method 1000 may be executed at a mobile device, HMD, desktop computer, laptop computer, server device, or by multiple devices communicating with each other. In some embodiments, method 1000 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, method 1000 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).

[0076] At block 1010, method 1000 obtains event camera data corresponding to light reflected from the physical environment and received at the event camera. In some embodiments, method 1000 obtains event camera data corresponding to light reflected from the physical environment and received at the event camera that is in a first wavelength band. In some embodiments, the event camera is specifically tuned to IR light to reduce noise (e.g., ambient light) in the event camera data.

[0077] In block 1020, method 1000 identifies event blocks associated with multiple times based on block criterion. In some implementations, each block in the event block is a region having a predetermined number of events occurring at a given time (e.g., a small time interval). In some implementations, the number of events determines the size or resolution of the event block. In some implementations, fewer events per block at each of the multiple times provide a smaller block or a finer resolution depiction of the physical environment. Similarly, more events per block at each of the multiple times provide a larger block or a lower resolution depiction of the physical environment. In some implementations, the number of event blocks at each of the multiple times is different. For example, there might be 100 event blocks associated with time T1 and 90 event blocks associated with a second time T2, etc.

[0078] At box 1030, method 1000 identifies entities (e.g., particles) at each of the plurality of times. In some implementations, the entities at each of the plurality of times comprise a subset of event blocks associated with the corresponding time. In the example described at box 1020, the entity could be 90 event blocks out of 100 event blocks associated with time T1, and so on. In some implementations, the subset of event blocks at each of the plurality of times is based on a grouping radius, a distance factor between the plurality of event blocks, a smoothing factor, or a time factor. In some implementations, clusters of individual event blocks are determined to represent one hand of a person in a physical environment.

[0079] At box 1040, method 1000 determines a path by tracking the position of an entity at multiple time points. In some implementations, a selected portion or position of the entity is tracked as it moves over time to determine the path. In some implementations, the path identifies a user's pose (e.g., a predetermined movement of a body part) based on event camera data. In some implementations, the pose is a hand gesture. In some implementations, at box 1040 (see box 840), ML or an ML model can be used to analyze the entity path to identify or detect the pose.

[0080] In some embodiments, at block 1010, the event camera data is a subset of event camera data identified by selectively including only event camera data corresponding to a part of the physical environment (e.g., a user's body part or hand). In some embodiments, at block 1010, method 1000 obtains frame-based camera data corresponding to light received at a frame-based camera and identifies a subset of event camera data based on this frame-based camera data. In some embodiments, the frame-based camera data corresponds to event camera data spatially or temporally. In some embodiments, a subset of event camera data is identified through color analysis of the frame-based camera data. In some embodiments, the frame-based camera data is used to remove the background of the physical environment based on factors considered in the event camera analysis.

[0081] In some embodiments, an event camera receiving data corresponding to IR light reflected from the physical environment is located at an event camera included in a first electronic device including a processor. In some embodiments, a second electronic device separate from the first electronic device includes the event camera. In some embodiments, at block 1010, the second electronic device receives IR light reflected from the physical environment and transmits it to the first electronic device. In some embodiments, the first electronic device is located where processing is performed to implement method 1000.

[0082] This document sets forth numerous specific details to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter can be practiced without these specific details. In other instances, methods, apparatus, or systems known to a person of ordinary skill have not been described in detail so as not to obscure the claimed subject matter.

[0083] Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as “processing,” “calculating,” “computing,” “determining,” and “identifying” refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, which manipulate or convert data representing physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0084] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific implementations of this subject. The teachings contained herein can be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.

[0085] Specific implementations of the methods disclosed herein can be performed in the operation of such a computing device. The order of the boxes presented in the above examples can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel.

[0086] The use of “applies to” or “configured to” in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Similarly, the use of “based on” implies openness and inclusivity, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated. The headings, lists, and numbering included in this document are for illustrative purposes only and are not intended to be restrictive.

[0087] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.

[0088] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising” as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0089] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.

[0090] The foregoing description and summary of the present invention should be understood as illustrative and exemplary in every respect, and not restrictive, and the scope of the invention disclosed herein is determined not only by the detailed description of the illustrative specific embodiments, but also by the full extent permitted by patent law. It should be understood that the specific embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention.

Claims

1. A method for pose recognition, the method comprising: In electronic devices that have processors; Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Obtain frame-based camera data corresponding to the light reflected from the physical environment and received at the frame-based camera; Based on the frame-based camera data, a region of interest is identified in the physical environment, and a subset of event camera data is identified within the region of interest, wherein the subset of event camera data is identified based on the identification of an event camera data corresponding to a part of a person; as well as Pose identification is based on the subset of the event camera data.

2. The method of claim 1, wherein the event camera uses a first optical wavelength band, and the frame-based camera uses a different second optical wavelength range.

3. The method according to any one of claims 1 to 2, further comprising illuminating the physical environment with an IR light source, wherein the event camera uses reflected IR light and the frame-based camera uses reflected visible light.

4. The method of claim 3, wherein the IR light source is variable in size, intensity, or maneuverability.

5. The method according to any one of claims 1 to 2, wherein the event camera uses reflected IR light within the intensity range.

6. The method according to any one of claims 1 to 2, wherein identifying the subset of the event camera data includes excluding event camera data corresponding to the background portion of the physical environment.

7. The method according to any one of claims 1 to 2, wherein identifying the subset of the event camera data includes color analysis of the frame-based camera data.

8. The method according to any one of claims 1 to 2, wherein identifying the subset of the event camera data includes temporally and spatially relating the field of view of the event camera to the field of view of the frame-based camera.

9. The method according to any one of claims 1 to 2, wherein the electronic device includes the event camera, or wherein the second electronic device includes the event camera, and the second electronic device is separate from the electronic device.

10. The method according to any one of claims 1 to 2, wherein the event camera data corresponds to a pixel event triggered based on a change in light intensity at the pixel sensor exceeding a comparator threshold, or wherein the event camera data excludes data corresponding to changes in visible light intensity (VL).

11. The method according to any one of claims 1 to 2, wherein the posture includes hand movement.

12. The method of any one of claims 1 to 2, wherein identifying the pose comprises evaluating the event camera data using a machine learning algorithm.

13. The method according to any one of claims 1 to 2, wherein recognizing the posture includes recognizing the path of one or more hands of a person in the physical environment.

14. A method for pose recognition, the method comprising: In electronic devices that have processors; Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Based on grouping criteria, event blocks associated with multiple times within the event camera data are identified; Identify entities at each of the plurality of times, wherein the entities at each of the plurality of times comprise a subset of the event blocks associated with the corresponding time; as well as The path is determined by tracking the location of the entities at the plurality of time points, wherein the entities comprising the subset of the event blocks at each of the plurality of time points are determined based on a grouping radius, a distance factor between the plurality of event blocks, a smoothing factor, or a time factor.

15. The method of claim 14, wherein tracking the position of the entity at the plurality of time points includes activating the entity, tracking the entity, and terminating the entity after an interval of inactivity.

16. The method of any one of claims 14 to 15, wherein the entity is one or more hands of a person in the physical environment, the method further comprising: Determine the first hand gesture at the starting point of the path of the entity; A second hand gesture is determined at the end of the path of the entity, wherein the gesture includes the first hand gesture, the path of the entity, and the second hand gesture, wherein the second hand gesture is different from the first hand gesture.

17. The method of any one of claims 14 to 15, wherein identifying the pose comprises evaluating the event camera data using a machine learning algorithm.

18. The method of any one of claims 14 to 15, further comprising using frame-based camera data corresponding to light reflected from the physical environment and received at a frame-based camera to identify a selected region of the physical environment.

19. The method of claim 18, wherein the event camera data includes light reflected from the selected area of ​​the physical environment.

20. The method of claim 18, wherein the event camera uses a first wavelength band and the frame-based camera uses a different second wavelength range.

21. The method of claim 18, further comprising illuminating the physical environment with an IR light source, wherein the event camera uses reflected IR light and the frame-based camera uses reflected visible light.

22. The method of claim 21, wherein the IR light is provided by the IR light source, the IR light source being variable in size, intensity, or direction, and wherein the event camera uses reflected IR light within the intensity range.

23. The method of claim 18, wherein the selected region corresponds to a portion of a person in the physical environment, and wherein the field of view of the event camera and the field of view of the frame-based camera are temporally and spatially correlated.

24. The method of any one of claims 14 to 15, wherein the electronic device includes the event camera, or wherein the second electronic device includes the event camera, and the second electronic device is separate from the electronic device.

25. The method according to any one of claims 14 to 15, wherein the event camera data corresponds to a pixel event triggered based on a change in light intensity at the pixel sensor exceeding a comparator threshold, or wherein the event camera data excludes data corresponding to changes in visible light intensity.

26. The method according to any one of claims 14 to 15, wherein the entity is a body part of a person in the physical environment, and the path includes the posture of the body part at the start point of the path or the end point of the path.

27. The method of any one of claims 14 to 15, wherein the weight of each block in the event block is related to the number of events, and wherein the events include positive events and negative events.

28. A system for pose recognition, comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the system to perform the following operations: Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Based on grouping criteria, event blocks associated with multiple times within the event camera data are identified; Identify entities at each of the plurality of times, wherein each entity at each of the plurality of times comprises a subset of the event blocks associated with the corresponding time; and The path is determined by tracking the location of the entities at the plurality of time points, wherein the entities comprising the subset of the event blocks at each of the plurality of time points are determined based on a grouping radius, a distance factor between the plurality of event blocks, a smoothing factor, or a time factor.

29. A non-transitory computer-readable storage medium storing computer-executable program instructions on a computer to perform the following operations: Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Based on grouping criteria, event blocks associated with multiple times within the event camera data are identified; Identify entities at each of the plurality of times, wherein the entities at each of the plurality of times comprise a subset of the event blocks associated with the corresponding time; as well as The path is determined by tracking the location of the entities at the plurality of time points, wherein the entities comprising the subset of the event blocks at each of the plurality of time points are determined based on a grouping radius, a distance factor between the plurality of event blocks, a smoothing factor, or a time factor.

30. A system for pose recognition, the system comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the system to perform the following operations: Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Obtain frame-based camera data corresponding to the light reflected from the physical environment and received at the frame-based camera; Based on the frame-based camera data, a region of interest (ROI) is identified in the physical environment; within the ROI, a subset of event camera data is identified, wherein the subset of event camera data is identified based on the identification of event camera data corresponding to a portion of a person; and Pose identification is based on the subset of the event camera data.

31. A non-transitory computer-readable storage medium storing computer-executable program instructions on a computer to perform the following operations: Event camera data corresponding to light reflected from the physical environment and received at the event camera is obtained, the event camera data being generated based on changes in light intensity detected at the pixel sensor of the event camera; Obtain frame-based camera data corresponding to the light reflected from the physical environment and received at the frame-based camera; Based on the frame-based camera data, a region of interest is identified in the physical environment, and a subset of event camera data is identified within the region of interest, wherein the subset of event camera data is identified based on the identification of an event camera data corresponding to a part of a person; as well as Pose identification is based on the subset of the event camera data.

Citation Information

Patent Citations

  • Method and Apparatus for Motion Recognition

    US20120257789A1