Human-machine interaction state sensing system and sensing method, human-machine interaction enhancement system

CN122111238BActive Publication Date: 2026-08-18SHENZHEN RUISHIZHIXIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610570762.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-08-18
Estimated Expiration
2046-04-27

AI Technical Summary

Technical Problem

该类方案通常依赖相干光干涉、散斑分析或微振动检测,对光源稳定性、佩戴位置及个体差异高度敏感,系统复杂度高、功耗大,不利于在消费级终端中长期运行

Benefits of technology

[0010]The human-computer interaction state perception system disclosed herein includes a processing component and a photosensitive component communicatively connected. Initially, the system operates in a first working mode. In this mode, the photosensitive component operates at a first power consumption and collects reflected light from a target area to form first optical change information. The processing component processes this first optical change information to determine whether movement has occurred in the target area (e.g., whether a user has made a vocal gesture). If a user has made a vocal gesture, the processing component switches the photosensitive component from the first working mode to a second working mode; otherwise, it maintains the photosensitive component in the first working mode. When the photosensitive component switches to the second working mode, a projection component begins operation, projecting an active light field with time-modulated characteristics onto the target area. The photosensitive component operates at a second power consumption and collects reflected light from the target area after being illuminated by the active light field to form second optical change information. The processing component determines the current human-computer interaction state based on this second optical change information. In response to determining that interaction exists, the processing component maintains the photosensitive component in the second working mode; otherwise, it switches the photosensitive component back to the first working mode. Therefore, the human-computer interaction state perception system disclosed herein employs hierarchical perception. It first uses a low-power motion-sensing component for initial perception, and then, after the processing component determines that motion has occurred in the target area, it controls the photosensitive component to switch to fine perception. When no interaction occurs, the system can maintain microwatt-level power consumption for an extended period. Furthermore, because the projection component projects an active light field with time-modulated characteristics onto the target area in the second operating mode, it offers higher bandwidth and lower latency compared to traditional frame-based structured light solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111238B_ABST
    Figure CN122111238B_ABST
Patent Text Reader

Abstract

The present disclosure provides a human-computer interaction state perception system and method, and a human-computer interaction enhancement system. The human-computer interaction state perception system provided by the present disclosure comprises a light sensing component, a projection component and a processing component. The light sensing component operates at a first power consumption to collect reflected light of a target area to form first optical change information in a first working mode, or operates at a second power consumption to collect reflected light of the target area after being irradiated by an active light field to form second optical change information in a second working mode, the first power consumption being less than the second power consumption. The projection component projects an active light field with time modulation characteristics to the target area in the second working mode. The processing component controls the light sensing component to switch from the first working mode to the second working mode according to the first optical change information, determines a current human-computer interaction state based on the second optical change information, and controls the light sensing component to remain in the second working mode or switch to the first working mode based on the human-computer interaction state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of optical sensing and human-computer interaction technology, and in particular to a human-computer interaction state perception system and perception method, and a human-computer interaction enhancement system. Background Technology

[0002] With the development of consumer electronics such as smartphones, wireless headphones, augmented reality (AR) devices, and smart glasses, voice interaction has gradually become a core human-computer interaction method. However, existing voice interaction systems mainly rely on microphones to collect audio signals, which has significant limitations in scenarios such as speaking in low voices or private interactions.

[0003] To address these issues, related technologies attempt to infer vocal intentions or speech content by optically detecting minute facial vibrations or muscle micro-movements. These solutions typically rely on coherent optical interferometry, speckle analysis, or micro-vibration detection, and are highly sensitive to light source stability, wearing position, and individual differences. They also suffer from high system complexity and power consumption, making them unsuitable for long-term operation in consumer-grade devices.

[0004] Therefore, there is an urgent need for a perception solution that can maintain the accuracy of structured light geometric measurement, significantly reduce power consumption, and be suitable for dynamic human-computer interaction scenarios. Summary of the Invention

[0005] Based on this, embodiments of this disclosure propose a human-computer interaction state perception system and perception method, a human-computer interaction enhancement system, and an electronic device.

[0006] In a first aspect, embodiments of this disclosure propose a human-computer interaction state perception system, including a photosensitive component, a projection component, and a processing component. The photosensitive component is configured to operate in a first working mode with a first power consumption and collect reflected light from a target area to form first optical change information, or to operate in a second working mode with a second power consumption and collect reflected light from the target area after being illuminated by an active light field to form second optical change information. The projection component is configured to project an active light field with time modulation characteristics onto the target area in the second working mode. The processing component is configured to control the switching of the photosensitive component from the first working mode to the second working mode based on the first optical change information; determine the current human-computer interaction state based on the second optical change information; and, in response to determining that the human-computer interaction state is one of interaction, control the photosensitive component to remain in the second working mode, otherwise control the photosensitive component to switch back to the first working mode; wherein the first power consumption is less than the second power consumption.

[0007] In a second aspect, embodiments of this disclosure propose a human-computer interaction enhancement system, including an interaction control module and a human-computer interaction state perception system described in any implementation of the first aspect. The interaction control module is configured to perform at least one of the following functions based on the current human-computer interaction state determined by the human-computer interaction state perception system described in any implementation of the first aspect: starting or stopping a speech recognition module, filtering out environmental noise, enhancing human voice, and providing control functions as a control signal.

[0008] Thirdly, this disclosure proposes a human-computer interaction state perception method, comprising: controlling a photosensitive component to switch from a first operating mode to a second operating mode based on first optical change information, and controlling a projection component to start operating, wherein the first optical change information is optical change information formed by the photosensitive component operating in the first operating mode and collecting reflected light from a target area. Based on second optical change information, determining the current human-computer interaction state, wherein the second optical change information is optical change information formed by the photosensitive component operating in the second operating mode and collecting reflected light from a target area after being irradiated by an active light field. In response to determining that the human-computer interaction state is one of interaction, controlling the photosensitive component to remain in the second operating mode; otherwise, controlling the photosensitive component to switch to the first operating mode, wherein the power consumption of the photosensitive component in the first operating mode is less than the power consumption in the second operating mode.

[0009] Fourthly, embodiments of this disclosure provide an electronic device including a memory and a processor. The memory stores computer-executable instructions, and the processor is configured to execute the stored instructions to implement a sensing method for a human-computer interaction state sensing system as described in any implementation of the third aspect.

[0010] The human-computer interaction state perception system disclosed herein includes a processing component and a photosensitive component communicatively connected. Initially, the system operates in a first working mode. In this mode, the photosensitive component operates at a first power consumption and collects reflected light from a target area to form first optical change information. The processing component processes this first optical change information to determine whether movement has occurred in the target area (e.g., whether a user has made a vocal gesture). If a user has made a vocal gesture, the processing component switches the photosensitive component from the first working mode to a second working mode; otherwise, it maintains the photosensitive component in the first working mode. When the photosensitive component switches to the second working mode, a projection component begins operation, projecting an active light field with time-modulated characteristics onto the target area. The photosensitive component operates at a second power consumption and collects reflected light from the target area after being illuminated by the active light field to form second optical change information. The processing component determines the current human-computer interaction state based on this second optical change information. In response to determining that interaction exists, the processing component maintains the photosensitive component in the second working mode; otherwise, it switches the photosensitive component back to the first working mode. Therefore, the human-computer interaction state perception system disclosed herein employs hierarchical perception. It first uses a low-power motion-sensing component for initial perception, and then, after the processing component determines that motion has occurred in the target area, it controls the photosensitive component to switch to fine perception. When no interaction occurs, the system can maintain microwatt-level power consumption for an extended period. Furthermore, because the projection component projects an active light field with time-modulated characteristics onto the target area in the second operating mode, it offers higher bandwidth and lower latency compared to traditional frame-based structured light solutions.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the structure of a human-computer interaction state perception system provided in an embodiment of the present disclosure; Figure 2a This is a schematic diagram of a human-computer interaction state perception system provided in this disclosure for use in a wireless headphone scenario; Figure 2b This is a schematic diagram of the human-computer interaction state perception system provided in this disclosure for use in smart glasses scenarios; Figure 3a A schematic diagram of a human-computer interaction state perception system for time-series encoded structured light provided in an embodiment of this disclosure; Figure 3bThis is a schematic diagram of a human-computer interaction state perception system for multi-channel time-series encoded structured light provided in an embodiment of the present disclosure. Figure 4 This is a schematic diagram of a human-computer interaction state perception system for pixel aggregation provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of a human-computer interaction state perception system for triangulation provided in an embodiment of this disclosure; Figure 6 A schematic diagram illustrating the high-pass / band-pass filtering of the original motion signal in the human-computer interaction state perception system provided in this embodiment of the disclosure; Figure 7 A schematic diagram of the entire signal chain for the human-computer interaction state perception system using vibration characteristics to perceive the interaction state, as provided in the embodiments of this disclosure. Figure 8 A schematic diagram illustrating the interaction confidence generated by the human-computer interaction enhancement system provided in this embodiment of the present disclosure for use in the audio processing link; Figure 9 A schematic diagram of a human-computer interaction enhancement system structure provided in this embodiment of the present disclosure; Figure 10 A flowchart of a human-computer interaction state perception method provided in this embodiment of the disclosure; Figure 11 A detailed flowchart of a human-computer interaction state perception method provided in this embodiment of the disclosure. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings. While various detailed descriptions are provided in these embodiments for ease of understanding, those skilled in the art should understand that these embodiments and their detailed descriptions are merely exemplary, and various modifications can be made to these embodiments without departing from the teachings of this disclosure. Similarly, for clarity and brevity, detailed descriptions of well-known functions and structures will be omitted in the following description. Furthermore, embodiments and features described in this disclosure can be combined with each other unless otherwise specified.

[0014] While traditional structured light 3D sensing technology is relatively mature in terms of geometric accuracy, it usually relies on high frame rate imaging, resulting in high power consumption and data bandwidth, making it difficult to meet the requirements of wearable devices and mobile terminals for battery life and real-time performance.

[0015] Based on the above, this disclosure provides a human-computer interaction state perception system. Through hierarchical perception, a low-power motion photosensitive component is first used for primary perception. After the processing component determines that motion has occurred in the target area, the photosensitive component is controlled to switch to fine perception. When no interaction occurs, the system can maintain a power consumption of microwatts for a long time.

[0016] This disclosure provides a human-computer interaction state perception system, such as Figure 1 As shown, the human-computer interaction state perception system includes a photosensitive component 11, a projection component 12, and a processing component 13. The photosensitive component 11 is configured to operate in a first working mode with a first power consumption and collect reflected light from a target area 14 to form first optical change information, or operate in a second working mode with a second power consumption and collect reflected light from the target area 14 after being illuminated by an active light field to form second optical change information. The projection component 12 is configured to project an active light field with time modulation characteristics onto the target area 14 in the second working mode. The processing component 13 is configured to control the switching of the photosensitive component 11 from the first working mode to the second working mode based on the first optical change information; determine the current human-computer interaction state based on the second optical change information; and, in response to determining that the human-computer interaction state is one where interaction exists, control the photosensitive component 11 to remain in the second working mode, otherwise control the photosensitive component 11 to switch back to the first working mode; wherein, the first power consumption is less than the second power consumption.

[0017] It should be noted that, Figure 1 The dashed arrow pointing to the photosensitive component 11 indicates that the photosensitive component 11 collects the reflected light from the target area 14 in the first working mode, and the solid arrow pointing to the photosensitive component 11 indicates that the photosensitive component 11 collects the reflected light from the target area 14 after being illuminated by the active light field in the second working mode.

[0018] It should be noted that the target area 14 may vary depending on the specific scenario. In this disclosure, it mainly refers to areas that may move due to the user's facial area, lip area, jaw area, throat area, etc., which may be areas that are actually not spoken or whose sound cannot be clearly recognized, such as those that are far away.

[0019] It should be noted that the projection component 12 projects an active light field with time-modulated characteristics onto the target area 14. This means that each luminescent unit (light spot / spot / stripe, etc.) in the projected active light field undergoes changes in brightness / darkness / flickering according to time-modulated characteristics. The photosensitive component 11 senses changes in light intensity and outputs event sequences at the corresponding pixels. These event sequences also possess time-modulated characteristics.

[0020] It is understood that the processing component 13 can be a system, such as a SoC (System-on-a-Chip). The processing component 13 can also include various discrete but collaborative hardware and software modules. The processing performed by the processing component 13 can be executed by a dedicated chip, a programmable chip, a general-purpose processor, etc.

[0021] The human-computer interaction state perception system provided in this disclosure has a processing component 13 and a photosensitive component 11 that are communicatively connected. In the initial state, the human-computer interaction state perception system is in a first working mode. In the first working mode, the photosensitive component 11 operates with a first power consumption and collects the reflected light of the target area 14 to form first optical change information. The processing component 13 processes the first optical change information output by the photosensitive component 11 to determine whether the target area 14 has moved (e.g., whether there is a user's vocalization action). If there is a user's vocalization action, the processing component 13 controls the photosensitive component 11 to switch from the first working mode to the second working mode; otherwise, the processing component 13 controls the photosensitive component 11 to remain in the first working mode. When the photosensitive component 11 switches to the second operating mode, the projection component 12 starts working, projecting an active light field with time-modulated characteristics onto the target area 14. The photosensitive component 11 operates at the second power consumption and collects the reflected light from the target area 14 after being illuminated by the active light field, forming second optical change information. The processing component 13 determines the current human-computer interaction state based on the second optical change information. In response to determining that there is interaction, the processing component 13 controls the photosensitive component 11 to remain in the second operating mode; otherwise, it controls the photosensitive component 11 to switch to the first operating mode. Therefore, the human-computer interaction state perception system of this disclosure uses hierarchical perception. First, it performs primary perception with the low-power motion photosensitive component 11. After the processing component 13 determines that the target area 14 has moved, it controls the photosensitive component 11 to switch to fine perception. When no interaction occurs, the system can maintain a power consumption of microwatts for a long time. In addition, since the projection component 12 projects an active light field with time-modulated characteristics onto the target area 14 in the second operating mode, it has higher bandwidth and lower latency compared to traditional frame-based structured light schemes.

[0022] The human-computer interaction state perception system disclosed herein can take different forms in various application scenarios, thereby enabling it to effectively collect optical change information of the target area 14.

[0023] For example, such as Figure 2a As shown, the human-computer interaction state perception system of this disclosure can be partially or wholly integrated into the wireless earphone 21. The wireless earphone 21 can be over-ear, ear-hook, or in-ear, etc. For over-ear earphones, the human-computer interaction state perception system of this disclosure can be partially or wholly disposed on a structure extending from the main body, such as a microphone structure or an additional extension structure. The photosensitive component 11 and the projection component 12 are placed on the extension structure, so that the photosensitive component 11 can collect optical change information of the target area 14, and the projection component 12 can project an active light field onto the target area 14.

[0024] For ear-hook or in-ear earphones where size control is critical, the human-computer interaction state perception system disclosed herein adopts a miniaturized integrated design, with some or all of its functional modules adapted to and embedded in the earphone body and ear-hook fitting parts of the reserved integration slots, achieving a compact arrangement of components.

[0025] For example, the human-computer interaction state perception system disclosed herein can also be partially or fully integrated into the camera of a smartphone, adapted and integrated into the camera module body, peripheral support structure or lens protection frame, with the photosensitive component 11 directly serving as the camera or part of the camera, and the projection component 12 correspondingly arranged in a reserved position next to the camera.

[0026] For example, such as Figure 2b As shown, the human-computer interaction state perception system of this disclosure can be partially or wholly integrated into the smart glasses 22, and is disposed in the temples, front bridge of the frame, nose pad structure, or lens periphery integration area of ​​the smart glasses 22. For example, a processing component 13 is disposed at the temples, and a photosensitive component 11 and a projection component 12 are correspondingly disposed on the side of the frame or lens.

[0027] For example, the human-computer interaction state perception system disclosed herein can also be partially or wholly integrated into a smart helmet, embedded in a pre-reserved slot in the helmet shell, an inner lining fitting structure, a goggle bracket, or a side communication module mounting position. For instance, the photosensitive component 11 and the projection component 12 are correspondingly arranged in the front, side, or forehead areas of the helmet to adapt to the perception needs of the target area 14.

[0028] It is understood that the human-computer interaction state perception system disclosed herein is not limited to the above scenarios.

[0029] In one embodiment, the processing component 13 calculates the motion indication of the target area 14 based on the first optical change information, and controls the photosensitive component 11 to switch from the first working mode to the second working mode if the motion indication exceeds the threshold.

[0030] It should be noted that a motion indicator exceeding the threshold can indicate that motion has occurred in the target area 14, such as indicating that the user has made a vocalization.

[0031] It should be noted that, in this disclosure, motion indication primarily refers to possible vocalization actions of the user. That is, movements of the facial area, lip area, jaw area, and throat caused by the user's audible or silent speech. The photosensitive component 11 can acquire optical change information of the target area 14 in the low-power state of the first operating mode, and obtain the motion indication after processing by the processing component 13.

[0032] Specifically, the event rate within a sliding time window can be used to determine motion events in the target area 14. When a person makes a vocalization, their facial muscles move, creating local undulations. Within a short time window, the ambient light can be considered essentially constant. These local undulations in facial muscle movement lead to changes in light and shadow, meaning the local light intensity changes with the movement. This change in light intensity is sensed by the photosensitive component 11 (e.g., an event sensor), generating optical change information. Similarly, the opening and closing of the lips, the movement of the jaw, and the movement of the throat can all generate optical change information.

[0033] Calculate motion indicators:

[0034] in: The set of pixels involved in motion detection; For pixel brightness changes; This is a threshold or normalization function. When When, it is determined that there is valid motion, among which, To determine the threshold for vocalization.

[0035] Compared to not making a vocalization, the rate of motion events increases, so the motion indication can be obtained based on the event rate within a sliding time window.

[0036] Furthermore, since the causes of movements in the facial, lip, jaw, and throat areas are not limited to vocalization, but may also include facial expressions, eating, and drinking behaviors, additional auxiliary conditions can be added to filter out invalid movements. For example, the frequency of facial expression changes may not be as consistent as that of speech; adding event rate criteria within multiple time windows can filter out simple facial expression changes. For instance, chewing has relatively stable characteristics; when the event rate in a specific area matches chewing characteristics, it can also be filtered out. Those skilled in the art can develop practical filtering strategies based on the basic ideas of this disclosure and the actual situation.

[0037] Alternatively, this disclosure can skip the screening at this stage and determine whether a vocalization action is being performed through more refined processing after switching to the second working mode. In the first working mode, the overall power consumption of the human-computer interaction state perception system can reach the microwatt (μW) level, making it suitable for continuous operation.

[0038] In this disclosure, the photosensitive component 11 is used to capture a sequence of events generated by brightness changes caused by the movement of an object within the target region 14. The "event sequence" in this disclosure refers to any type of time-stamped sampling data generated based on changes in light intensity, phase, or periodic changes in modulated light. This data may include, but is not limited to: brightness change events output by an event camera, brightness pulse signals output by a pulse camera, photon trigger timestamps output by a photodetector array, transition moments obtained by a high-speed image sensor based on light intensity sampling, and other non-frame-based sampling data capable of recording the timing or phase information of changes in reflected light signals.

[0039] Therefore, "event sequence" is a general term for all sampling methods that can record temporal changes in reflected light, and is not limited to a specific hardware type or data format. Among these, event cameras are a commonly used solution, and this disclosure uses event cameras as an example for illustration.

[0040] Event cameras can include standalone EVS (Event-based Vision Sensor) or hybrid sensors (HVS) with event channels.

[0041] HVS is a sensor that integrates EVS and APS (Active Pixel Sensor). HVS simultaneously outputs event data and traditional frame images. This disclosure primarily uses EVS to acquire event data and analyzes it to perceive whether a corresponding target area is moving. However, it does not preclude simultaneously acquiring frame images and fusing event data analysis and frame image analysis to perceive whether a corresponding target area is moving.

[0042] An event camera may include an optical system, an event vision sensor, an analog front-end, and a processor.

[0043] The optical system projects the light signal from the target area 14 onto the photosensitive surface of the event vision sensor. The optical system can determine the field of view, resolution, and dynamic response performance. The optical system may include a lens assembly and an aperture. The lens assembly is used for optical imaging of the target area 14 and can be either fixed-focus or zoom; the aperture is used to adjust the amount of light entering the camera, balancing overexposure in strong light conditions and signal strength in low light conditions.

[0044] The event vision sensor triggers data generation only when a pixel detects a change in brightness, thus recording only pixels in the scene whose brightness has changed. Each pixel can operate independently; when the brightness detected by a pixel changes and the change exceeds a threshold, the corresponding pixel will trigger data generation. This data is called event data.

[0045] Event data is typically represented as a quadruple (x, y, t, p). Here, (x, y) represents the position or coordinates of a pixel; p represents the polarity of the event, i.e., a change in brightness or light intensity. Generally, p = +1 indicates an increase in brightness, p = 0 indicates no change in brightness, and p = -1 indicates a decrease in brightness; t represents the time the event occurred, stored as a timestamp, achieving microsecond-level precision.

[0046] The output of an event camera can be either a 1-bit or 2-bit matrix representing all pixels. When using a 1-bit matrix, the value of any element indicates whether the corresponding pixel has generated an event; an element is 1 if an event occurs, and 0 if no event occurs. When using a 2-bit matrix, the value of any element indicates whether the corresponding pixel generated a positive, negative, or no event. For example, a positive event is represented by a +1 value, a negative event by a -1 value, and no event by a 0 value. The first bit of the 2-bit matrix can represent the sign.

[0047] In this disclosure, the projection component 12 projects an active light field with time-modulated characteristics onto the target area. This means that each luminescent unit (light spot / spot / stripe, etc.) in the projected active light field undergoes changes in brightness, darkness, or flickering according to time-modulated characteristics. The photosensitive component 11 senses changes in light intensity and outputs an event sequence at the corresponding pixel. These event sequences also possess time-modulated characteristics.

[0048] The analog front end connects the event vision sensor and the processor, which can preprocess the analog electrical signals output by the event vision sensor, improve signal quality, and achieve analog-to-digital conversion.

[0049] The processor is used for real-time processing, timing calibration, and format conversion of event data. It can be built on high-performance chips such as FPGA (Field Programmable Gate Array), MCU (Microcontroller Unit), DSP (Digital Signal Processor), GPU (Graphics Processing Unit), or NPU (Neural Network Processing Unit), possessing parallel computing capabilities and enabling low-latency processing of massive asynchronous event data output from event vision sensors. In dot matrix structured light time-series encoding applications, the processor can perform processing including but not limited to: timing calibration of event data, correcting timestamp deviations caused by hardware transmission by combining the clock reference provided by the synchronization module, ensuring strict alignment of event timestamps with the encoding timing of projection component 12; filtering event data to remove erroneous events caused by ambient light interference and extracting the event stream of valid encoded light points. The processor can also integrate a parameter configuration interface to dynamically adjust sensor thresholds, filtering parameters, etc., from external sources to adapt to different encoding schemes and application scenarios.

[0050] In one embodiment, the photosensitive component 11 includes a plurality of pixels arranged in an array; in a first operating mode, the photosensitive component 11 operates in at least one of a pixel binning mode, a downsampling mode, and a mode that only enables a portion of the region of interest. The photosensitive component 11 operates with reduced power consumption and acquires optical change information of the target region 14 by using at least one of the pixel binning mode, the downsampling mode, and the mode that only enables a portion of the region of interest.

[0051] In the first implementation, the photosensitive component 11 reduces power consumption by using a pixel-merging mode. The pixel-merging mode merges multiple adjacent pixels on the photosensitive component 11, making them work together as a single pixel, thereby reducing the overall data processing volume and computational load of the photosensitive component 11, and ultimately achieving an effective reduction in power consumption.

[0052] Specifically, the photosensitive component 11 can be composed of a large array of pixels, each with an independent photoelectric conversion function, capable of converting received light signals into electrical signals, which are then analyzed and transmitted by the subsequent processing component 13. In normal operation mode, all pixels are in working state, continuously performing photoelectric conversion and signal output even in scenarios where detailed information is not required, resulting in relatively high power consumption. However, in pixel merging mode, adjacent 2×2, 4×4, or even more pixels can be divided into a merging unit according to actual sensing needs. Multiple pixels within this merging unit simultaneously receive light signals and perform photoelectric conversion, then the converted electrical signals are superimposed, averaged, or summed, ultimately outputting only a single merged electrical signal. This processing method not only reduces the number of signals transmitted and lowers power consumption during signal transmission, but also reduces the computational load of the subsequent processing component 13, avoiding unnecessary computational power consumption, thereby significantly reducing the overall power consumption of the photosensitive component 11.

[0053] In practical applications, the number and range of merged pixels can be flexibly adjusted according to the perception accuracy requirements of the target area. For example, when the human-computer interaction state perception system of this disclosure is integrated into an in-ear wireless earphone, due to the limited size of the earphone and the high requirements for power consumption control, when detecting whether there is a sound action in the first working mode, a high resolution is not required. At this time, the photosensitive component 11 can adopt a 2×2 pixel merging mode, merging four adjacent pixels into one unit. While ensuring accurate acquisition of optical change information within the target area 14, power consumption is minimized and the earphone's battery life is extended. At the same time, the pixel merging mode also has the advantage of fast response speed. The merged pixel unit can quickly complete the acquisition and conversion of light signals without complex calculation processing, and can quickly output the perception results to meet the real-time requirements in the human-computer interaction process.

[0054] In the second implementation, the photosensitive component 11 operates with reduced power consumption through a downsampling mode. Downsampling mode, also known as the sampling mode, involves the photosensitive component 11 acquiring complete raw image data and then, using a specific sampling algorithm, extracting a portion of the image pixels according to a preset ratio, discarding redundant pixel data. This reduces resource consumption during subsequent data processing, transmission, and storage, thus lowering power consumption. Unlike pixel merging mode, which reduces the number of signals during the photoelectric conversion stage, downsampling mode filters data after event data acquisition. By retaining key pixel information and discarding redundant pixel information, it optimizes power consumption without affecting the core sensing effect. Under normal operating conditions, the photosensitive component 11 acquires event data from the target area 14 at the highest resolution. Even if the key information required for human-computer interaction is concentrated in only a portion of the area, complete high-resolution data is generated. This not only consumes a significant amount of power from the photosensitive component 11 but also increases the computational burden on the processing component 13 and data transmission power consumption. Especially for small-sized smart devices with limited battery life, excessive power consumption can severely impact the user experience.

[0055] In practice, the photosensitive component 11 first acquires complete event data of the target area 14 at a preset original resolution. Then, based on the actual needs of human-computer interaction state perception, it sets a reasonable downsampling ratio, such as 1 / 2, 1 / 4, or 1 / 8, and performs downsampling processing on the original data using common algorithms such as nearest neighbor sampling, bilinear interpolation, and bicubic interpolation. For example, when the downsampling ratio is set to 1 / 4, one pixel is extracted every three pixels, converting the original high-resolution image into a low-resolution image with a resolution of 1 / 4 of the original, discarding the remaining 3 / 4 of redundant pixel data. These discarded redundant pixel data do not require subsequent signal analysis, transmission, and storage, thereby significantly reducing the computational load and data transmission bandwidth of the processing component 13 and lowering the corresponding power consumption.

[0056] In the third implementation, the photosensitive component 11 reduces power consumption by enabling only a portion of the region of interest (ROI). Based on the specific needs of human-computer interaction state perception, key areas within the photosensitive component 11's acquisition range, such as the ROI, can be clearly defined. This allows the photosensitive component 11 to collect and process data only within the ROI, thus avoiding power waste in unnecessary areas and achieving overall power reduction. In human-computer interaction state perception scenarios, the acquisition range of the photosensitive component 11 is often larger than the actual required perception range. For example, when the system is integrated into a smartphone camera, the camera's acquisition range can cover the entire face, but human-computer interaction perception only needs to capture changes in specific areas of the face. When the system is integrated into a smart helmet, the photosensitive component 11's acquisition range can cover the head and surrounding environment, but only the user's facial state needs to be captured. In this case, the continuous operation of pixel units in non-ROI areas would result in wasted power.

[0057] Specifically, multiple configurable regions of interest (ROIs) can be preset. Users or the system can select the corresponding ROI through software configuration or hardware triggering based on the actual sensing scenario and needs, and activate the pixel units, photoelectric conversion modules, and signal processing modules in that region. Simultaneously, all components in non-ROI regions are turned off, putting them into a low-power sleep state. For example, when the human-computer interaction state sensing system of this disclosure is integrated into a smart helmet to detect the user's facial movement state, the ROI can be set to the user's facial area. The photosensitive component 11 only activates the pixel units in that region, continuously collecting optical change information of the face to determine whether the user is performing a vocalization action.

[0058] Furthermore, the setting of the region of interest (ROI) offers the advantage of dynamic adjustment. The photosensitive component 11 can switch and adjust the range and position of the ROI in real time according to changes in the human-computer interaction scenario. This implementation method has a particularly significant effect on power consumption reduction because it reduces unnecessary component operation from the source of data acquisition. Compared with pixel merging mode and downsampling mode, it can reduce redundant operations and resource consumption to a greater extent. It is suitable for various smart terminal devices with extremely high power consumption control requirements and relatively fixed sensing scenarios, such as in-ear headphones, smartwatches, and smart helmets. It can ensure the accuracy and real-time nature of human-computer interaction state perception while effectively extending the device's battery life and improving the user experience.

[0059] It should be noted that the above three implementation methods can be used in combination.

[0060] In the second operating mode, the projection component 12 is activated to project an active light field with time modulation characteristics onto the target area 14, and the photosensitive component 11 operates in the normal power consumption state in the second operating mode, so that it can output the optical change information of all pixels after being illuminated by the active light field, so as to perform more refined perception.

[0061] The projection component 12 and the photosensitive component 11 are arranged in such a way that the target area 14 is optically accessible. Optical accessibility means that the projection component 12 can project an active light field onto the surface of the target area 14, and the photosensitive component 11 can capture light intensity changes in the target area 14. The most direct method of optical accessibility is relative arrangement. It is understood that at least one mirror can also be used to adjust the optical path, making the projection component 12 and the photosensitive component 11 optically accessible to the target area 14.

[0062] In one embodiment, the projection assembly 12 may include an optical element, a drive controller, and a synchronization module. The optical element is configured to project an active light field onto the target area 14, the drive controller is configured to generate a multi-channel time-coded pulse sequence and drive the active light field in sequence, and the synchronization module is configured to synchronize the timing of the time-coded pulse sequence with the timing of the reflected light collected by the photosensitive assembly after the target area 14 is illuminated by the active light field.

[0063] In one specific embodiment, the optical elements of the projection assembly 12 may include a DOE (Diffractive Optical Element), a diffuser, a collimating lens, etc., to create a multi-spot array (e.g., a dot matrix), speckle, or stripes. The DOE is used to generate random speckles or regular dot matrices; its interior or surface has micro- or nano-structures that can diffract light, forming a certain number of tiny spots with unique distribution characteristics. The diffuser is used to homogenize the light beam and widen the illumination angle, ensuring a uniform illumination field. The collimating lens is used to convert the diverging light source output into parallel light, ensuring that the projected pattern remains stable and undistorted within the effective distance.

[0064] In one specific embodiment, the drive controller of projection component 12 is used to generate a multi-channel time-coded pulse sequence. A multi-channel time-coded pulse sequence refers to multiple different codes within an encoding duration, where each bit in the code is assigned a value sequentially according to time. For example, in channel 1, in the time sequence... In channel 1, the encoding is [0101100…1] (a total of M bits). In channel 2, for the same time series, the encoding is [1010011…0] (a total of M bits). Other channels follow the same pattern. The encodings for different channels are distinct.

[0065] In one specific embodiment, the synchronization module of the projection component 12 is used to synchronize with the photosensitive component 11 or to establish a time alignment relationship with the projection component 12 through calibration. In order to acquire the event sequence reflecting the aforementioned channel encoding, the photosensitive component 11 must maintain synchronization with the timing of the projection pulses of the projection component 12. If the timing is mismatched, the photosensitive component 11 will be unable to accurately capture the time nodes of the on / off state transitions of the luminous units, resulting in the loss of encoded sequence features and consequently causing decoding confusion. A stable time alignment relationship between the two can be established through hardware clock synchronization or software calibration compensation. At the hardware level, a precise time protocol or synchronization trigger signal can be used to directly unify the clock references of the photosensitive component 11 and the projection component 12. At the software level, a reference timing signal can be acquired to calculate and compensate for the time deviation between the photosensitive component 11 and the projection component 12, ensuring synchronization accuracy.

[0066] In one embodiment, the active light field is a structured light with temporal coding formed by encoding a multi-channel time-coded pulse sequence. The optical element forms at least two distinguishable illumination regions in space. The structured light with temporal coding from different channels is projected into different illumination regions, and the coding Hamming distance between different channels is not less than 2.

[0067] In this disclosure, the time-coded pulse sequence of the drive controller is used to sequentially drive the projected structured light spots / speckles / stripes to exhibit bright / dark changes according to the coded pulses. For ease of description, each spot / speckle / stripe is referred to as a luminous unit. One or more luminous units can be programmed into a channel, meaning that the luminous units within that channel all exhibit the same luminous rhythm according to the same time-coded pulse sequence.

[0068] In this disclosure, each channel's encoding can be controlled independently or in groups. Independent channel encoding control means that each channel is driven and controlled independently, enabling fully parallel signal output. Group channel encoding control means dividing multiple channels into several groups, using a unified control strategy within each group and a differentiated strategy between groups. For example, channels that are close together should have different encodings to distinguish them. Channels that are more than a certain distance apart can use the same encoding. In this case, channels with the same encoding can be grouped together for unified driving and control.

[0069] In this disclosure, the encoding can be periodic encoding, Gray code encoding, binary code encoding, pseudo-random code encoding, etc. Periodic encoding uses a fixed period as a unit, allowing the emitting unit to switch between on / off or bright / dark states according to a repetitive temporal pattern. The temporal sequence within the period serves as a unique identifier, and redundancy verification is achieved through multi-period acquisition. Gray code encoding, based on the characteristic that adjacent codes differ by only one bit, maps multiple temporal states to Gray code values, achieving high-precision sequence matching and reducing the bit error rate through frame-by-frame comparison. Binary code encoding directly maps the temporal state of the emitting unit to binary numbers, with each frame corresponding to one binary bit, and multiple binary bits combined to form a unique code. Sequence matching can be quickly completed through bitwise operations. Pseudo-random code encoding, based on a pseudo-random sequence generation algorithm, assigns a temporal sequence with no obvious pattern to each emitting unit, utilizing the autocorrelation of the sequence to achieve highly robust matching and avoid confusion with the encoded sequences of adjacent channels.

[0070] In this disclosure, the timing codes of different channels maintain a sufficient Hamming distance threshold (e.g., ≥2) on the ideal event response sequence to ensure differentiation. Hamming distance refers to the number of different bits at corresponding positions in two equal-length coded sequences. For example, the timing codes "1011" and "1101" differ in two corresponding positions, so the Hamming distance is 2. The Hamming distance threshold can be used to determine whether two timing codes are distinguishable and whether there are valid errors. For example, setting the Hamming distance threshold to 2 means that if the Hamming distance between the two channel codes is greater than 2, they can be significantly distinguished. Maintaining a sufficient Hamming distance prevents different channel codes from being assigned to the same channel due to decoding errors caused by slight interference.

[0071] To enhance robustness, one or more methods can be used for encoding, including but not limited to: increasing the ±1 event ratio, replicating the encoding cycle, and increasing the number of encoding steps.

[0072] "±1 events" refer to the transition behavior of the light-emitting unit between its on and off states (similar to the dark state) (on → off is a -1 event, off → on is a +1 event). Increasing the proportion of ±1 events, i.e., by optimizing the coding sequence design, increases the number of state transitions per unit coding length. State transitions allow the event camera to capture enough events. A higher transition ratio makes the features of the coded sequence more prominent.

[0073] A replication coding cycle refers to repeatedly projecting a complete coding sequence multiple times, forming a continuous time sequence of "coding cycle 1 → coding cycle 2 → ... → coding cycle n". It utilizes the consistency verification of redundant data to offset random interference.

[0074] Increasing the number of encoding steps is equivalent to extending the length of the encoded sequence. The number of encoding steps is the total number of bits in the encoded sequence; increasing the number of encoding steps is the same as extending the length of the encoded sequence. A longer encoded sequence can carry more uniqueness. The more encoding steps there are, the greater the number of differences between any two channel codes if the minimum Hamming distance between different channels is increased simultaneously, resulting in stronger anti-aliasing capabilities.

[0075] The projection component 12 of this disclosure forms at least two spatially distinguishable illumination regions, which differ from each other in temporal modulation characteristics. Temporal modulation methods include, but are not limited to: intensity modulation, frequency modulation, phase modulation, duty cycle modulation, or combinations thereof. The illumination regions are not required to be regular dot or stripe structures, only that they can be spatially correlated through temporal modulation.

[0076] This disclosure uses structured light with an active light field based on a multi-channel time-coded pulse sequence and a photosensitive component 11 as an example for illustration.

[0077] like Figure 3a As shown, assuming the encoding step size M=6, that is, 6 bits are used to represent an encoding. One encoding cycle The structured light projection component is configured according to code [011010]. Within a given time period, light spots are projected in a sequence of "off-on-on-off-on-off".

[0078] The event sequence acquisition component can detect events sequentially at the same pixel location where the aforementioned light spot is projected: 0 (the event polarity is set to 0 in the first frame), +1, 0, -1, 1, -1. It can be seen that the event sequence and the encoding do not directly correspond numerically, but as long as the initial state of the pixel is known—for example, if all event pixels are reset once before an encoding cycle (to zero, or the encoding starts with 0)—the encoding perceived by each pixel can be easily deduced.

[0079] Furthermore, such as Figure 3b As shown, structured light from different channels is projected into different regions, and the encoding of each channel is different. Furthermore, the Hamming distance between the encodings of different channels is not less than 2. This disclosure employs structured light with brightness abrupt changes, therefore, at least one abrupt change occurs within one encoding cycle. Locations not affected by structured light projection will not generate events for the event camera; pixels at these locations can be excluded to reduce the amount of data processed.

[0080] In one embodiment, the processing component 13 of this disclosure is further configured to: decode and match the event sequence output by all pixels of the photosensitive component 11 on the time axis to obtain the encoding channel where the pixel is located, and the event sequence is used to record the temporal variation of the second optical change information; aggregate multiple adjacent pixels belonging to the same encoding channel and whose spatial distance meets a set value into a spot pixel region; and perform geometric inference or motion estimation based on the coordinates of the center pixel of the spot pixel region to determine the current human-computer interaction state.

[0081] In one specific embodiment, after structured light with time-series encoding is projected onto the target area 14, the event sensor can collect the event sequence. The processing component 13 decodes and matches the event sequences output by all pixels of the event sensor on the time axis to obtain the encoding channel where the pixel is located. In this disclosure, one flash corresponds to one whole frame, and time decoding and matching are performed on event sequences of multiple frames (such as 4 frames or 6 frames) on the same pixel.

[0082] The first j The ideal encoding sequence for each projection channel is represented as follows: The encoding is an M-bit sequence of 0s and 1s. If M=6, the example encodings could be [011010], [110110], [100100], etc. Therefore, there are 64 different encodings in the encoding space. Different encodings should be used for different channels, and as described earlier, the Hamming distance between encodings of different channels needs to be significantly distinguishable.

[0083] Pixels The observed binary sequence is represented as The values ​​in the sequence are determined by whether the event is triggered at each encoding step. See Table 1. Channel Encoding For [011010], the position of the modulated structured light illuminating the target area corresponds to the pixel coordinates of the event camera ( a, b At point ), the flickering of the structured light causes the event camera to record pixel coordinates ( a, b The event sequence at () will be identical to the channel encoding if recorded accurately. At another location on the target area, corresponding to the coordinates of the event camera () c, d At point ), the flickering of the structured light causes the event camera to record pixel coordinates ( c, d The event sequence at position [001010] is related to the channel encoding. different.

[0084] Table 1

[0085] In one embodiment, regarding pixel coordinates ( x, y Is it through the channel?j Structured light projection patterns, also known as pixels ( x, y Does it belong to the first category? j For each channel, the Hamming distance can be used for decision-making:

[0086] in, Indicates the first j Preset time series encoding for each channel, This represents the Hamming distance calculation function. Indicates the first j Calculate the Hamming distance for each channel, and take the channel number with the smallest distance as the optimal matching result.

[0087] This place will also Used to represent the binary code after decoding, i.e., in Table 1 The event sequence is [0,+1,0,-1,+1,-1]. After being reset before encoding, the event change begins with the light spot being off. Therefore, the binary code [011010] can be recovered from the event sequence. For simplicity, in this formula... This can refer to the recovered binary code [011010]. Similarly, The event sequence is [0,0,+1,-1,+1,-1], which can be deduced from the binary code [001010].

[0088] One can compare pixels one by one. x, y ) observation binary sequence With channel encoding The Hamming distance between them is used. For example, if there are N channels in total, let j = 1, 2, ..., N, and calculate them one by one. Take The value of j when it is at its minimum is the pixel value. x, y The channel to which it belongs.

[0089] in, This means subtracting the code values ​​at the same position from two codes of the same length, taking the absolute value, and then adding them all back together.

[0090] The above With encoding They are exactly the same, therefore the Hamming distance is 0. The pixel point can be confirmed ( a, b This belongs to channel 1. The above... With encoding Since one digit is different, the Hamming distance is 1.

[0091] A decision threshold can be set to reduce noise:

[0092] in, This indicates that Hamming is far from the decision threshold. Represents the observation coding sequence of pixels. Indicates the optimal matching channel The preset time series encoding. That is, after calculating the Hamming distance, the minimum Hamming distance should be less than or equal to a threshold. Only then can the current value of j be used to classify the pixel into the channel.

[0093] For example, when After calculating the Hamming distance with other channel codes, if there is also a channel with a Hamming distance of 0 (e.g., channel 3), then the pixel ( c, d ) belongs to channel 3; conversely, if the calculated Hamming distance is 1 for That is the minimum value, if Then you can also use pixels ( c, d It belongs to channel 1. If Then you cannot put pixels ( c, d It belongs to channel 1.

[0094] In one embodiment, regarding pixel coordinates ( x, y Whether the structured light projection pattern of channel j is used to determine the pixel ( x, y Whether a string belongs to the j-th channel can be determined using binary decoding. Binary decoding directly interprets a binary sequence into an integer value and then compares these integer values. For example, the decimal integer value of [011010] is 26, and the decimal integer value of [001010] is 10. When using binary decoding, if the difference occurs in different positions, the decoded decimal values ​​may differ significantly. Generally, the two values ​​need to be completely equal to determine the channel affiliation. For example, the two codes above differ only in the second bit from the left, and the two decimal values ​​differ by 16.

[0095] In one embodiment, regarding pixel coordinates ( x, y Whether the structured light projection pattern of channel j is used to determine the pixel ( x, y Whether a pixel belongs to the j-th channel can be determined using Gray code decoding. If the decoded Gray code sequence belongs to a certain channel, then the pixel can be classified as belonging to that channel.

[0096] In one specific embodiment, the processing component 13 aggregates multiple adjacent pixels belonging to the same encoding channel and with a spatial distance satisfying a set value into a light spot pixel region. The light-emitting units (speckle / spot / stripes, etc.) projected by the projection component 12 will affect the generation events of multiple related pixels within the light-emitting region. For example, a speckle may cover a diameter of 0.5 mm, and the event camera may record dozens of pixels that cause the event by the flickering of this speckle.

[0097] As described above, after filtering, decoding, and channel matching, the channel to which each pixel belongs has been obtained. Pixels with the same channel encoding and close spatial distance can be aggregated to reconstruct the spot pixel region. The aggregation distance threshold can be, for example, 10 pixels, meaning that in the aggregated pixel set, besides having the same channel encoding, the difference between the horizontal and vertical coordinates of any two pixels does not exceed 10 pixels. Alternatively, a Euclidean distance metric can be used, where the distance between any two pixels does not exceed 10 pixel spacing. The above aggregation distance threshold is only an example and can be adjusted according to actual needs.

[0098] Specifically, for the candidate pixel set of the same channel j , This represents the optimal matching channel calculated using the Hamming distance. The calculation formula has been described above. Indicates pixels Assigned to channel The i-th spot region is obtained by clustering connected components. Connected component clustering involves two main steps: first, determining the neighboring pixels of each pixel; and second, forming clusters based on the neighborhood relationships between pixels.

[0099] like Figure 4 As shown, in one embodiment, two adjacent pixels are considered to be neighboring pixels. That is, either the horizontal coordinate or the vertical coordinate of the two pixels differs by 1. For example, pixel ( x, y The neighboring pixels of (non-boundary pixels) are In addition, the range of neighboring pixels can be expanded, and the eight pixels surrounding a non-boundary pixel can be considered as neighboring pixels. Alternatively, the Euclidean distance between pixels can be used as a measure; when the Euclidean distance between two pixels is less than a set threshold, they are considered neighboring pixels.

[0100] After determining the neighboring pixels of each pixel, clusters can be formed based on the neighborhood relationships between pixels. The core of connected component clustering is to determine the connectivity between points. For example, the flood fill method (flood algorithm) or the Union-Find algorithm can be used.

[0101] In one specific embodiment, the processing component 13 calculates the coordinates of the center pixel of the aforementioned spot pixel region.

[0102] Define weights: That is, the pixels already described above. In an encoding cycle The event count (or event rate) within.

[0103] The center of the light spot (centroid) is calculated based on the following formula using weights:

[0104] in, For the first The first channel The center coordinates of each light spot region; The x-axis coordinate of the center of the non-homogeneous spot in the image coordinate system; The y-axis coordinate of the center of the non-homogeneous spot in the image coordinate system; For the first The first channel The set of pixels corresponding to each spot area; For pixels The weights; Sum the weights × coordinates of all pixels within the spot area; Sum the weights of all pixels within the spot area.

[0105] That is, the number of events per pixel within the light spot area. The weighted sum of the coordinate values, divided by the total number of events within the spot area, gives the coordinates of the spot's center point. The pixel with the most events indicates that the projected spot's flicker events have been collected most thoroughly, preserving the encoded information to the greatest extent. This calculation method assigns a higher weight to the pixel with the most events, thus biasing the center point towards the pixel that retains the most encoded information.

[0106] This disclosure obtains the mapping of structured light onto the event sensor pixel through the timing decoding, spot aggregation, and calculation of the coordinates of the center pixel of the spot pixel region performed by the aforementioned processing component 13. Based on the fixed positional relationship between the projection component 12 and the event sensor, by calibration parameters, the spatial coordinates representing the structured light pattern of the same channel and the pixel coordinates reflecting the structured light pattern on the event sensor can be unified into a coordinate system. By the difference in their spatial coordinates, the relative displacement, relative motion trend, or relative geometric relationship between the local areas within the target area 14 can be determined.

[0107] In one embodiment, the target region 14 is divided into multiple sub-regions. After the active light field is projected, there are multiple illuminated areas in space. Each illuminated area is mapped to a set of pixels on the event sensor. This can be done based on region. The acquisition signal of the j-th pixel in To calculate regional motion index :

[0108] in: It is a region The set of pixel indices.

[0109] The above formula represents: traversing the region Obtain the acquisition signal of all pixels in the image. , for its set Apply To obtain regional motion indicators . This is a geometric inference or motion estimation function.

[0110] It should be noted that the geometric inference or motion estimation performed by the processing component 13 described above, after time-series decoding, spot aggregation, and calculation of the coordinates of the center pixel of the spot pixel region, can be used to determine the current human-computer interaction state. The geometric inference or motion estimation function may not be based on the calculation of three-dimensional coordinates or depth values. Of course, this disclosure does not exclude the use of three-dimensional coordinates or depth values ​​for geometric inference or motion estimation.

[0111] Specifically, this disclosure can employ two methods to perform geometric inference or motion estimation to determine the current human-computer interaction state.

[0112] The first method can be achieved by generating a disparity map, which compares the coordinates of the center pixels of the same coded channel with the original projection coordinates to obtain the disparity map. Specifically, based on the coordinates of the center pixels of the spot pixel region and their channel encoding, the emission coordinates with the same channel encoding as the transmitter are obtained. The disparity is then obtained by comparing the coordinates of the center pixels of the spot pixel region with the emission coordinates. The above operation is performed on all spots to obtain the disparity map.

[0113] The second method can be to use triangulation to form camera rays and projected rays in the event camera coordinate system by combining the center pixel coordinates of the same encoded channel with the original projection coordinates, and then find the intersection in three-dimensional space to obtain a three-dimensional point.

[0114] Specifically, the above-obtained center coordinates of the light spot The coordinates corresponding to the event camera coordinate system are:

[0115] in, yes The image coordinate representation. This indicates that the image coordinates are projected onto the plane at z=1. The image coordinates of the center of the projected light spot are... The point is converted to the event camera coordinate system using the camera intrinsic parameter K.

[0116] like Figure 5 As shown, the camera ray can be represented using the point method as follows:

[0117] in It is the optical center of the event camera, which is usually (0, 0, 0). It is a point on the plane with z=1 in the event camera coordinate system (corresponding to the center coordinates of the spot after the intrinsic parameter transformation of the event camera). It is a proportionality coefficient, which can be based on Adjust and generate a series of points in the camera coordinate system. This formula means that the camera ray is formed by a ray drawn from the camera's optical center to these points.

[0118] Similarly, the projection ray in the event camera coordinate system can be represented as:

[0119] in It is the representation of the projection center in the event camera coordinate system. The two can be obtained by converting through calibration parameters based on the fixed positional relationship between the event camera and the projection component 12. This represents the projection coordinates in the event camera coordinate system. Similarly, the coordinates of the projection component 12 can be transformed to the event camera coordinate system through calibration parameters. This is a scaling factor used to adjust the position of a point along the direction of the projected light beam. Greater than 0.

[0120] like Figure 5 As shown, two light rays can be calculated in space. and The distance. When The distance between the two light rays changes when different values ​​are taken. This can be determined by finding the shortest distance between the two spatial lines. (Least squares closed-form solution) yields three-dimensional points: This is a 3D representation of the center point of the i-th spot in the j-th channel. By combining the encoding time of the spot, all 3D point clouds can be represented as 3D point clouds with time information. .in, These are the coordinates of the optimal position point on the projected ray in the event camera coordinate system. These are the coordinates of the optimal position point on the event camera ray in the event camera coordinate system.

[0121] The three-dimensional point cloud data obtained in this disclosure can be used to determine whether the surface of the target area 14 has undergone local undulation, and the current human-computer interaction state can be determined by performing pattern recognition on the three-dimensional point cloud data.

[0122] The hierarchical perception and temporal modulation active light field technology disclosed herein can be further used to acquire the temporal and frequency domain characteristics of the user's facial skin vibration during speech, thereby enhancing voice interaction judgment and interaction state recognition. When a user speaks, the vocal cords vibrate and couple to the facial area through the bones and soft tissues, causing periodic minute deformations and motion changes in the skin near the jaw, cheeks, and corners of the mouth.

[0123] It should be noted that this disclosure does not rely on precise measurement of the absolute displacement of the skin, but rather utilizes the changes in the optical response time structure caused by skin movement for state analysis. This minute movement can be manifested in the optical acquisition signal as: micro-perturbations in reflection intensity, temporal variations in local brightness, and changes in event trigger density.

[0124] To extract vibration-related components, the signal is high-pass filtered, such as... Figure 6 As shown.

[0125] In one embodiment, the target region 14 in this disclosure may include a region that moves due to spoken or silent speech behavior. The processing component 13 is configured to: calculate a motion index of the target region 14 based on second optical change information; and perform high-pass filtering or band-pass filtering on the motion index to obtain high-frequency components of the motion index related to the vibration of the target region 14, and determine the current human-computer interaction state based on the high-frequency components of the motion index.

[0126] Among them, the motion index of target region 14 is calculated. The method has already been introduced above and will not be repeated here. Figure 6 As shown, for the motion index High-pass or band-pass filtering is performed to obtain the high-frequency components of motion parameters related to the vibration of the target region 14. :

[0127] in, This indicates a high-pass filter or a band-pass filter.

[0128] In one implementation, such as Figure 7As shown, the processing component 13 is further configured to: calculate a high-frequency vibration intensity index based on the high-frequency components of the motion index; generate an interaction confidence score based on the high-frequency vibration intensity index, the interaction confidence score being used to characterize the credibility of the current vocal action of the user in the target area; and fuse the audio speech activity detection probability of the target area with the interaction confidence score to obtain the final speech probability, and determine the current human-computer interaction state based on the final speech probability.

[0129] In the first embodiment, generating interactive confidence based on the high-frequency vibration intensity index may include: the processing component 13 performing frequency domain transformation based on the high-frequency vibration intensity index to obtain the vibration dominant frequency, and generating interactive confidence based on the vibration dominant frequency and the high-frequency vibration intensity index.

[0130] In the second embodiment, generating interactive confidence based on the high-frequency vibration intensity index may include: the processing component 13 performing frequency domain transformation based on the high-frequency vibration intensity index to obtain the vibration dominant frequency, calculating the spectral peak significance based on the high-frequency vibration intensity index, and generating interactive confidence based on the spectral peak significance and the high-frequency vibration intensity index.

[0131] The following is in conjunction with the appendix Figure 7 This paper details the method for determining the current human-computer interaction state based on the final speech probability.

[0132] like Figure 7 As shown, based on the high-frequency components of the motion index It can calculate high-frequency vibration intensity index Among them, high-frequency vibration intensity index It must meet at least one of the following forms:

[0133] in, It is the vibration detection cycle; or,

[0134] Among them, RMS (Root Mean Square) is used to calculate the root mean square of the high-frequency components in the vibration signal and is used to represent the vibration intensity.

[0135] or,

[0136] in, This represents frequency domain transformation (FFT / DFT / sliding spectrum analysis). This is the preset frequency band.

[0137] High-frequency vibration intensity index The physical meanings include: high-frequency motion energy, skin vibration activity, and dynamic indicators related to vocalization.

[0138] In one embodiment, the dominant vibration frequency can be obtained through frequency domain transformation. :

[0139] in: This represents frequency domain transformation (FFT / DFT / sliding spectrum analysis). Represents multiple frequency values ​​after frequency domain transformation The largest frequency value is taken as the dominant vibration frequency. .

[0140] In one embodiment, based on the dominant vibration frequency and high-frequency vibration intensity index Generate interactive confidence ,

[0141] in, For the Sigmoid function, These are parameters that can be adjusted according to actual conditions. To prevent division by zero of constants.

[0142] Alternatively, a threshold-based method can be used to generate interactive confidence scores. , in, For indicator functions, The vibration intensity threshold. The threshold for spectral peak significance. In It refers to logical AND, in other words, it refers to... and .

[0143] In another embodiment, to obtain a reliability measure of the vibration frequency, the significance of the spectral peaks can be further calculated. :

[0144] in, To prevent division by zero constants, This refers to multiple frequency values ​​after frequency domain transformation. The maximum value in, This refers to multiple frequency values ​​after frequency domain transformation. The median.

[0145] In another implementation, based on spectral peak significance and high-frequency vibration intensity index Generate interactive confidence ,

[0146] in, For the Sigmoid function, These are parameters that can be adjusted according to actual conditions. To prevent division by zero of constants.

[0147] Alternatively, a threshold-based method can be used to generate interactive confidence scores. ,

[0148] in, For indicator functions, The vibration intensity threshold. This is the threshold for the significance of the spectral peak.

[0149] It should be noted that this disclosure does not require micron-level measurement of absolute displacement on the skin surface, but rather utilizes "vibration-induced optical time-varying patterns" to construct interactive confidence levels. .

[0150] In one embodiment, the probability of detecting audio speech activity on the audio side is... With interaction confidence The final speech probability is obtained by fusion. :

[0151] in, For weight fusion. It can be used to control noise estimation updates and noise suppression strength, thereby avoiding the wearer's voice from being misjudged as noise and weakened for wireless headphone applications.

[0152] In one implementation, such as Figure 7 As shown, processing component 13 is further configured to: calculate a high-frequency vibration intensity index based on the high-frequency components of the motion index, and generate an interaction confidence score based on the high-frequency vibration intensity index; gate or weight the frequency domain gain function according to the interaction confidence score to obtain a gated or weighted frequency domain gain function, where the frequency domain gain function is the gain function between the audio frequency domain signal of the target region and the audio frequency domain signal of the target region after noise suppression; perform frequency domain transformation based on the high-frequency vibration intensity index to obtain the vibration dominant frequency; construct a harmonic protection or enhancement mask using the vibration dominant frequency; process the gated or weighted frequency domain gain function based on the harmonic protection or enhancement mask to obtain a refined frequency domain gain function; and determine the current human-computer interaction state based on the refined frequency domain gain function.

[0153] like Figure 7 As shown, audio frequency domain signal The output is obtained after noise suppression:

[0154] in This is the frequency domain gain function. Furthermore, it can be determined based on the cross-confidence level. Gating or weighting the frequency domain gain function can be implemented as follows:

[0155] in A gain strategy that is friendly to voice fidelity, This represents a more aggressive gain strategy for noise suppression. Therefore, for wireless headphone applications, when the interaction confidence level... When the interaction confidence level is high, the system tends to protect the wearer's voice; when the interaction confidence level is high... At lower levels, the system tends to enhance its suppression of external noise.

[0156] In an alternative implementation, the dominant vibration frequency is utilized. Constructing harmonic protection or enhancement masks:

[0157] Based on this, the frequency domain gain of the gated frequency domain gain function is refined:

[0158] in For harmonic order, Gaussian bandwidth (engineering setting value, used for bandwidth control). For coefficient parameters, The vibration frequency is [value missing]. By protecting or moderately enhancing the frequency structure related to the wearer's voice, speech intelligibility in noisy environments can be further improved.

[0159] In one implementation, such as Figure 8 As shown, embodiments of this disclosure can utilize interactive confidence. The audio processing links can be gated or weighted, and the audio processing links may include at least one of the following: microphone, voice activity detection link, beamforming link, noise reduction link, and echo cancellation link.

[0160] This disclosure can generate interactive confidence scores based on the geometric deformation or motion indicators of the target region 14. ,

[0161] in: It is the weight of the i-th illuminated region. It is the geometric deformation or motion index of the i-th illuminated region. This represents the weighted sum of motion metrics to the interaction confidence level. The mapping function can be a suitable normalized mapping function, such as the Sigmoid family of functions. Interactive confidence. It can be used to: trigger or suppress voice interaction, serve as a priori for voice recognition reliability, and assist in human-computer interaction control processes, etc. (See reference) Figure 8 As shown.

[0162] The human-computer interaction state perception system disclosed herein uses hierarchical perception. The photosensitive component 11 first operates at low power consumption. After the processing component 13 determines that the object in the target area has moved (i.e., when an interaction occurs), the photosensitive component 11 switches to normal power consumption operation for fine perception. Therefore, the human-computer interaction state perception system can maintain a microwatt-level power consumption for a long time when no interaction occurs.

[0163] In the fine perception stage, a human-computer interaction state perception scheme based on event data stream can be formed by combining EVS with active light field projection. Compared with traditional frame-based structured light and other schemes, this disclosure has the following advantages: First, this disclosure enables higher bandwidth and lower latency. The event camera is only sensitive to change; it only outputs events when pixels detect change, and does not output events when there is no change. It can achieve a frame rate of over 1000 frames per second, but with low redundant data, making it suitable for sensing facial muscle movements.

[0164] Second, this disclosure maintains the Hamming distance through multi-channel coding and introduces redundancy to enhance noise resistance, thus achieving stronger robustness.

[0165] Third, this disclosure does not directly perform semantic recognition, but it can quickly enter the interaction mode by recognizing the initial stage of the interaction. In scenarios involving voice interaction, it eliminates the need for steps such as activating a voice assistant, making it faster and more efficient.

[0166] In addition, the active light field disclosed herein does not limit the structure to a regular pattern, but only requires that the spatially distinguishable region and the temporally modulated region be distinguishable. The geometric calculation is not limited to displaying 3D depth, and can be relative displacement / trend / relationship.

[0167] This disclosure also provides a human-computer interaction enhancement system, such as Figure 9 As shown, the human-computer interaction enhancement system includes the aforementioned human-computer interaction state perception system and the interaction control module 15. The interaction control module 15 is configured to perform at least one of the following functions based on the current human-computer interaction state determined by the aforementioned human-computer interaction state perception system: starting or stopping the speech recognition module, filtering out environmental noise, enhancing human voice, and providing control functions as control signals.

[0168] The human-computer interaction enhancement system disclosed herein adds an interaction control module 15 to the aforementioned human-computer interaction state perception system. When the human-computer interaction state perception system determines the human-computer interaction state, it can be further utilized to enhance the human-computer interaction experience. This further utilization includes, but is not limited to, starting or stopping the speech recognition module, filtering environmental noise, enhancing human voice, and providing control functions as a control signal.

[0169] Specifically, the activation or deactivation of the speech recognition module based on the human-computer interaction state may include: the interaction control module 15 analyzes the data of the human-computer interaction state in real time. When it detects speech-related geometric changes (such as the lip opening and closing amplitude reaching a preset threshold) or motion information (such as the regular movement of facial muscles corresponding to pronunciation actions) on the user's face, it determines that the user is about to start voice interaction and immediately and automatically activates the speech recognition module to ensure that the recognition function is ready synchronously when the user speaks. Conversely, when the detected facial geometric changes and motion information disappear and there is no related movement for a preset duration, it determines that the user's voice interaction has ended and automatically deactivates the speech recognition module, reducing the device's computing power consumption and avoiding false triggering of recognition by environmental noise, thereby improving the intelligence level of the interaction.

[0170] Specifically, filtering environmental noise based on the human-computer interaction state can include: combining local facial features when the user speaks to filter environmental noise and optimize the clarity of voice interaction. Since the human-computer interaction state is determined based on the user's facial geometric changes and motion information when speaking, the interaction control module 15 can determine the user's pronunciation period and pronunciation state accordingly. When the user's face shows pronunciation-related movements (such as lip opening and closing, jaw movement), it is determined to be a valid pronunciation period. At this time, the human voice frequency band signal is mainly retained, and environmental noise (such as wind noise, background conversation noise) is specifically filtered out.

[0171] Specifically, enhancing human voice based on human-computer interaction status may include: the interaction control module 15 first assists in language recognition by analyzing the geometric deformation of the user's face (such as the angle of lip opening and closing, and the degree of facial muscle tension) and motion information (such as the frequency and amplitude of facial movements during pronunciation). For example, by analyzing the pattern of lip opening and closing and the characteristics of movement, it can accurately determine whether the user is making a sound, the approximate time and place of pronunciation, and even help distinguish similar pronunciations. After completing the initial language recognition, it enhances the human voice by combining the speech information of the speech system with the human voice and performing cross-verification.

[0172] Specifically, providing control functions based on human-computer interaction states as control signals may include: the interaction control module 15 quantifies the geometric changes and motion information of the user's face when speaking, and assigns different control commands to different facial movement features. For example, when a specific geometric deformation (such as the lips opening and closing rapidly twice) or movement trajectory is detected on the user's face, it is converted into a "pause" control signal, and the system suspends the current operation; when continuous and regular facial movements are detected, it is converted into a "switch" control signal to achieve function switching. This control method does not require the user to issue explicit voice commands; control can be completed solely through pronunciation-related facial movements, simplifying the interaction steps, improving interaction efficiency, and is suitable for scenarios where hands are inconvenient to operate.

[0173] This disclosure also provides a method for human-computer interaction state perception, such as Figure 10 As shown, the sensing method includes: S101. Based on the first optical change information, control the photosensitive component to switch from the first working mode to the second working mode, and control the projection component to start working. The first optical change information is the optical change information formed by the photosensitive component operating in the first working mode and collecting reflected light from the target area.

[0174] S102. Based on the second optical change information, determine the current human-computer interaction state, wherein the second optical change information is the optical change information formed by the photosensitive component operating in the second working mode and collecting the reflected light after the target area is irradiated by the active light field.

[0175] S103. In response to determining that the human-computer interaction state is that there is interaction, control the photosensitive component to remain in the second working mode; otherwise, control the photosensitive component to switch to the first working mode, wherein the power consumption of the photosensitive component in the first working mode is less than the power consumption in the second working mode.

[0176] In some implementations, controlling the photosensitive component to switch from a first operating mode to a second operating mode based on first optical change information includes: calculating a motion indication amount of the target area based on the first optical change information, and controlling the photosensitive component to switch from the first operating mode to the second operating mode when the motion indication amount exceeds a threshold.

[0177] The following is combined Figure 11 This disclosure provides a detailed description of the human-computer interaction state perception method.

[0178] The photosensitive component 11 disclosed herein is used to receive reflected light from the target area 14 and capture optical change information therein. The photosensitive component 11 can operate in a first operating mode or a second operating mode. The projection component 12 projects an active light field with time modulation characteristics onto the target area 14 only in the second operating mode. The processing component 13 is communicatively connected to the photosensitive component 11 and is configured to perform a series of processes on the optical change information output by the photosensitive component 11, and finally output relevant information reflecting the user's vocalization actions.

[0179] like Figure 11 As shown, in the initial state, the human-computer interaction state perception system is in the first working mode. In the first working mode, the photosensitive component 11 operates with a low first power consumption and collects reflected light from the target area 14 to form first optical change information; that is, the human-computer interaction state perception system executes S201. Then, the processing component 13 calculates the motion indication amount of the target area 14 based on the first optical change information; that is, the human-computer interaction state perception system executes S202. Next, the processing component 13 determines whether the motion indication amount exceeds a threshold; that is, the human-computer interaction state perception system executes S203. If the motion indication amount exceeds the threshold, S204 is executed; otherwise, S201 is executed.

[0180] like Figure 11 As shown, in S204, the processing component 13 can obtain the second optical change information. Based on the second optical change information, the processing component 13 determines the geometric deformation or motion state of the target area 14, i.e., S205 is executed. Then, based on the geometric deformation or motion state of the target area 14, the processing component 13 determines whether there is human-computer interaction, i.e., S206 is executed. If there is human-computer interaction, S207 is executed; otherwise, S208 is executed.

[0181] In some implementations, determining the current human-computer interaction state based on the second optical change information includes: calculating the motion index of the target area based on the second optical change information; performing high-pass filtering or band-pass filtering on the motion index to obtain the high-frequency components of the motion index related to the vibration of the target area, and determining the current human-computer interaction state based on the high-frequency components of the motion index.

[0182] In one implementation, determining the current human-computer interaction state based on the high-frequency components of motion indicators includes: calculating a high-frequency vibration intensity index based on the high-frequency components of motion indicators; generating an interaction confidence score based on the high-frequency vibration intensity index, wherein the interaction confidence score is used to characterize the credibility of the user's current vocal action in the target area; and fusing the audio speech activity detection probability of the target area with the interaction confidence score to obtain the final speech probability, and determining the current human-computer interaction state based on the final speech probability.

[0183] In another implementation, determining the current human-computer interaction state based on the high-frequency components of motion indicators includes: calculating a high-frequency vibration intensity index based on the high-frequency components of motion indicators, and generating an interaction confidence score based on the high-frequency vibration intensity index; gating or weighting the frequency domain gain function according to the interaction confidence score to obtain a gated or weighted frequency domain gain function, wherein the frequency domain gain function is the gain function between the audio frequency domain signal of the target area and the audio frequency domain signal of the target area after noise suppression; performing a frequency domain transformation based on the high-frequency vibration intensity index to obtain the vibration dominant frequency; constructing a harmonic protection or enhancement mask using the vibration dominant frequency, and processing the gated or weighted frequency domain gain function based on the harmonic protection or enhancement mask to obtain a refined frequency domain gain function, and determining the current human-computer interaction state based on the refined frequency domain gain function.

[0184] In some implementations, determining the current human-computer interaction state based on the second optical change information includes: decoding and matching the event sequences output by all pixels of the photosensitive component on the time axis to obtain the encoding channel where the pixel is located; the event sequence is used to record the temporal changes of the second optical change information; aggregating multiple adjacent pixels belonging to the same encoding channel and whose spatial distance meets a set value into a spot pixel region; and performing geometric inference or motion estimation based on the coordinates of the center pixel of the spot pixel region to determine the current human-computer interaction state.

[0185] In one optional implementation, geometric inference or motion estimation is performed based on the coordinates of the center pixel of the spot pixel region, including: comparing the coordinates of the center pixels of the same encoding channel with the original projection coordinates to obtain a disparity map, and performing geometric inference or motion estimation based on the disparity map. Alternatively, the coordinates of the center pixels of the same encoding channel and the original projection coordinates are used to form camera rays and projection rays in the photosensitive component coordinate system, and the intersection of the camera rays and projection rays is obtained in three-dimensional space to obtain a three-dimensional point, and geometric inference or motion estimation is performed based on the three-dimensional point.

[0186] The following provides directly implementable application scenarios and exemplary computational models from a product implementation perspective. These computational models are exemplary implementations and do not limit the scope of protection of this disclosure. Specifically, they can include three main product scenarios: The first scenario can be a wireless earphone scenario, where the target area can be the jaw / cheek / corner of the mouth region, outputting the confidence level of "the wearer is speaking / preparing to speak" as a gating / weighted prior for the audio link. The second scenario can be AI glasses / AR glasses, and the third scenario can be a mobile phone front camera / front sensor.

[0187] In the first implementation, the application scenario is a wireless earphone scenario. In this scenario, a small near-infrared (NIR) transmitter and receiver (which can be a miniature event sensor or a timestamp imaging sensor) are integrated into the earphone stem or ear hook, pointing towards the jaw / cheek / corner of the mouth area. The human-computer interaction state perception system resides in the first working mode to perform motion pre-detection; after detecting possible sound-related motion, it enters the second working mode, activates the time-modulated active light field, and performs fine perception.

[0188] In this scenario, the interaction confidence C(t) of "the wearer is speaking / preparing to speak" is output, which is used to: (i) gate and weight the audio link in noisy environments; (ii) suppress false triggers (others speaking / ambient noise); and (iii) provide priors for noise reduction / echo cancellation.

[0189] Example of motion pre-detection in this scenario: within a time window Internal statistical events or brightness changes:

[0190] in: Indicates the number of pixel events or the count of brightness changes; This refers to the set of low-resolution pixels that are being detected.

[0191] To enhance stability, an exponential moving average (EMA) is used to reduce data volatility.

[0192] in, α Smoothing coefficient (value range 0 < α < 1), This refers to the time window obtained after exponential moving average processing. Internal statistical events or brightness changes, This refers to the raw, unprocessed time window. Internal statistics on events or brightness changes.

[0193]

[0194] The human-computer interaction state perception system enters the second working mode and starts the projection of the active light field. To determine the threshold for vocalization.

[0195] In this scenario, an example of mandibular movement / skin deformation indices (obtained in the second working mode) is shown: the target area is divided into multiple sub-regions. Calculate sub-regions Central Event Center :

[0196] in, sub-region In the event pixel x The sum of coordinates sub-region In the event pixel y The sum of coordinates sub-region The count of events in the middle.

[0197] Calculate deformation / motion intensity: The physical meaning can include: the intensity of jaw movement, the amount of speech gesture indication, and the amount of facial geometric changes.

[0198] In this scenario, an example of speech filtering / enhancement through audio fusion is shown: building interaction confidence. C ( t This is used to gate audio features or acoustic model outputs and construct interactive confidence scores.

[0199] in, for The weighting coefficients, Yes The offset Sigmoid confidence mapping will Mapped to the range of 0 to 1 This refers to the scaling factor, which controls the steepness of the curve. This refers to the bias term, which is used to eliminate baseline offset, background noise, etc.

[0200] Will Used for audio link control:

[0201] in, It can provide beamforming output or voice channel estimation for the user-facing direction. It can be used for background channel estimation, and weights can be assigned using interactive confidence to control the audio link. and Both are results from audio system processing, which will not be elaborated upon here. In terms of implementation, such as... Figure 10 As shown, the interaction confidence level C ( tIt can also be used as a prior for Voice Activity Detection (VAD) or as a conditional input for noise reduction networks.

[0202] In this scenario, the interaction state perception based on skin vibration characteristics has been discussed above. Figure 8 The details have been explained in detail, so I will not repeat them here.

[0203] In the second implementation, the application scenario is AI glasses / AR glasses. In this scenario, one or more sets of projection components 12 and photosensitive components 11 are arranged on the bridge / temples of the glasses frame, pointing towards the face (eyebrows, around the eyes, mouth) or the side of the face. The human-computer interaction state perception system achieves all-weather low-power operation through a hierarchical architecture.

[0204] In this scenario, the target outputs are: (i) facial expression / emotion related facial action unit (AU) estimation; (ii) silent interaction: outputting discrete commands (e.g., 10–30 gestures / lip shapes) through mouth / jaw movements to control the AI ​​in a silent scenario; and (iii) continuous identity verification (optional).

[0205] Example of facial expression / AU regression in this scenario: with multiple sub-regions Deformation vector As input, estimate the AU intensity.

[0206] Specifically, a linear model can be built based on the input vector and predictions can be made. :

[0207] in, This refers to the coefficient. This refers to the intercept.

[0208] It is also possible to train a nonlinear model and make predictions based on the input vector. Nonlinear model:

[0209] in These can be MLP (Multilayer Perceptron), LSTM (Long Short-Term Memory), Transformer, CNN (Convolutional Neural Network), etc., all of which can be trained using labeled data. This disclosure is not limited to a specific model structure.

[0210] Example of silent interaction (lip-reading / gesture commands) in this scenario: Extract short-term features Φ(t) (such as D(t), dD / dt, frequency band energy, etc.) and classify them: Feature vector:

[0211] in, This refers to frequency band energy.

[0212] Classification model:

[0213] in: For classification models, This refers to the characteristics of time t. Input parameters are Classification model ,go through The process is converted into predicted probabilities for each category, and the category c with the highest probability is output. Category c corresponds to a predefined set of silent commands (e.g., "confirm / cancel / next / record / mute"). This approach avoids direct decoding of continuous semantics and is more suitable for low-power, low-latency human-computer interaction. Silent interaction can be used in scenarios such as silent environment control, gesture / lip-reading commands, and sign language control input.

[0214] In the third implementation, the application scenario is the front-facing camera / front-facing sensor of a mobile phone. In this scenario, it can work in conjunction with existing front-facing NIR (Near Infrared) emitters / receivers or under-display sensors. By time-modulated active light field and timestamp acquisition, dynamic deformation and motion priors of key facial regions can be obtained without relying on high frame rate video.

[0215] In this scenario, the target outputs are: (i) triggering and filtering of low-voice / privacy voice interaction; (ii) facial expression-driven enhancement of selfie videos / video conferences (facial expression capture, Avatar-driven); and (iii) robust interaction under complex lighting / motion conditions (motion priors are used for image stabilization / frame selection, etc.).

[0216] In this scenario, the spatial correspondence with the temporally modulated active light field is as follows: assuming the active light field contains multiple illumination regions. Each region has a different modulation code For sensor signals and Perform correlation analysis to obtain the response strength.

[0217] Define the active light field region Corresponding modulation code :

[0218] in, Refers to sensor signals For the illuminated area Corresponding modulation code In the time window The response intensity within, It refers to the largest value. label The corresponding region with the strongest response can be obtained. .

[0219] Establish spatial correspondence. Based on this correspondence, relative displacement or geometric changes can be calculated further using methods such as triangulation, relative parallax, and response differences.

[0220] In this scenario, the "vibration / high-frequency micro-motion" energy index (example, avoiding coherent interference path): For Perform a high-pass filter and calculate the energy.

[0221] High-pass filtering can be achieved using an engineering-level model:

[0222] Vibrational energy:

[0223] V(t) can be used as an index of "high-frequency micro-motion / vibration intensity" to distinguish between the rapid motion components of a stationary face and those related to vocalization, or to detect fine-grained changes in facial expressions. This refers to the vibration detection cycle. The above implementation does not require micrometer-level vibration measurement, nor does it rely on coherent interferometry or speckle phase decoding.

[0224] The physical meaning of V(t) includes high-frequency motion components, speech action enhancement indicators, and fine-grained changes in facial expressions.

[0225] This disclosure also provides an electronic device including a memory and a processor. The memory stores computer-executable instructions, and the processor is configured to execute the instructions to implement the perception method of the human-computer interaction state perception system described above.

[0226] It should be understood that, based on the teachings of the embodiments of this disclosure and according to actual needs, the above-described process steps can be reordered, or additional steps can be added, or some of the steps already discussed can be deleted. Furthermore, depending on the actual situation, the steps described in this disclosure can be performed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0227] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Various modifications can be made to the above embodiments, including the mutual substitution or replacement of features, depending on design requirements and other factors. Any modifications made within the teachings of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A human-machine interaction state sensing system, characterized by, include: The photosensitive component is configured to operate in a first working mode with a first power consumption and collect reflected light from a target area to form first optical change information, or to operate in a second working mode with a second power consumption and collect reflected light from the target area after being irradiated by an active light field to form second optical change information, wherein the first power consumption is less than the second power consumption. The projection component is configured to project an active light field with time-modulated characteristics onto the target region in the second operating mode; and The processing component is configured as follows: Based on the first optical change information, control the switching of the photosensitive component from a first working mode to a second working mode; Based on the second optical change information, the current human-computer interaction state is determined; In response to determining that the human-computer interaction state is an interaction, the photosensitive component is controlled to remain in the second working mode; otherwise, the photosensitive component is controlled to switch to the first working mode. The target region includes areas that move due to spoken or silent speech, and the processing component is configured to: Based on the second optical change information, the motion index of the target area is calculated; The motion index is subjected to high-pass or band-pass filtering to obtain the high-frequency components of the motion index related to the vibration of the target region; A high-frequency vibration intensity index is calculated based on the high-frequency components of the motion index, and an interactive confidence score is generated based on the high-frequency vibration intensity index. The interactive confidence score is used to characterize the credibility of the user's current vocalization action in the target area. The final speech probability is obtained by fusing the audio speech activity detection probability of the target area with the interaction confidence, and the current human-computer interaction state is determined based on the final speech probability.

2. The perception system of claim 1, wherein, in, The processing component calculates the motion indication of the target area based on the first optical change information, and controls the photosensitive component to switch from the first working mode to the second working mode if the motion indication exceeds a threshold.

3. The perception system of claim 1, wherein, The processing component performs frequency domain transformation based on the high-frequency vibration intensity index to obtain the dominant vibration frequency, and generates the interactive confidence score based on the dominant vibration frequency and the high-frequency vibration intensity index; or The processing component performs frequency domain transformation based on the high-frequency vibration intensity index to obtain the vibration dominant frequency, calculates the spectral peak significance based on the high-frequency vibration intensity index, and generates the interactive confidence score based on the spectral peak significance and the high-frequency vibration intensity index.

4. The perception system of claim 1, wherein, The processing component is further configured to: The frequency domain gain function is gated or weighted according to the interactive confidence level to obtain the gated or weighted frequency domain gain function. The frequency domain gain function is the gain function between the audio frequency domain signal of the target region and the audio frequency domain signal of the target region after noise suppression. as well as The vibration dominant frequency is obtained by performing frequency domain transformation based on the high-frequency vibration intensity index; a harmonic protection or enhancement mask is constructed using the vibration dominant frequency, and the gated or weighted frequency domain gain function is processed based on the harmonic protection or enhancement mask to obtain a refined frequency domain gain function, and the current human-computer interaction state is determined based on the refined frequency domain gain function.

5. The perception system of claim 1, wherein, The photosensitive component includes multiple pixels arranged in an array; In the first operating mode, the photosensitive component operates in the first operating mode by at least one of the following methods: pixel binning mode, downsampling mode, and enabling only a portion of the region of interest.

6. The perception system of claim 1, wherein, The projection component includes: Optical elements are configured to project the active light field onto the target region; A drive controller is configured to generate a multi-channel time-coded pulse sequence and drive the active optical field in a time sequence; and The synchronization module is configured to synchronize the timing of the time-coded pulse sequence with the timing of the reflected light collected by the photosensitive component after the target area is irradiated by the active light field.

7. The perception system of claim 6, wherein, The active light field is a structured light with temporal coding formed by encoding the multi-channel time-coded pulse sequence; The optical element forms at least two distinguishable illumination regions in space, and the time-coded structured light from different channels is projected into the different illumination regions, with the coded Hamming distance between the different channels not less than 2.

8. The sensing system according to claim 7, characterized in that, The processing component is further configured to: The event sequence output by all pixels of the photosensitive component is decoded and matched on the time axis to obtain the encoding channel where the pixel is located. The event sequence is used to record the time change of the second optical change information. Multiple adjacent pixels belonging to the same encoding channel and whose spatial distance meets the set value are aggregated into a spot pixel region; as well as Geometric inference or motion estimation is performed based on the coordinates of the center pixel of the light spot pixel region to determine the current human-computer interaction state.

9. A human-computer interaction enhancement system, characterized in that, include: The human-computer interaction state perception system according to any one of claims 1 to 8; as well as The interactive control module is configured to perform at least one of the following functions based on the current human-computer interaction state determined by the human-computer interaction state perception system: starting or stopping the voice recognition module, filtering out environmental noise, enhancing human voice, and providing control functions as a control signal.

10. A method for perceiving the state of human-computer interaction, characterized in that, include: Based on the first optical change information, the photosensitive component is controlled to switch from the first working mode to the second working mode, and the projection component is controlled to start working. The first optical change information is the optical change information formed by the photosensitive component operating in the first working mode and collecting the reflected light of the target area. The power consumption of the photosensitive component in the first working mode is less than the power consumption in the second working mode. Based on the second optical change information, the current human-computer interaction state is determined, wherein the second optical change information is the optical change information formed by the photosensitive component operating in the second working mode and collecting the reflected light after the target area is irradiated by the active light field; In response to determining that the human-computer interaction state is one of interaction, the photosensitive component is controlled to remain in the second operating mode; otherwise, the photosensitive component is controlled to switch to the first operating mode. The step of determining the current human-computer interaction state based on the second optical change information includes: Based on the second optical change information, the motion index of the target area is calculated; The motion index is subjected to high-pass or band-pass filtering to obtain the high-frequency components of the motion index related to the vibration of the target region; A high-frequency vibration intensity index is calculated based on the high-frequency components of the motion index, and an interactive confidence score is generated based on the high-frequency vibration intensity index. The interactive confidence score is used to characterize the credibility of the user's current vocalization action in the target area. The final speech probability is obtained by fusing the audio speech activity detection probability of the target area with the interaction confidence, and the current human-computer interaction state is determined based on the final speech probability.

11. The sensing method according to claim 10, characterized in that, The step of controlling the photosensitive component to switch from a first operating mode to a second operating mode based on the first optical change information includes: Based on the first optical change information, the motion indication of the target area is calculated. If the motion indication exceeds a threshold, the photosensitive component is controlled to switch from the first working mode to the second working mode.

12. The sensing method according to claim 10, characterized in that, The step of determining the current human-computer interaction state based on the second optical change information further includes: Based on the interactive confidence level, the frequency domain gain function is gated or weighted to obtain a gated or weighted frequency domain gain function, wherein the frequency domain gain function is the gain function between the audio frequency domain signal of the target region and the audio frequency domain signal of the target region after noise suppression; and The vibration dominant frequency is obtained by performing frequency domain transformation based on the high-frequency vibration intensity index; a harmonic protection or enhancement mask is constructed using the vibration dominant frequency, and the gated or weighted frequency domain gain function is processed based on the harmonic protection or enhancement mask to obtain a refined frequency domain gain function, and the current human-computer interaction state is determined based on the refined frequency domain gain function.

13. The sensing method according to claim 10, characterized in that, The determination of the current human-computer interaction state based on the second optical change information includes: The event sequence output by all pixels of the photosensitive component is decoded and matched on the time axis to obtain the encoding channel where the pixel is located. The event sequence is used to record the time change of the second optical change information. Multiple adjacent pixels belonging to the same encoding channel and with a spatial distance meeting a set value are aggregated into a spot pixel region; and Geometric inference or motion estimation is performed based on the coordinates of the center pixel of the light spot pixel region to determine the current human-computer interaction state.

14. The sensing method according to claim 13, characterized in that, The step of performing geometric inference or motion estimation based on the coordinates of the center pixel of the light spot pixel region includes: A disparity map is obtained by comparing the coordinates of the center pixels in the same encoded channel with the original projected coordinates, and geometric inference or motion estimation is performed based on the disparity map; or The center pixel coordinates of the same encoding channel and the original projection coordinates are used to form camera rays and projection rays in the coordinate system of the photosensitive component. The intersection of the camera rays and the projection rays is obtained in three-dimensional space to obtain a three-dimensional point. Geometric inference or motion estimation is performed based on the three-dimensional point.

15. An electronic device, characterized in that, include: Memory, which stores instructions that a computer can execute; The processor is configured to execute the instructions to implement the human-computer interaction state perception method as described in any one of claims 10-14.

Citation Information

Patent Citations

  • Tactile perception system and method, electronic device, storage medium, and program product

    CN121614037A

  • Speech transcription from facial skin movements

    US20230215437A1

  • Optical smart trunk opener

    WO2017102573A1