A control system and method for IoT devices based on steady-state visual evoked potentials and augmented reality.

By integrating SSVEP brain-computer interface technology into augmented reality devices, and utilizing virtual control objects and EEG signal recognition, the control challenges of IoT devices in specific scenarios have been solved, achieving seamless, intuitive, and efficient interactive control.

CN122086245APending Publication Date: 2026-05-26CHENGDU WABO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU WABO TECHNOLOGY CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing IoT device control methods suffer from physical contact limitations, environmental noise interference, and a sense of separation in interaction under certain scenarios. Traditional SSVEP system devices are bulky and difficult to meet the needs of mobile scenarios. AR interaction has failed to completely solve the core pain points of hands-occupancy and noise interference.

Method used

By integrating SSVEP-based brain-computer interface technology with augmented reality technology, virtual control objects are overlaid in the user's field of vision through wearable AR devices. The EEG signal acquisition device identifies the user's gaze intention, and combined with eye tracking and context awareness, efficient and intuitive control of IoT devices is achieved.

Benefits of technology

It enables hands-free control without the need for physical gestures or voice input, provides intuitive interaction with spatial anchoring, improves the robustness and security of control, and reduces the cognitive load on users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086245A_ABST
    Figure CN122086245A_ABST
Patent Text Reader

Abstract

This invention discloses a control system and method for Internet of Things (IoT) devices based on steady-state visual evoked potentials (SSVEP) and augmented reality (AR). The system is applied to wearable AR devices and includes an SSVEP signal acquisition module, an AR display and stimulus presentation module, a processing and recognition module, a wireless communication and control module, and an AR feedback presentation module. The AR module renders a virtual control object with a specific flickering frequency in the user's field of vision, while the acquisition module acquires EEG signals from the user's occipital cortex in real time. The processing and recognition module decodes the user's gaze intent and generates control commands using CCA or deep learning algorithms. Furthermore, the system integrates eye tracking for dual intent verification and utilizes a context-aware module to adaptively switch scene modes. This invention achieves "what you see is what you get" hands-free, silent interaction, effectively solving the control challenges of IoT devices under conditions of hand-occupancy and environmental noise interference, and possesses the advantages of high robustness and low cognitive load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of human-computer interaction technology, wearable computing and Internet of Things (IoT) technology, and in particular to a system and method for controlling Internet of Things (IoT) devices by deeply integrating steady-state visual evoked potential (SSVEP) brain-computer interface (BCI) technology with augmented reality (AR) technology. Background Technology

[0002] With the rapid development of Internet of Things (IoT) technology, smart devices have widely penetrated into fields such as smart homes, industrial automation, smart healthcare, and assisted rehabilitation. Currently, the mainstream methods for controlling these IoT devices include smartphone applications (APPs), voice assistants, and traditional physical switches.

[0003] However, these interaction methods have significant shortcomings in certain scenarios. First, there are limitations related to physical contact: in situations where users' hands are occupied or cannot move freely, such as surgeons during operations, maintenance engineers working at heights, or chefs cooking, operating via a mobile phone or physical switch is extremely inconvenient and may even lead to safety accidents. Second, there is interference from environmental noise: in acoustically complex environments such as factory workshops or noisy streets, the recognition accuracy of voice assistants drops significantly, leading to wake-up failures or incorrect commands. Finally, there is a sense of separation between the interaction and the user experience: existing technologies often require users to frequently switch their attention between the physical device, the phone screen, and the real environment, lacking an immersive and intuitive "what you see is what you get" experience.

[0004] Brain-computer interface (BCI) technology, especially BCI based on steady-state visual evoked potentials (SSVEP), offers a promising solution for achieving "hands-free" and "silent" interactive control. SSVEP refers to the stable, detectable electrical response signal generated by the visual cortex of the brain when the human eye focuses on a visual stimulus flashing at a specific frequency. However, traditional SSVEP systems typically rely on external, fixed displays or LED arrays as stimulus sources, resulting in bulky devices that are disconnected from the user's real-world environment, making them unsuitable for mobile applications.

[0005] Meanwhile, while augmented reality (AR) glasses can overlay virtual information onto the real world, current AR interactions mainly rely on gestures or voice, failing to fundamentally address the core pain points of "hands-occupying" and "noise interference." Therefore, there is an urgent need in this field for a lightweight solution that seamlessly integrates high-precision EEG recognition, contextualized AR display, and IoT control. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an IoT device control system and method based on SSVEP and AR. This solution aims to create a highly integrated, lightweight, and wearable interactive system that allows users to efficiently, intuitively, and reliably control IoT devices in the physical world simply by "gazing" at virtual controls within the AR field of view.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an IoT device control system based on steady-state visual evoked potentials and augmented reality, applied to a wearable augmented reality device, the system comprising:

[0008] An augmented reality (AR) interactive terminal is configured to acquire an environmental image within the user's current field of view and overlay at least one virtual control object in the field of view;

[0009] The virtual control object establishes a mapping relationship with the target Internet of Things (IoT) device in the physical space based on spatial coordinates or visual tags, and the virtual control object dynamically flashes with preset light modulation parameters;

[0010] An electroencephalogram (EEG) signal acquisition device is installed on the main body or wearable component of the AR interactive terminal to acquire the electroencephalogram (EEG) signals generated when the user gazes at the virtual control object.

[0011] A processor, connected to the AR interactive terminal and the EEG signal acquisition device, is configured to perform the following steps: perform frequency domain or time domain feature analysis on the EEG signal to extract steady-state visual evoked potential (SSVEP) features; match the SSVEP features with the light modulation parameters of the virtual control object to determine the target virtual control object that the user is looking at; generate control commands corresponding to the target IoT device according to the mapping relationship; and a communication module is used to send the control commands to the target IoT device.

[0012] Furthermore, the AR interactive terminal also includes an eye-tracking sensor;

[0013] The processor is also configured to execute intent confirmation logic: when the SSVEP feature matching is detected to be successful, it further determines whether the gaze point coordinates collected by the eye tracking sensor are located within the display area of ​​the target virtual control object, and whether the gaze duration exceeds a preset time threshold.

[0014] The control command is triggered only when both the SSVEP feature matching and intent confirmation logic pass. Upon receiving the SSVEP feature signal, it is determined whether the user's gaze is located within the corresponding virtual control object area and whether the gaze duration exceeds a preset threshold (e.g., 1.5 seconds). Command generation is triggered only when both conditions are met simultaneously.

[0015] Furthermore, the system also includes an environment perception module for collecting contextual data of the environment in which the AR interactive terminal is located; the contextual data includes geographical location, inertial measurement data, or object recognition results in environmental images; the processor is configured to identify the current scene mode based on the contextual data, and adaptively filter and display a set of virtual control objects associated with the current scene based on the scene mode.

[0016] Furthermore, the optical modulation parameters include scintillation frequency, scintillation phase, or a combination thereof; the feature analysis algorithm used by the processor is selected from at least one of the following: power spectral density analysis (PSDA), canonical correlation analysis (CCA), filter bank canonical correlation (FBCCA), long short-term memory network (LSTM), or convolutional neural network (CNN).

[0017] Furthermore, the AR interactive terminal renders a visual feedback layer in the field of view; the communication module is also configured to receive a status feedback signal from the target IoT device; the processor controls the visual feedback layer to change the visual attributes of the target virtual control object according to the status feedback signal, the visual attributes including color, brightness, border highlighting or dynamic micro-animation.

[0018] Furthermore, the EEG signal acquisition device includes several dry electrodes integrated on the inner side of the temple or nose pad of the AR interactive terminal; the dry electrodes are configured to fit the user's occipital visual cortex region, which covers at least one EEG measurement point including O1, Oz or O2 in the International 10-20 system.

[0019] This invention also provides a method for controlling IoT devices based on augmented reality and steady-state visual evoked potentials, comprising: displaying a virtual control object overlaid in the user's field of vision through an augmented reality (AR) interactive terminal, and driving the virtual control object to blink according to preset light modulation parameters; establishing a mapping relationship between the virtual control object and the location of IoT devices in the real environment based on spatial coordinates or visual tags; collecting electroencephalogram (EEG) signals generated when the user gazes at the virtual control object; analyzing the EEG signals using a decoding algorithm to identify implicit frequency or phase features; comparing the identified features with the light modulation parameters to lock the target device intended to be controlled by the user; generating control commands according to the mapping relationship, and sending them to the target device through a communication module.

[0020] Furthermore, before acquiring EEG signals, the method also includes performing a user calibration process: guiding the user to sequentially gaze at calibration stimulus sources of different frequencies, acquiring baseline SSVEP response data; and using the baseline SSVEP response data to perform personalized fine-tuning of the model parameters of the decoding algorithm.

[0021] Compared with existing technologies, the present invention has the following significant advantages: true hands-free control, requiring no physical movements or voice input; intuitive interaction with spatial anchoring, realizing "what you see is what you control"; high robustness and security, introducing eye-tracking assisted verification and closed-loop visual feedback; scene adaptability, reducing the user's cognitive load. Attached Figure Description

[0022] Figure 1 This is the overall architecture logic block diagram of the system of the present invention.

[0023] Figure 2 This is a schematic diagram of the hardware structure layout of the wearable AR device of the present invention.

[0024] Figure 3 This is a schematic diagram of a visual interface in which a virtual control object within the AR field of view is integrated with the real environment in one embodiment of the present invention.

[0025] Figure 4 This is a data flow diagram of the signal processing and recognition algorithm of this invention.

[0026] Figure 5 This is a flowchart of the interactive steps of the control method of the present invention.

[0027] It should be noted that the symbols for the main components used in the accompanying drawings of this specification are explained as follows:

[0028] 100 - Wearable AR device; 110 - SSVEP signal acquisition module (dry electrode); 120 - AR optical display system; 130 - Eye tracking sensor; 200 - Processing unit (SoC); 300 - Virtual control object (control); 400 - Target IoT device. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Example 1: Hardware Architecture

[0031] refer to Figure 2 This embodiment provides an integrated wearable device 100, specifically an augmented reality (AR) glasses. The SSVEP signal acquisition module 110 includes 3-5 dry electrodes integrated into the inner side of the temples and the nose pads of the glasses. The layout of the dry electrodes corresponds to the O1, Oz, and O2 regions of the occipital visual cortex of the user's head, as well as the mastoid region (A1 / A2) or the forehead region (Fpz) as reference electrodes. The glasses integrate a high-performance, low-power system-on-a-chip (SoC), including a central processing unit (CPU), a graphics processing unit (GPU), memory (RAM / Flash), and one or more wireless communication modules (supporting protocols such as Wi-Fi, Bluetooth, Zigbee, or Matter). The dry electrodes use conductive polymers or gold-plated comb-like structures to ensure good contact with the scalp without conductive paste, with impedance controlled below 50kΩ. The glasses also integrate a high-performance SoC 200, serving as the hardware carrier for the processing and recognition modules.

[0032] Example 2: Visual Stimulus and Signal Processing

[0033] See Figure 3This embodiment discloses a virtual-real fusion encoding mapping mechanism. When a user wears the wearable AR device 100 and observes a target IoT device (such as a smart fan 400) in a real environment through its optical display system 120, the AR display and stimulus presentation module overlays and renders a virtual control object 300 associated with the target IoT device above its physical spatial location or within its corresponding visual salience area. The virtual control object 300 undergoes periodic brightness or contrast modulation at a specific frequency (such as 12Hz) in a preset orthogonal frequency sequence. To ensure the stability of the stimulus frequency and eliminate beat frequency interference between the display refresh rate and the stimulus frequency, the system captures the V-Sync (vertical synchronization) signal of the display system through the processing unit 200, locking the rendering cycle of each frame to an integer multiple of the specific frequency. In specific applications, the system presets a device-frequency mapping table, for example: smart fan icon: assigned frequency 12Hz, initial phase 0; air conditioner icon: assigned frequency 10Hz, initial phase 0; desk lamp icon: assigned frequency 8.5Hz, initial phase 0.5π.

[0034] When a user's gaze is focused on the "smart fan icon", their retina is modulated by the 12Hz periodic visual stimulation, which in turn induces a steady-state visual evoked potential (SSVEP) signal containing the 12Hz fundamental frequency and its higher harmonic components in the occipital visual cortex.

[0035] The processing and recognition module performs as follows Figure 4 The algorithm flow is as follows: Preprocessing: The raw EEG signal is bandpass filtered from 4-40Hz to remove electrooculography (EOG) artifacts and electromyography (EMG) noise, and a 50Hz notch filter is used to remove power frequency interference. Feature Decoding: This embodiment provides two optional decoding schemes. Scheme A (based on statistics): The canonical correlation analysis (CCA) algorithm is used to calculate the canonical correlation coefficient between the real-time acquired multi-channel EEG signal and the system-pre-generated standard reference signal (including sine and cosine waves). Scheme B (based on deep learning): A neural network model containing a spatiotemporal feature extraction layer is constructed. This model combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). The CNN layer is used to extract the spatial distribution features of the EEG signal between each electrode channel, and the LSTM layer is used to capture the dynamic change features of the signal in the time series. This model is pre-installed in the processor through offline training and can directly output the probability distribution of each preset frequency category. Decision: The system compares the correlation coefficients (Scheme A) or classification probabilities (Scheme B) corresponding to each frequency. If the value corresponding to 12Hz is the largest and exceeds the preset confidence threshold (e.g., 0.6), it is determined that the user is looking at the target corresponding to 12Hz.

[0036] Example 3: Security Mechanisms and Context Awareness

[0037] To prevent accidental operation, this system incorporates a user status monitoring and safety module. This module utilizes an infrared eye-tracking camera 130 located inside the glasses to track the pupil position in real time. The decision logic is as follows: the system will only generate a control command when "the decoding algorithm detects obvious SSVEP features" and "eye-tracking data shows that the gaze point falls within the corresponding icon area for more than 1.5 seconds." At this time, a green progress ring will appear around the icon in the AR interface, filling up as the gaze duration increases, providing the user with clear time feedback.

[0038] The context-aware module works as follows: The system detects the user's motion state through an IMU (Inertial Measurement Unit) and determines the location using GPS or indoor positioning beacons. Scenario A: If the GPS location is "home" and the IMU data shows that the user's three-axis acceleration variance is below a preset threshold (determined as "stationary seated posture"), the system automatically loads the "living room / bedroom control mode". Scenario B: If the GPS location is "factory" and the camera recognizes a specific machine's QR code, the system automatically loads the "industrial control mode", prioritizing the display of machine tool start / stop and parameter adjustment icons.

[0039] Example 4: Interactive Control Method

[0040] like Figure 5 As shown, the method flow of this invention includes: S1 system initialization and scene scanning; S2 multimodal intent capture; S3 joint decoding and verification; S4 instruction distribution; S5 closed-loop feedback. Through the above steps, closed-loop control of IoT devices is achieved.

[0041] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A control system for an Internet of Things (IoT) device based on augmented reality and steady-state visual evoked potentials, characterized in that, include: An augmented reality (AR) interactive terminal is configured to acquire an environmental image within the user's current field of view and overlay at least one virtual control object in the field of view; The virtual control object establishes a mapping relationship with the target IoT device in the physical space based on spatial coordinates or visual tags, and the virtual control object dynamically flashes with preset light modulation parameters; An electroencephalogram (EEG) signal acquisition device is installed on the main body or wearable component of the AR interactive terminal to acquire EEG signals generated when the user gazes at the virtual control object. The processor, connected to the AR interactive terminal and the EEG signal acquisition device, is configured to perform the following steps: perform frequency domain or time domain feature analysis on the EEG signal to extract steady-state visual evoked potential (SSVEP) features; match the SSVEP features with the light modulation parameters of the virtual control object to determine the target virtual control object being gazed at by the user; generate control commands corresponding to the target IoT device according to the mapping relationship; and a communication module is used to send the control commands to the target IoT device.

2. The system according to claim 1, characterized in that, The AR interactive terminal also includes an eye-tracking sensor (130); the processor is further configured to execute intent confirmation logic: when the SSVEP feature matching is detected to be successful, it further determines whether the gaze point coordinates collected by the eye-tracking sensor are located within the display area of ​​the target virtual control object, and whether the gaze duration exceeds a preset time threshold; the generation of the control command is triggered only when both the SSVEP feature matching and intent confirmation logic are successful.

3. The system according to claim 1, characterized in that, The system also includes an environment perception module for collecting contextual data of the environment in which the AR interactive terminal is located; the contextual data includes geographical location, inertial measurement data, or object recognition results in environmental images; The processor is configured to identify the current scene mode based on the context data, and adaptively filter and display a set of virtual control objects associated with the current scene based on the scene mode.

4. The system according to claim 1, characterized in that, The optical modulation parameters include scintillation frequency, scintillation phase, or a combination thereof; the feature analysis algorithm used by the processor is selected from at least one of the following: power spectral density analysis (PSDA), canonical correlation analysis (CCA), filter bank canonical correlation analysis (FBCCA), long short-term memory network (LSTM), or convolutional neural network (CNN).

5. The system according to claim 1, characterized in that, The AR interactive terminal renders a visual feedback layer in the field of view; the communication module is also configured to receive a status feedback signal from the target IoT device; the processor controls the visual feedback layer to change the visual attributes of the target virtual control object according to the status feedback signal, the visual attributes including color, brightness, border highlighting or dynamic micro-animation.

6. The system according to claim 1, characterized in that, The EEG signal acquisition device includes several dry electrodes integrated on the inner side of the temple or nose pad of the AR interactive terminal; the dry electrodes are configured to fit the user's occipital visual cortex region, which covers at least one EEG measurement point including O1, Oz or O2 in the international 10-20 system.

7. A method for controlling Internet of Things (IoT) devices based on augmented reality and steady-state visual evoked potentials, characterized in that, include: Step S1: The virtual control object is superimposed and displayed in the user's field of vision through the augmented reality (AR) interactive terminal, and the virtual control object is driven to flash dynamically according to the preset light modulation parameters as a steady-state visual evoked stimulus source. Step S2: Use an EEG signal acquisition device to collect EEG signals generated by the user when gazing at the virtual control object in real time; Step S3: Perform frequency domain or time domain feature analysis on the EEG signal to extract steady-state visual evoked potential (SSVEP) features; Step S4: Match the extracted SSVEP features with the optical modulation parameters of each virtual control object to determine the target virtual control object currently being gazed at by the user; Step S5: Generate control commands corresponding to the target IoT device based on the preset mapping relationship; Step S6: Send the control command to the target IoT device through the communication module to realize interactive control of the device in the physical space.

8. The method according to claim 7, characterized in that, Before acquiring EEG signals, the process also includes performing a user calibration procedure: guiding the user to sequentially gaze at calibration stimulus sources of different frequencies, acquiring baseline SSVEP response data; and using the baseline SSVEP response data to perform personalized fine-tuning of the model parameters of the decoding algorithm.