Brain-computer interface system and brain-computer interface interaction method

CN121095544BActive Publication Date: 2026-09-08TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511225598.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-09-08
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

然而,用户的视野都受限于自身视角或第一人称摄像头视角,无法获取全局场景信息,这限制了其在实际应用中的实用性

Benefits of technology

1.本发明提供的基于第三人称视角的目标检测模块,通过在环境中部署独立摄像头,以第三人称视角捕获周围环境的信息,配置改进型YOLOv5网络模型,设定置信度阈值,设定交并比阈值;减轻了用户佩戴设备的负担和不适感,克服了视野受限,避免了因头戴设备位移导致的刺激位置偏移问题,避免了视觉卡顿问题,同时保留了最优预选框。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095544B_ABST
    Figure CN121095544B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of brain-computer interfaces, in particular to a brain-computer interface system and a brain-computer interface interaction method. A target detection module comprises an image acquisition unit and a target detection unit; wherein the image acquisition unit acquires image data comprising a controlled object and global environment information in real time; the target detection unit is configured with a pre-trained improved YOLOv5 network model, the number of C3 modules is simplified, and a non-maximum suppression algorithm is optimized; the improved YOLOv5 network model performs target detection after receiving the image data output by the image acquisition unit; the frame coordinates of the optimal preselected frame are selected to complete target detection. The application realizes reduction of jumps caused by environmental changes, improvement of timing synchronicity, guarantee of the stability of stimulation frequency output, keeping of the stability of the position of stimulation superimposed on the target position of the controlled object, and reduction of the visual lag phenomenon caused by image refresh delay or detection lag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of brain-computer interface technology, and more particularly to a brain-computer interface system and a brain-computer interface interaction method. Background Technology

[0002] Brain-computer interfaces (BCIs), as a technology that can establish a real-time communication and control bridge between the biological brain and external devices or environments, thereby enabling direct interaction between the brain and external devices, are gradually becoming a research hotspot across multiple disciplines. It cleverly combines methods, ideas, and concepts from neurophysiology, computer science, and engineering, aiming to build a real-time bidirectional connection between the brain and machine devices, opening up a completely new communication and control pathway that bypasses peripheral nerves and muscles. It is a highly transformative human-computer interaction technology. BCI technology encompasses non-invasive, semi-invasive, and invasive methods; based on the differences in input signals, it can also be divided into BCIs based on motor imagery, P300, and steady-state visual evoked potentials.

[0003] However, current brain-computer interface (BCI) technology still faces a series of serious challenges in its development. In existing technologies, presenting visual evoked potentials (VEP) stimuli using AR devices typically requires users to wear head-mounted displays. However, these devices generally suffer from drawbacks such as large size and weight. Prolonged wear can lead to discomfort, visual fatigue, and affect natural head movements, thus limiting the user's freedom of movement. This not only reduces the user's interactive experience but also limits the device's application in dynamic scenarios, making it difficult to support long-term BCI interactive tasks. Furthermore, VEP stimulation usually requires users to gaze at a fixed area, but in a first-person perspective, the user's gaze naturally shifts, making it difficult to maintain a stable focus on the stimulus source. Wearing AR devices too tightly can affect user comfort, while wearing them too loosely can cause head movements to shift the device, leading to a shift in the stimulus position, affecting the stability and evoked effect of the VEP signal, and consequently impacting the accuracy of the BCI system. When presenting Visual Evoked Potentials (VEP) stimuli and a first-person perspective environment on a PC, it is necessary to use a camera on the controlled object or place a camera on the controlled object, allowing the user to perceive the environment through the camera's perspective. The environmental image information captured by the camera is then displayed on the PC, and the necessary virtual stimulus elements (such as images and text) are superimposed on the image information. However, the user's field of vision is limited to their own perspective or the first-person camera's perspective, making it impossible to obtain global scene information, which limits its practicality in real-world applications.

[0004] Therefore, there is an urgent need to propose a brain-computer interface solution that reduces user dependence on wearing devices, overcomes the instability and limited field of vision of first-person perspective stimulation, avoids stimulation position shift caused by user movement, reduces computational load, improves temporal synchronization, reduces visual stuttering, and ensures the stability of augmented reality stimulation. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by providing a brain-computer interface system and brain-computer interface interaction method. By deploying a third-person perspective camera in the environment to capture information about the surrounding environment, configuring an improved YOLOv5 network model, smoothly transitioning the optimal pre-selection box of the current frame of the target detection image, and combining computer vision algorithms and augmented reality technology to perform real-time adsorption and adjustment of virtual stimuli, the invention avoids the problems of stimulus position shift, visual lag, and limited field of view in existing technologies, thereby improving the comfort of users wearing the device.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a target detection module based on a third-person perspective, comprising an image acquisition unit and a target detection unit; wherein, the image acquisition unit acquires image data including information of the controlled object and the global environment in real time; the target detection unit is configured with a pre-trained improved YOLOv5 network model; the improved YOLOv5 network model is based on the original YOLOv5 network model by removing the P5 branch and its related FPN / PAN layers, removing the SPP module, retaining only the two detection scales P3 and P4, and simplifying the number of C3 modules to reduce network complexity; the improved YOLOv5 network model also improves the non-maximum suppression algorithm, i.e., by setting a confidence threshold and an intersection-union ratio threshold; The improved YOLOv5 network model receives image data output from the image acquisition unit and performs target detection. Specifically, it extracts image features through forward inference and performs target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores. The optimal box coordinates are then selected by using confidence thresholds and intersection-union (IUU) thresholds to complete target detection.

[0007] As one possible implementation, improvements to the nonmaximum suppression algorithm specifically include: Set a confidence threshold and retain the pre-selected boxes that are greater than or equal to the confidence threshold; Set an intersection-union ratio (IU) threshold. From the preselected boxes that are greater than or equal to the confidence threshold, filter out the preselected boxes that are greater than or equal to the IU threshold. Then, select the preselected box with the higher confidence among the two preselected boxes with the largest IU as the optimal preselected boxes.

[0008] As one possible implementation, the confidence threshold is 0.6 and the crossover ratio (CUP) threshold is 0.4.

[0009] As one possible implementation, the latency of the target detection module is less than or equal to 100ms.

[0010] In a second aspect, the present invention provides an augmented reality stimulation module for target detection and aVEP stimulation adsorption, including a preprocessing unit, a weak asymmetric visual stimulation unit, and an adsorption stimulation image generation unit. The preprocessing unit receives the target detection image output by the target detection module provided by the first aspect, and performs smooth transition processing on the bounding box coordinates of the optimal preselected box in the current frame to obtain the bounding box coordinates of the preprocessed optimal preselected box. The weak asymmetric visual stimulation unit is used to provide weak asymmetric visual stimulation; The adsorption stimulus image generation unit constructs a two-way temporal constraint mechanism based on Shannon sampling theorem, which includes target detection and weak asymmetric visual stimuli. That is, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli, and the superposition position of the weak asymmetric visual stimuli is adjusted in real time so that it is adsorbed around the box coordinates of the preprocessed optimal preselected box, that is, around the controlled object.

[0011] As one possible implementation, the coordinates of the optimal preselected box in the current frame are smoothly transitioned, specifically including the following steps: Preserve the coordinates of the optimal preselected box from the previous frame; After obtaining the coordinates of the optimal preselection box in the current frame, a smoothing factor is used. The coordinates of the optimal preselected box in the current frame and the coordinates of the optimal preselected box in the previous frame are weighted and summed to obtain the coordinates of the optimal preselected box in the current frame after smooth transition processing.

[0012] As one possible implementation, the sampling frequency of the target detection image is greater than or equal to 32Hz, and the frequency of the weak asymmetric visual stimulus is 4Hz.

[0013] Thirdly, the present invention provides a brain-computer interface system, including a target detection module, an augmented reality stimulation module, an EEG data acquisition module, and an EEG data analysis module; The target detection module performs the following steps: acquiring continuous image frames of the controlled object; preprocessing the continuous image frames to match the input requirements of the pre-trained improved YOLOv5 network model, including at least scaling, normalization, and color space conversion; using the pre-trained improved YOLOv5 network model to perform target detection on the preprocessed continuous image frames, i.e., extracting image features through forward inference and performing target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores; and using confidence thresholds and intersection-over-union (IoU) thresholds to select the optimal box coordinates of the pre-selected boxes to complete the target detection. The augmented reality stimulus module performs the following steps: receiving the target detection image output by the target detection module; performing smooth transition processing on the bounding box coordinates of the optimal preselected box in the current frame to obtain the preprocessed bounding box coordinates; providing weak asymmetric visual stimuli; constructing a two-way temporal constraint mechanism based on Shannon sampling theorem, including target detection and weak asymmetric visual stimuli, i.e., the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli; adjusting the superposition position of the weak asymmetric visual stimuli in real time so that it is attached to the bounding box coordinates of the preprocessed optimal preselected box; and presenting the augmented reality stimulus to the user. The EEG data acquisition module records the EEG signals induced by augmented reality stimuli in real time and sends them to the EEG data analysis module. The EEG data analysis module decodes and analyzes the EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

[0014] Fourthly, the present invention provides a brain-computer interface interaction method, comprising the following steps: Acquire consecutive image frames of the controlled object; preprocess the consecutive image frames to match the input requirements of the pre-trained improved YOLOv5 network model. Preprocessing includes at least scaling, normalization, and color space conversion; use the pre-trained improved YOLOv5 network model to perform target detection on the preprocessed consecutive image frames, i.e., extract image features through forward inference and perform target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores; use confidence thresholds and intersection-over-union (IoU) thresholds to filter out the box coordinates of the optimal pre-selected boxes to complete target detection. The system receives the target detection image output by the target detection module, performs smooth transition processing on the coordinates of the optimal preselected box in the current frame to obtain the preprocessed coordinates of the optimal preselected box; provides weak asymmetric visual stimuli; constructs a two-way temporal constraint mechanism based on Shannon sampling theorem, including target detection and weak asymmetric visual stimuli, that is, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli, and adjusts the superposition position of the weak asymmetric visual stimuli in real time so that it is attached to the area around the preprocessed optimal preselected box coordinates; and presents augmented reality stimuli to the user. The system records EEG signals induced by augmented reality stimuli in real time, decodes and analyzes the EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

[0015] Compared with the prior art, the beneficial effects of this invention are as follows: 1. The target detection module based on third-person perspective provided by the present invention captures information about the surrounding environment from a third-person perspective by deploying an independent camera in the environment, configuring an improved YOLOv5 network model, setting a confidence threshold, and setting an intersection-over-union (IoU) threshold; it reduces the burden and discomfort of users wearing the device, overcomes the limitation of field of view, avoids the problem of stimulus position shift caused by head-mounted device displacement, avoids visual lag, and retains the optimal preselected box.

[0016] 2. The augmented reality stimulus module for target detection and aVEP stimulus adsorption provided by this invention improves temporal synchronization and ensures the stability of augmented reality stimuli through a dual-thread architecture that combines target detection and aVEP stimulation. It performs smooth transition processing on the bounding box coordinates of the optimal pre-selected box in the current frame of the target detection image, and combines computer vision algorithms and augmented reality technology to perform real-time adsorption and adjustment of virtual stimuli, ensuring that the stimuli are always within the user's visual range.

[0017] 6. The brain-computer interface system provided by this invention includes a target detection module, an augmented reality stimulation module, an EEG data acquisition module, and an EEG data analysis module. In the target detection module, continuous image frames are preprocessed to match the input requirements of a pre-trained improved YOLOv5 network model. The preprocessing includes at least scaling, normalization, and color space conversion. The pre-trained improved YOLOv5 network model is used to perform target detection on the preprocessed high-quality continuous image frames, which can effectively control the quality of the acquired images. In the augmented reality stimulation module, virtual visual stimuli are superimposed on the optimal pre-selected box of the target detection image through smooth transition processing, thereby achieving a more immersive and realistic interactive experience. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a schematic diagram of the SSVEP brain-computer interface system structure proposed in an embodiment of the present invention; Figure 2 This is a schematic diagram of adsorption stimulation from a third-person perspective, as proposed in an embodiment of the present invention. Figure 3 This is a schematic diagram of the third-person perspective adsorption stimulation scheme architecture proposed in an embodiment of the present invention; Figure 4 This is a schematic diagram of a typical brain-computer interface structure proposed in an embodiment of the present invention; Figure 5 This is a schematic diagram of the various performance indicators for model training proposed in this embodiment of the invention. Figure 6 This is a schematic diagram of the confusion matrix data proposed in an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating the timing synchronization optimization principle proposed in an embodiment of the present invention; Figure 8 This is a schematic diagram of the original YOLOv5 model structure proposed in the embodiments of the present invention; Figure 9 This is a schematic diagram of the improved YOLOv5 model structure proposed in an embodiment of the present invention; Figure 10 This is a schematic diagram of the AR superimposed stimulation interface proposed in an embodiment of the present invention; Figure 11 This is a schematic diagram of AR-based SSVEP fusion visual stimulation proposed in an embodiment of the present invention; Figure 12 This is a schematic diagram of PC-based SSVEP fusion visual stimulation proposed in an embodiment of the present invention; Figure 13 This is a schematic diagram of the spatial arrangement of visual stimuli and guiding points proposed in an embodiment of the present invention; Figure 14 This is a flowchart of the time smoothing algorithm proposed in an embodiment of the present invention. Detailed Implementation

[0019] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0020] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0021] In this invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, "at least one of a, b, or c" can represent: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0022] With the rapid development of non-invasive BCI technology, BCI based on visually evoked potentials (VEPs) has attracted widespread attention due to its high information transmission rate and decoding reliability. Among them, the SSVEP-BCI paradigm, with its advantages of stable evoked features and high signal-to-noise ratio, has become one of the mainstream BCI paradigms. (See [link to relevant documentation]). Figure 1 However, prolonged direct viewing of flickering stimuli can lead to eye dryness, soreness, and other discomfort. Researchers have proposed solutions from multiple perspectives. These improvements have reduced visual fatigue caused by direct viewing to some extent, but have not eliminated visual occupation. Furthermore, some improvements still rely on large-sized visual stimuli or expensive induction methods.

[0023] In practical control scenarios, the overlay of the environment and virtual stimuli can better adapt to complex dynamic changes and provide real-time feedback. Overlay images can be presented through both AR devices and PCs. AR devices use cameras and sensors to acquire data from the real environment and combine this with computer vision technology to analyze the scene, dynamically adjusting virtual stimuli and displaying the actual scene. PCs, on the other hand, capture images through the camera of the controlled object or by placing a camera on the controlled object, and use computer vision algorithms to overlay virtual elements onto the screen image.

[0024] Existing technologies face multiple limitations when presenting virtual stimuli and environmental images overlaid using AR devices or PCs. When using AR devices for virtual stimulus presentation, the physical characteristics of the head-mounted device directly impact the user experience. Due to the integration of modules such as cameras, sensors, and displays, the device is large and heavy, leading to physiological discomfort and visual fatigue with prolonged wear. This discomfort affects user comfort, weakens the user's willingness to continue operating, and limits the natural range of head movement, making it difficult for users to flexibly adjust their body movements in dynamic environments. When the device shifts due to improper fit, the spatial correspondence between the virtual stimulus source and the actual visual focus is disrupted, resulting in inaccurate spatial characteristics of the induced EEG signals and ultimately reducing the decoding accuracy of the BCI system. At the interaction level, VEP stimulation typically requires users to maintain a stable gaze at a specific area. However, in real-world scenarios, users need to continuously observe dynamically changing environmental information from a first-person perspective, and their gaze inevitably moves with the task requirements. AR devices also struggle to ensure that users maintain focus on the stimulus source while moving, leading to unstable stimulus positions, affecting the induced VEP signal effect, and consequently impacting the system's accuracy.

[0025] This invention aims to provide a brain-computer interface system and brain-computer interface interaction method. By deploying an independent camera in the environment to capture information about the surrounding environment from a third-person perspective, and by configuring a pre-trained improved YOLOv5 network model, setting confidence thresholds and intersection-over-union (IoU) thresholds, a dual-thread architecture combining target detection and aVEP stimulation is used to smoothly transition the optimal pre-selected box of the current frame of the target detection image. Furthermore, computer vision algorithms and augmented reality technology are combined to perform real-time adsorption and adjustment of virtual stimuli. For third-person perspective stimuli adsorption, see [link to documentation]. Figure 2 This avoids the problems of stimulus position displacement, visual lag, and limited field of view that exist in existing technologies, improves the temporal synchronization, ensures the stability of augmented reality stimuli, and ensures that the stimulus is always within the user's field of vision, greatly reducing the user's burden and improving the comfort of long-term use.

[0026] In a first aspect, the present invention provides a brain-computer interface system, see [link to previous document]. Figure 3It includes a target detection module, an augmented reality stimulation module, an EEG data acquisition module, and an EEG data analysis module.

[0027] The target detection module performs the following steps: acquiring continuous image frames of the controlled object; preprocessing the continuous image frames to match the input requirements of the pre-trained improved YOLOv5 network model, including at least scaling, normalization, and color space conversion; using the pre-trained improved YOLOv5 network model to perform target detection on the preprocessed continuous image frames, i.e., extracting image features through forward inference and performing target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores; and using confidence thresholds and intersection-over-union (IoU) thresholds to select the optimal box coordinates of the pre-selected boxes to complete the target detection. The augmented reality stimulus module performs the following steps: receiving the target detection image output by the target detection module; performing smooth transition processing on the bounding box coordinates of the optimal preselected box in the current frame to obtain the preprocessed bounding box coordinates; providing weak asymmetric visual stimuli; constructing a two-way temporal constraint mechanism based on Shannon sampling theorem, including target detection and weak asymmetric visual stimuli, i.e., the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli; adjusting the superposition position of the weak asymmetric visual stimuli in real time so that it is attached to the bounding box coordinates of the preprocessed optimal preselected box; and presenting the augmented reality stimulus to the user. The EEG data acquisition module records the EEG signals induced by augmented reality stimuli in real time and sends them to the EEG data analysis module; The EEG data analysis module decodes and analyzes EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

[0028] In this system, visual perception and brain-controlled feedback are key to achieving control. This solution constructs an augmented reality stimulus module as its core, acquiring environmental information through cameras and detecting the real-time position of the controlled object. YOLOv5 object detection technology is used to achieve real-time tracking of the autonomous vehicle, and PsychoPy rendering is combined to generate augmented reality stimuli that are stably attached to the controlled object.

[0029] Brain-computer interface (BCI) systems based on electroencephalography (EEG) can non-invasively establish an information transmission channel between the brain and external devices, thereby helping people with severe movement disorders or other healthy people in need to communicate amicably with the external environment. Figure 4A complete brain-computer interface (BCI) system comprises a stimulation module, an acquisition module, and a processing module. The stimulation module uses a display to encode visual stimuli at a fixed frequency. When the subject focuses on a flashing target, specific visual steady-state potentials are induced in the occipital region. The acquisition module obtains electroencephalographic (EEG) signals representing brain activity and cognitive processes from the user's cerebral cortex or from within the cranial cavity via implanted electrodes and acquisition devices. The processing module extracts relevant EEG characteristic signals through a series of signal processing methods and performs pattern recognition to convert these signals into verbal commands.

[0030] Currently, when presenting VEP stimuli on a PC, although the camera captures images of the real environment and combines them with virtual elements for rendering, the user's field of view is limited by the angle of the first-person camera. This fixed-viewpoint presentation mode restricts the dimensions of environmental information acquisition, preventing users from obtaining information about the overall scene. This limitation prevents users from having a comprehensive perception of the environment, affecting its practicality in real-world applications.

[0031] The EEG data acquisition module (acquisition device name: Neuroscan SynAmps2) is used to record EEG signals induced by aVEP stimulation in real time, record high-precision signals, and support data storage and export. This module receives serial port tags sent from the augmented reality stimulation module, captures EEG data by capturing cutoff bits, and transmits the acquired EEG signals to the EEG data analysis module for further processing via TCP / IP protocol. It includes: (1) Input: EEG signal; (2) Communication: TCP / IP communication, serial port communication; (3) Output: EEG signal.

[0032] The EEG data analysis module is responsible for processing and analyzing the EEG signals acquired by the EEG signal acquisition module, extracting features and classifying them, establishing offline and online processing models, and then outputting brain control commands. This module receives the EEG signals transmitted by the EEG data acquisition module via TCP / IP protocol, decodes and analyzes them to extract the user's control commands, and sends the results to the augmented reality stimulation module of the first computer via UDP protocol to achieve control over the controlled object. It includes: (1) Input: EEG signal; (2) Communication: UDP communication, TCP / IP communication; (3) Output: control commands.

[0033] The pre-trained improved YOLOv5 network model includes: S101. Configure the network environment: Select the Windows operating system and train the model based on the PyTorch deep learning framework; install the necessary dependency libraries, including numpy, opencv-python, torch, torchvision, tqdm and other tools for image preprocessing and data loading; configure CUDA and cuDNN hardware acceleration to make full use of GPU computing resources and improve training efficiency.

[0034] S102. Configure network model parameters: Set the batch size appropriately based on GPU memory availability to balance training speed and memory usage; set the input image size to suit the network structure requirements while ensuring detection accuracy. Adjust the anchor box size according to the data distribution to better match the true size of the target, thereby improving detection performance; set training hyperparameters, such as the initial learning rate, learning rate decay strategy (e.g., cosine decay or step decay), and weight decay, to optimize the model's convergence speed and stability. Furthermore, set the maximum number of training iterations (epochs), typically choosing 100–300 epochs, and decide whether to terminate early based on the loss reduction during training.

[0035] S103 Model Training: YOLOv5 is trained using the training set, and the parameters are updated using the stochastic gradient descent (SGD) optimizer. The network weights are continuously optimized through the backpropagation algorithm to reduce the loss function value and improve the model's recognition accuracy. Data augmentation and regularization strategies (such as Dropout and L2 regularization) are used to prevent overfitting and improve generalization ability.

[0036] S104. Model Testing and Evaluation: Load the test dataset and input it into the trained model. Use the test set to evaluate the model performance and calculate metrics such as mAP (mean accuracy), precision, and recall. See [link to relevant documentation]. Figure 5 Adjust hyperparameters based on evaluation results to improve model stability and accuracy; further analysis of error categories can be conducted using a confusion matrix, and the model can be optimized for misidentification. See the confusion matrix documentation. Figure 6 .

[0037] S105. Model Optimization and Saving: A learning rate decay strategy is adopted to improve convergence performance; YOLOv5 does not save the model in each epoch by default during training, but only retains the model parameters and weights of the latest training last.pt and the model parameters and weights of the best performing model on the validation set best.pt for subsequent deployment and application.

[0038] Secondly, this invention provides a target detection module based on a third-person perspective, including an image acquisition unit and a target detection unit; wherein, the image acquisition unit acquires image data including information about the controlled object and the global environment in real time; the target detection unit is configured with a pre-trained improved YOLOv5 network model; the improved YOLOv5 network model is based on the original YOLOv5 network model by removing the P5 branch and its related FPN / PAN layers, removing the SPP module, retaining only the two detection scales P3 and P4, and simplifying the number of C3 modules to reduce network complexity; the improved YOLOv5 network model also improves the non-maximum suppression algorithm, that is, by setting a confidence threshold and an intersection-over-union (IoU) threshold; The improved YOLOv5 network model receives image data output from the image acquisition unit and performs target detection. Specifically, it extracts image features through forward inference and performs target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores. The optimal box coordinates are then selected by using confidence thresholds and intersection-union (IUU) thresholds to complete target detection.

[0039] As one possible implementation, improvements to the nonmaximum suppression algorithm specifically include: Set a confidence threshold and retain the pre-selected boxes that are greater than or equal to the confidence threshold; Set an intersection-union ratio (IU) threshold. From the preselected boxes that are greater than or equal to the confidence threshold, filter out the preselected boxes that are greater than or equal to the IU threshold. Then, select the preselected box with the higher confidence among the two preselected boxes with the largest IU as the optimal preselected boxes.

[0040] As one possible implementation, the confidence threshold is 0.6 and the crossover ratio (CUP) threshold is 0.4.

[0041] As one possible implementation, the latency of the target detection module is less than or equal to 100ms.

[0042] The image acquisition unit acquires global environmental information in real time through a camera and captures the position of the controlled object, facilitating subsequent position tracking. The image acquisition unit transmits image data to the augmented reality stimulus module via a USB interface for the presentation of augmented reality stimuli and the generation of control commands. It includes: (1) Input: environmental information; (2) Communication: USB communication; (3) Output: image data.

[0043] After data preparation is complete, model building and optimization, model training and testing begin. Finally, the trained model is applied to object detection and localization regression. The specific implementation process is as follows: S201 Data Acquisition: The camera in the image acquisition unit records video of the scene where the controlled object (unmanned vehicle) is located, including various directions and angles, both static and dynamic processes. This video is then decomposed into frames and sampled at equal intervals using OpenCV to obtain image data containing the unmanned vehicle. A large amount of data needs to be collected to allow the model to overfit, laying the foundation for the stability of subsequent stimulus position changes.

[0044] S202 Screening and Organizing: The collected images are screened to ensure that the data covers different angles and positions to improve the model's generalization ability. During the organizing process, duplicate, unclear, or irrelevant images must be removed, and the images are classified and organized according to the autonomous vehicle's position and posture in the scene to ensure the diversity of the dataset and avoid bias during training.

[0045] S203 Data Labeling: Data labeling is a crucial step in training deep learning models. The quality of the labeling directly impacts the model's performance. Tools like LabelImg are used to manually label the selected images. During labeling, targets (autonomous vehicles) in the images must be bounded according to a unified naming convention, and corresponding label files must be generated. Each label file records the location information of the autonomous vehicle in the image, including its coordinates, category, and other attributes. Label files are typically in TXT format, with each line recording one type of target and its corresponding location information. Accuracy in labeling is essential to avoid omissions or mislabeling, ensuring the model learns accurate object recognition and localization capabilities.

[0046] S204 Data Augmentation: Enhances images by performing operations such as flipping, contrast enhancement, cropping, and mirroring to expand the size of the dataset, enabling the model to learn richer features during training and thus improve its performance on unknown data.

[0047] S205 Data Partitioning: To evaluate the model's training performance, the dataset needs to be divided into training and test sets. The training set is used for model training, while the test set is used to evaluate the model's performance on unseen data. Typically, the training set comprises 80% of the total dataset, and the test set comprises 20%. The data partitioning method can be further optimized using methods such as cross-validation. When partitioning the data, it is necessary to ensure that the distribution of the training and test sets is similar to avoid data leakage. The purpose of data partitioning is to ensure that the model can learn general features during training and provide reliable evaluation results during testing.

[0048] The improved YOLOv5 network model includes: S206. Determine the Detection Network Model: The feature extraction network structure of the deep neural network was determined to be the YOLOv5 object detection network. YOLOv5 is a lightweight and efficient object detection network, using CSPDarknet as the backbone for feature extraction. Its neck uses a PANet structure to enhance feature fusion and improve detection accuracy. The detection head adopts the YOLO approach, predicting the object class and bounding box through detection layers of different scales. YOLOv5 uses a focus structure to improve feature extraction efficiency. In version v5.x, the SPP module is used for multi-scale feature fusion, while in version v6.x, the SPPF module replaces the SPP module to enhance the receptive field. Furthermore, YOLOv5 supports various model sizes (such as YOLOv5s, YOLOv5m, YOLOv5l, YOLOv5x), allowing selection of an appropriate model size based on available computing resources. During training, optimization strategies such as automatic anchor box learning, adaptive image scaling (Mosaic Augmentation), and label smoothing were also employed.

[0049] S207. Improved Network Structure: Based on the original YOLOv5 network, the P5 branch and its associated FPN / PAN layers are removed, along with the SPP module, retaining only the P3 and P4 detection scales. The number of C3 modules is also simplified to reduce network complexity and computational complexity. This significantly reduces the computational load during the inference phase, improving detection speed while ensuring model stability. Furthermore, this structural adjustment helps improve the real-time performance of object detection and lays the foundation for the stability of subsequent stimulus timing.

[0050] S208. Improved Non-Maximum Suppression (NMS) Algorithm: The Non-Maximum Suppression (NMS) process is optimized to ensure the reliability and stability of detection results, while avoiding interference from redundant detection boxes to target tracking and control commands. Detection boxes meeting certain criteria are selected by setting a confidence threshold. The box with the highest confidence among these is chosen as the final detection result. If no box meets the confidence requirement, an empty result is returned, indicating no target was detected, and no control command is output. This ensures that only the pre-selected boxes with the highest confidence are retained, laying the foundation for the stability of subsequent stimulus positions. The specific steps are as follows: (1) Set confidence threshold Only confidence level is retained. The detection box is designed to filter out low-confidence targets and reduce false detections. (2) Set the IoU (Intersection over Union) threshold In the remaining detection boxes, if two boxes Then only the box with the highest confidence is retained, further reducing the interference caused by duplicate detection; (3) If the confidence scores of all boxes are equal If the target is lost, no detection result will be output, and no control command will be sent to ensure the stability of subsequent target detection results and reduce the risk of erroneous control.

[0051] S209. Real-time Image Acquisition and Preprocessing: Continuous image frames of the controlled object are acquired through the camera of the image acquisition unit. The input images are preprocessed to match the input requirements of the YOLOv5 model. The images are adjusted to the size required by the YOLOv5 input layer (e.g., 640×640) using a proportional scaling method to avoid target distortion. Simultaneously, the images are normalized to adapt to the input specifications of the deep neural network. Furthermore, color space conversion (e.g., RGB to BGR) is used to enhance image contrast and improve detection performance.

[0052] S210. Deep learning model object detection: Use the trained YOLOv5 model to perform object detection on the acquired images, extract image features through forward inference and perform object classification and localization, and obtain the coordinates (x, y, w, h), category (there is only one category here, i.e., the controlled object), and confidence of the preselected bounding box of the controlled object.

[0053] S211. Non-maximum suppression (NMS) processing (improved NMS algorithm): The improved non-maximum suppression (NMS) is used to filter redundant target boxes with high overlap, retaining only the best detection results; an IoU (Intersection over Union) threshold is set to reduce false detections and duplicate detections, thereby improving target localization accuracy.

[0054] S212. Detection result output: Obtain the coordinates (x, y, w, h) of the pre-selected bounding box of the controlled object and transmit the detection result to the augmented reality stimulus module.

[0055] S213. Localization Regression and Stimulus Overlay: The YOLOv5 network predicts the target's position in the image by regressing the coordinates of the preselected bounding box of the controlled object, and adjusts the overlay position of the aVEP stimulus in real time to make it adhere to the controlled object.

[0056] Among them, temporal synchronization optimization; in order to ensure the frame rate and stimulus frequency of the returned images, Shannon sampling theorem is incorporated to optimize the network model. Combined with the design of a dual-thread architecture, temporal synchronization is improved, thereby ensuring the stability of augmented reality stimuli.

[0057] S214. In the process of combining object detection with augmented reality visual stimuli, ensuring temporal synchronization is crucial. A bidirectional temporal constraint mechanism is constructed based on Shannon's sampling theorem, requiring the sampling frequency to be at least twice the highest frequency of the signal to fully reconstruct it. By modifying the structure of the YOLOv5 model, removing the P5 branch and its associated FPN / PAN layers, and also removing the SPP module, only retaining the P3 and P4 detection scales, and simplifying the number of C3 modules, the computational load is reduced, thereby decreasing the model inference time and improving temporal synchronization. With the original image frame rate at 30fps, after adding the YOLOv5 model, the time interval between processed return image frames will not exceed 30ms, and the image frequency will not be less than 32Hz, which is more than twice the stimulus frequency (4Hz) (8Hz), thus ensuring complete signal sampling and meeting the temporal synchronization requirements. The temporal synchronization optimization principle diagram is shown below. Figure 7 As shown, through a dual-thread architecture combining object detection and aVEP stimulation, thread 1 uses a camera to acquire images, obtaining the raw video stream at a rate of 30fps (frame time approximately 33ms). The system combines CPU affinity management to bind the improved YOLOv5 model inference task to the high-performance computing core. After model processing, the frame rate of the returned object detection image is greater than 32Hz. Meanwhile, thread 2 performs aVEP visual stimulation at a frequency of 4Hz. The frame rate of the object detection image (>32Hz) is more than twice that of the aVEP stimulation (4Hz) (8Hz), satisfying the Shannon sampling theorem. Therefore, the visual stimulus signal can be completely sampled.

[0058] S215. By improving the Non-Maximum Suppression (NMS) algorithm, a confidence threshold is set to filter out detection boxes that meet the criteria. The box with the highest confidence among the qualified boxes is selected as the final detection result. If no box meets the confidence requirement, an empty result is returned, indicating no target was detected, and no control command is output. This ensures that only the pre-selected boxes with the highest confidence are retained, further reducing positional fluctuations. The specific steps are as follows: (a) Setting a confidence threshold Only confidence level is retained. The detection box is designed to filter out low-confidence targets and reduce false detections. (b) Set the IoU (Intersection over Union) threshold In the remaining detection boxes, if two boxes Then only the box with the highest confidence is retained, further reducing the interference caused by duplicate detection; (c) If the confidence scores of all boxes are... If the target is lost, no detection result will be output, and no control command will be sent to ensure the stability of subsequent target detection results and reduce the risk of erroneous control.

[0059] Visual smoothness optimization aims to reduce visual stuttering caused by image refresh or detection lag. Based on the spatiotemporal perception threshold of human vision (100ms~200ms), visual smoothness optimization ensures that the latency of image processing and object detection is controlled within 100ms after optimizing the YOLOv5 model structure. Combined with overfitting training based on scene features, it makes the detection boxes more stable and reduces jumps caused by the environment.

[0060] S216. Within the 100ms-200ms perception interval of human vision, if the displacement of the controlled object within a 100ms time window is lower than the spatiotemporal perception threshold of human vision (100ms-200ms), and the system cannot complete sufficient displacement feature updates within the visual persistence period, it will lead to the superposition of residual images between consecutive frames, specifically manifested as the trailing phenomenon of visual persistence, i.e., the target appears as a ghost image or blur. To solve this problem, the key is to ensure that the latency of image processing and target detection is controlled within 100ms. By optimizing the YOLOv5 model structure, removing the P5 branch and its related FPN / PAN layers, and removing the SPP module, only retaining the P3 and P4 detection scales, and simplifying the number of C3 modules, the computational load is effectively reduced, the inference speed is improved, and the inference time per frame is less than 32ms, ensuring the temporal synchronization of the system. At the same time, overfitting training based on scene features, combined with a large amount of scene data for model fine-tuning, makes the detection boxes more stable, reduces environmental jumps, and improves detection confidence. These optimization measures work together to reduce visual trailing, improve the system's real-time response capability, and enable the target displacement between adjacent frames to exceed the visual perception threshold, thereby achieving a clear spatiotemporal representation of moving targets within a continuous visual perception cycle.

[0061] S217. Original YOLOv5 model structure, see [link / reference]. Figure 8 This includes: Backbone, Neck, and Head.

[0062] (1) The backbone is responsible for extracting useful features from the input image. It is usually a convolutional neural network (CNN) trained on large-scale image classification tasks, such as ImageNet. The backbone network captures hierarchical features at different scales, extracting low-level features (such as edges and textures) in earlier layers and high-level features (such as object parts and semantic information) in deeper layers.

[0063] In Backbone, P1 / 2 indicates that the size of the output feature map of this layer is reduced to 1 / 2 of the original input size; In the Backbone layer, layer P2 is responsible for detecting small-scale targets. Due to the large feature map size (e.g., 160×160), it preserves the features of small targets well, but deep convolutions may lead to feature loss.

[0064] In the Backbone, layer P4 is a relatively deep detection head, typically used for detecting medium-scale targets. Its corresponding feature map size is 40x40, suitable for detecting medium-sized targets larger than 16x16.

[0065] In the Backbone layer, P3 / 8 is responsible for extracting feature information from an 80x80 region in the input image. The feature map output by this layer is 80x80 in size, which is suitable for detecting small targets. Each grid cell predicts three bounding boxes (bboxes) for localization and classification in target detection tasks.

[0066] In Backbone, P4 / 16 is a feature layer level with a downsampling factor of 16x, used to detect medium-sized targets.

[0067] In Backbone, P5 / 32 is a feature layer level with a downsampling factor of 32, used for detecting large targets.

[0068] (2) The Neck is an intermediate component connecting the Backbone and the Head. It aggregates and refines the features extracted from the backbone, typically focusing on enhancing spatial and semantic information at different scales. The neck may include additional convolutional layers, feature pyramids (FPNs), or other mechanisms to improve the representativeness of the features.

[0069] (3) The Head is the final component of the object detector. It is responsible for making predictions based on the features provided by the Backbone and Neck. It typically consists of one or more task-specific subnetworks that perform classification, localization, and the most recent instance segmentation and pose estimation. The Head processes the features provided by the Neck to generate a prediction for each candidate. Finally, a post-processing step, such as Non-Maximum Suppression (NMS), filters out overlapping predictions and retains only the detections with the highest confidence.

[0070] In the Head layer, P5 / 32-large is the feature layer level used for detecting large targets. In the YOLOv5 architecture, P5 represents the output of the feature map after 5 downsampling steps, with a size of 1 / 32 of the original image size (i.e., 32 times smaller), while -large indicates that this layer is specifically used for detecting large targets.

[0071] Among them, the P4 / 16-medium in the Head is a detection head in the YOLOv5 network structure that is specifically designed to handle medium-sized targets. It improves the detection performance of medium-sized targets by fusing feature map information from different levels.

[0072] In the Head layer, P3 / 8-small integrates feature map information at different scales through PANet (Path Aggregation Network) to ultimately generate predicted bounding box coordinates, class probabilities, and confidence scores for detecting small targets. In the YOLOv5 CSPDarknet53 backbone network, P3 / 8-small corresponds to the small-scale feature map of FPN (Feature Pyramid Network), which is used to improve the accuracy of small target detection.

[0073] S218. Improved YOLOv5 model structure, see [link / reference]. Figure 9 This includes: Backbone, Neck, and Head.

[0074] (1) The backbone is responsible for extracting useful features from the input image. It is usually a convolutional neural network (CNN) trained on large-scale image classification tasks, such as ImageNet. The backbone network captures hierarchical features at different scales, extracting low-level features (such as edges and textures) in earlier layers and high-level features (such as object parts and semantic information) in deeper layers.

[0075] In Backbone, P1 / 2 indicates that the size of the output feature map of this layer is reduced to 1 / 2 of the original input size; In the Backbone layer, layer P2 is responsible for detecting small-scale targets. Due to the large feature map size (e.g., 160×160), it preserves the features of small targets well, but deep convolutions may lead to feature loss.

[0076] In the Backbone, layer P4 is a relatively deep detection head, typically used for detecting medium-scale targets. Its corresponding feature map size is 40x40, suitable for detecting medium-sized targets larger than 16x16.

[0077] In the Backbone layer, P3 / 8 is responsible for extracting feature information from an 80x80 region in the input image. The feature map output by this layer is 80x80 in size, with each grid predicting 3 bounding boxes (bboxes), which are used for localization and classification in object detection tasks.

[0078] In Backbone, P4 / 16 is a feature layer level with a downsampling factor of 16x, used to detect medium-sized targets.

[0079] (2) The Neck is an intermediate component connecting the Backbone and the Head. It aggregates and refines the features extracted from the backbone, typically focusing on enhancing spatial and semantic information at different scales. The neck may include additional convolutional layers, feature pyramids (FPNs), or other mechanisms to improve the representativeness of the features.

[0080] (3) The Head is the final component of the object detector. It is responsible for making predictions based on the features provided by the Backbone and Neck. It typically consists of one or more task-specific subnetworks that perform classification, localization, and the most recent instance segmentation and pose estimation. The Head processes the features provided by the Neck to generate a prediction for each candidate. Finally, a post-processing step, such as Non-Maximum Suppression (NMS), filters out overlapping predictions and retains only the detections with the highest confidence.

[0081] Among them, the P4 / 16-medium in the Head is a detection head in the YOLOv5 network structure that is specifically designed to handle medium-sized targets. It improves the detection performance of medium-sized targets by fusing feature map information from different levels.

[0082] In the Head layer, P3 / 8-small integrates feature map information at different scales through PANet (Path Aggregation Network) to ultimately generate predicted bounding box coordinates, class probabilities, and confidence scores for detecting small targets. In the YOLOv5 CSPDarknet53 backbone network, P3 / 8-small corresponds to the small-scale feature map of FPN (Feature Pyramid Network), which is used to improve the accuracy of small target detection.

[0083] Thirdly, the present invention provides an augmented reality stimulation module for target detection and aVEP stimulation adsorption, including a preprocessing unit, a weak asymmetric visual stimulation unit, and an adsorption stimulation image generation unit. The preprocessing unit receives the target detection image output by the target detection module provided by the second aspect, and performs smooth transition processing on the bounding box coordinates of the optimal preselected box in the current frame to obtain the bounding box coordinates of the preprocessed optimal preselected box. The weak asymmetric visual stimulation unit is used to provide weak asymmetric visual stimulation; The adsorption stimulus image generation unit constructs a two-way temporal constraint mechanism based on Shannon sampling theorem, which includes target detection and weak asymmetric visual stimuli. That is, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli, and the superposition position of the weak asymmetric visual stimuli is adjusted in real time so that it is adsorbed around the box coordinates of the preprocessed optimal preselected box, that is, around the controlled object.

[0084] As one possible implementation, the coordinates of the optimal preselected box in the current frame are smoothly transitioned, specifically including the following steps: Preserve the coordinates of the optimal preselected box from the previous frame; After obtaining the coordinates of the optimal preselection box in the current frame, a smoothing factor is used. The coordinates of the optimal preselected box in the current frame and the coordinates of the optimal preselected box in the previous frame are weighted and summed to obtain the coordinates of the optimal preselected box in the current frame after smooth transition processing.

[0085] As one possible implementation, the sampling frequency of the target detection image is greater than or equal to 32Hz, and the frequency of the weak asymmetric visual stimulus is 4Hz.

[0086] In practical control scenarios, the overlay of environmental and virtual stimuli can enhance the system's perception capabilities, enabling it to better adapt to complex and dynamic changes. By supplementing the system with virtual information or the real environment, users can obtain more real-time feedback, thereby improving operational efficiency. This overlay of environmental and virtual stimuli can be presented through two methods: Augmented Reality (AR) devices and Personal Computers (PCs). AR devices can acquire images and data of the real environment through cameras (first-person perspective) and sensors, and combine this with computer vision technology to analyze the scene in real time. The system uses graphics rendering technology to overlay virtual stimulus elements (such as images and text) onto the real world. These virtual elements typically adjust dynamically according to the user's position, viewpoint, and actions. See [link to documentation]. Figure 10 and Figure 11 On a PC, the system can capture images of the real-world environment through the camera (first-person perspective) of the controlled object. Using computer vision algorithms, the system can analyze the scene and render virtual elements (such as images and text) onto the screen. (See also...) Figure 12 .

[0087] Weak asymmetrical visual stimuli transitioned subjects from direct visual stimulation to direct visual cues. Simultaneously, a fixed-frequency, unilaterally weaker asymmetrical stimulus was used. The combined effect of these two methods significantly alleviated visual fatigue in the subjects. The spatial arrangement of the visual stimuli and guiding points is detailed in [reference needed]. Figure 13 Weak asymmetric visual stimuli consist of a visual stimulus displayed at a fixed frequency (such as a flickering light source or a graphic on a screen) and a guide point. When a subject fixates on a specific guide point, different EEG signals are induced, with different guide points corresponding to different categories.

[0088] The implementation of third-person perspective stimulus adsorption includes: building the original network architecture > improving the network architecture > training the model > target detection and object tracking > fusing the aVEP paradigm > improving the system.

[0089] The weak asymmetric visual stimulus consists of 6 different colors or shapes, with a single stimulus cycle of 250ms, including a bright state for the first 60ms (100% brightness) and a dark state for the last 190ms (0% brightness). The stimulus size is 2°. Because the stimulus interval is 250ms, the stimulus flashes at a frequency of 4Hz. The stimulus is surrounded by 24 visual cue points at a distance of 3° from the stimulus center.

[0090] This invention overcomes the limitations of existing technologies in virtual stimulus presentation on AR devices and PCs by introducing a third-person perspective adsorption stimulation scheme based on aVEP. This scheme deploys independent cameras in the environment to capture information about the surrounding environment from a third-person perspective, and combines computer vision algorithms and augmented reality technology to adjust the virtual stimulus in real time, ensuring a stable spatial position even when the controlled object moves or the user's line of sight changes. Compared to the first-person perspective stimulation methods of existing AR devices, this invention avoids the problem of stimulus position displacement caused by head-mounted device movement, while reducing the burden and discomfort of wearing the device and improving the interactive experience.

[0091] Compared to the first-person perspective approach, this invention utilizes a third-person perspective approach to absorb stimuli, providing a wider range of environmental information and enabling users to observe their surroundings from a third-person perspective, thus allowing for more natural interaction.

[0092] This not only reduces reliance on wearable devices and improves user comfort, but also provides a wider field of view, enhancing the applicability of the BCI system in practical control scenarios.

[0093] The augmented reality stimulation module runs on the PC of the first computer. It is responsible for receiving image data transmitted by the image acquisition unit, detecting the position of the controlled object, generating superimposed stimulation images based on this, and forwarding control commands in real time. The module adopts deep learning-based target detection and tracking technology (YOLOv5), which can efficiently and accurately analyze environmental features and target positions. Based on Psychopy graphics rendering technology, the module generates a third-person perspective image and uses YOLOv5 to detect the real-time position of the target. It adjusts the superimposed position of the aVEP stimulation in real time so that it adheres to the area around the controlled object, presenting augmented reality stimulation to the user to enhance the user's environmental perception. At the same time, the module receives control commands sent by the EEG data analysis module and forwards them to the controlled object via WiFi to achieve remote control. It includes: (1) Input: image data, control commands; (2) Communication: USB communication, UDP communication, WIFI communication; (3) Output: augmented reality stimulation images, control commands.

[0094] As the core control subject of the entire system, the user generates specific EEG signals by gazing at the augmented reality stimulation presented by the augmented reality stimulation module. These signals are recorded by the EEG data acquisition module and processed and converted into control commands by the EEG data analysis module, thereby controlling the controlled object. The user's intention is ultimately realized through brain control commands, enabling them to remotely control the behavior of the controlled object and obtain real-time feedback through augmented reality stimulation. This includes: (1) Input: Augmented reality stimulation image; (2) Output: EEG signal.

[0095] The controlled object can receive WiFi commands sent by the augmented reality stimulation module and perform corresponding operations according to these commands. This includes: (1) Input: control commands; (2) Communication: WiFi communication; (3) Output: corresponding actions after receiving commands.

[0096] An augmented reality stimulation module combining target detection and aVEP stimulation adsorption acquires global scene information and captures changes in the surrounding environment in real time to achieve a more immersive and realistic interactive experience. This enhances the user's environmental awareness of the controlled object while reducing user burden and improving comfort during extended use. Furthermore, the module integrates with an EEG data analysis module to analyze and identify the user's intentions and transmit brain-controlled commands to the controlled object, enabling remote control.

[0097] To ensure stable presentation of augmented reality stimuli, the system needs to accurately detect and track the autonomous vehicle's position, ensuring that the stimuli remain attached to the target. Therefore, the YOLOv5 target detection technology used in this solution requires large-scale data training to achieve efficient and accurate target detection capabilities. To ensure YOLOv5 has good detection accuracy and robustness, the system first needs to collect and label a large amount of training data containing autonomous vehicles, and then perform preprocessing.

[0098] Spatial stability is optimized by improving the network structure and the NMS (non-maximum suppression) algorithm, and introducing a weighted moving average algorithm to reduce stimulus position jumps and ensure the stability of the position of the stimulus superimposed on the controlled object.

[0099] By collecting extensive scene data and fine-tuning the model, the S301's overfitting significantly improves the stability of detection boxes and reduces jumps caused by subtle environmental changes. Furthermore, it enhances detection confidence, leading to more accurate target recognition. Simultaneously, overfitting reduces the model's dependence on unknown environments, making its performance more reliable and consistent in specific scenarios.

[0100] S302 adds temporal smoothing between frames. During the detection of each frame, it retains the position information of the previous frame and performs smoothing processing with the current frame. Based on a user-defined smoothing factor (α), it adjusts the new position of the current frame (…). , Historical position relative to the previous frame ( , Perform a weighted average calculation. Smoothed coordinate values ​​( , It is given by the following formula: in, Smoothing factor ( It is used to control the position weights between the current frame and historical frames.

[0101] (1) When When the value is close to 1, it becomes more dependent on the current frame, and the detection box responds quickly, but there may be jumps.

[0102] (2) When When the value is close to 0, it relies more on historical frames, and the detection box transitions smoothly, but the response speed may be slower.

[0103] (3) Recommended settings The value should be between 0.6 and 0.8 to balance smoothness and response speed.

[0104] This method ensures a smooth transition between rapidly moving or changing stimulus positions, avoiding visual stuttering caused by frame skipping or sudden positional changes. When the previous frame position ( , When empty, the method directly returns the position of the current frame. , This ensures error-free processing during the initial calculation or initialization. This weighted moving average method maintains the stability of stimulus positions in fast-moving environments, thereby optimizing the smoothness of visual presentation. See the flowchart for the time smoothing algorithm. Figure 14 .

[0105] Fourthly, the present invention provides a brain-computer interface interaction method, comprising the following steps: Acquire consecutive image frames of the controlled object; preprocess the consecutive image frames to match the input requirements of the pre-trained improved YOLOv5 network model. Preprocessing includes at least scaling, normalization, and color space conversion; use the pre-trained improved YOLOv5 network model to perform target detection on the preprocessed consecutive image frames, i.e., extract image features through forward inference and perform target classification and localization to obtain multiple pre-selected boxes of the controlled object and their corresponding box coordinates, categories, and confidence scores; use confidence thresholds and intersection-over-union (IoU) thresholds to filter out the box coordinates of the optimal pre-selected boxes to complete target detection. The system receives the target detection image output by the target detection module, performs smooth transition processing on the coordinates of the optimal preselected box in the current frame to obtain the preprocessed coordinates of the optimal preselected box; provides weak asymmetric visual stimuli; constructs a two-way temporal constraint mechanism based on Shannon sampling theorem, including target detection and weak asymmetric visual stimuli, that is, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli, and adjusts the superposition position of the weak asymmetric visual stimuli in real time so that it is attached to the area around the preprocessed optimal preselected box coordinates; and presents augmented reality stimuli to the user. The system records EEG signals induced by augmented reality stimuli in real time, decodes and analyzes the EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

Claims

1. A brain-computer interface system, characterized in that, It includes a target detection module, an augmented reality stimulation module, an EEG data acquisition module, and an EEG data analysis module; The target detection module includes an image acquisition unit and a target detection unit. The image acquisition unit acquires image data in real time, including information about the controlled object and the global environment. The target detection unit is equipped with a pre-trained improved YOLOv5 network model. The improved YOLOv5 network model removes the P5 branch and its related FPN / PAN layers from the original YOLOv5 network model, and also removes the SPP module, retaining only the P3 and P4 detection scales. The number of C3 modules is also simplified to reduce network complexity. The improved YOLOv5 network model also improves the non-maximum suppression algorithm by setting confidence thresholds and intersection-over-union (IoU) thresholds. The augmented reality stimulation module includes a preprocessing unit, a weak asymmetric visual stimulation unit, and an adsorption stimulation image generation unit. The preprocessing unit receives the target detection image output by the target detection module and performs smooth transition processing on the coordinates of the optimal preselected bounding box in the current frame to obtain the preprocessed coordinates of the optimal preselected bounding box. The weak asymmetric visual stimulation unit provides weak asymmetric visual stimulation. The adsorption stimulation image generation unit constructs a bidirectional temporal constraint mechanism based on Shannon's sampling theorem, including target detection and weak asymmetric visual stimulation. Specifically, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimulation, and the superposition position of the weak asymmetric visual stimulation is adjusted in real time so that it adsorbs around the coordinates of the preprocessed optimal preselected bounding box, i.e., around the controlled object. The EEG data acquisition module records the EEG signals induced by the augmented reality stimulation in real time and sends them to the EEG data analysis module. The EEG data analysis module decodes and analyzes EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

2. The brain-computer interface system according to claim 1, characterized in that, The improvements to the nonmaximum suppression algorithm specifically include: Set a confidence threshold and retain the pre-selected boxes that are greater than or equal to the confidence threshold; Set an intersection-union ratio (IU) threshold. From the preselected boxes that are greater than or equal to the confidence threshold, filter out the preselected boxes that are greater than or equal to the IU threshold. Then, select the preselected box with the higher confidence among the two preselected boxes with the largest IU as the optimal preselected boxes.

3. The brain-computer interface system according to claim 1, characterized in that, The confidence threshold is 0.6, and the crossover ratio (CUP) threshold is 0.

4.

4. The brain-computer interface system according to claim 1, characterized in that, The latency of the target detection module is less than or equal to 100ms.

5. The brain-computer interface system according to any one of claims 1 to 4, characterized in that, The sampling frequency of the target detection image is greater than or equal to 32Hz, and the frequency of the weak asymmetric visual stimulus is 4Hz.

6. A brain-computer interface interaction method, characterized in that, Includes the following steps: Acquire consecutive image frames of the controlled object; The process involves preprocessing consecutive image frames to match the input requirements of a pre-trained improved YOLOv5 network model. Preprocessing includes at least scaling, normalization, and color space conversion. The pre-trained improved YOLOv5 network model is then used to perform target detection on the preprocessed consecutive image frames. This involves extracting image features through forward inference and classifying and locating targets, obtaining multiple pre-selected bounding boxes of the controlled object along with their corresponding box coordinates, categories, and confidence scores. Optimal box coordinates are then selected using confidence and intersection-over-union (IoU) thresholds to complete target detection. The improved YOLOv5 network model removes the P5 branch and its associated FPN / PAN layers from the original YOLOv5 network model, while also removing the SPP module, retaining only the P3 and P4 detection scales, and simplifying the number of C3 modules to reduce network complexity. The improved YOLOv5 network model also improves the non-maximum suppression algorithm by setting confidence and IoU thresholds. The system receives the target detection image output by the target detection module, performs smooth transition processing on the coordinates of the optimal preselected box in the current frame to obtain the preprocessed coordinates of the optimal preselected box; provides weak asymmetric visual stimuli; constructs a two-way temporal constraint mechanism based on Shannon sampling theorem, including target detection and weak asymmetric visual stimuli, that is, the sampling frequency of the target detection image is at least twice the highest frequency of the weak asymmetric visual stimuli, and adjusts the superposition position of the weak asymmetric visual stimuli in real time so that it is attached to the area around the preprocessed optimal preselected box coordinates; and presents augmented reality stimuli to the user. The system records EEG signals induced by augmented reality stimuli in real time, decodes and analyzes the EEG signals to extract the user's control commands, and sends the control commands to the controlled object through the augmented reality stimulation module to achieve control over the controlled object.

7. The brain-computer interface interaction method according to claim 6, characterized in that, The coordinates of the optimal preselected box in the current frame are smoothed, which includes the following steps: Preserve the coordinates of the optimal preselected box from the previous frame; After obtaining the coordinates of the optimal preselection box in the current frame, a smoothing factor is used. The coordinates of the optimal preselected box in the current frame and the coordinates of the optimal preselected box in the previous frame are weighted and summed to obtain the coordinates of the optimal preselected box in the current frame after smooth transition processing.

Citation Information

Patent Citations

  • PCB surface defect detection method based on improved YOLOv5

    CN115719338A

  • AR-based intelligent target detection, identification and tracking system

    CN116883814A