Brain-eye-computer data fusion target detection method

By combining computer vision and brain-computer interface technology, using eye trackers and EEG signal fusion methods, high-precision, strong and robust weak hidden object detection in drone aerial images is achieved, solving the problems of low detection accuracy and poor robustness in the existing technology, and is suitable for intelligent monitoring and early warning.

CN116524381BActive Publication Date: 2025-08-19NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310507104.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-08-19
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

In drone aerial images, weak hidden object detection accuracy is low and poor robustness is poor. Existing computer vision and single-modal methods of EEG are difficult to achieve efficient and accurate object detection, especially under small sample data conditions.

Method used

Combining computer vision and brain-computer interface technology, the target area of interest is captured in real time through eye trackers, synchronizes the EEG signal and image data, and target detection is performed using brain-eye-computer data fusion method, and the event-related potential signals are used to induce eye tracking data, and target positioning is performed by combining EEG and image features.

Benefits of technology

Under small sample conditions, the detection accuracy and robustness of weak hidden targets in drone aerial images are improved, and they have strong generalization performance, which is suitable for fields such as intelligent monitoring and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524381B_ABST
    Figure CN116524381B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for target detection using brain-computer interface (BCI) data fusion, comprising the following steps: constructing a sample image dataset; conducting an eye-movement-based slow-sequence visual presentation experiment on a subject based on the sample image dataset, collecting the subject's first eye movement data and first EEG data; obtaining a first candidate region frame of the subject's gaze based on the first eye movement data, and extracting the corresponding first candidate image; segmenting the first EEG data based on the timestamp of the first eye movement data to obtain corresponding first EEG data segments; constructing a training sample set to train a BCI signal fusion classification model, and performing target detection based on the trained BCI signal fusion classification model. This method, which has applications in the fields of brain-computer interface technology and target detection technology, can improve the detection accuracy and robustness of weakly hidden targets in drone aerial images while ensuring good real-time performance. It exhibits strong generalization performance and is not constrained by small sample data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of brain-computer interface technology and target detection technology, and specifically to a slow sequence visual presentation experimental paradigm design based on eye movement and a brain-eye-computer data fusion target detection method construction. Background Art

[0002] With the rapid development of artificial intelligence (AI), target detection technology has found widespread application in intelligent monitoring, early warning, facial recognition, and other fields. It is also a key component in the development and application of more advanced technologies in the era of big data intelligence. In recent years, computer vision-based target detection methods have achieved significant success in target detection tasks. Powered by large datasets, they can effectively detect a wide range of targets. Applications such as facial recognition and unmanned obstacle avoidance have also been adopted in various civilian fields. However, due to the complex backgrounds and high noise levels in drone aerial images, and the often faint and hidden nature of targets, detection accuracy and efficiency remain limited in certain areas, requiring improvement.

[0003] Compared to computer vision technology, brain-computer interface technology can build a bridge of communication between the human brain and computers, leveraging the human brain's advanced cognitive abilities to accomplish various tasks that are difficult for computers. Among them, event-related potential (ERP)-based target detection methods can leverage the human brain's cognitive analysis capabilities to generate neural responses to key sensitive information in a scene within a few hundred milliseconds, and have achieved many important achievements in target detection tasks. However, because targets in aerial imagery are difficult to detect, single-modality detection methods based solely on EEG are susceptible to interference from environmental noise, have low signal-to-noise ratios, and require long target search times, which increases the difficulty of EEG signal analysis. Therefore, it is difficult to achieve robust and accurate target detection performance in real-world scenarios relying solely on EEG data.

[0004] Therefore, by combining the powerful information processing capabilities of computer vision with the human brain's ability to recognize complex situations and sensitive information, it is hoped that the challenge of quickly and accurately detecting faintly hidden targets in massive amounts of drone aerial imagery will be solved. However, existing object detection technologies that combine computer vision and human vision are still in the exploratory stage, and there is little research on object detection in drone aerial imagery. Summary of the Invention

[0005] To address these technical issues, we propose a slow-motion sequence visual presentation experimental paradigm based on eye movements. This uses computer vision techniques to select suspicious areas in an image and an eye tracker to capture the target of interest in real time. This synchronizes EEG signals with image data, laying the foundation for target detection using brain-computer signal fusion. A brain-eye-computer data fusion target detection method is proposed. After acquiring synchronized EEG and image data, a brain-computer signal fusion classification method is used to detect regional target images and locate the target position.

[0006] The present invention can effectively broaden the coverage of information contained in the input data, and improve the detection accuracy and robustness of weak and hidden targets in drone aerial images while ensuring good real-time performance. It has strong generalization performance and is not constrained by small sample data. It can be applied to intelligent monitoring, reconnaissance and early warning and other fields.

[0007] To achieve the above object, the present invention provides a brain-eye-computer data fusion target detection method, comprising the following steps:

[0008] Step 1: construct a sample image dataset, wherein the sample image dataset includes a plurality of sample images, each of which has at least one first candidate region frame containing a suspicious object, and the sample image dataset also includes position data and true label data of each first candidate region frame;

[0009] Step 2: Conduct an eye movement-based slow sequence visual presentation experiment on the subject based on the sample image dataset, and collect the subject's first eye movement data and first EEG data in real time;

[0010] Step 3: obtaining a first candidate region frame that the subject is looking at based on the first eye movement data, and defining it as a second candidate region frame; and extracting a corresponding first candidate image from the corresponding sample image based on the second candidate region frame;

[0011] Step 4: segmenting the first EEG data based on the timestamp of the first eye movement data to obtain first EEG data segments corresponding to the first candidate image;

[0012] Step 5: construct a training sample set, wherein the training sample set includes a plurality of sample data, each of which includes a first candidate image and its corresponding true label data and a first EEG data segment;

[0013] Step 6: Train the brain-computer signal fusion classification model based on the training sample set, and perform target detection based on the trained brain-computer signal fusion classification model.

[0014] In one embodiment, constructing the sample image dataset comprises:

[0015] capturing a plurality of target images with targets and a plurality of non-target images without targets, and using both the target images and the non-target images as the sample images;

[0016] At least one first candidate area frame containing a suspected target is selected in the sample image by a manual labeling method and / or a computer vision extraction method, and the position data of the first candidate area frame in the corresponding sample image and the real label data of whether the first candidate area frame contains the target are saved.

[0017] In one embodiment, in step 2, the eye movement-based slow sequence visual presentation experiment is conducted on the subject based on the sample image dataset, specifically:

[0018] Randomly arranging the sample images in the sample image data set to form N groups of image stimulation sequences, each group of image stimulation sequences contains 50 to 100 sample images, where N ≥ 1;

[0019] Playing the image stimulus sequence to the subjects at a playback rate of 3 to 5 seconds per frame. During the playback of the image stimulus sequence, the subjects were required to subjectively search for the target and maintain their gaze for more than 0.3 seconds after finding the target.

[0020] After each set of image stimulation sequences is played, the next set of image stimulation sequences is played after a preset period of time.

[0021] In one embodiment, among N groups of image stimulation sequences, the target images in group a of image stimulation sequences only include target 1, and the target images in group b of image stimulation sequences only include target 2, where a+b=N;

[0022] The first candidate image and the first EEG data segment generated by the image stimulation sequence with target 1 are used as a training set and a validation set for training the brain-computer signal fusion classification model;

[0023] The first candidate image and the first EEG data segment generated by the image stimulation sequence with target 1 are used as a sample set for transfer training of the brain-computer signal fusion classification model.

[0024] In one embodiment, Gaussian noise with a mean of 0 and a variance of 0.2 and salt and pepper noise with a variance of 0.2 are added to the image data used as the validation set to simulate noise interference that may exist in a real environment.

[0025] In one embodiment, in a set of image stimulation sequences, the probability of a target image appearing is 40% to 50%, and 1 to 2 non-target images are randomly spaced between every two target images.

[0026] In one embodiment, when the image stimulation sequence is played, the subject faces the display playing the image stimulation sequence and is 50 cm to 70 cm away from the display.

[0027] In one embodiment, in step 6, target detection is performed based on the trained brain-computer signal fusion classification model, specifically:

[0028] Step 6.1, obtaining an image to be tested, selecting third candidate region frames containing suspected targets in the image to be tested, and determining position data of each third candidate region frame in the image to be tested;

[0029] Step 6.2: Conduct an eye movement-based slow sequence visual presentation experiment on the subject based on the image to be tested, and collect the subject's second eye movement data and second EEG data in real time;

[0030] Step 6.3: obtaining a third candidate region frame that the subject is looking at based on the second eye movement data, and defining it as a fourth candidate region frame; and extracting the corresponding second candidate image from the image to be tested based on the fourth candidate region frame;

[0031] Step 6.4, segmenting the second EEG data based on the timestamp of the second eye movement data to obtain second EEG data segments corresponding to the second candidate images;

[0032] Step 6.5: Input a set of corresponding second candidate images and the second EEG data segment into the trained brain-computer signal fusion classification model to obtain a target detection result corresponding to the second candidate image. If the target detection result indicates that there is a target, the position of the target in the image to be tested can be located by following the corresponding second candidate image;

[0033] Step 6.6: Repeat step 6.5 until all second candidate images and second EEG data segments are traversed, thus completing the target detection of the image to be tested.

[0034] In one embodiment, in step 6.1, a computer vision extraction method is used to extract all third candidate region frames containing suspected targets in the image to be tested.

[0035] Compared with the prior art, the present invention has the following beneficial technical effects:

[0036] The present invention studies the brain-computer interface experimental paradigm to solve the problems of low generalization, large sample data requirements, and low detection accuracy in weak hidden target detection in image data such as drone aerial images, and constructs a brain-eye-computer data fusion target detection system to combine the advanced cognitive functions of the human brain with the powerful information processing capabilities of computer vision. The slow sequence visual presentation experimental paradigm based on eye movements can effectively induce the subject's event-related potential signal (ERP) and synchronously store the candidate area image and related position data of the subject's attention, providing data input for the subsequent brain-computer signal fusion classification algorithm. It effectively combines the dual-modality characteristics of the human brain and the computer, designs the experimental paradigm and target detection system, and makes full use of the characteristics of each modality to achieve high-precision, strong robustness, and high generalization target detection of drone weak hidden target images under small sample conditions, providing new technical ideas for subsequent research, and has important significance and practical value for the weak hidden target detection task of drone aerial images. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0038] Figure 1 Flowchart of the brain-eye-computer data fusion target detection method according to an embodiment of the present invention;

[0039] Figure 2 Schematic diagram of the weakly hidden target image and non-target image of a drone in an embodiment of the present invention, where: (a) is a target image containing target 1, (b) is a background image excluding target 1, (c) is a target image containing target 2, and (d) is a background image excluding target 2;

[0040] Figure 3 Schematic diagram of a specific process framework of a brain-eye-computer data fusion target detection system according to an embodiment of the present invention;

[0041] Figure 4 Schematic diagram of the hardware architecture of the brain-eye-computer data fusion target detection system in an embodiment of the present invention;

[0042] Figure 5 Schematic diagram of the software architecture of the brain-eye-computer data fusion target detection system in an embodiment of the present invention.

[0043] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0045] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0046] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referenced. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "plurality" means at least two, such as two or three, unless otherwise specifically defined.

[0047] In the present invention, unless otherwise specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; it can mean mechanical connection, electrical connection, physical connection, or wireless communication connection; it can mean direct connection or indirect connection through an intermediate medium; it can mean internal communication between two elements or interaction between two elements, unless otherwise specified. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0048] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0049] like Figure 1 The present invention discloses a method for detecting a target using brain-eye-computer data fusion, which mainly includes the following steps:

[0050] Step 1: Construct a sample image dataset, which includes several sample images. Each sample image has at least one first candidate region frame containing a suspicious object. The sample image dataset also includes position data and true label data of each first candidate region frame.

[0051] Step 2: Design an eye-movement-based slow sequence visual presentation paradigm, build a target detection system environment for brain-eye-computer data fusion, and clarify the relevant experimental requirements; conduct an eye-movement-based slow sequence visual presentation experiment on the subjects based on the sample image dataset, and collect the subjects' first eye movement data and first EEG data in real time;

[0052] Step 3: obtaining a first candidate region frame that the subject is looking at based on the first eye movement data, and defining it as a second candidate region frame; and extracting a corresponding first candidate image from the corresponding sample image based on the second candidate region frame;

[0053] Step 4: segmenting the first EEG data based on the timestamp of the first eye movement data to obtain first EEG data segments corresponding to the first candidate image;

[0054] Step 5: construct a training sample set, the training sample set includes a plurality of sample data, each sample data includes a first candidate image and its corresponding true label data and a first EEG data segment;

[0055] Step 6: Train the brain-computer signal fusion classification model based on the training sample set, and perform target detection based on the trained brain-computer signal fusion classification model.

[0056] In step 1, the specific implementation method of constructing the sample image dataset is as follows:

[0057] Step 1.1: Use a drone to shoot low-altitude and high-altitude videos in different scenes, and extract frames to obtain several images with various object models and background images. After setting a specific object as the target according to the experimental requirements, the images can be divided into target images and non-target images. Both target images and non-target images are used as sample images;

[0058] Step 1.2: Select at least one first candidate region frame containing a suspected target in the sample image, and save the position data of the first candidate region frame in the corresponding sample image and the real label data of whether the first candidate region frame contains the target for subsequent classification model training and performance evaluation.

[0059] In this embodiment, suspicious targets in an image can be framed using either manual labeling or computer vision extraction. Manual labeling allows for recording real-world label data during the selection process, which can be used for subsequent classification model training and optimization, as well as for performance evaluation of the brain-eye-computer data fusion target detection system. Computer vision extraction methods can suffer from missed detections and false detections, but after simple training, they can ensure the selection of suspicious target areas, enabling practical detection tasks within the brain-eye-computer data fusion target detection system.

[0060] In the specific implementation process, the first candidate region in the sample image is selected by manual annotation to obtain the target candidate region position data and true label data for training and optimizing the subsequent classification model. The position data of the first candidate region frame in the sample image can also be extracted by referencing the pre-trained Faster-RCNN method. Figure 2 These are the weakly hidden target images and non-target images captured by the drone in this example. The sample image dataset includes 191 target images containing target 1, 114 target images containing target 2, and 62 non-target images excluding targets 1 and 2. In addition, some of the images used to create the validation set were added with Gaussian noise with a mean of 0 and a variance of 0.2, and salt and pepper noise of 0.2 to simulate noise interference that may be present in a real environment.

[0061] The specific implementation of designing the eye-movement-based slow sequence visual presentation paradigm is as follows:

[0062] The sample images selected in the sample image dataset are randomly arranged to form image stimulation sequences. The experimental summary contains a total of 10 groups of image stimulation sequences, each group of image stimulation sequences contains a total of 50 images. The first 8 groups of image stimulation sequences are composed of target images containing target 1 and non-target images, and the last 2 groups of image stimulation sequences are composed of target images containing target 2 and non-target images. To verify the generalization performance of the target detection system, the first 8 groups of image stimulation sequences and the corresponding data collected during the experiment will be used as training sets and validation sets, and the last two groups of image stimulation sequences and the corresponding data collected during the experiment will be used as sample sets for transfer training.

[0063] The subjects were required to stand 50 to 70 cm away from the monitor, look directly at the center of the monitor, wear an EEG cap, and calibrate the eye tracker before the experiment began. Before the experiment began, the subjects were informed that they needed to detect the specific features of the target and view an image containing the target. The subjects subjectively searched for the target in the sample image and stared at the target after finding it (staring for more than 0.3 seconds). The image stimulus sequence was presented at a rate of one image per 3 seconds. The probability of the sample image containing the target appearing in the image stimulus sequence was 40%. Every two target images were randomly separated by 1 to 2 non-target images. After each set of image stimulus sequences was played, the next set of image stimulus sequences was played after a preset period of time. The subjects could also choose the rest time.

[0064] refer to Figure 3 ,The specific detection process framework for building a brain-eye-machine data fusion target detection system is as follows:

[0065] First, the drone sensor image is processed using a computer vision-based candidate region extraction technique to annotate candidate region frames (possibly multiple frames) that may be the target in the image. Second, by designing a slow-motion sequential visual presentation paradigm based on eye movements, the subject's EEG data and the image frames of the region (one or several) that the subject is focusing on at the corresponding moment are simultaneously extracted as input to the EEG feature extraction and image feature extraction modules. After the EEG and computer signal processing modules obtain the corresponding EEG features and image features, brain-computer data fusion methods are used to achieve target image recognition. Finally, based on the candidate frame data extracted from the eye movement data, the location of the region image in the original image is found to locate the target, thereby achieving target detection based on brain-eye-computer multimodal data fusion.

[0066] like Figure 4Figure 2 shows the hardware architecture of the brain-eye-computer data fusion target detection system. The system hardware primarily handles sensor data acquisition and algorithm execution, including a 30-inch monitor with 2K resolution, a computer equipped with an NVIDIA GeForce GTX 2070s graphics card, a 64-channel EEG cap, a signal amplifier, a multi-parameter synchronization box, a router, an eye tracker, and several data cables. The monitor and computer are key components for playing back slow-motion visual presentations and storing and processing EEG, eye movement, and image data in real time. The monitor must be positioned 50-70 cm from the subject to maximize ERP signal induction and collect more accurate eye movement data. The eye tracker is placed directly below the monitor, tilted 15 degrees toward the subject's eyes. It effectively captures the subject's gaze position and transmits it to the computer at a 50 Hz frequency. EEG data is collected by the cap's 64 wet electrodes, amplified by an amplifier connected to the cap's signal port, and transmitted to the computer. The synchronization box can record various types of events, such as sound, light, and program output, synchronously with neurophysiological data with sufficiently high event accuracy (<1ms). It is used to synchronize EEG data with event-triggered tag signals (i.e., timestamps), providing a core basis for subsequent data analysis. A router is used to build a wireless local area network for wireless data transmission between hardware devices; the computer transmits image data to the monitor via an HDMI cable; the eye tracker connects to the computer via a USB data cable, transmitting the subject's gaze point on the monitor in real time; the EEG cap records EEG data generated while the subject views visual sequences, and transmits it to the computer and the synchronization box via the local area network using the WiFi module in the amplifier. The synchronization box communicates with the computer via a USB data cable, aligning the EEG data received by the WiFi module with the tag synchronization data. The communication interface implements the hexadecimal Dynamic Configuration Protocol (DCP).

[0067] like Figure 5The following figure shows the software architecture of a brain-eye-computer data fusion object detection system. Research on object detection systems based on multimodal data fusion is highly comprehensive, involving the acquisition and processing of data from various modalities, and placing high demands on code reusability and modularity. The software architecture primarily encompasses two aspects: paradigm design and result prediction. The paradigm design involves first extracting candidate regions from the input system image, processing the image based on the data from the candidate region boxes, and transmitting the image sequence to a display to induce EEG and eye movement data from the subject. After extracting eye movement data using the eye tracker's built-in C++ program module, it is transferred to a Python environment via a dynamic library. Based on the eye movement data, the subject's focus region and the corresponding label signal are derived. The fixation region is then combined with the original candidate regions to determine the fixation region, and the image is segmented. This regional image is then fed into the result prediction module. The EEG data acquisition system collects the subject's EEG data in real time and preprocesses the data using the C-based EEGLAB toolkit. EEG signal segments corresponding to the label signal and regional image are then fed into the result prediction module. In terms of result prediction, EEG and image features are extracted respectively through EEG feature extraction algorithm and image feature extraction algorithm, so that the brain-computer data fusion classification method can fuse brain-computer features and predict the results, and finally judge the overall performance of the system.

[0068] The specific implementation process of the eye-movement-based slow sequence visual presentation experiment on subjects based on the sample image dataset is as follows:

[0069] After the subject sits down, a 64-electrode EEG cap is fitted for the subject and EEG paste is applied to each electrode channel. The cap is fitted only when the relative impedance of all electrode channels relative to the reference electrodes (FCZ and REF) is less than 50. The subject faces the screen for eye tracker calibration to ensure that the position the subject is looking at is consistent with the position captured by the eye tracker. The subject is explained the experimental task and the target features to be searched for, and the target image is displayed. After the experiment begins, the image stimulus sequence is played, and the eye tracker captures the subject's eye movement data in real time. When the subject is found to have framed the candidate area in the image for more than 0.3 seconds, the eye movement tag is sent to the EEG signal acquisition system. Synchronized with the EEG signal, the computer also records the position data of the candidate area that the subject is focusing on at that moment.

[0070] After the experiment, the EEG signal was segmented according to the eye movement label data on the EEG signal to obtain the corresponding EEG segment data. The EEG data was then preprocessed using the EEGLab toolkit to remove noise, perform re-referencing, bandpass filtering, downsampling, and baseline correction operations. Finally, the candidate region position data was used to crop the corresponding image to obtain the candidate region image data. This data was then aligned with the processed EEG segment data and combined with the manually calibrated candidate region real label data to obtain a sample set for training and optimizing the subsequent brain-computer signal fusion classification model. The EEG segment data was combined with the regional image data obtained by the candidate region extraction method based on computer vision as a test set to verify the detection performance of the brain-eye-computer data fusion target detection system.

[0071] The specific process of preprocessing EEG data in this embodiment is as follows: use the C-language-based EEGLAB toolkit to preprocess the EEG data in the public dataset and the self-collected dataset. First, the electrodes of the EEG data are positioned according to the international 10-20 standard system, and the electrode position information is assigned to each EEG channel; the EEG data of useless channels are excluded according to the electrode channels of the known EEG cap; a 2-30Hz bandpass filter is used to remove high-frequency noise and low-frequency drift in the data; the EEG signal with the original 1000Hz sampling rate is downsampled to 250Hz to increase the subsequent calculation speed. Finally, the time segment between -200 and 1000ms with or without the target label in the RSVP public dataset is intercepted to obtain an EEG data segment with 300 sampling points and the corresponding label as the EEG signal sample. For the EEG data in the ESSVP self-collected data set, the EEG data is divided according to the labels given when the subjects see the target and non-target annotation boxes for more than 0.3 seconds, and the EEG signal data segments between -500 and 700ms with and without target candidate box labels are taken to obtain EEG signal samples with and without target labels. The EEG signal samples in the above two data sets need to be baseline corrected to subtract the 200ms mean value when the target stimulus does not appear at the beginning to avoid data drift. In this embodiment, EEG data and corresponding image area data of 10 subjects were collected. All subjects were college students aged between 22 and 26 years old, with normal vision and without any mental illness.

[0072] After processing the eye movement and EEG data, a training sample set is constructed. This sample set consists of several sample data sets, each of which is composed of the first candidate image that the subject focused on during the experiment and its corresponding first EEG data segment. The first candidate image and its corresponding first EEG data segment are input into the brain-computer signal classification model, and the model is trained and optimized. The optimal classification model is obtained based on the AUC value.

[0073] After training is completed, target detection can be performed by inputting the candidate area image and EEG signal fragment data obtained by computer vision methods into the brain-computer signal classification model to obtain prediction results. According to the candidate area results of the area image, the target position in the original image is located, the target detection of the drone aerial image is completed, and the target detection results of the candidate area image that the subject focused on in the experiment are obtained. Based on this result and the candidate area position data, the target position in the original aerial image is located to realize the detection and positioning of weak hidden targets in drone images. In the specific application process, the target detection process is as follows:

[0074] Step 6.1: Acquire the image to be tested, select third candidate region frames containing suspected targets in the image to be tested using a computer vision extraction method, and determine the position data of each third candidate region frame in the image to be tested;

[0075] Step 6.2: Conduct an eye movement-based slow sequence visual presentation experiment on the subject based on the image to be tested, and collect the subject's second eye movement data and second EEG data in real time;

[0076] Step 6.3: obtaining a third candidate region frame that the subject is looking at based on the second eye movement data, and defining it as a fourth candidate region frame; and extracting the corresponding second candidate image from the image to be tested based on the fourth candidate region frame;

[0077] Step 6.4, segmenting the second EEG data based on the timestamp of the second eye movement data to obtain second EEG data segments corresponding to the second candidate images;

[0078] Step 6.5: Input a set of corresponding second candidate images and the second EEG data segment into the trained brain-computer signal fusion classification model to obtain a target detection result corresponding to the second candidate image. If the target detection result indicates that there is a target, the position of the target in the image to be tested can be located by following the corresponding second candidate image;

[0079] Step 6.6: Repeat step 6.5 until all second candidate images and second EEG data segments are traversed, thus completing the target detection of the image to be tested.

[0080] In this embodiment, the corresponding performance testing method is set according to the characteristics of the input data, so as to effectively evaluate the performance of the brain-eye-computer data fusion target detection system. After the experiment starts, the candidate region extraction algorithm will obtain a large number of candidate frames after processing the image data. After further screening the candidate frames using eye movement data, the candidate region images and the corresponding pre-processed EEG data segments are input into the brain-computer data fusion model to obtain the predicted candidate frames. and classification results To predict the results, the true annotation boxes of all input image label files And the corresponding tags is the real result. By setting the hyperparameter IOU (intersection-over-union ratio of the predicted candidate box and the real labeled box) threshold, the predicted box in each image is matched with the real box. If the overlap between the predicted box and the labeled box is greater than the IOU threshold, the match is successful, and the classification label of the predicted sample is the same as the classification result in the prediction result; if the overlap between the predicted box and the labeled box is less than the IOU threshold, the match fails, and the classification label of the predicted sample is 0. Finally, using the results of the predicted sample Compared with the real sample results Compare and calculate the corresponding performance evaluation indicators.

[0081] In this embodiment, the AUC, F1-score, accuracy, recall and precision of target recognition are used as performance verification indicators of various algorithms in the simulation test system. The coverage area (Area Under the ROC Curve, AUC) under the Receiver Operating Characteristic Curve (ROC) is used as the main measurement indicator of detection performance to solve the problem of imbalance in the number of positive and negative sample data in the data set. The ROC curve uses the false positive rate (False Positive Rate, FPR) as the horizontal axis, which represents the proportion of actual negative instances in the positive class predicted by the classifier to all positive instances, and the true positive rate (True Positive Rate, TPR) as the vertical axis, which represents the proportion of actual positive instances in the positive class predicted by the classifier to all positive instances.

[0082] AUC is the sum of the areas under the curve, which can be used to intuitively evaluate the quality of the classifier. The larger the value, the better. It is an important evaluation indicator used by current researchers for imbalanced classification problems. In addition, it can also be directly calculated using the probability of positive and negative samples:

[0083]

[0084] M and N represent a dataset with M positive samples and N negative samples, where:

[0085]

[0086] F1-score can also evaluate the classification problem under the condition of sample imbalance to a certain extent. The specific calculation formula is:

[0087]

[0088] The False Negative Rate (FNR) is the ratio of actual negative instances to all negative instances in the negative class predicted by the classifier. Furthermore, the True Negative Rate (TNR) is included in the confusion matrix, which is the ratio of actual positive instances to all negative instances in the negative class predicted by the classifier.

[0089] The expression of accuracy is:

[0090]

[0091] It can judge the overall accuracy, but its judgment ability fails when the sample is unbalanced.

[0092] The precision rate can effectively determine the probability of all samples predicted to be positive being actually positive samples, and its expression is:

[0093]

[0094] The recall rate is for the original sample. Its meaning is the probability of predicting a positive sample among the actually positive samples, that is, judging how many positive samples are recalled. The specific expression is:

[0095]

[0096] In this embodiment, the candidate boxes extracted by the Faster R-CNN model and the subject's eye movement data during the experiment are mainly used to screen the candidate boxes to obtain the final target candidate region. This module (the eye-computer combined candidate region extraction algorithm) provides strong support for the processing of visual presentation images in the eye movement-based slow sequence visual presentation experimental paradigm and the extraction of image data in the brain-computer data fusion module. In addition, the eye movement data recorded by the subject during the target search process also screens the target candidate region to a certain extent.

[0097] To verify the performance of this module, assuming that the subsequent modules correctly predict the candidate region images, the final detection results of the entire process are used to evaluate the performance of target candidate region extraction. Because the candidate boxes obtained by the candidate region extraction algorithm differ from the actual labeled boxes in the image label data, the algorithm is verified based on different IOU threshold conditions. The evaluation results are shown in the table below:

[0098]

[0099] The analysis results show that due to the assumption that subsequent predictions are correct, the accuracy and precision values are high but not statistically significant. Recall effectively reflects the model's ability to capture positive samples and can be used to validate module performance. When the IOU threshold is set to 0.6, the combined eye-machine candidate region extraction algorithm is able to capture the target candidate region boxes to the greatest extent possible. However, increasing the IOU threshold significantly decreases the recall rate, leading to missed detection of target regions. Overall, the eye-machine signal combination module effectively leverages human cognitive capabilities and computer image data processing capabilities to capture target candidate regions and classify some non-target candidate regions, achieving effective screening and localization of target regions.

[0100] In this embodiment, the EEG data generated by the subjects during the eye movement-based slow sequence visual presentation experimental paradigm is processed based on the EEG signal algorithm to predict the target area image and verify the performance of the single-modal target recognition algorithm based on EEG signals. This module mainly uses the EEG signal classification algorithm to detect the ERP signal in the input EEG data. In order to effectively verify the target recognition performance of this module, the eye-machine data fusion target candidate area extraction module is ablated, and only this module is used to determine whether there is a problem with the target in the image. The specific detection performance is shown in the following table:

[0101]

[0102]

[0103] The analysis results show that the unimodal target recognition algorithm based only on EEG signals can detect targets without pre-training. The accuracy, precision, recall rate, F1-score and AUC are still at a high level, and it has strong generalization ability, but the detection accuracy still needs to be improved.

[0104] In this example, a brain-computer data fusion method was used to simultaneously process the image data of the target candidate region and the corresponding EEG data segments to obtain the target image classification results. To effectively verify the module's performance, the eye-computer data fusion target candidate region extraction module was ablated. The specific detection performance is shown in the table below.

[0105]

[0106] As shown in the table, the target recognition algorithm based on brain-computer data fusion shows significant improvements in evaluation metrics compared to single-modal algorithms, with the F1-score increasing by 8.87% and the AUC value increasing by 12.97%. This demonstrates that the fusion model has higher detection performance and generalization capabilities than single-modal algorithms under conditions of imbalanced sample data. Furthermore, the system uses a cross-subject model trained on a self-collected dataset to achieve cross-scene and cross-environment target detection. This demonstrates that the eye-movement-based slow-sequence visual presentation experimental paradigm proposed in this embodiment can effectively acquire synchronized EEG and image data for target detection tasks. The brain-eye-computer data fusion target detection system exhibits strong robustness and generalization performance, and has broad application prospects in the field of target image recognition.

[0107] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made by using the contents of the present invention description and drawings under the inventive concept of the present invention, or direct / indirect application in other related technical fields are included in the patent protection scope of the present invention.

Claims

1. A brain-eye-computer data fusion target detection method, characterized in that: The steps include: Step 1: construct a sample image dataset, wherein the sample image dataset includes a plurality of sample images, each of which has at least one first candidate region frame containing a suspicious object, and the sample image dataset also includes position data and true label data of each first candidate region frame; Step 2: Conduct an eye movement-based slow sequence visual presentation experiment on the subject based on the sample image dataset, and collect the subject's first eye movement data and first EEG data in real time; The eye movement-based slow sequence visual presentation experiment is conducted on the subjects based on the sample image dataset, specifically: Randomly arrange the sample images in the sample image data set to form N groups of image stimulation sequences, each group of image stimulation sequences contains 50 to 100 sample images, where N ≥ 1; The image stimulus sequence is played to the subjects at a playback rate of 3 to 5 seconds per frame. During the playback of the image stimulus sequence, the subjects are required to subjectively search for the target and maintain their gaze for more than 0.3 seconds after finding the target; After each set of image stimulation sequences is played, the next set of image stimulation sequences will be played after a preset period of time. Step 3: obtaining a first candidate region frame that the subject is looking at based on the first eye movement data, and defining it as a second candidate region frame; and extracting a corresponding first candidate image from the corresponding sample image based on the second candidate region frame; Step 4: segmenting the first EEG data based on the timestamp of the first eye movement data to obtain first EEG data segments corresponding to the first candidate image; Step 5: construct a training sample set, wherein the training sample set includes a plurality of sample data, each of which includes a first candidate image and its corresponding true label data and a first EEG data segment; Step 6: Train the brain-computer signal fusion classification model based on the training sample set, and perform target detection based on the trained brain-computer signal fusion classification model.

2. The brain-eye-computer data fusion target detection method according to claim 1, characterized in that: The construction of the sample image dataset is specifically as follows: capturing a plurality of target images with targets and a plurality of non-target images without targets, and using both the target images and the non-target images as the sample images; At least one first candidate area frame containing a suspected target is selected in the sample image by a manual labeling method and / or a computer vision extraction method, and the position data of the first candidate area frame in the corresponding sample image and the real label data of whether the first candidate area frame contains the target are saved.

3. The brain-eye-computer data fusion target detection method according to claim 1, characterized in that: In N groups of image stimulation sequences, the target images in group a of image stimulation sequences only contain target 1, and the target images in group b of image stimulation sequences only contain target 2, where a+b=N; The first candidate image and the first EEG data segment generated by the image stimulation sequence with target 1 are used as a training set and a validation set for training the brain-computer signal fusion classification model; The first candidate image and the first EEG data segment generated by the image stimulation sequence with target 2 are used as a sample set for transfer training of the brain-computer signal fusion classification model.

4. The brain-eye-computer data fusion target detection method according to claim 3, characterized in that: In the image data used as the validation set, Gaussian noise with a mean of 0 and a variance of 0.2 and salt and pepper noise of 0.2 were added to simulate the noise interference that may exist in the actual environment.

5. The method for target detection using brain-eye-computer data fusion according to any one of claims 1 to 4, characterized in that: In a set of image stimulation sequences, the probability of the target image appearing is 40% to 50%, and 1 to 2 non-target images are randomly spaced between every two target images.

6. The method for target detection using brain-eye-computer data fusion according to any one of claims 1 to 4, characterized in that: When the image stimulation sequence is played, the subject faces the display playing the image stimulation sequence and is 50 cm to 70 cm away from the display.

7. The method for target detection using brain-eye-computer data fusion according to any one of claims 1 to 4, characterized in that: In step 6, target detection is performed based on the trained brain-computer signal fusion classification model, specifically: Step 6.1, obtaining an image to be tested, selecting third candidate region frames containing suspected targets in the image to be tested, and determining position data of each third candidate region frame in the image to be tested; Step 6.2: Conduct an eye movement-based slow sequence visual presentation experiment on the subject based on the image to be tested, and collect the subject's second eye movement data and second EEG data in real time; Step 6.3: obtaining a third candidate region frame that the subject is looking at based on the second eye movement data, and defining it as a fourth candidate region frame; and extracting the corresponding second candidate image from the image to be tested based on the fourth candidate region frame; Step 6.4, segmenting the second EEG data based on the timestamp of the second eye movement data to obtain second EEG data segments corresponding to the second candidate images; Step 6.5: Input a set of corresponding second candidate images and the second EEG data segment into the trained brain-computer signal fusion classification model to obtain a target detection result corresponding to the second candidate image. If the target detection result indicates that there is a target, the position of the target in the image to be tested can be located by following the corresponding second candidate image; Step 6.6: Repeat step 6.5 until all second candidate images and second EEG data segments are traversed, thus completing the target detection of the image to be tested.

8. The brain-eye-computer data fusion target detection method according to claim 7, characterized in that: In step 6.1, a computer vision extraction method is used to extract all third candidate region frames containing suspected targets in the image to be tested.

Citation Information

Patent Citations

  • Electroencephalogram and eye movement fusion method and device for remote sensing image target detection

    CN109255309A

  • Method for detecting image target in smart home environment

    WO2021244079A1