Human-machine cooperation camouflage object detection method based on uncertainty perception
By adopting a human-computer collaboration method based on uncertainty perception in camouflage object detection, combined with computer vision model and brain-computer interface technology, the problem of insufficient feature extraction in traditional models in complex scenarios is solved, which significantly improves the reliability and accuracy of detection and reduces manual intervention.
Patent Information
- Application Number
- CN202510243060.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art In complex backgrounds, low contrast or highly dynamic scenarios, traditional computer vision models may miss detection or misjudgment due to insufficient feature extraction capabilities, and humans need to consume a lot of time and energy when performing camouflage object detection tasks, and EEG signals are susceptible to individual differences and physiological fatigue.
The human-computer collaborative camouflage object detection method based on uncertainty perception is used to output the predicted uncertainty through the computer vision model as a human-computer bridge. The CV model can use the ability to process a large amount of data in parallel to output the preliminary prediction probability, and the samples screened based on uncertainty are further detected through the brain-computer interface.
It significantly improves the reliability and accuracy of camouflage object detection. By deeply integrating computer vision model and brain-computer interface technology, a multi-view backbone network is built, the confidence of the model is dynamically evaluated, uncertain cases in the detection are accurately identified, and the frequency of manual intervention is reduced.
Smart Images

Figure CN120219707A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of brain-computer collaboration, relates to a target detection method, and particularly relates to a human-machine collaboration camouflaged object detection method based on uncertainty perception. Background Art
[0002] The core task of camouflaged object detection (COD) is to identify camouflaged objects hidden in the environment, which has been widely applied in multiple fields, including defect detection in manufacturing, pest monitoring in agriculture, lesion segmentation in medicine, pedestrian detection in night environments, etc. The COD task usually involves two key steps: firstly, judging whether a camouflaged object exists; secondly, when a camouflaged object exists, performing object segmentation processing. Currently, most research focuses on the segmentation stage, assuming that the model can already generate a continuous mask map, and using a black saliency map to indicate the absence of camouflaged objects. However, considering that camouflaged objects do not always exist, identifying the existence of objects is a key step for subsequent segmentation. Especially in tasks such as rescue operations and pest monitoring, accurately detecting the existence of camouflaged objects is the primary task.
[0003] Although computer vision (CV) technology shows potential in camouflaged object detection, it still faces significant challenges in practical applications. In complex backgrounds, low contrast, or highly dynamic scenarios (such as night environments, agricultural monitoring with dense vegetation cover), traditional CV models may miss detections or make misjudgments due to insufficient feature extraction capabilities. For example, when the camouflaged object is highly homogeneous with the background texture (such as a camouflage target in a desert environment) or there are drastic changes in illumination, models based on convolutional neural networks may not be able to effectively capture the edge and semantic information of the object. If the performance is improved solely by increasing the model complexity or data augmentation means, the computational cost will increase exponentially.
[0004] The human brain demonstrates extraordinary abilities in adapting to changing environments, recognizing subtle patterns, and finding hidden objects in complex scenarios. However, manually performing this task usually requires a large amount of time and effort. Non-invasive brain-computer interface (BCI) technology provides a novel solution to solve this problem. However, electroencephalogram (EEG) signals are easily affected by individual differences, resulting in a decrease in the stability of the results. At the same time, the attention drift caused by the operator's physiological fatigue will also weaken the classification accuracy.
[0005] Human-machine cooperation has become an important trend in the development of the digital society, especially in application scenarios where humans and artificial intelligence complement each other and collaborate to complete tasks. The key to building a trustworthy artificial intelligence system lies in enabling the AI model to evaluate and communicate its own uncertainty, which can not only provide important guidance for the decision-making process but also guide human operators when needed or request more data in complex scenarios. However, how to effectively utilize uncertainty quantification to promote this cooperation and its impact remains an under-explored research area. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention proposes an uncertainty-aware human-machine collaborative camouflaged object detection method, which uses the uncertainty of the prediction output by the CV model as a bridge between humans and machines. First, it utilizes the ability of the CV model to process a large amount of data in parallel to output preliminary prediction probabilities, and then further detects the samples screened based on uncertainty through a brain-computer interface (BCI), thereby improving the performance of the final target detection.
[0007] An uncertainty-aware human-machine collaborative camouflaged object detection method, the specific steps are as follows:
[0008] Step 1: Collect picture sample data
[0009] Collect pictures under different environments, scenes, and lighting conditions as sample data. Some pictures include camouflaged objects of different types and different degrees of camouflage. Mark the sample data according to whether it contains a camouflaged target to obtain labeled data.
[0010] Step 2: Train the computer vision model
[0011] Perform weak data augmentation operations and strong data augmentation operations on the sample data collected in Step 1 respectively, and then input the augmented samples and the original samples into the computer vision model together to train the computer vision model's detection ability for whether there is a camouflaged target in the picture.
[0012] The weak data augmentation operation is one of random horizontal flipping, slight rotation, or random cropping, aiming to simulate common slight changes in natural scenes to help the model remain robust when facing small image changes.
[0013] The strong data augmentation operation is one of cropping, color transformation, image quality adjustment, occlusion, or composite augmentation. The strong data augmentation operation involves a larger range of perturbations to simulate complex scene changes and help the model better adapt to complex situations in the real world.
[0014] Preferably, the computer vision model is SwinT, DenseNet-161, or ResNet-18.
[0015] Step 3: Electroencephalogram (EEG) data acquisition
[0016] Using the rapid serial visual presentation (RSVP) paradigm in a brain-computer interface (BCI), sample data in the input computer vision model is played to the subject, and at the same time, the subject's EEG data is collected. After preprocessing the EEG data, it is used as an EEG model sample. The event-related potentials in the EEG signals are marked as EEG model labels.
[0017] Step 4: EEG model training
[0018] The preprocessed EEG signals in Step 3 are input into the EEGNet model for feature extraction to capture the spatio-temporal patterns of the EEG signals, and the EEGNet model is trained to recognize the event-related potentials in the EEG signals, realizing the ability of detecting camouflage targets based on EEG signals.
[0019] Step 5: Obtaining CV prediction probability and uncertainty estimation
[0020] For the pictures to be detected for camouflage targets, first, they are input into the computer vision model trained in Step 2 to obtain the preliminary prediction probability of containing camouflage targets. Then, a weak data augmentation operation and multiple strong data augmentation operations are performed respectively to obtain a weakly augmented sample and multiple strongly augmented samples, which are respectively input into the computer vision model trained in Step 2, and then the distribution differences between the weakly augmented sample and each strongly augmented sample are calculated as the uncertainty of the sample.
[0021] Step 6: Obtaining EEG prediction probability
[0022] According to the uncertainty calculation result in Step 5, the samples with uncertainty higher than the set threshold are played to the subject in the RSVP paradigm, and at the same time, the subject's EEG data is collected. After preprocessing, it is input into the EEGNet model trained in Step 4 for inference to obtain the corresponding sample category and probability distribution.
[0023] Step 7: Decision fusion for camouflage target detection
[0024] For the samples with uncertainty lower than the set threshold, the preliminary prediction probability output by the computer vision model in Step 5 is directly used as the final camouflage target detection result. For the samples with uncertainty higher than the set threshold, the probability distribution output by the EEGNet model in Step 6 is used as the camouflage target detection result.
[0025] The present invention has the following beneficial effects:
[0026] 1. By deeply integrating computer vision models with brain-computer interface technology based on rapid serial visual presentation, an innovative framework is proposed, constructing a multi-view backbone network that adopts a dual-path parallel inference mechanism of strong and weak data augmentation, dynamically evaluating model confidence through the differences in prediction results, thereby accurately identifying uncertain cases in detection. For low-confidence samples, leveraging humans' keen recognition ability of camouflaged targets at the visual cognitive level, cognitive feedback re-evaluation is performed on difficult cases of the CV model, significantly improving the reliability and accuracy of camouflaged object detection.
[0027] 2. A learning system of "machine main judgment - human-machine collaboration" is constructed. Through a confidence screening mechanism, a small number of uncertain cases are handed over to humans for processing. Utilizing humans' semantic understanding and context reasoning abilities in complex visual scene analysis, it not only gives full play to the high-efficiency data processing ability of the CV model but also integrates humans' cognitive advantages in marginal cases, significantly reducing the frequency of human intervention while ensuring a 5% increase in the overall detection accuracy of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a flowchart of model training for human-machine collaborative camouflaged object detection based on uncertainty perception;
[0029] Figure 2 is a schematic diagram of a sample containing a camouflaged target in the embodiment;
[0030] Figure 3 is a schematic diagram of a sample without a camouflaged target in the embodiment;
[0031] Figure 4 is a schematic diagram of the computer vision model structure used in the embodiment;
[0032] Figure 5 is a schematic diagram of the rapid serial visual presentation paradigm in the embodiment;
[0033] Figure 6 is a flowchart of the method for human-machine collaborative camouflaged object detection based on uncertainty perception. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The method of the present invention will be described in detail below with reference to the accompanying drawings. A method for human-machine collaborative camouflaged object detection based on uncertainty perception comprises the following specific steps:
[0035] Step 1. Collect picture sample data
[0036] Such as Figure 1As shown, pictures are collected under different environments, scenarios, and lighting conditions as sample data, including indoor and outdoor scenes, different meteorological conditions (such as sunny, cloudy, rainy, foggy weather, etc.), and different background complexities (such as uniform backgrounds, backgrounds with rich textures, dynamic environments, etc.) to ensure the diversity and representativeness of the data. Some pictures include different types of camouflaged objects, such as natural camouflage (such as the protective color of animals or insects), artificial camouflage (such as military camouflage equipment), and other complex camouflage situations. At the same time, the pictures also contain different degrees of camouflage effects, such as partial occlusion of the target object, low-contrast camouflage, background interference (such as noise, shadows, obstacles), etc., to simulate the visual challenges in real scenarios.
[0037] The sample data is labeled according to whether there is a camouflaged target to obtain labeled data.
[0038] In this embodiment, 1250 camouflaged target images provided in the publicly available dataset CAMO are selected, and 1250 background images are generated by cropping these 1250 images. Finally, 2000 images are selected as the training set and 500 images as the test set, and the number of positive and negative samples in the training set and the test set is the same. Figure 2 is a positive sample with a camouflaged target, Figure 3 is a negative sample with a camouflaged target.
[0039] Step 2: Training of the computer vision model
[0040] In this embodiment, SwinT as shown in Figure 4 is selected as the computer vision model. During the training process, through strong and weak data augmentation strategies, it is ensured that the model can adapt to diverse input data and complex environmental changes, improve the robustness of the model, and evaluate the uncertainty based on multiple perspectives. Specifically, weak data augmentation operations and strong data augmentation operations are respectively performed on the sample data in the training set, and then the augmented samples and the original samples are input into the computer vision model together to train the computer vision model's detection ability of whether there is a camouflaged target in the picture.
[0041] The weak data augmentation operation is one of random horizontal flipping, slight rotation, or random cropping, aiming to simulate common minor changes in natural scenes to help the model remain robust when facing minor image changes.
[0042] The strong data augmentation operation is one of cropping, color transformation, image quality adjustment, occlusion, or composite augmentation. The strong data augmentation operation involves a larger range of perturbations to simulate complex scene changes and help the model better adapt to complex situations in the real world.
[0043] Step 3: EEG data acquisition
[0044] The training set sample data before and after augmentation is played to the subjects using the rapid serial visual presentation paradigm in the brain-computer interface, while the electroencephalogram (EEG) data of the subjects is recorded while viewing the picture content.
[0045] The experiment was conducted in an enclosed environment to reduce external light interference and ensure the accuracy of the experimental data. The distance between the subject and the screen was maintained at approximately 75 cm, and the line of sight was aligned with the center of the screen. As Figure 5 shown, the experiment was divided into 7 sessions, each session contained 11 blocks, and each block contained 55 pictures. The resolution of the pictures was 512×512, and they were presented sequentially at a frequency of 1 Hz. To ensure the rhythm control of the experiment, the presentation time of each picture was fixed, and at least three non-target pictures were inserted between consecutive target pictures to reduce the influence of potential serial effects on the EEG signals.
[0046] Before the start of each block, a "+" symbol appeared in the center of the screen as a fixation point to guide the subject to concentrate. After the end of each block, the subject could decide the rest time according to their own situation to avoid affecting the experimental results due to fatigue.
[0047] During the experiment, a 64-channel Neuroscan EEG acquisition system was used to record EEG data. The electrode placement followed the international 10-20 system, and the data sampling frequency was set to 1000 Hz to ensure high-time-resolution data acquisition. The impedance of all electrodes was maintained below 15 kΩ to reduce noise interference and improve data quality.
[0048] Step 4: Data preprocessing
[0049] The collected EEG data was preprocessed to remove possible artifact interferences and ensure the accuracy and reliability of subsequent analysis results.
[0050] First, the interference sources such as electrooculogram (EOG) and electromyogram (EMG) signals were identified and removed to reduce the influence of these non-brain-derived signals on the EEG data. Then, the EEG signals were filtered to remove low-frequency drift and high-frequency noise. After filtering, the signals were downsampled, and the sampling frequency was adjusted to 250 Hz to optimize storage and computational efficiency while ensuring sufficient time resolution to capture the key changes in EEG activities. Then, the missing signals or abnormal signals caused by electrode detachment, poor electrode contact, or other technical problems were removed to avoid negative impacts on subsequent analysis. Finally, the preprocessed EEG signals were added with labels corresponding to the experimental paradigm. According to the time markers of the stimuli in the rapid serial visual presentation (RSVP) paradigm, the EEG signal segments were precisely aligned with the stimulus events, and category labels were assigned to the EEG data within each time window.
[0051] Step 5: Pretraining of EEG model
[0052] Use the EEGNet model to efficiently extract features from the preprocessed EEG data in Step 4, train the EEGNet model to recognize event-related potentials in EEG signals, and achieve the ability to detect camouflage targets based on EEG signals.
[0053] EEGNet, as a deep learning model specifically designed for EEG data, can effectively extract discriminative features from raw EEG signals through a multi-layer convolutional network structure. These features not only reflect the spatial distribution of EEG signals but also capture the dynamic patterns of signal changes over time. In this way, the model can more comprehensively learn the intrinsic feature distribution of EEG data, thereby improving its classification performance when dealing with complex tasks.
[0054] Step 6: Obtaining CV prediction probability and uncertainty estimation
[0055] As shown in Figure 6 , for the sample x in the test set, first input it into the computer vision model trained in Step 2 to obtain the preliminary prediction probability of the camouflage target. Then, perform a weak data augmentation operation and n strong data augmentation operations on the sample x respectively to obtain a weakly augmented sample x w and n strongly augmented samples {x s1 , x s2 ,..., x sn}, input them into the computer vision model trained in Step 2 to obtain the corresponding prediction distributions p w and p sj , j = 1, 2,... n. Then, calculate the uncertainty Uncertrainty based on the cross-entropy (CE) to measure the difference between the prediction distribution p w of the weakly augmented sample and the prediction distribution p sj of the strongly augmented samples:
[0056]
[0057] where C represents the number of sample categories, p w (i) is the prediction probability that the weakly augmented sample p w belongs to category i, and p sj (i) is the prediction probability that the strongly augmented sample p sj belongs to category i.
[0058] By calculating the difference in the prediction distributions of weakly augmented samples and strongly augmented samples, the uncertainty of the samples can be effectively quantified, and the stability of the model's prediction results from different perspectives can be evaluated. Combining the prediction probability and uncertainty estimation can improve the robustness and adaptability of the model in complex scenarios, providing a more reliable decision-making basis for subsequent camouflaged target detection tasks.
[0059] Step 7: Obtain the electroencephalogram prediction probability
[0060] Samples with high uncertainty usually involve factors such as the presence or absence of camouflaged targets, partial occlusion of targets, low contrast, background interference (such as complex textures, noise, or shadows), etc., resulting in significant differences in the prediction results of the model from different augmentation perspectives.
[0061] In this embodiment, samples x with Uncertrainty≥0.55 are selected as samples with high uncertainty, played to the subjects in the RSVP paradigm, and at the same time, the electroencephalogram data of the subjects are collected. After preprocessing, they are input into the EEGNet model trained in step 5 for inference to obtain the corresponding sample categories and probability distributions.
[0062] Step 8: Decision fusion for camouflaged target detection
[0063] For samples with uncertainty lower than the set threshold, the preliminary prediction probability output by the computer vision model in step 6 is directly used as the final camouflaged target detection result. For samples with uncertainty higher than the set threshold, the probability distribution output by the EEGNet model in step 7 is used as the camouflaged target detection result.
[0064] To illustrate the effectiveness of this method, several common computer vision object classification models are used to test the detection performance of complex targets in various situations, and the experimental results are measured using the metrics F1 Score and Balanced Accuracy (BA). F1 Score can take into account both the precision and recall metrics, and BA is used to characterize the average of the precisions between classes:
[0065]
[0066] Among them, precision represents the precision rate, which is the proportion of true positive samples among the samples predicted as positive. Recall represents the recall rate, which is the proportion of positive samples that are correctly predicted. TP represents the number of positive classes predicted as positive classes, FN represents the number of positive classes predicted as negative classes, FP represents the number of negative classes predicted as positive classes, and TN represents the number of negative classes predicted as negative classes.
[0067] The specific experimental results are shown in the following table:
[0068]
[0069] According to the results of the comparative experiment, it can be seen that introducing strong and weak data augmentation strategies during the training process of the computer vision model can enhance the model's adaptability to changes in input data and effectively improve the robustness of the model. And this method is based on an uncertainty-aware human-machine collaboration method, which can fully combine the advantages of computer vision and electroencephalogram signals to further improve the recognition accuracy of camouflaged targets and enhance the overall reliability and adaptability of the model.
Claims
1. A human-machine collaborative disguised object detection method based on uncertainty perception, collecting pictures under different environments, scenes and lighting conditions as sample data, marking the sample data according to whether they contain disguised targets, and obtaining an original data set; inputting the pictures in the original data set into a computer vision model, and training the computer vision model to detect whether the pictures contain disguised targets; the characteristics are: The specific steps include: Step 1: Play the pictures in the original data set to the subjects and collect the EEG data of the subjects at the same time; pre-process the EEG data and use it as the EEG model sample; The event-related potentials in the EEG signals are marked as EEG model labels, and then input into the EEGNet model. The EEGNet model is trained to identify the event-related potentials in the EEG signals, thus realizing camouflaged target detection based on EEG signals. Step 2: For the image to be detected, first use the trained computer vision model to obtain the preliminary prediction probability of containing the camouflaged target; Then, different degrees of data enhancement operations are performed to obtain a weakly enhanced sample and multiple strongly enhanced samples. Then, the difference in the initial predicted probability distribution between the weakly and strongly enhanced samples is calculated as the uncertainty of the image to be detected. Step 3: Select the images to be tested whose uncertainty is higher than the set threshold, play them to the subjects, and collect the EEG data of the subjects at the same time. After preprocessing, input them into the trained EEGNet model to obtain the corresponding sample categories and probability distribution; Step 4: For samples whose uncertainty is lower than the set threshold, the preliminary prediction probability output by the computer vision model is directly used as the final camouflaged target detection result; for samples whose uncertainty is higher than the set threshold, the probability distribution output by the EEGNet model in step 3 is used as the camouflaged target detection result.
2. A method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1, characterized in that: The computer vision model is SwinT, DenseNet-161 or ResNet-18.
3. The method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1, characterized in that: Perform different degrees of data augmentation operations on the images in the original dataset to expand the scale of the original dataset for training computer vision models and EEGNet models.
4. A method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1 or 3, characterized in that: The different degrees of data enhancement operations include weak data enhancement operations and strong data enhancement operations; the weak data enhancement operation is one of random horizontal flipping, slight rotation or random cropping, and the strong data enhancement operation is one of cropping, color transformation, image quality adjustment, occlusion or composite enhancement.
5. The method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1, characterized in that: The pictures were played to the subjects using the rapid serial visual presentation paradigm in brain-computer interface.
6. The method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1, characterized in that: The sampling frequency of the EEG signal is set to 1000 Hz. The non-brain-derived signals of the collected raw EEG data are identified and removed. The low-frequency drift and high-frequency noise are filtered out and then downsampled to 250 Hz. Then the missing and abnormal signals are removed. Finally, the EEG signals are divided according to the time stamps of the stimuli in the rapid serial visual presentation paradigm, and the category labels are marked and input into the EEGNet model for training.
7. The method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1, characterized in that: The uncertainty is calculated as: Among them, Uncertrainty represents the uncertainty of the image to be detected, p w 、p sj Weakly enhanced samples x output by computer vision models w And the strongly enhanced sample x sj The initial predicted probability distribution of n represents the number of strongly enhanced samples, j = 1, 2, ... n, C represents the number of sample categories, p w (i) is a weakly enhanced sample p w The predicted probability of belonging to category i, p sj (i) is a strongly enhanced sample p sj The predicted probability of belonging to class i.
8. A method for detecting camouflaged objects by human-machine collaboration based on uncertainty perception as claimed in claim 1 or 7, characterized in that: Set the uncertainty threshold to 0.55.