Intelligent feature cognition detection device for visual detection
By combining multimodal data acquisition and dynamic stimulus generation with the Transformer-LSTM model, the problem of insufficient detection accuracy in existing visual detection technologies is solved, enabling in-depth analysis and personalized evaluation of the visual cognitive process, and improving the adaptability and accuracy of the detection results.
Patent Information
- Application Number
- CN202511568233.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-27
AI Technical Summary
Existing visual inspection technologies lack sufficient accuracy in detecting human visual feature cognition in complex scenarios, have incomplete multimodal data acquisition, and their stimulus generation and evaluation models cannot adapt to different cognitive levels, resulting in weak correlation between detection results and actual job capabilities.
The system employs a multimodal perception module to simultaneously collect eye movement, pupil, and EEG data. Combined with a dynamic stimulus generation unit, it generates highly adaptive visual detection targets. A Transformer-LSTM hybrid deep learning model is used for deep feature analysis, and personalized cognitive levels are assessed based on three-dimensional indicators.
It achieves a comprehensive reflection of the visual cognitive process, enhances the actual relevance and personalization of detection results, supports real-time adjustment of complexity and duration, and improves the accuracy of feature parsing and the adaptability of evaluation.
Smart Images

Figure CN121400829A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of feature recognition detection technology, specifically to an intelligent feature recognition detection device for visual detection. Background Technology
[0002] In fields such as industrial quality inspection, precision instrument testing, and medical imaging diagnosis, the visual feature recognition ability of workers (including the ability to quickly identify target features, distinguish interference, and analyze complex features) directly determines the efficiency and accuracy of inspection. With the development of technology, automated inspection systems based on machine vision have been widely used, but human visual inspection still has irreplaceable value in complex scenarios (such as the identification of minute defects and the collaborative judgment of multiple features). Therefore, the accurate detection and evaluation of human visual feature recognition ability has become a key requirement.
[0003] Existing visual cognition detection technologies mainly suffer from the following shortcomings: Data acquisition is limited: Most systems rely only on eye-tracking data or simple behavioral data (such as button response), ignoring key modal data such as dynamic pupil changes (reflecting cognitive load) and electroencephalogram (EEG) signals (reflecting neurocognitive processes), resulting in an incomplete characterization of cognitive processes and limited detection accuracy.
[0004] Static stimulus generation: The stimuli are mostly single-feature images of fixed complexity, and the presentation position and sequence are designed in a fixed way. It is impossible to dynamically adjust the difficulty and type according to the subject's real-time cognitive state, making it difficult to adapt to subjects with different cognitive levels, and it is impossible to simulate the dynamic feature changes in the actual detection scenario.
[0005] Superficial feature analysis: Traditional machine learning algorithms (such as SVM and random forest) can only extract static features (such as gaze coordinates and total gaze duration) and cannot capture the spatiotemporal dynamic correlation features in the cognitive process, resulting in insufficient depth of analysis of cognitive mechanisms.
[0006] Fixed assessment model: Using a uniform assessment model (such as scoring based solely on accuracy) does not take into account individual differences in the cognitive abilities of the subjects, nor does it achieve dynamic updates of model parameters, making it impossible to generate personalized assessment reports, and the assessment results have a weak correlation with actual testing work.
[0007] To address these issues, some technologies have attempted to incorporate multimodal data, but these suffer from problems such as poor data synchronization and complex calibration processes (requiring multiple independent calibrations). Furthermore, the lack of a linkage mechanism between stimulus generation and cognitive assessment results in insufficient system adaptability and practicality. Therefore, there is an urgent need for an intelligent detection system and method capable of accurately acquiring multimodal data, intelligently generating dynamic stimuli, analyzing deep features, and providing personalized cognitive assessment. Summary of the Invention
[0008] The purpose of this invention is to provide an intelligent feature recognition detection device for visual inspection, so as to solve the problems existing in the prior art.
[0009] To achieve the above objectives, the present invention provides the following technical solution: an intelligent feature recognition detection device for visual inspection, comprising the following modules: The multimodal perception module is configured to simultaneously acquire multi-source visual cognition-related data from subjects. Specifically, the eye-tracking data sampling rate is no less than 600Hz, with measurement accuracy controlled within the range of 0.15° to 0.3°; the pupil diameter data sampling rate is no less than 300Hz, with accuracy ≤0.01mm; and the EEG signal sampling rate is no less than 1000Hz, with a common-mode rejection ratio ≥110dB. It also supports adaptive perception and calibration for ambient light levels of 200-1200 lumens and ambient temperatures of 15-40℃. The dynamic stimulus generation unit is configured to generate four types of visual detection target stimuli, including structured features, fuzzy features, interference features and composite features. The stimulus resolution is not less than 2K, the refresh rate is ≥240Hz, the stimulus presentation duration is continuously adjustable within the range of 20ms-500ms, and it supports dynamic stimulus difficulty grading based on the subject's real-time cognitive state. The difficulty grading includes beginner, intermediate and advanced levels. The intelligent feature parsing unit is configured to perform the following operations: a) The raw data collected by the multimodal perception module is spatiotemporally segmented according to a 50ms time window to extract motion features of eye movement trajectory (such as saccadic speed and fixation stability), dynamic changes of pupil (such as dilation rate and contraction amplitude), and cognitive-related features of EEG signal (such as P300 amplitude and alpha wave power). b) A three-stage dynamic cognitive analysis model of feature perception, cognitive association and decision output is constructed using the Transformer-LSTM hybrid deep learning model. Adaptive feature weight allocation sub-models are constructed for different stimulus types. The dynamic cognitive assessment unit is configured to provide a real-time cognitive state visualization interface, supporting personalized parameter configuration based on the subject's visual cognitive ability level, ranging from C1 to C5, and generating a feature cognitive efficiency comparison map across stimulus types and time stages, including three-dimensional indicators such as cognitive accuracy, response delay, and cognitive load index.
[0010] Preferably, in the three-stage dynamic cognitive analysis model: The feature perception stage lasts from 0 to 1.2 seconds. An attention-enhanced CNN sub-model is used, with convolutional kernel sizes of 3×3, 5×5, and 7×7. The output feature map has 256 channels. The cognitive association phase lasts from 1.2 to 3.5 seconds. A Transformer encoder sub-model is used, with 8 heads for multi-head attention and 1024 dimensions for the hidden layers of the Feed-Forward network. The decision output phase lasts 3.5-5.0 seconds and uses an LSTM decoder sub-model with 512 hidden layer units and a dropout rate of 0.2.
[0011] Preferably, the four types of visual detection target stimuli of the dynamic stimulus generation unit are designed with the presentation position and sequence balance according to the dual dimensions of target feature density and interference item complexity; each type of stimulus is presented 3-5 times in a single experimental process, dynamically adjusted according to the cognitive ability level, and the presentation interval of adjacent stimuli adopts an adaptive adjustment mode with an adjustment range of 30ms-200ms, and the adjustment is based on the cognitive accuracy of the previous stimulus.
[0012] Preferably, the dynamic cognitive assessment unit includes a visual cognitive ability grading module, which is configured as follows: Based on a three-dimensional index of feature recognition accuracy (weight 40%), cognitive response delay (weight 30%), and cognitive load index (weight 30%), the subjects were divided into five cognitive levels: C1 (basic level), C2 (intermediate level), C3 (proficient level), C4 (professional level), and C5 (expert level). Differentiated model parameters were assigned to different cognitive levels. Among them, C4 / C5 level subjects showed a 0.4-0.6-fold increase in the attention weight of the Transformer encoder during the cognitive association phase (1.2-3.5s), while C1 and C2 level subjects showed a 2-3-fold increase in the number of feature extraction iterations of the CNN sub-model during the feature perception phase (0-1.2s).
[0013] Preferably, the intelligent feature parsing unit is further configured to: set a dynamic threshold for the cognitive accuracy of composite feature stimuli during the decision output phase of 3.5-5.0s; when the cognitive accuracy of C1 and C2 level subjects is detected to be lower than 60% for three consecutive times, trigger the dynamic stimulus generation unit to switch the subsequent stimuli to low-complexity structured feature stimuli and extend the presentation time by 1.5 times; when the cognitive accuracy of C4 and C5 level subjects is detected to be higher than 95% for three consecutive times, trigger the dynamic stimulus generation unit to switch the subsequent stimuli to ultra-high complexity composite features and random interference stimuli and shorten the presentation time to 70% of the original time.
[0014] Preferably, a multi-source data synchronization calibration device is provided between the multimodal sensing module and the dynamic stimulus generation unit. This device is configured to: execute a 12-point calibration procedure (eye movement / pupil) and EEG signal baseline calibration before the experiment begins. The 12-point calibration procedure is used to calibrate the eye movement data and pupil diameter data in the multimodal sensing module, and the calibration error is controlled within eye movement ≤0.1°, pupil ≤0.005mm, and EEG signal noise ≤2μV; and dynamically adjust the timestamp alignment algorithm according to the actual refresh rate (240Hz) of the dynamic stimulus generation unit and the multimodal data sampling rate.
[0015] A detection method for a visual inspection intelligent feature recognition detection device includes the following steps: A. Activate the multimodal perception module to complete the initial calibration of eye movement, pupil, EEG and environmental parameters, and simultaneously collect multi-source visual cognition related data of the subjects; B. Input the collected multi-source raw data into the intelligent feature parsing unit, perform spatiotemporal joint segmentation according to a 50ms time window, and extract multi-dimensional cognitive features; C. The dynamic stimulus generation unit generates four types of visual detection target stimuli of corresponding complexity based on the subject's initial cognitive level, which is C3 by default, and presents them in sequence. D. The intelligent feature parsing unit adopts the Transformer-LSTM hybrid deep learning model, and analyzes the subject's feature cognition process and pattern of different types of stimuli through three-stage modeling of feature perception, cognitive association and decision output. E. Based on the model analysis results, the dynamic cognitive assessment unit calculates the subject's cognitive accuracy, reaction delay, and cognitive load index, generates a visual cognitive ability level assessment report and a feature cognitive efficiency comparison map, and triggers the dynamic stimulus generation unit to adjust the parameters of subsequent stimuli (complexity, presentation duration, and interval).
[0016] Preferably, step B further includes: performing fusion preprocessing on the segmented multi-source data, wherein eye-tracking data is filtered by Kalman filtering to remove noise, pupil data is transformed by wavelet transform (db4 wavelet basis, decomposition level 3) to eliminate blinking interference, EEG signals are separated by independent component analysis (ICA) to separate EEG artifacts, and then multi-dimensional features are mapped to the [0,1] interval by feature normalization (Z-score normalization), and finally a three-dimensional feature matrix of time, features and modality is constructed for model input.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This invention integrates eye-tracking, pupillary, and electroencephalogram (EEG) data, achieving superior sampling rate and accuracy compared to existing technologies. Combined with environmental adaptive calibration, data integrity is enhanced, comprehensively reflecting the neural and behavioral processes of visual feature cognition. Based on a linkage mechanism between cognitive state and stimulus parameters, four types of stimuli closely resemble actual testing scenarios, supporting real-time adjustment of complexity, duration, and interval to suit subjects with different cognitive levels. Compared to fixed stimulus schemes, the correlation between test results and actual job capabilities is improved. Employing a Transformer-LSTM hybrid network, it captures the spatiotemporal dynamic features of the cognitive process in stages. Compared to traditional machine learning algorithms, feature parsing accuracy is improved, revealing the complete cognitive chain of perception, association, and decision-making. Based on three-dimensional indicators, five cognitive levels are defined, generating an assessment report including job suitability suggestions. It supports online model fine-tuning, evaluation, and closed-loop control, enhancing personalization compared to uniform assessment models and providing precise data support for enterprise talent selection and employee training. Attached Figure Description
[0018] Figure 1 This is a system schematic diagram of the present invention; Figure 2 This is a schematic diagram of the multimodal sensing module of the present invention; Figure 3 This is a schematic diagram of the dynamic stimulus generation unit of the present invention; Figure 4 This is a schematic diagram of the dynamic cognitive assessment unit of the present invention; Figure 5 This is a flowchart of the method of the present invention. Detailed Implementation
[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figure 1-5 The present invention provides an intelligent feature recognition detection device for visual detection, including a multimodal perception module, a dynamic stimulus generation unit, an intelligent feature parsing unit, and a dynamic recognition evaluation unit; The multimodal perception module integrates multiple types of sensors to achieve synchronous acquisition and adaptive calibration of multi-source data related to visual cognition. Specific configuration: Eye-tracking submodule: It adopts a desktop high-precision eye tracker with a sampling rate of 600Hz, a measurement accuracy of 0.15°-0.3°, a horizontal tracking range of ±40°, a vertical tracking range of ±30°, and supports dual-mode positioning of corneal reflection and pupil center. Pupil Dynamic Monitoring Submodule: Works in conjunction with the eye tracking submodule, independently optimizes the pupil diameter measurement algorithm, with a sampling rate of 300Hz and an accuracy of 0.01mm, and can capture minute pupil changes within 10ms (such as pupil dilation when cognitive load increases). EEG signal acquisition submodule: adopts dry electrode EEG cap (8 channels, covering the frontal, central, and parietal lobe regions), sampling rate 1000Hz, common-mode rejection ratio 110dB, input impedance ≥100MΩ, and supports real-time removal of electrooculography and electromyography artifacts; Environmental perception submodule: integrates light sensor and temperature sensor, collects environmental parameters in real time and feeds them back to the calibration unit to achieve environmental adaptation of data acquisition.
[0021] The dynamic stimulus generation unit, based on cognitive science and the needs of visual detection scenarios, generates four types of visual stimuli with clear detection significance and supports dynamic adjustment. Specific configurations include: Types of irritants: Structured feature stimuli: such as regular feature images of standard-sized holes and slots in industrial products (defect rate 0%). Blurred feature stimuli: such as product defect images with blurred edges due to wear and uneven lighting (medium difficulty in defect recognition). Interference stimuli: Images containing a large number of interference items similar to the target defect (such as a mixture of scratches and stains on a metal surface) (interference items account for 30%-60%). Composite feature stimuli: Complex product images that simultaneously contain multiple defects (such as cracks, deformation, color difference) (2-4 types of defects). Presentation parameters: 2K resolution (2560×1440), 240Hz refresh rate (to ensure no ghosting in dynamic stimuli), presentation duration continuously adjustable from 20ms to 500ms, presentation position randomly distributed based on the screen's nine-grid layout (to avoid the influence of positional preferences). Dynamic adjustment mechanism: The built-in cognitive state and stimulus parameter mapping model automatically adjusts the complexity of subsequent stimuli (e.g., reducing the proportion of interference items by 20% when the accuracy is <70%), presentation duration (extended by 20%), and interval (extended by 15%) based on the subject's cognitive accuracy and response delay of the previous stimulus.
[0022] The intelligent feature parsing unit is responsible for deep processing of multi-source data and parsing of the feature recognition process. Specific configuration: Data preprocessing: Multi-source data are spatiotemporally segmented in 50ms time windows (ensuring precision in the time dimension). Noise and artifacts are removed using Kalman filtering (eye-tracking data), wavelet transform (pupil data), and ICA (electroencephalogram data). Then, a three-dimensional feature matrix is constructed through feature normalization (time dimension: 50ms / window; feature dimensions: 12-dimensional eye-tracking + 8-dimensional pupil + 10-dimensional EEG; modality dimension: 3 categories). Deep learning model: Employs a Transformer-LSTM hybrid network, consisting of three sub-models: Feature perception stage (0-1.2s): Attention-enhanced CNN (containing 3×3, 5×5, 7×7 convolutional kernels) extracts local features and outputs a 256-channel feature map, focusing on capturing the subject's initial perception process of stimulus features; Cognitive association stage (1.2-3.5s): An 8-head Transformer encoder constructs the spatiotemporal association between features (such as the linkage between eye movement trajectory and P300 amplitude), and the Feed-Forward network has a dimension of 1024 to achieve cross-modal feature fusion; Decision output phase (3.5-5.0s): The 512-unit LSTM decoder outputs the cognitive decision result (e.g., "crack defect exists" or "no defect"), with a dropout rate of 0.2 to prevent overfitting; Dynamic feature library: Stores the subject's cognitive feature data in real time, including feature patterns of different stimulus types and different cognitive stages, supports online model fine-tuning, and updates model parameters every 100 trials.
[0023] The dynamic cognitive assessment unit enables personalized assessment and visualization of cognitive abilities. Specific configuration: Cognitive level classification: Based on three-dimensional indicators (weight: recognition accuracy 40%, reaction delay 30%, cognitive load index 30%), the K-means clustering algorithm was used to classify the subjects into five levels, C1-C5, with each level corresponding to a clear visual inspection job suitability standard; Evaluation index calculation: Cognitive accuracy: the percentage of trials in which features are correctly identified; Response delay: the time from stimulus presentation to decision-making, averaged after removing outliers; Cognitive Load Index: Calculated based on pupil dilation rate (weight 50%) and EEG alpha wave power (weight 50%) (index range 0-1, 0 indicates low load, 1 indicates high load); Visualization and Feedback: Provides a real-time cognitive status interface (displaying the current stimulus, multimodal data waveforms, and cognitive load curves), generates comparative graphs across stimulus types and time periods (such as the cognitive accuracy change curves of subjects at different levels under composite characteristic stimuli), and supports exporting assessment reports (including cognitive strengths / weaknesses analysis and job suitability suggestions).
[0024] A multi-source data synchronization calibration device is positioned between the multimodal sensing module and the dynamic stimulus generation unit to solve the problems of time synchronization and accuracy calibration of multi-source data. Initial calibration: Before the experiment, perform 12-point eye movement / pupil calibration (covering the entire screen area) and 10-second baseline calibration of EEG signals (recording resting EEG). The calibration error is controlled within the range of eye movement ≤0.1°, pupil ≤0.005mm, and EEG noise ≤2μV. Dynamic synchronization: Based on the 240Hz refresh rate of the dynamic stimulus generation unit, the timestamps of multimodal data are adjusted in real time (one data point every 1.67ms for eye-tracking at 600Hz, and one data point every 1ms for EEG at 1000Hz), and precise synchronization between data and stimulus presentation is achieved through hardware trigger signals.
[0025] The intelligent feature recognition detection method for visual detection of the present invention is implemented based on the above system, and the specific steps are as follows: Step A: System Startup and Initial Calibration The multimodal perception module was activated, and the subjects wore EEG caps and adjusted the position of the eye tracker (chin rest fixed, eyes 65cm from the center of the screen). Perform 12-point eye-tracking / pupil calibration and 10-second EEG baseline calibration. The environmental perception submodule collects current illumination (e.g., 500 lumens) and temperature (25°C) parameters, and automatically adjusts the eye-tracking sampling gain and EEG signal amplifier parameters to ensure data acquisition accuracy. Subjects complete basic information filling and cognitive ability pre-test (10 structured feature stimuli trials). The system initially classifies the cognitive level based on the pre-test results (default level C3, if the pre-test accuracy is >90%, the initial level is set to level C4, and if <60%, it is set to level C2).
[0026] Step B: Multi-source data acquisition and preprocessing The multimodal perception module synchronously collects subject data at 600Hz (eye movement), 300Hz (pupil), and 1000Hz (EEG), while the environmental perception submodule updates environmental parameters every 100ms. The intelligent feature parsing unit segments the raw data in 50ms time windows and sequentially performs noise removal (Kalman filtering, wavelet transform, ICA), feature extraction (such as eye movement saccade velocity, pupil dilation rate, and EEG P300 amplitude), and feature normalization (Z-score standardization), generating a three-dimensional feature matrix of time, features, and modality, which is transmitted to the model calculation unit in real time.
[0027] Step C: Dynamic Stimulus Generation and Presentation The dynamic stimulus generation unit randomly selects four types of stimuli from the stimulus library based on the subject's initial cognitive level (5-8 trials per type, for a total of 20-32 trials), and initially presents them in the order of "structured → ambiguous → interference → complex", and then dynamically adjusts them according to the cognitive state. When the stimulus is presented, the hardware trigger signal of the multimodal perception module is triggered synchronously to ensure that the data acquisition and stimulus presentation are synchronized. During the presentation, gaze point prompts are displayed in real time to guide the subject to concentrate.
[0028] Step D: Modeling and Analysis of Feature Cognition Process The intelligent feature parsing unit inputs the 3D feature matrix into the Transformer-LSTM hybrid model and performs calculations in stages: Feature perception stage (0-1.2s): Attention-enhanced CNN extracts local features and outputs a feature perception probability map, reflecting the subject's attention to each region of the stimulus; Cognitive association stage (1.2-3.5s): The Transformer encoder constructs the spatiotemporal association of eye movement-pupil-EEG features and outputs the feature association strength matrix, such as the correlation coefficient between P300 amplitude and fixation duration of target features; Decision output stage (3.5-5.0s): The LSTM decoder outputs the cognitive decision result (e.g., "There are two types of defects: cracks + color difference") and the decision confidence level (0-1). The dynamic feature library automatically fine-tunes model parameters after receiving 100 parsing results, such as adjusting the Transformer attention head weights, to improve the accuracy of subsequent parsing.
[0029] Step E: Cognitive Ability Assessment and Dynamic Regulation The dynamic cognitive assessment unit calculates the three-dimensional assessment indicators of the subjects based on the model output results, including cognitive accuracy, reaction delay, and cognitive load index. It uses the K-means clustering algorithm to update the cognitive level. For example, the subjects who were initially at level C3 had a composite feature stimulus accuracy of 92%, a reaction delay of 0.8s, and a cognitive load index of 0.3, and were upgraded to level C4. Generate a personalized assessment report, including: Cognitive ability radar chart, three-dimensional performance of four types of stimuli; A temporal analysis diagram of the cognitive process, showing the changes in cognitive efficiency at each stage within 0-5 seconds. Job suitability suggestions, such as "To be suitable for a precision quality inspection position for electronic components, it is necessary to strengthen training in fuzzy feature recognition"; The dynamic stimulus generation unit is triggered to adjust the parameters of subsequent stimuli. For example, the number of defective types of subsequent composite characteristic stimuli for C4 level subjects is increased to 4, and the presentation time is shortened to 80% of the original time, thus achieving a closed loop of assessment and control.
[0030] Example: An intelligent feature recognition detection method for visual inspection, comprising the following steps: Recruitment and Screening of Participants Recruitment targets: 100 quality inspection staff for an auto parts manufacturing company, aged 20-45, with corrected vision ≥1.0 and no history of neurological diseases; Grouping: Based on current work performance (quality inspection accuracy), the group is divided into three groups: Excellent Group (accuracy ≥ 95%, 30 people), Satisfactory Group (accuracy 85%-95%, 40 people), and Group Needing Improvement (accuracy < 85%, 30 people). Experimental preparation: Subjects should acclimatize to the laboratory environment (500 lumens light, 25°C) 10 minutes in advance, wear an EEG cap and complete the initial calibration, and read the experimental instructions (understanding the task requirements of "identifying the type of defect in the stimulus").
[0031] Experimental Procedure Pre-test: Ten structured stimuli were presented, and the system initially classified the cognitive levels (the excellent group was mostly at level C4 / C5, and the group that needed improvement was mostly at level C1 / C2). Formal experiment: 20 trials were presented (5 of each of the four types of stimuli). The dynamic stimulus generation unit adjusted the parameters according to the cognitive state of the previous trial (e.g., in the third trial of the excellent group, the number of defective composite feature stimuli increased to 4). Data recording: The multimodal perception module collects data synchronously, the intelligent feature analysis unit processes it in real time, and the dynamic cognitive assessment unit generates a real-time cognitive state curve; Post-experiment: Output personalized evaluation reports, and participants fill out an experiment experience questionnaire (to evaluate the system's usability).
[0032] Analysis of Experimental Results Data integrity: The multimodal data loss rate and time synchronization error decreased during the experiment, demonstrating the stability of the system's data acquisition. Analysis accuracy: The accuracy of the intelligent feature analysis unit in cognitive decision-making for all four types of stimuli has been improved; Level matching degree: The improved matching rate between the cognitive levels classified by the system and the actual work performance of the subjects proves the effectiveness of the evaluation results.
[0033] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A visual inspection intelligent feature recognition detection device, characterized in that, Includes the following modules: The multimodal perception module is configured to simultaneously acquire multi-source visual cognition-related data from subjects. Specifically, the eye-tracking data sampling rate is no less than 600Hz, with measurement accuracy controlled within the range of 0.15° to 0.3°; the pupil diameter data sampling rate is no less than 300Hz, with accuracy ≤0.01mm; and the EEG signal sampling rate is no less than 1000Hz, with a common-mode rejection ratio ≥110dB. It also supports adaptive perception and calibration for ambient light levels of 200-1200 lumens and ambient temperatures of 15-40℃. The dynamic stimulus generation unit is configured to generate four types of visual detection target stimuli, including structured features, fuzzy features, interference features and composite features. The stimulus resolution is not less than 2K, the refresh rate is ≥240Hz, the stimulus presentation duration is continuously adjustable within the range of 20ms-500ms, and it supports dynamic stimulus difficulty grading based on the subject's real-time cognitive state. The difficulty grading includes beginner, intermediate and advanced levels. The intelligent feature parsing unit is configured to perform the following operations: a) The raw data collected by the multimodal perception module is spatiotemporally segmented according to a 50ms time window to extract motion features of eye movement trajectory, dynamic change features of pupil, and cognitive-related features of EEG signal; b) A three-stage dynamic cognitive analysis model of feature perception, cognitive association and decision output is constructed using the Transformer-LSTM hybrid deep learning model. Adaptive feature weight allocation sub-models are constructed for different stimulus types. The dynamic cognitive assessment unit is configured to provide a real-time cognitive state visualization interface, supporting personalized parameter configuration based on the subject's visual cognitive ability level, ranging from C1 to C5, and generating a feature cognitive efficiency comparison map across stimulus types and time stages, including three-dimensional indicators such as cognitive accuracy, response delay, and cognitive load index.
2. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that, In the aforementioned three-stage dynamic cognitive analysis model: The feature perception stage lasts from 0 to 1.2 seconds. An attention-enhanced CNN sub-model is used, with convolutional kernel sizes of 3×3, 5×5, and 7×7. The output feature map has 256 channels. The cognitive association phase lasts from 1.2 to 3.5 seconds. A Transformer encoder sub-model is used, with 8 heads for multi-head attention and 1024 dimensions for the hidden layers of the Feed-Forward network. The decision output phase lasts 3.5-5.0 seconds and uses an LSTM decoder sub-model with 512 hidden layer units and a dropout rate of 0.
2.
3. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that, The four types of visual detection target stimuli in the dynamic stimulus generation unit are designed with a dual-dimensional approach of target feature density and interference complexity to balance their presentation positions and sequences. In a single experimental procedure, each type of stimulus is presented 3-5 times, dynamically adjusted according to the level of cognitive ability. The interval between presentations of adjacent stimuli adopts an adaptive adjustment mode, with an adjustment range of 30ms-200ms, and the adjustment is based on the cognitive accuracy of the previous stimulus.
4. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that, The dynamic cognitive assessment unit includes a visual cognitive ability grading module, which is configured as follows: Based on three-dimensional indicators—feature recognition accuracy, cognitive response delay, and cognitive load index—the subjects were divided into five cognitive levels: C1, C2, C3, C4, and C5. Differentiated model parameters were assigned to different cognitive levels. Among them, the attention weight of the Transformer encoder was increased by 0.4-0.6 times in the cognitive association stage for C4 / C5 level subjects, and the number of feature extraction iterations of the CNN sub-model was increased by 2-3 times in the feature perception stage for C1 and C2 level subjects.
5. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that, The intelligent feature parsing unit is further configured to: set a dynamic threshold for the cognitive accuracy of composite feature stimuli during the decision output phase (3.5-5.0s); when the cognitive accuracy of C1 and C2 level subjects is detected to be below 60% for three consecutive times, trigger the dynamic stimulus generation unit to switch subsequent stimuli to low-complexity structured feature stimuli and extend the presentation time by 1.5 times; when the cognitive accuracy of C4 and C5 level subjects is detected to be above 95% for three consecutive times, trigger the dynamic stimulus generation unit to switch subsequent stimuli to ultra-high-complexity composite features and random interference stimuli and shorten the presentation time to 70% of the original time.
6. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that, A multi-source data synchronization calibration device is provided between the multimodal sensing module and the dynamic stimulus generation unit. This device is configured to: execute a 12-point calibration procedure and EEG signal baseline calibration before the experiment begins. The 12-point calibration procedure is used to calibrate the eye movement data and pupil diameter data in the multimodal sensing module; the calibration error is controlled within the range of eye movement ≤0.1°, pupil ≤0.005mm, and EEG signal noise ≤2μV; and the timestamp alignment algorithm is dynamically adjusted according to the actual refresh rate of the dynamic stimulus generation unit and the multimodal data sampling rate.
7. The detection method of the intelligent feature recognition detection device for visual detection according to any one of claims 1-6, characterized in that, Includes the following steps: A. Activate the multimodal perception module to complete the initial calibration of eye movement, pupil, EEG and environmental parameters, and simultaneously collect multi-source visual cognition related data of the subjects; B. Input the collected multi-source raw data into the intelligent feature parsing unit, perform spatiotemporal joint segmentation according to a 50ms time window, and extract multi-dimensional cognitive features; C. The dynamic stimulus generation unit generates four types of visual detection target stimuli of corresponding complexity based on the subject's initial cognitive level, which is C3 by default, and presents them in sequence. D. The intelligent feature parsing unit adopts the Transformer-LSTM hybrid deep learning model, and analyzes the subject's feature cognition process and pattern of different types of stimuli through three-stage modeling of feature perception, cognitive association and decision output. E. Based on the model analysis results, the dynamic cognitive assessment unit calculates the subject's cognitive accuracy, reaction delay, and cognitive load index, generates a visual cognitive ability level assessment report and a feature cognitive efficiency comparison map, and triggers the dynamic stimulus generation unit to adjust the parameters of subsequent stimuli.
8. The detection method of the intelligent feature recognition detection device for visual inspection according to claim 7, characterized in that, Step B further includes: performing fusion preprocessing on the segmented multi-source data, wherein eye-tracking data is filtered by Kalman filtering to remove noise, pupil data is filtered by wavelet transform to eliminate blinking interference, EEG signals are filtered by independent component analysis (ICA) to separate EEG artifacts, and then multi-dimensional features are mapped to the [0,1] interval through feature normalization. Finally, a three-dimensional feature matrix of time, features and modality is constructed for model input.
Citation Information
Cited By
Method for assessing cognitive performance under environmental light stimulation based on multi-modal physiological monitoring
CN122423816A