Intelligent Feature Cognition Detection Device for Visual Inspection
Patent Information
- Application Number
- NL2041437
- Authority / Receiving Office
- NL · NL
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2025-10-30
- Filing Date
- 2025-11-14
- Publication Date
- 2026-08-20
- Estimated Expiration
- 2045-11-14
AI Technical Summary
Existing visual cognition detection technologies rely solely on incomplete data modalities, static stimulus generation, shallow feature analysis, and fixed evaluation models, failing to adapt to individual cognitive levels and dynamic inspection scenarios, leading to inaccurate and non-personalized evaluations.
A system integrating multimodal data acquisition, adaptive stimulus generation, and deep feature analysis using a Transformer-LSTM hybrid model for personalized cognitive evaluation, with synchronized data collection and dynamic stimulus adjustment based on real-time cognitive states.
Enhances data integrity and adaptability, improves feature analysis accuracy, and provides personalized evaluation reports matching actual inspection tasks, supporting real-time complexity adjustment and individualized training recommendations.
Abstract
Description
Technical Field The present invention relates to the technical eld of feature cognition detection, and specically to an . Background Art In elds such as industrial quality inspection, precision instrument testing, and medical image diagnosis, the visual feature cognition ability of personnel (including the abilities of rapid recognition of target features, anti-interference discrimination, and complex feature analysis) directly determines the efciency and accuracy of inspection. With technological advancement, automated detection systems based on machine vision have been widely applied. However, human visual inspection still holds irreplaceable value in complex scenarios (such as identifying subtle defects and making multi-feature collaborative judgments). Therefore, accurate detection and evaluation of human visual feature cognition ability have become a critical requirement. The existing visual cognition detection technologies mainly suffer from the following deciencies: Single-Modality Data Acquisition: Most systems rely solely on eye movement data or simple behavioral data (e.g., key press response time), while ignoring key modal data such as dynamic changes in pupil diameter (reecting cognitive load) and EEG signals (reecting neural cognitive processes). This leads to an incomplete characterization of cognitive processes and limits detection accuracy. Static Stimulus Generation: Stimuli mostly consist of single-feature images with xed complexity, and the presentation positions and sequences are designed statically. These systems cannot dynamically adjust the difculty and type of stimuli based on the subject's realtime cognitive state, thus failing to adapt to subjects of different cognitive levels and being incapable of simulating the dynamic feature variations found in actual inspection scenarios. Shallow Feature Analysis: Traditional machine learning algorithms (such as SVM or random forest) are used for feature analysis, which can only extract static features (e.g., gaze point coordinates, total xation duration). These approaches are unable to capture the spatiotemporal dynamic correlation features during cognitive processes, resulting in insufcient depth in the analysis of cognitive mechanisms. Fixed Evaluation Models: Unied evaluation models (e.g., scoring based solely on accuracy) are adopted, without consideration of individual differences in cognitive ability. There is also no dynamic updating of model parameters, making it impossible to generate personalized evaluation reports. The relevance between evaluation results and actual inspection tasks is therefore weak. In order to solve the above problems, some technologies have attempted to introduce multimodal data, but issues such as poor data synchronization and complex calibration procedures (requiring multiple independent calibrations) remain. Moreover, the linkage mechanism between stimulus generation and cognitive evaluation is lacking, which reduces the systems adaptability and practicality. Therefore, there is an urgent need for an intelligent detection system and method capable of precise acquisition of multimodal data, intelligent dynamic generation of stimuli, deep feature analysis, and personalized cognitive evaluation. Summary of the Invention The objective ofthe present invention is to provide an , in order to solve the problems existing in the prior art. To achieve the above objective, the present invention provides the following technical solution: an , comprising the following modules: a multimodal perception module, congured to synchronously collect multisource data related to the visual cognition of the subject, wherein the sampling rate of eye movement data is not less than 600 Hz, with measurement accuracy controlled within the range of 0.15° to 03°; the sampling rate of pupil diameter data is not less than 300 Hz, with accuracy s 0.01 mm; the sampling rate of electroencephalogram (EEG) signals is not less than 1000 Hz, with a common mode rejection ratio 2 110 dB; meanwhile, adaptive perception and calibration of ambient light ranging from 200 to 1200 lumens and ambient temperature ranging from 15°C to 40°C are supported; a dynamic stimulus generation unit, congured to generate four types of visual inspection target stimuli including structured features, fuzzy features, interference features, and composite features, with stimulus resolution not less than 2K, refresh rate 2 240 Hz, presentation duration of stimuli continuously adjustable within the range of 20 ms to 500 ms, and support for dynamic stimulus difculty grading based on the subject's real-time cognitive state, wherein difculty levels include elementary, intermediate, and advanced; an intelligent feature analysis unit, congured to perform the following operations: a) perform spatiotemporal joint segmentation of the raw data collected by the multimodal perception module in 50 ms time windows, and extract motion features of eye movement trajectories (e.g., saccade velocity, xation stability), dynamic variation features ofthe pupil (e.g., dilation rate, contraction amplitude), and cognition-related features of EEG signals (e.g., P300 amplitude, alpha wave power); b) adopt a Transformer-LSTM hybrid deep learning model to construct a three-stage dynamic cognitive analysis model comprising feature perception, cognitive association, and decision output, and construct respective feature weight allocation sub-models adapted to different types of stimuli; a dynamic cognitive evaluation unit, congured to provide a real-time cognitive state visualization interface, support personalized parameter conguration based on the subjects visual cognitive ability level ranging from level C1 to C5, and generate cross-stimulus-type and cross-time-phase comparative maps of feature cognitive efciency, comprising threedimensional indicators of cognitive accuracy, response delay, and cognitive load index. Preferably, in the threestage dynamic cognitive analysis model: the feature perception stage ranges from 0 to 1.2 s, employing an attention mechanism- enhanced CNN sub-model, wherein the convolution kernel sizes include 3x3, 5x5, and 7><7, and the number of output feature map channels is 256; the cognitive association stage ranges from 1.2 to 3.5 s, employing a Transformer encoder sub-model, with the number of multi-head attention heads being 8 and the hidden layer dimension of the FeedForward network being 1024; the decision output stage ranges from 3.5 to 5.0 s, employing an LSTM decoder sub-model with 512 hidden units and a dropout rate of 0.2. Preferably, the four types of visual inspection target stimuli in the dynamic stimulus generation unit are presented based on a dual-dimensional design of target feature density and interference item complexity to balance presentation position and sequence; each type of stimulus is presented 3 to 5 times in a single experimental process, dynamically adjusted according to the cognitive ability level, and the presentation interval between adjacent stimuli adopts an adaptive adjustment mode within a range of 30 ms to 200 ms, based on the cognitive accuracy ofthe preceding stimulus. Preferably, the dynamic cognitive evaluation unit comprises a visual cognitive ability grading module, congured to: classify subjects into ve cognitive levels of C1 (basic level), C2 (intermediate level), C3 (procient level), C4 (professional level), and C5 (expert level) based on the three-dimensional indicators of feature recognition accuracy (weight 40%), cognitive response delay (weight 30%), and cognitive load index (weight 30%); assign differentiated model parameters to different cognitive levels, wherein for subjects at levels C4 / C5, the attention weights of the Transformer encoder in the cognitive association stage (1.2 to 3.5 s) are increased by 0.4 to 0.6 times, and for subjects at levels C1 and C2, the number of iterations of feature extraction in the CNN sub-model during the feature perception stage (0 to 1.2 s) is increased by 2 to 3 times. Preferably, the intelligent feature analysis unit is further congured to: in the decision output stage from 3.5 to 5.0 s, set dynamic thresholds for the cognitive accuracy of composite feature stimuli, such that when the cognitive accuracy of subjects at levels C1 and C2 is detected to be continuously below 60% for three times, the dynamic stimulus generation unit is triggered to switch the subsequent stimuli to low-complexity structured feature stimuli and extend the presentation duration by 1.5 times; when the cognitive accuracy of subjects at levels C4 and C5 is detected to be continuously above 95% for three times, the dynamic stimulus generation unit is triggered to switch the subsequent stimuli to ultrahigh complexity composite features and randomly interfering stimuli, and shorten the presentation duration to 70% of the original duration. Preferably, a multisource data synchronization calibration device is arranged between the multimodal perception module and the dynamic stimulus generation unit, congured to: perform a 12-point calibration procedure (eye movement / pupil) and EEG baseline calibration before the experiment, wherein the 12point calibration procedure is used to calibrate the eye movement data and pupil diameter data in the multimodal perception module; the calibration errors are controlled to be S 01° for eye movement, 5 0.005 mm for pupil, and EEG signal noise s 2 uv; dynamically adjust the timestamp alignment algorithm based on the actual refresh rate (240 Hz) ofthe dynamic stimulus generation unit and the sampling rate of multimodal data. A detection method for the , it comprises the following steps: A. activate the multimodal perception module, complete initial calibration of eye movement, pupil, EEG, and environmental parameters, and synchronously collect multisource data related to the subjects visual cognition; B. input the collected multisource raw data into the intelligent feature analysis unit, perform spatiotemporal joint segmentation in 50 ms time windows, and extract multidimensional cognitive features; C. the dynamic stimulus generation unit generates corresponding complexity levels of the four types of visual inspection target stimuli based on the subjects initial cognitive level, defaulted to level C3, and presents them sequentially; D. the intelligent feature analysis unit employs the TransformerLSTM hybrid deep learning model to perform three-stage modeling of feature perception, cognitive association, and decision output, analyzing the subject's feature cognition processes and patterns in response to different types of stimuli; E. the dynamic cognitive evaluation unit calculates the subjects cognitive accuracy, response delay, and cognitive load index based on the model analysis results, generates a visual cognitive ability level evaluation report and a feature cognitive efciency comparative map, and triggers the dynamic stimulus generation unit to adjust subsequent stimulus parameters (complexity, presentation duration, interval). Preferably, step B further comprises: performing fusion preprocessing of the segmented multisource data, wherein Kalman ltering is used to remove noise from the eye movement data, wavelet transform (using the db4 wavelet basis, with three decomposition levels) is used to eliminate blink interference in the pupil data, independent component analysis (ICA) is used to separate electrooculographic artifacts from the EEG signals, followed by feature normalization (Z- score standardization) to map the multidimensional features to the [0,1] interval, and nally constructing a three-dimensional feature matrix oftime, features, and modalities for model input. Compared with the prior art, the benecial effects of the invention are: The invention integrates eye movement, pupil, and EEG tri-modal data, with sampling rate and accuracy both superior to the prior art. Combined with environmentadaptive calibration, data integrity is enhanced, enabling comprehensive reection of the neural and behavioral processes involved in visual feature cognition. Based on a linkage mechanism between cognitive states and stimulus parameters, four types of stimuli close to actual inspection scenarios are generated, supporting realtime adjustment of complexity, duration, and interval, and adapting to subjects of different cognitive levels. Compared with xed stimulus schemes, the correlation between detection results and actual job competencies is improved. A TransformerLSTM hybrid network is used to capture the spatiotemporal dynamic features ofthe cognitive process in stages. Compared with traditional machine learning algorithms, feature analysis accuracy is improved, and the complete cognitive chain of perception, association, and decision-making is revealed. Based on threedimensional indicators, ve cognitive levels are dened, and evaluation reports includingjob-matching recommendations are generated. The system supports online model ne- tuning and a closed loop of evaluation and regulation. Compared with unied evaluation models, the level of personalization is enhanced, providing precise data support for enterprise talent selection and employee training. Brief Description of the Drawings FIG. 1 is a system schematic diagram of the present invention; FIG. 2 is a schematic diagram ofthe multimodal perception module of the present invention; FIG. 3 is a schematic diagram of the dynamic stimulus generation unit of the present invention; FIG. 4 is a schematic diagram of the dynamic cognitive evaluation unit of the present invention; FIG. 5 is a owchart of the method of the present invention. Detailed Description The following will clearly and fully describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments. It is evident that the described embodiments are only part of the embodiments of the present invention and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without any creative effort shall fall within the scope of protection of the present invention. Please refer to FIGs. 1 to 5. The present invention provides an , comprising a multimodal perception module, a dynamic stimulus generation unit, an intelligent feature analysis unit, and a dynamic cognitive evaluation unit. The multimodal perception module integrates multiple types of sensors to achieve synchronized acquisition and adaptive calibration of multisource data related to visual cognition, it is specically congured as follows: Eye movement tracking submodule: adopts a desktop high-precision eye tracker with a sampling rate of 600 Hz and measurement accuracy of 0.15°-0.3°, horizontal tracking range of i40°, vertical tracking range of 30°, and supports dual-mode positioning via corneal reection and pupil center; Pupil dynamics monitoring submodule: cooperates with the eye movement tracking submodule, with a separately optimized pupil diameter measurement algorithm, sampling rate of 300 Hz, accuracy of 0.01 mm, capable of capturing minute pupil changes within 10 ms (e.g., pupil dilation during increased cognitive load); EEG signal acquisition submodule: employs a dry electrode EEG cap (8 channels covering frontal, central, and parietal regions), sampling rate of 1000 Hz, common mode rejection ratio of 110 dB, input impedance 2 100 IVIQ, and supports real-time removal of ocular and muscular artifacts; Environmental perception submodule: integrates light sensors and temperature sensors to collect environmental parameters in real time and feed them back to the calibration unit, achieving environment-adaptive data acquisition. The dynamic stimulus generation unit is based on cognitive science and the demands of visual inspection scenarios to generate four types of visual stimuli with clear detection signicance, and supports dynamic adjustment. It is specically congured as follows: Stimulus types: Structured feature stimuli: images with regular features such as standard holes or grooves in industrial products (defect rate 0%); Fuzzy feature stimuli: images of product defects with blurred edges caused by wear or uneven lighting (moderate defect identication difculty); Interference feature stimuli: images containing a large number of interference items similar to target defects (e.g., mixed scratches and stains on metal surfaces), with an interference ratio of 30%-60%; Composite feature stimuli: complex product images containing multiple types of defects (e.g., cracks, deformation, color differences), with 2-4 types of defects; Presentation parameters: resolution 2K (256OX1440), refresh rate 240 Hz (ensuring no ghosting during dynamic stimulus presentation), presentation duration continuously adjustable from 20 ms to 500 ms, and presentation positions randomly distributed based on a nine-grid screen layout (to avoid location preference bias); Dynamic adjustment mechanism: an internal cognitive state and stimulus parameter mapping model automatically adjusts the complexity of subsequent stimuli (e.g., reducing interference ratio by 20% when accuracy < 70%), presentation duration (extended by 20%), and interval (extended by 15%) based on the cognitive accuracy and response delay of the previous stimulus. The intelligent feature analysis unit is responsible for deep processing of multisource data and analysis of the feature cognition process, it is specically congured as follows: Data preprocessing: performs spatiotemporal joint segmentation of multisource data in 50 ms time windows (ensuring time dimension granularity), applies Kalman ltering (eye movement data), wavelet transform (pupil data), and ICA (EEG data) to remove noise and artifacts, and constructs a three-dimensional feature matrix via feature normalization (time dimension: 50 ms / window; feature dimension: 12 dimensions for eye movement + 8 dimensions for pupil + 10 dimensions for EEG; modality dimension: 3 types); Deep learning model: adopts a TransformerLSTM hybrid network, divided into three sub model stages: Feature perception stage (01.2 s): attentionenhanced CNN (with 3x3, 5><5 and 7><7 convolution kernels) to extract local features, outputting a 256-channel feature map, focusing on the subjects initial perception of stimulus features; Cognitive association stage (1.2-3.5 s): 8-head Transformer encoder constructs spatiotemporal associations between features (e.g., correlation between eye movement trajectories and P300 amplitude), Feed-Forward network dimension 1024, achieving cross-modal feature fusion; Decision output stage (3.5-5.0 s): 512-unit LSTM decoder outputs cognitive decision results (e.g. crack defect present" or no defect), dropout rate 0.2 to prevent overtting; Dynamic feature library: stores the subjects feature cognition data in real time, including feature patterns from different stimulus types and cognitive stages, supports online model ne tuning, with model parameters updated every 100 trials. The dynamic cognitive evaluation unit enables personalized assessment and visualization of cognitive ability, it is specically congured as follows: Cognitive level classication: based on threedimensional indicators (weights: recognition accuracy 40%, response delay 30%, cognitive load index 30%), applies K-means clustering algorithm to classify subjects into ve levels C1-C5, each corresponding to specic visual inspection job suitability standards; Evaluation indicator calculation: Cognitive accuracy: the proportion oftrials with correct feature identication; Response delay: the time from stimulus presentation to decisionmaking, with outliers removed and average taken; Cognitive load index: calculated based on pupil dilation rate (weight 50%) and EEG alpha wave power (weight 50%) (index range: 0-1, where 0 indicates low load and 1 indicates high load); Visualization and feedback: provides a real-time cognitive state interface (displaying current stimulus, multimodal data waveforms, cognitive load curve), generates cross-stimulus-type and cross-time-phase comparison maps (e.g., variation curves of cognitive accuracy under composite feature stimuli across different subject levels), and supports exporting evaluation reports (including analysis of cognitive strengths / weaknesses and job suitability recommendations). A multisource data synchronization calibration device is arranged between the multimodal perception module and the dynamic stimulus generation unit to address issues of time synchronization and accuracy calibration of multisource data: Initial calibration: Prior to the experiment, a 12point eye movement / pupil calibration (covering the entire screen range) and a 10-second EEG baseline calibration (recording resting- state EEG) are performed. Calibration error is controlled within: eye movement S 0.1", pupil S 0.005 mm, EEG noise S 2 uV; Dynamic synchronization: Based on the 240 Hz refresh rate of the dynamic stimulus generation unit, timestamps of multimodal data are adjusted in real time (eye movement at 600 Hz, one data point every 1.67 ms; EEG at 1000 Hz, one data point every 1 ms), achieving precise synchronization between data and stimulus presentation through hardware trigger signals. The intelligent feature cognition detection method for visual inspection of the present invention is implemented based on the aforementioned system and comprises the following steps: Step A: System Initialization and Initial Calibration The multimodal perception module is activated, the subject wears the EEG cap, and the eye tracker is adjusted (chin rest xed, distance from eyes to screen center is 65 cm). 12-point eye movement / pupil calibration and 10-second EEG baseline calibration are performed, the environmental perception submodule collects current light intensity (e.g., 500 lumens) and temperature (25°C) parameters and automatically adjusts the gain of the eye tracker and the amplier parameters of the EEG signals to ensure data acquisition accuracy. The subject completes basic information entry and a cognitive ability pre-test (10 trials with structured feature stimuli). Based on pretest results, the system preliminarily classies the initial cognitive level (default level C3; if pre-test accuracy > 90%, initial level is set to C4; if < 60%, set to C2). Step B: Multisource Data Acquisition and Preprocessing The multimodal perception module synchronously collects data from the subject at 600 Hz (eye movement), 300 Hz (pupil), and 1000 Hz (EEG). The environmental perception submodule updates environmental parameters every 100 ms. The intelligent feature analysis unit performs segmentation of the raw data in 50 ms time windows and sequentially executes noise removal (Kalman ltering, wavelet transform, ICA), feature extraction (e.g., saccade velocity from eye movement, pupil dilation rate, P300 amplitude from EEG), and feature normalization (Zscore standardization) to generate a threedimensional feature matrix of time, feature, and modality, which is transmitted in real time to the model computing unit. Step C: Dynamic Stimulus Generation and Presentation The dynamic stimulus generation unit, based on the subjects initial cognitive level, randomly selects four types of stimuli from the stimulus library (5-8 trials per type, totaling 20-32 trials), initially presented in the sequence structured % fuzzy % interference % composite", and subsequently adjusted dynamically according to cognitive state. During stimulus presentation, a hardware trigger signal from the multimodal perception module is triggered simultaneously to ensure time synchronization between data acquisition and stimulus presentation. A gaze point prompt is displayed in real time to guide the subject's attention. Step D: Feature Cognition Process Modeling and Analysis The intelligent feature analysis unit inputs the three-dimensional feature matrix into the TransformerLSTIVI hybrid model, performing phased computations: Feature perception stage (0-1.2 5): The attention-enhanced CNN extracts local features and outputs a feature perception probability map, reecting the subject's attention distribution over different areas of the stimulus; Cognitive association stage (1.2-3.5 5): The Transformer encoder constructs spatiotemporal associations between eye movement, pupil, and EEG features, outputting a feature association strength matrix, such as the correlation coefcient between P300 amplitude and gaze duration on target features; Decision output stage (3.55.0 s): The LSTM decoder outputs the cognitive decision result (e.g., two types of defects detected: crack + color difference") and decision condence (0-1); The dynamic feature library automatically ne-tunes model parameters every 100 trials of parsed results, such as adjusting the attention head weights of the Transformer to improve subsequent analysis accuracy. Step E: Cognitive Ability Evaluation and Dynamic Regulation The dynamic cognitive evaluation unit calculates the subject's threedimensional evaluation indicators based on model outputs, including cognitive accuracy, response delay, and cognitive load index. It applies the K-means clustering algorithm to update the cognitive level. For example, if a subject initially at level C3 achieves a 92% accuracy rate for composite feature stimuli, 0.8 s response delay, and cognitive load index of 0.3, the level is upgraded to C4. A personalized evaluation report is generated, including: A cognitive ability radar chart, showing the subjects performance on the three indicators under four types of stimuli; A cognitive process timeseries analysis chart, showing changes in feature cognition efciency across stages within 0-5 5; Job matching recommendations, such as: Suitable for precision inspection of electronic components; fuzzy feature cognition training is recommended"; The report triggers the dynamic stimulus generation unit to adjust parameters ofsubsequent stimuli. For example, for a C4-level subject, the number of defect types in subsequent composite stimuli is increased to four, and the presentation duration is shortened to 80% of the original, thus achieving a closed loop of evaluation and regulation. Embodiment: An intelligent feature cognition detection method for visual inspection, comprising the following steps: Subject Recruitment and Screening Recruitment targets: 100 employees from a quality inspection post in an automotive parts manufacturing enterprise, aged 2045, corrected vision 2 1.0, no history of neurological disorders; Grouping: Based on existing work performance (quality inspection accuracy), subjects are divided into three groups: excellent group (accuracy 2 95%, 30 persons), qualied group (accuracy 85%-95%, 40 persons), and improvement group (accuracy < 85%, 30 persons); Experiment preparation: Subjects adapt to the laboratory environment for 10 minutes in advance (lighting 500 lumens, temperature 25°C), wear the EEG cap and complete initial calibration, and read the experiment instruction manual (clearly dening the task: identify the types of defects in the stimulus). Experimental Procedure Pre-test: Ten structured feature stimuli are presented, and the system preliminarily classies the initial cognitive level (subjects in the excellent group are mostly classied as level C4 / C5, while those in the improvement group are mostly classied as level C1 / C2); Formal experiment: Twenty trials are presented (ve for each of the four types of stimuli), and the dynamic stimulus generation unit adjusts parameters based on the cognitive state of the previous trial (e.g., for the excellent group, in the third trial with composite feature stimuli, the number of defect types is increased to four); Data recording: The multimodal perception module synchronously collects data, the intelligent feature analysis unit processes it in real time, and the dynamic cognitive evaluation unit generates a realtime cognitive state curve; Post-experiment: A personalized evaluation report is generated, and the subject completes an experimental experience questionnaire (evaluating system usability). Experimental Result Analysis Data integrity: During the experiment, the loss rate of multimodal data is reduced, and time synchronization error is decreased, demonstrating the stability of system data acquisition; Analysis accuracy: The intelligent feature analysis unit improves the cognitive decision accuracy for all four types of stimuli; Level matching: The match rate between the cognitive levels assigned by the system and the subjects' actual work performance is improved, verifying the effectiveness of the evaluation results. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that modications may still be made to the technical solutions described in the foregoing embodiments, or equivalent replacements may be made to certain technical features. Any modications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall fall within the scope of protection of the present invention.
Claims
1. An intelligent feature recognition detection device for visual inspection, characterized because it includes the following modules: a multimodal perception module, configured to simultaneously process multisource data with to collect information about the subject's visual cognition, where the sampling rate of eye movement data is not less than 600 Hz, with a measurement accuracy within the range of 0.15° to 0.3°; the sampling rate of pupil diameter data is not less than 300 Hz, with an accuracy of S 0.01 mm; the sampling rate of EEG signals (electroencephalogram) is not less than 1000 Hz, with a common mode rejection ratio of 2 110 dB; adaptive perception and calibration are also available of ambient light (between 200 and 1200 lumens) and ambient temperature (between 15°C and 40°C) supported; a dynamic stimulus generation unit, configured to generate four types of visual to generate inspection target stimuli, including structured features, vague features, interference features and composite features, with a stimulus resolution of not less than 2K, a refresh rate of 2240Hz, a presentation duration that is continuous is adjustable within the range of 20ms to 500ms, and supports dynamic difficulty classification of stimuli based on the real-time cognitive state of the subject, with difficulty levels including elementary, intermediate and advanced; an intelligent feature analysis unit, configured to perform the following operations to perform: a) spatiotemporal joint segmentation of the raw data collected by the perform multimodal perception module in 50 ms time windows, and motion features of eye movement trajectories, dynamic variation characteristics of the pupil and cognition-related extract features from EEG signals; b) apply a hybrid deep learning model based on Transformer-LSTIVI to a to build a three-part dynamic cognitive analysis module, consisting of feature perception, cognitive association and decision output, and corresponding construct feature weight distribution submodels adapted to different types stimuli; a dynamic cognitive evaluation unit, configured to provide a real-time visualization to provide a cognitive state interface, to support personalized parameter configuration based on the subject's visual cognitive ability level (level C1 to C5), and comparison maps of trait cognition efficiency between to generate stimulus types and time phases consisting of three-dimensional indicators of cognitive accuracy, reaction latency and cognitive load index.
2. The intelligent feature recognition detection device for visual inspection according to conclusion 1, characterized in that in the three-part dynamic cognitive analysis module: the feature perception phase runs from 0 to 1.2 s and uses a CNN submodel enhanced with an attention mechanism, where the dimensions of the convolution kernels are 3><3, 5><5 and 7><7 are, and the number of channels of the output feature card is 256; the cognitive association phase runs from 1.2 to 3.5 s and uses a Transformer encoder submodel, where the number of multi-head attention heads is 8 and the dimension of the hidden layer in the feed-forward network is 1024; the decision execution phase runs from 3.5 to 5.0 5 and uses an LSTIVl decoder- submodel with 512 hidden units and a dropout rate of 0.
2.
3. The intelligent feature recognition detection device for visual inspection according to conclusion 1, characterized in that the four types of visual inspection target stimuli in the dynamic stimulus generation unit are presented based on a two-dimensional design of target feature density and interference complexity to determine presentation position and order in to balance; each stimulus type is presented 3 to 5 times during a single experiment, dynamically adapted to the cognitive skill level, and the presentation interval between adjacent stimuli is adaptively adjusted within a range from 30 ms to 200 ms, based on the cognitive accuracy of the preceding stimulus.
4. The intelligent feature recognition detection device for visual inspection according to conclusion 1, characterized in that the dynamic cognitive evaluation unit is a module for visual cognitive ability classification includes, configured to: to divide subjects into five cognitive levels C1, C2, C3, C4 and C5 based on the three-dimensional indicators of feature recognition accuracy, cognitive reaction delay and cognitive load index; to assign differentiated model parameters to different cognitive levels, where for subjects at level C4 / C5 the attention weights of the Transformer- encoder in the cognitive association stage are increased by 0.4 to 0.6 times, and for subjects at level C1 and C2 the number of iterations for feature extraction in the CNN- submodel during the feature perception phase is increased by 2 to 3 times.
5. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that the intelligent feature analysis unit is further configured to: set dynamic thresholds in the decision execution phase from 3.5 to 5.0 s for the cognitive accuracy of compound feature stimuli, so that when found that the cognitive accuracy of subjects at levels C1 and C2 is below 60% for three consecutive times, the dynamic stimulus generation- unit is activated to change the following stimuli to structured features with low complexity and extend the presentation time by 1.5 times; when found that the cognitive accuracy of subjects at levels C4 and C5 is above 95% for three consecutive times, the dynamic stimulus generation unit is activated to change the following stimuli to compound features with ultra-high complexity and random interference stimuli, and the presentation duration to shorten to 70% of the original duration.
6. The intelligent feature recognition detection device for visual inspection according to claim 1, characterized in that a multi-source data synchronization calibration device is placed between the multimodal perception module and the dynamic stimulus generation unit, configured to perform a 12-point calibration procedure and an EEG baseline calibration before the experiment, using the 12-point calibration procedure to eye movement data and pupil diameter data in the multimodal perception module calibration; calibration errors are controlled to S 0.1° for eye movement, S 0.005 mm for pupil, and EEG signal noise S 2 uv; dynamically adjust the time-stamp alignment algorithm based on the actual refresh rate of the dynamic stimulus generation unit and the sampling rate of multimodal data.
7. A detection method for the intelligent feature recognition detection device for visual inspection according to any one of claims 1 to 6, characterised in that it comprises the following steps include: A. activate the multimodal perception module, complete the initial eye movement calibration, pupil, EEG and environmental parameters, and collect multisource data simultaneously with relating to the subject's visual cognition; B. Input the collected multisource raw data into the intelligent feature analysis- unit, perform spatiotemporal joint segmentation in 50 ms time windows, and extract multidimensional cognitive features; C. the dynamic stimulus generation unit generates corresponding complexity levels of the four types of visual inspection target stimuli based on the initial subject's level of cognition, set to level C3 by default, and presents this consecutive; D. The intelligent feature analysis unit applies the hybrid deep learning model Transformer-LSTM to perform three-phase modeling of feature perception, cognitive association and decision output, involving the subject's feature recognition process and pattern for different types of stimuli are analyzed; E. the dynamic cognitive evaluation unit calculates cognitive accuracy, reaction delay and cognitive load index of the subject based on the analysis results of the model, generates an evaluation report of the visual cognitive ability level and a comparison map of trait cognition efficiency, and activates the dynamic stimulus generation unit to adjust the parameters of subsequent stimuli.
8. The detection method for the intelligent feature recognition detection device for visual inspection according to claim 7, characterised in that step B further comprises: performing fusion preprocessing of the segmented multisource data, using Kalman filtering is used to remove noise from eye movement data, wavelet transform is used to eliminate flicker interference in the pupil data, independent component analysis (ICA) is used to separate electrooculographic artifacts from EEG signals followed by feature normalization to capture the multidimensional features in the interval [0,1], and finally a three-dimensional feature matrix of time, features and modalities are constructed for input into the model. ultimodal perception module Intelligent feature analysis unit Dynamic cognitive evaluation unit ynamic stimulus generation unit Embedded linkage or multisource data synchronization calibration device FIG. 1 EEG acquisition submodule e movement tracking sub-module Pupil monitoring sub-module Data synchronization interface Adaptive calibration unit mental perception submodule FIG. 2 Stimulus generation submodule Presentation parameter control Dynamic adjustment submodule Position balancing submodule FIG. 3 Three-dimensional indicator tive level classification module Visualization and feedback submodule calculation submodule Parameter adaptation submodule FIG. 4 Step A Activate the multimodal perception module Step B Input the collected multisource raw data into the intelligent feature analysis unit Step C The dynamic stimulus generation unit generates stimuli based on the subject’s initial cognitive level Step D intelligent feature analysis unit employs a Transformer-LSTM hybrid deep learning model Step E The dynamic cognitive evaluation unit processes the model analysis results FIG. 5