Attention monitoring method and related equipment

By using target detection and EEG decoding technology, a list of key targets is generated and attention status is monitored in real time, which solves the problems of insufficient real-time performance and accuracy of attention monitoring in existing technologies and enables reliable attention monitoring in multi-screen layouts and complex environments.

CN120899252APending Publication Date: 2025-11-07启元实验室
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511076834.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing attention monitoring technologies cannot monitor attention status in real time and objectively in multi-screen layouts and complex environments. They are easily affected by emotions and cognitive biases, and their operation latency and accuracy are insufficient in long-distance monitoring, making it difficult to meet complex monitoring needs.

Method used

The target detection convolutional network is used to extract target information, and a list of key targets is generated by combining a multi-factor risk assessment model. The target of attention is determined by visual stimulus signals and EEG decoding technology. The SSVEP response frequency is collected and decoded in real time to generate attention deviation signals to activate the guidance strategy.

Benefits of technology

It enables real-time and reliable attention monitoring in multi-screen layouts and complex environments, reducing the risk of missing key targets and improving the efficiency and reliability of human-machine collaborative monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120899252A_ABST
    Figure CN120899252A_ABST
Patent Text Reader

Abstract

The invention discloses an attention monitoring method and related equipment. Monitoring is realized through coordination of five steps of target detection, risk assessment, visual stimulation, electroencephalogram decoding and attention guidance. The method specifically solves the defects in the prior art: electroencephalogram signal decoding is adopted to replace a subjective report, so that emotion and cognition deviation is avoided, instantaneous attention monitoring is realized, and the hysteresis and non-objectivity of a subjective method are overcome; visual stimulation and electroencephalogram response association is used for replacing operation behavior inference, a mouse or touch equipment is not needed, remote monitoring is adapted, and insufficient real-time performance caused by operation delay is eliminated; the focus target is decoded through the SSVEP response frequency, the limitation of the eye tracker on the monitoring angle and the fixed posture is eliminated, and the method is suitable for the complex scene with multiple screens, multiple angles and multiple postures. Meanwhile, key target list dynamic generation and guide strategy automatic activation form closed-loop management, the key target omission risk is significantly reduced, and the man-machine cooperative monitoring efficiency and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent monitoring, and more particularly to an attention monitoring method and related device. BACKGROUND

[0002] In real life, many scenarios have an urgent need for attention monitoring. The monitoring center of a traffic hub needs to track abnormal behavior in densely populated areas in real time to avoid large-scale accidents; in the security system, the monitoring personnel at the border checkpoint need to continuously identify suspicious personnel and prohibited items, and attention slackness may lead to security loopholes; the production monitoring room of a large factory needs to monitor the equipment operation status of multiple production lines simultaneously, and any abnormality in a key parameter that is not captured in time may cause a production accident; when handling emergencies, the command personnel of the emergency command platform need to quickly lock the rescue target from a large number of monitoring screens, and improper attention allocation may delay the rescue opportunity. These scenarios all rely on the continuous attention of the monitoring personnel, and human attention is easily affected by factors such as fatigue and multitasking interference, therefore, building a precise and real-time attention monitoring technology has become a necessary prerequisite for ensuring the stable operation of complex systems.

[0003] Existing attention monitoring technologies have obvious limitations: subjective judgment-based methods assess attention states through questionnaire surveys, self-reports, and other means, which are simple to operate, but the results are affected by emotional and cognitive biases, lack objectivity, and need to interrupt the task for evaluation, which cannot reflect the instantaneous fluctuations of attention, and are difficult to adapt to scenarios that require continuous monitoring; operation behavior-based methods rely on mouse clicks, touch operations, and other means to infer attention focus, which have insufficient operation delay and precision in long-distance monitoring or multi-screen systems, for example, it is difficult for the operators of a large monitoring center to quickly locate the target in multiple screens through a mouse, resulting in lag in attention monitoring; eye movement feature-based methods capture fixation points through an eye tracker, but the effective monitoring angle is limited and cannot cover all areas of a multi-screen layout, and when the operator changes the observation posture, such as turning around to view the side screen, the device cannot accurately capture the eye movement data, making it difficult to adapt to the needs of complex monitoring environments. These technical defects have a high risk of missing key targets, and a more reliable solution is urgently needed. SUMMARY

[0004] The present application provides an attention monitoring method and related device, which solves the defects of existing subjective judgment, operation behavior, and eye movement feature monitoring methods through target detection, risk assessment, visual stimulation, electroencephalogram decoding, and attention guidance, reduces the risk of missing key targets, and improves the efficiency and reliability of human-machine collaborative monitoring.

[0005] An attention monitoring method, comprising:

[0006] The target detection convolutional network detects targets in images or videos in a monitoring interface, outputs target information containing target categories, location information and confidence, and marks the targets in the interface;

[0007] The target type features, spatial location information and behavior trajectory data are extracted from the target information, a risk score is calculated by a multi-factor risk assessment model and sorted, and a key target list is generated;

[0008] A visual stimulation signal with independent frequency and preset phase interval is assigned to each target in the key target list, and a periodic brightness change is superimposed on the corresponding target position in the monitoring interface;

[0009] The time domain features, frequency domain features and spatial domain features of the brain electrical signals of the monitored person are collected and extracted in real time, the SSVEP response frequency is decoded by a canonical correlation analysis and task-related component analysis fusion algorithm, and the attention target number of the monitored person is determined;

[0010] The key target list is compared with the attention target number, and if there is a key target that is not paid attention to, an attention deviation signal is generated to activate the attention prompt guidance strategy.

[0011] Optionally, the target detection convolutional network adopts a single parameter training strategy and an incremental training mechanism;

[0012] The single parameter training strategy comprises:

[0013] Data enhancement operations are performed on the training set images, and image mixing technology is used to proportionally fuse multiple images of different scenes;

[0014] The input image size is dynamically adjusted for multi-scale training, the learning rate is dynamically adjusted by a cosine annealing strategy, and the number of frozen layers and the number of preheating rounds of learning rate in the initial training stage are controlled by a hot start mechanism;

[0015] The incremental training mechanism comprises:

[0016] After preliminary training, the training loss, accuracy and recall rate under different parameter combinations are compared, and the optimal parameter combination is selected;

[0017] Based on the optimal parameter combination, additional training samples are gradually added and the network structure is fine-tuned for incremental training until the model reaches the preset recognition accuracy threshold.

[0018] Optionally, the target type features, spatial location information and behavior trajectory data extracted from the target information comprise:

[0019] Based on the target category of the target information, a key target category with potential risk properties is determined through risk attribute analysis processing, and corresponding basic risk weight parameters are preset for different types to generate target type features;

[0020] Based on the location information of the target information, the distance of the target from the boundary of the predefined safety area and the relative position in the monitoring interface are determined through spatial distance and area attribution analysis processing to generate spatial location information;

[0021] Based on the location information of the target information in consecutive frames, the appearance frequency and duration of the target are determined through time series analysis processing, and abnormal behavior patterns are identified by combining motion speed, direction change, and dwell time features to generate behavior trajectory data.

[0022] Optionally, a visual stimulation signal with an independent frequency and a preset phase interval is assigned to each target in the key target list, and a periodic brightness change is superimposed on the corresponding target position in the monitoring interface, including:

[0023] A visual stimulation signal with an independent frequency in the range of 8Hz-15Hz is assigned to each target in the key target list, and a preset phase interval is maintained between each frequency to ensure that the signals do not interfere with each other;

[0024] Adjust the contrast of the visual stimulation according to the size of the target, including using low contrast for targets larger than a threshold and high contrast for targets smaller than a threshold;

[0025] With the adjusted visual stimulation contrast, a periodic brightness change is superimposed on the corresponding target position in the monitoring interface, and the visual stimulation is kept synchronized with the brain electrical signal acquisition frame level through hardware triggering or software timestamping.

[0026] Optionally, the time domain features, frequency domain features, and spatial domain features of the brain electrical signals of the monitored person are collected and extracted in real time, and the SSVEP response frequency is decoded through a typical correlation analysis and task-related component analysis fusion algorithm to determine the attention target number of the monitored person, including:

[0027] The brain electrical signals of the monitored person are collected in real time, and the time domain features, frequency domain features, and spatial domain features of the brain electrical signals are extracted using peak valley analysis, power spectral density analysis, and electroencephalogram distribution topographic map analysis, respectively;

[0028] The typical correlation coefficient between the multi-channel brain electrical signals and the reference stimulation signal is calculated through typical correlation analysis to generate the first SSVEP frequency identification result, and the second SSVEP frequency identification result is determined by extracting the stable neural response features across trials through task-related component analysis;

[0029] The first SSVEP frequency recognition result and the second SSVEP frequency recognition result are combined with confidence to make a decision-level fusion judgment, and an SSVEP response frequency is decoded to determine the attention target number of the monitored person.

[0030] Optionally, after collecting the brain electrical signals of the monitored person in real time, the method further comprises:

[0031] Based on the environmental power frequency interference, a corresponding notch frequency is set, and the power frequency interference periodic noise in the brain electrical signals is filtered out through notch filtering;

[0032] Based on the frequency characteristics of the SSVEP signal, a band-pass filtering passband range is set, and the effective brain electrical signals related to the SSVEP in the brain electrical signal frequency band are purified through band-pass filtering;

[0033] The effective brain electrical signals after the notch filtering and the band-pass filtering are decomposed into a plurality of independent components, and the artifact components are identified and separated to obtain the brain electrical signals after removing the artifacts.

[0034] An attention monitoring device comprises:

[0035] A target recognition module is configured to perform target detection on images or videos in a monitoring interface through a target detection convolutional network, output target information containing target categories, location information and confidence, and mark targets in the interface;

[0036] A risk level evaluation module is configured to extract target type features, spatial location information and behavior trajectory data from the target information, calculate risk scores and sort them through a multi-factor risk evaluation model, and generate a key target list;

[0037] A target visual coding module is configured to assign independent frequencies to each target in the key target list and superimpose periodic brightness changes on the corresponding target positions in the monitoring interface;

[0038] An attention analysis module is configured to collect and extract time domain features, frequency domain features and spatial domain features of brain electrical signals of a monitored person in real time, decode SSVEP response frequencies through a typical correlation analysis and task-related component analysis fusion algorithm, and determine the attention target number of the monitored person;

[0039] An attention transfer module is configured to compare the key target list with the attention target number, and if there is a key target that is not being paid attention to, generate an attention deviation signal to activate an attention prompt guidance strategy.

[0040] An attention monitoring device comprises a memory and a processor;

[0041] The memory is configured to store a program;

[0042] The processor is configured to execute the program to implement each step of the attention monitoring method according to any one of the preceding embodiments.

[0043] A readable storage medium, having a computer program stored thereon, the computer program, when executed by a processor, implements each step of the attention monitoring method according to any one of the preceding embodiments.

[0044] A computer program product, comprising a computer program, the computer program, when executed by a processor, implements each step of the attention monitoring method according to any one of the preceding embodiments.

[0045] As can be seen from the technical solutions described above, the attention monitoring method and related device provided by the embodiments of the present application are implemented through five steps of cooperation: first, target information is extracted and labeled by a target detection convolutional network; second, a key target list is generated by a multi-factor risk assessment model after feature extraction; third, a visual stimulation signal with an independent frequency and a preset phase interval is assigned to the list target, and a periodic brightness change is superimposed; fourth, electroencephalogram signals are collected in real time, time-frequency domain and spatial domain features are extracted, and the SSVEP response frequency is decoded by a typical correlation analysis and task-related component analysis fusion algorithm to determine the number of the target of interest; and finally, the list and the target of interest are compared, and if the target of interest is not focused on, an attention deviation signal is generated and a guidance strategy is activated.

[0046] The scheme solves the defects of the prior art in a targeted manner: first, electroencephalogram signal decoding is used to replace subjective reporting, avoiding the influence of emotional and cognitive bias, and real-time collection and analysis enable instantaneous monitoring of the attention state, overcoming the lag and subjectivity of subjective methods; second, visual stimulation signals and electroencephalogram responses are used to replace operation behavior inference, eliminating the need for mouse or touch devices, adapting to long-distance monitoring scenarios, and eliminating the real-time deficiency problem caused by operation delay; third, the target of interest is decoded by the SSVEP response frequency, breaking free from the restrictions of the eye tracker on monitoring angles and fixed postures, and regardless of multi-screen layout, complex observation angles, or changes in posture, the target of interest can be stably identified, breaking through the application limitations of eye tracking technology. At the same time, the dynamic generation of the key target list and the automatic activation of the attention guidance strategy enable closed-loop management from target priority sorting to attention correction, significantly reducing the risk of missing key targets and improving the efficiency and reliability of human-machine collaborative monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creating any inventive labor.

[0048] Figure 1 A flow chart of a method for monitoring attention disclosed in an embodiment of the present application;

[0049] Figure 2 A schematic diagram of a method for monitoring attention disclosed in an embodiment of the present application;

[0050] Figure 3 A schematic diagram of a device for monitoring attention disclosed in an embodiment of the present application;

[0051] Figure 4 A hardware structure block diagram of a device for monitoring attention disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0053] The present application can be used in many general or special-purpose computing device environments or configurations. For example: personal computers, server computers, handheld or laptop devices, tablet devices, multiprocessor devices, distributed computing environments that include any of the above devices or devices, and the like.

[0054] Next, the technical solutions of the present application are introduced. The present application proposes the following technical solutions, which are described below.

[0055] Figure 1 A flow chart of a method for monitoring attention disclosed in an embodiment of the present application;

[0056] Figure 2 A schematic diagram of a method for monitoring attention disclosed in an embodiment of the present application;

[0057] As shown in Figure 1 and Figure 2 , the method can include:

[0058] Step S1, performing target detection on images or videos in a monitoring interface through a target detection convolutional network, outputting target information containing target categories, position information and confidence, and marking the target in the interface.

[0059] Specifically, first, the image or video data input by the monitoring interface is preprocessed, including image enhancement (such as adjusting contrast, brightness, and chroma), denoising, and data enhancement processing such as image rotation, mirror flipping, and random cropping, to simulate shooting scenes under different times, lighting conditions, camera angles, or heights, expand the coverage of the training data, and improve the model's recognition ability for complex situations such as boundary blur, occlusion, and reflections.

[0060] Subsequently, the target detection convolutional network improves the model performance through the system's fine-tuning training strategy for specific monitoring task datasets. During training, a multi-scale training strategy is used to dynamically adjust the size of the input image, allowing the model to adapt to the recognition needs of targets of different sizes. Different warm-up rounds are set to explore parameters such as the number of frozen layers at the initial stage of the model, the learning rate warm-up method, and other parameter configurations to improve training stability. A cosine annealing strategy is used to dynamically adjust the learning rate, avoiding the model from falling into local optima and speeding up convergence. Based on single-parameter training, an incremental training mechanism is further designed, which first compares key performance indicators such as training loss, precision, and recall rate under different parameter combinations to select the optimal parameter combination. Then, additional training samples are gradually added, the training rounds are adjusted, or the network structure is fine-tuned to continuously optimize the model weights until the detection performance meets the expected standards.

[0061] The trained model recognizes and locates various targets in the monitoring screen, outputting target information including target categories (such as suspicious individuals, fast-moving targets, and targets entering the warning zone), location information (represented by bounding box coordinates), and corresponding confidence levels. Combining the non-maximum suppression algorithm to eliminate overlapping or repeated detection boxes, the final recognition results are displayed in the form of bounding boxes and labels on the monitoring interface, providing stable and efficient target data support for subsequent modules.

[0062] Step S2: Extract target type features, spatial location information, and behavior trajectory data from the target information, calculate risk scores and sort them through a multi-factor risk assessment model, and generate a key target list.

[0063] Specifically, three types of core features are extracted from the target information output by the target detection:

[0064] Target type features: identify the risk attribute category of the target, such as fast-moving targets, abnormal behavior individuals (such as unconscious floating personnel in water monitoring), and targets entering the warning zone, etc. Different types are assigned a basic risk weight (for example, the basic weight of unconscious floating personnel is higher than that of conscious swimmers);

[0065] Spatial location information: calculate the relative distance between the target and the core monitoring area (such as the shore or safety boundary). The farther the distance from the shore or the farther the distance outside the safety range, the higher the risk, and the higher the spatial risk weight assigned.

[0066] Behavioral trajectory data: Analyze the motion characteristics of the target through time series, including duration, motion speed, direction change, and appearance frequency (such as long-term static, sudden acceleration, or frequent appearance in the monitoring field), extract abnormal behavior patterns, and convert them into trajectory risk weights.

[0067] Subsequently, the multi-factor risk assessment model fuses the above features by weighted summation to generate a comprehensive risk score, where the weight parameters can be flexibly configured according to the application scenario. The comprehensive risk score is calculated by fusing the above features through the multi-factor risk assessment model, including: the model uses weighted summation to accumulate the target type weight, spatial position weight, and behavior trajectory weight according to the preset proportion (which can be flexibly adjusted according to the monitoring scene) to generate the risk score of each target. For example, in the water monitoring scene, the score of a certain person falling into the water may consist of "unconscious floating" type weight (0.3), "within 5 meters from the shore" spatial weight (0.2), and "continuous static for 10 seconds" trajectory weight (0.2), with a comprehensive score of 0.7.

[0068] Finally, according to the risk score, all targets are sorted in descending order, and targets with scores higher than a preset threshold (such as 0.6) are selected to generate a key target list. The list contains target number, category, and risk level, which are the core objects for subsequent visual coding and attention monitoring, ensuring that high-risk targets are given priority.

[0069] Further, the target type feature, spatial position information, and behavior trajectory data are extracted from the target information, including:

[0070] ① Based on the target category of the target information, determine the key target category with potential risk properties through risk attribute analysis processing, and preset corresponding basic risk weight parameters for different types to generate target type features.

[0071] Based on the target category output by the target detection (such as "falling person", "intruder", "fast-moving vehicle", etc.), the key target categories with potential risk properties are determined through risk attribute analysis. Specifically, risk attribute judgment rules are preset, for example, "unconscious floating falling person", "stranger entering the warning area", "moving object with speed exceeding the threshold" are defined as high-risk potential categories, and "normal walking staff", "compliantly driving vehicle" are defined as low-risk categories. For categories with different risk attributes, corresponding basic risk weight parameters (weight value range 0-1.0) are pre-configured, for example, the basic weight of "unconscious floating person" is set to 0.35, the basic weight of "intruder" is set to 0.3, the basic weight of "fast-moving vehicle" is set to 0.25, and the basic weight of "normal activity person" is set to 0.1. In this way, the target category is converted into a quantifiable target type feature, providing a basic parameter for subsequent risk assessment.

[0072] ②Based on the position information of the target information, the distance of the target from the boundary of the predefined safety area and the relative position in the monitoring interface are determined through spatial distance and area ownership analysis processing to generate spatial position information.

[0073] Based on the position information of the target detection output (represented by the bounding box coordinates), spatial position information is generated through spatial distance and area ownership analysis. First, the boundary of the safety area (such as "shore safety line" in water monitoring, "core area boundary" in forbidden area, "fence boundary" in park monitoring, etc.) is predefined in the monitoring system, and its coordinate parameters are recorded into the system. Then, the shortest Euclidean distance between the center point of the target bounding box and the safety area boundary is calculated, and the closer the distance, the higher the risk (for example, when the distance is less than 5 meters, it is marked as "near boundary", when the distance is 5-10 meters, it is marked as "medium distance", and when the distance is more than 10 meters, it is marked as "far distance"). At the same time, the relative position of the target in the monitoring interface is determined (such as "upper left corner area", "center area", "lower right corner area"), and combined with the key attention area of the monitoring scene (such as the center area is usually the core monitoring area), an additional spatial weight adjustment factor is given. Finally, the distance parameter and the relative position parameter are integrated to form spatial position information that can participate in risk calculation.

[0074] ③Based on the position information of the target information in consecutive frames, the occurrence frequency and duration of the target are determined through time series analysis processing, and the abnormal behavior pattern is recognized by combining the motion speed, direction change and dwell time features to generate behavior trajectory data.

[0075] Based on the position information (bounding box coordinates arranged in time sequence) of consecutive frames in the target information, the behavior trajectory data is generated through time sequence analysis. First, the target tracking algorithm (such as Kalman filter) is used to associate the same target in consecutive frames to ensure the continuity of the trajectory. Then, the time features are extracted: the frequency of the target appearing in the monitoring field (the proportion of the number of frames appearing in a unit of time) and the duration (the cumulative length of time from the first appearance to the current frame) are calculated. The target with a duration exceeding a preset threshold (such as 30 seconds) or a frequency higher than a threshold (such as appearing ≥5 times per minute) is marked as a "high-frequency / continuous appearance target". Next, the motion features are analyzed: the motion speed is calculated by the position difference between adjacent frames (speed = distance / frame interval time), and the target with a speed exceeding the scene threshold (such as personnel speed >3m / s in the park) is marked as "fast moving"; the direction angle change of three consecutive frames is calculated, and the target with a direction change exceeding 90° is marked as "direction mutation"; the target's stay time in the same area (the number of consecutive frames with a position change amplitude less than a preset pixel threshold) is calculated, and the target with a stay time exceeding the threshold (such as 15 seconds) is marked as "abnormal stay". Finally, the time features and motion features are integrated to identify abnormal behavior patterns such as continuous high-frequency appearance, fast movement, direction mutation, etc., and complete behavior trajectory data is formed.

[0076] Step S3, assign a visual stimulus signal with independent frequency and preset phase interval to each target in the key target list, superimpose periodic brightness changes at the corresponding target position in the monitoring interface.

[0077] Specifically, a visual stimulus signal is assigned to each target in the key target list, and a preset phase interval is set according to the number of targets to ensure that the multi-target signals do not interfere with each other in the time domain through phase separation. The technical personnel can flexibly adjust the phase allocation rules according to the actual number of targets. For target size differences, an adaptive contrast strategy is used, with a preset size threshold. For targets larger than the threshold (such as close-range clear targets), low contrast is used, and for targets smaller than the threshold (such as distant small targets), high contrast is used. The technical personnel can adjust the threshold parameters according to the resolution and target characteristics of the monitoring scene. Finally, the brightness changes are superimposed and synchronized, and periodic brightness changes are superimposed at the assigned frequency at the target corresponding position, with the brightness following a sinusoidal modulation to ensure that the original image is not obscured. Frame-level synchronization between the EEG signal and the stimulus signal is achieved through hardware triggering or software timestamp alignment. The stability of the visual encoding signal is optimized through optimization of the display parameter settings and program design. If a frequency deviation exceeding ±0.5Hz is detected, an alarm is triggered to ensure the effectiveness of the stimulus signal.

[0078] Further, a visual stimulus signal with independent frequency and preset phase interval is assigned to each target in the key target list, and periodic brightness changes are superimposed at the corresponding target position in the monitoring interface, including:

[0079] assigning a visual stimulation signal of an independent frequency in the range of 8-15 Hz to each target in the list of critical targets, and maintaining a preset phase interval between each frequency to ensure that the signals do not interfere with each other;

[0080] adjusting the visual stimulation contrast according to the target size, including using low contrast for targets larger than a threshold value, and using high contrast for targets smaller than the threshold value;

[0081] superimposing periodic brightness changes on the corresponding target positions in the monitoring interface with the adjusted visual stimulation contrast, and maintaining the visual stimulation in synchronization with the brain electrical signal acquisition frame level through hardware triggering or software timestamping.

[0082] First, an independent visual stimulation frequency and an initial phase are assigned to each target in the list of critical targets. The selected frequency needs to be in the range of 8-15 Hz to avoid signal confusion; at the same time, the initial phase corresponding to each frequency needs to maintain a preset interval (such as distributing 360° phase equally according to the number of targets, and the phase interval is 120° for 3 targets), to ensure the spatiotemporal separation of signals when multiple targets are encoded in parallel.

[0083] Second, a differentiated flicker contrast strategy is adopted for target size differences: for larger targets in the monitoring interface (such as targets occupying more than 5% of the screen area), a low-contrast flicker method (luminance fluctuation amplitude is 10-15% of the reference luminance) is used to avoid excessive visual stimulation interfering with normal observation by the monitor; for smaller targets (such as targets occupying less than 1% of the screen area), a high-contrast flicker method (luminance fluctuation amplitude is 20-30% of the reference luminance) is used to enhance visual saliency and ensure that the monitor's gaze can stably induce SSVEP (Steady-State Visual Evoked Potential) response.

[0084] Subsequently, periodic brightness changes are superimposed on the corresponding target positions in the monitoring interface to achieve visual stimulation. Specifically, the brightness of the pixels in the target area is modulated to achieve this, i.e., the brightness value of the pixels within the target bounding box is periodically changed according to the assigned frequency (such as changing according to the sine law: brightness value = reference brightness x contrast x [1+sin(2πft+θ)] / 2, where θ is the phase, f is the assigned frequency, and t is the time), and the brightness modulation process does not block the original image content of the target, and only transmits the encoding signal through brightness changes.

[0085] In addition, a timing synchronization mechanism needs to be integrated to ensure that the visual stimulation signal and the brain electrical signal acquisition system are frame-level synchronized through hardware triggering or software timestamp alignment, to avoid decoding errors caused by time offset; at the same time, the system monitors the stability of the visual encoding signal in real time, and if the deviation exceeds the threshold, the system will issue a system warning, prompting to check the display refresh rate or environmental light interference, to ensure the effectiveness of the stimulation signal.

[0086] Step S4, real-time acquisition and extraction of the time domain features, frequency domain features and spatial domain features of the brain electrical signals of the monitored person, decoding of the SSVEP response frequency through a typical correlation analysis and a task-related component analysis fusion algorithm, and determination of the attention target number of the monitored person.

[0087] Specifically, real-time acquisition and extraction of the time domain features, frequency domain features and spatial domain features of the brain electrical signals of the monitored person, decoding of the SSVEP response frequency through a typical correlation analysis and a task-related component analysis fusion algorithm, and determination of the attention target number of the monitored person. Specifically, the implementation process of this step is as follows:

[0088] A multi-channel portable electroencephalogram device is used, and the signals of the occipital electrodes (such as O1, O2, Oz, POz, etc.) are collected, with a sampling rate not less than 250 Hz to completely retain the time-frequency features. During collection, the electrodes are ensured to be in good contact with the scalp, and motion artifacts and environmental noise are reduced, and at the same time, the starting time of the key target visual stimulation in the monitoring interface is recorded synchronously through hardware triggering or software timestamp, to realize frame-level synchronization of the brain electrical signals and the stimulation signals.

[0089] The collected raw brain electrical signals are preprocessed, including: 50 Hz notch filtering is used to remove power frequency interference; band-pass filtering is used to retain the SSVEP signals and their low-order harmonics, generally 0.5-30 Hz; independent component analysis (ICA) is used to identify and strip out eye movement, electromyogram and other artifact components, to ensure signal purity.

[0090] Three types of features are extracted from the preprocessed signals: time domain features (such as peak-trough amplitude, peak-trough period, signal mean and variance), frequency domain features (such as power spectral density, focusing on analyzing the energy distribution in the 8-15 Hz frequency band), and spatial domain features (such as signal correlation of different electrode channels, and brain electrical distribution topography).

[0091] The correlation coefficient between the multi-channel electroencephalogram signal and the target reference stimulus frequency signal is calculated by canonical correlation analysis (CCA) to preliminarily locate the possible attention frequency; meanwhile, the neural response characteristics stable across trials are extracted by task-related component analysis (TRCA) to strengthen the significance of the target frequency. The results of the two algorithms are fused at the decision level, and the target number corresponding to the frequency with the highest correlation coefficient is selected as the real-time output of the monitoring person's attention target number. The technical personnel can adjust the fusion weight or introduce a trial accumulation mechanism (such as confirming the target number only when the results of three consecutive time windows are consistent) according to the actual scene to improve the decoding stability.

[0092] Step S5, compare the key target list with the attention target number, if there is a key target not being paid attention to, generate an attention deviation signal to activate the attention prompting guidance strategy.

[0093] Specifically, first, the system compares the target number in the key target list with the attention target number output in step S4 in real time according to a fixed time window, and checks whether all key targets in the list are included in the attention range. The judgment standard is: if a key target does not appear in the attention target number in a continuous preset number of time windows (such as 3 time windows, i.e. 1.5 seconds), it is determined that the target is not being paid attention to.

[0094] For the key target not being paid attention to, the system generates an attention deviation signal, which contains the target number, risk level and the duration of not being paid attention to, etc. as a trigger instruction to activate the attention prompting guidance strategy.

[0095] The activated guidance strategy adopts a multi-modal prompting method decoupled from the SSVEP frequency to avoid interfering with the decoding of the electroencephalogram signal: visually, a static highlight outline (such as a low-saturation yellow border, 3 pixels wide) is displayed on the edge of the target area or a short color emphasis (such as the border changing from the default color to orange) is performed without obscuring the original image content of the target; aurally, a short prompt sound (such as a 1kHz tone, 0.3 seconds long) or a voice reminder (such as "please confirm the status of target X") is played. At the same time, the system supports dynamically adjusting the prompt intensity according to the target risk level, for example, high-risk targets use a joint prompting method of static highlight and prompt sound, and low-risk targets only use static highlight prompting.

[0096] In addition, the strategy can also integrate an online feedback mechanism to monitor the change of attention after guidance in real time: if the attention target number output in step S4 contains the key target within the set time limit, it is determined that the guidance is successful, and the system automatically terminates the prompting process; if it is still not paid attention to after the time limit, the system will enhance the prompt level (such as increasing the brightness of the highlight border, prolonging the duration of the prompt sound), and record the relevant parameters for subsequent strategy optimization until the target is paid attention to or a manual intervention prompt is triggered.

[0097] It can be seen from the technical solution that the attention monitoring method and the related device provided by the embodiment of the application are realized through five steps of cooperation: first, target information is extracted by a target detection convolutional network and is labeled; then, features are extracted to generate a key target list by a multi-factor risk assessment model; subsequently, a visual stimulation signal with an independent frequency and a preset phase interval is allocated to the list target, and a periodic brightness change is superimposed; then, electroencephalogram signals are collected in real time, time-frequency domain and spatial domain features are extracted, a SSVEP response frequency is decoded by a typical correlation analysis and a task-related component analysis fusion algorithm, and a target number of attention is determined; finally, the list and the target of attention are compared, and if the target is not the target of attention, an attention deviation signal is generated and a guiding strategy is activated.

[0098] The scheme solves the defects of the prior art in a targeted manner: first, electroencephalogram signal decoding is used to replace subjective reporting, avoiding the influence of emotional and cognitive bias, and real-time collection and analysis realize instantaneous monitoring of the attention state, overcoming the lag and non-objectivity of subjective methods; second, visual stimulation signals and electroencephalogram responses are used to replace operation behavior inference, without relying on a mouse or a touch device, suitable for long-distance monitoring scenes, and eliminating the real-time deficiency problem caused by operation delay; third, the target of attention is decoded by the SSVEP response frequency, breaking free from the restrictions of the eye tracker on the monitoring angle and the fixed posture, and regardless of multi-screen layout, complex observation angles or posture changes, the target of attention can be stably identified, breaking through the application limitations of eye tracking technology. At the same time, the dynamic generation of the key target list and the automatic activation of the attention guiding strategy realize closed-loop management from target priority sorting to attention correction, significantly reduce the risk of missing key targets, and improve the efficiency and reliability of human-machine collaborative monitoring.

[0099] In some embodiments of the application, the training process of the target detection convolutional network is introduced, and the target detection convolutional network adopts a single parameter training strategy and an incremental training mechanism.

[0100] The single parameter training strategy comprises:

[0101] Data enhancement operations are performed on the training set images, and image mixing technology is used to fuse multiple different scene images in proportion;

[0102] The input image size is dynamically adjusted for multi-scale training, the learning rate is dynamically adjusted by using a cosine annealing strategy, and the number of frozen layers and the number of preheating rounds of the learning rate in the initial training stage are controlled by a hot start mechanism.

[0103] Specifically, multi-dimensional data augmentation operations are performed on the training set images, including adjusting visual attributes such as contrast, brightness, and chrominance, and simultaneously performing appropriate geometric transformations such as rotation, flipping, and cropping, to simulate image features under different environmental conditions. On this basis, image mixing technology is used to fuse multiple images of different scenes in a certain proportion, thereby improving the model's adaptability to complex backgrounds and multi-target coexistence scenes.

[0104] During model training, the size of the input image is dynamically adjusted for multi-scale training, enabling the model to adapt to the recognition needs of targets of different sizes. The learning rate adjustment uses a cosine annealing strategy, that is, as the training process progresses, the learning rate gradually changes according to the law of the cosine function, to facilitate better convergence of the model. At the same time, the initial training stage is controlled through a warm start mechanism, including setting a certain number of frozen layers (i.e., fixing part of the network layer parameters to not participate in initial training), and preheating the learning rate. By gradually adjusting the initial stage parameters of the learning rate, the stability in the early stage of training is ensured.

[0105] The incremental training mechanism includes:

[0106] After preliminary training, the training loss, accuracy, and recall rate under different parameter combinations are compared to select the optimal parameter combination;

[0107] Based on the optimal parameter combination, additional training samples are gradually added and the network structure is fine-tuned for incremental training until the model reaches the preset recognition accuracy threshold.

[0108] Specifically, after completing the preliminary training, the model performance under different parameter combinations is evaluated, mainly comparing training loss, recognition accuracy, and recall rate, etc., to select the optimal parameter combination.

[0109] Based on the optimal parameter combination selected, incremental training is performed. In this process, additional training samples are gradually added to the training set, which typically contain target features in edge scenarios, such as occlusion, small targets, and targets in complex backgrounds. At the same time, the network structure is fine-tuned, such as adjusting the parameters of some network layers or unfreezing some previously frozen network layers, to further optimize the model performance. Incremental training continues until the model's recognition accuracy reaches the preset threshold requirement.

[0110] In some embodiments of the present application, the process of step S4, real-time acquisition and extraction of the time domain features, frequency domain features, and spatial domain features of the brain electrical signals of the monitored person, decoding the SSVEP response frequency through canonical correlation analysis and task-related component analysis fusion algorithm, and determining the attention target number of the monitored person, can specifically include:

[0111] Step S41, real-time acquisition of the brain electrical signals of the monitored person, and extraction of time domain features, frequency domain features and space domain features of the brain electrical signals by using peak-valley analysis, power spectral density analysis and electroencephalogram distribution topography analysis respectively;

[0112] Step S42, generation of a first SSVEP frequency recognition result by calculating the canonical correlation coefficient between the multi-channel brain electrical signals and the reference stimulation signals through canonical correlation analysis, and determination of a second SSVEP frequency recognition result by extracting the neural response features stable across trials through task-related component analysis;

[0113] Step S43, decision-level fusion determination of the first SSVEP frequency recognition result and the second SSVEP frequency recognition result combined with confidence, decoding of the SSVEP response frequency, and determination of the attention target number of the monitored person.

[0114] Specifically, the multi-channel electroencephalogram device is used to collect the electroencephalogram signals of the relevant channels in the occipital region in real time at a proper sampling rate, and the visual stimulation frequency and phase information corresponding to the key target are recorded synchronously. The raw electroencephalogram signals are preprocessed first, including removal of power frequency interference, preservation of signal components in a specific frequency band, and stripping of eye movement, electromyogram and other artifact interference. Subsequently, the time domain features are extracted through peak-valley analysis to capture the amplitude change characteristics of the signals in the time dimension; the frequency domain features are extracted through power spectral density analysis to identify the energy distribution of the frequency band related to the visual stimulation; and the space domain features are extracted through electroencephalogram distribution topography analysis to reflect the spatial distribution pattern of the signals on the scalp surface.

[0115] In the canonical correlation analysis, the reference signals corresponding to the visual stimulation frequencies of each key target (including the fundamental frequency and related harmonic components) are constructed first, and then the canonical correlation coefficients of the preprocessed multi-channel electroencephalogram signals and each reference signal are calculated, and the first SSVEP frequency recognition result is generated according to the correlation coefficient size. In the task-related component analysis, based on the electroencephalogram trial data under the visual stimulation of the same target for multiple times, the neural response components stable across trials are extracted, the neural response template corresponding to the target is constructed, and then the matching degree between the real-time signals and the template is determined to determine the corresponding second SSVEP frequency recognition result.

[0116] The first SSVEP frequency recognition result obtained by canonical correlation analysis (CCA) and the second SSVEP frequency recognition result obtained by task-related component analysis (TRCA) are acquired, and confidence indicators corresponding to the two results are extracted simultaneously, wherein the confidence of CCA is determined based on the size of the canonical correlation coefficient of the multi-channel electroencephalogram signal and the reference stimulation signal, and the confidence of TRCA is calculated according to the matching degree of the real-time signal and the stable neural response template across trials. Subsequently, the system performs decision-level fusion based on the two confidences: if the two recognition results are consistent, the frequency is directly taken as the SSVEP response frequency; if there is a difference between the results, the result with higher confidence is preferentially selected according to the confidence weight, so as to effectively suppress noise interference and individual differences. After obtaining the final SSVEP response frequency, the target number currently focused on by the monitored person is decoded through a preset frequency-to-target number mapping relationship.

[0117] In addition, to adapt to different scene requirements, the system supports flexible mode selection: in simple tasks with higher real-time requirements, a lightweight and fast recognition mode based on CCA can be enabled; in complex task scenarios such as multiple targets and high noise, an individualized model initialization mechanism based on TRCA template learning can be enabled to optimize the recognition performance by pre-constructing a neural response template for the monitored person, ensuring the accuracy and adaptability of the fusion decision. To ensure the stability of the results, the recognition results can be confirmed through consistency verification of multiple time windows or majority voting mechanism, and finally the reliable focus target number is output.

[0118] On the basis of the above, to improve the quality of electroencephalogram signals to ensure the accuracy of subsequent feature extraction and frequency decoding, an electroencephalogram signal preprocessing process can also be added, specifically including the following steps after real-time acquisition of the electroencephalogram signals of the monitored person:

[0119] ①Based on the setting of the corresponding notch frequency of the environmental power frequency interference, the periodic noise of the power frequency interference in the electroencephalogram signals is filtered out through notch filtering. Specifically, according to the power frequency standard of the area where the monitoring scene is located (such as the common power frequency), a matching notch frequency is set in the electroencephalogram signal processing link, and a notch filtering circuit or algorithm is used to specifically attenuate the periodic interference signal at this frequency and nearby, reducing the influence of power grid fluctuations on the electroencephalogram signals.

[0120] ②Based on the frequency characteristics of the SSVEP signal, the passband range of the band-pass filter is set, and the effective electroencephalogram signal related to the SSVEP in the frequency band of the electroencephalogram signals is obtained through band-pass filtering. Since the SSVEP signal has a specific frequency range (usually related to the visual stimulation frequency), the upper and lower limits of the passband of the band-pass filter are set accordingly, so that only the components in this frequency band are retained after filtering, and high-frequency noise (such as electronic device interference) and low-frequency drift (such as baseline fluctuation) are suppressed, highlighting the neural electrical activity signal related to visual attention.

[0121] ③Decompose the effective brain electrical signal after the notch filtering and the band-pass filtering into a plurality of independent components, identify and separate a pseudo-trace component, and obtain a brain electrical signal after removing the pseudo-trace. Specifically, the preprocessed brain electrical signal can be decomposed into a plurality of statistically independent components by using an independent component analysis algorithm, the waveform characteristics, the frequency spectrum distribution, and the spatial distribution mode of each component are analyzed, the pseudo-trace components generated by non-neural activities such as eye movement, blinking, and electromyographic contraction are identified and stripped from the signal, and finally a purer brain electrical signal is obtained, which provides a reliable data basis for subsequent time domain, frequency domain, and spatial domain feature extraction.

[0122] Next, a kind of attention monitoring device provided by the embodiment of the application is described, and the following description of a kind of attention monitoring device can be mutually corresponding with the description of a kind of attention monitoring method in the foregoing.

[0123] Referring to Figure 3 , Figure 3 a schematic diagram of a kind of attention monitoring device disclosed by the embodiment of the application.

[0124] As Figure 3 indicated, the kind of attention monitoring device can include:

[0125] Target identification module 110, for target detection of image or video in monitoring interface by target detection convolution network, output target information containing target category, location information and confidence, and mark target in interface;

[0126] Risk level evaluation module 120, for extracting target type feature, spatial position information and behavior trajectory data from the target information, calculating risk score and sorting by multi-factor risk evaluation model, and generating key target list;

[0127] Target visual coding module 130, for assigning independent frequency and maintaining preset phase interval visual stimulation signal for each target in the key target list, and superimposing periodic brightness change in the corresponding target position in the monitoring interface;

[0128] Attention analysis module 140, for real-time acquisition and extraction of time domain feature, frequency domain feature and spatial domain feature of brain electrical signal of monitoring person, decoding SSVEP response frequency by canonical correlation analysis and task-related component analysis fusion algorithm, and determining the attention target number of the monitoring person;

[0129] Attention transfer module 150, for comparing the key target list with the attention target number, if there is a key target not being paid attention to, generating attention deviation signal to activate attention prompt guidance strategy.

[0130] From the above technical solution, the attention monitoring method and related device provided by the embodiment of the application is realized through five-step cooperation: first, the target information is extracted by the target detection convolutional network and labeled; then the feature is extracted to generate a key target list by a multi-factor risk assessment model; subsequently, the list target is assigned an independent frequency and a visual stimulation signal with a preset phase interval is superimposed with periodic brightness changes; then the electroencephalogram signal is collected in real time, the time-frequency domain and spatial domain features are extracted, the SSVEP response frequency is decoded by the canonical correlation analysis and task-related component analysis fusion algorithm, and the attention target number is determined; finally, the list and the attention target are compared, and if the attention target is not focused, an attention deviation signal is generated and the guidance strategy is activated.

[0131] The scheme solves the defects of the prior art in a targeted manner: first, the electroencephalogram signal decoding is used to replace the subjective report, avoiding the influence of emotional and cognitive bias, and the real-time collection and analysis realize the instantaneous monitoring of the attention state, overcoming the lag and non-objectivity of the subjective method; second, the visual stimulation signal and the electroencephalogram response are associated to replace the operation behavior inference, without relying on the mouse or touch device, adapting to the remote monitoring scene, and eliminating the real-time deficiency problem caused by the operation delay; third, the attention target is decoded by the SSVEP response frequency, breaking free from the restrictions of the eye tracker on the monitoring angle and fixed posture, and regardless of the multi-screen layout, complex observation angle or posture change, the attention target can be stably identified, breaking through the application limitations of the eye tracking technology. At the same time, the dynamic generation of the key target list and the automatic activation of the attention guidance strategy realize the closed-loop management from the target priority sorting to the attention correction, significantly reducing the risk of missing key targets and improving the efficiency and reliability of human-machine collaborative monitoring.

[0132] Optionally, the target detection convolutional network adopts a single parameter training strategy and an incremental training mechanism;

[0133] The single parameter training strategy comprises:

[0134] Performing a data enhancement operation on the training set image, and fusing multiple different scene images by image mixing technology according to a proportion;

[0135] Dynamically adjusting the input image size for multi-scale training, dynamically adjusting the learning rate by using a cosine annealing strategy, and controlling the number of frozen layers and the number of learning rate warm-up rounds in the initial training stage by using a hot start mechanism;

[0136] The incremental training mechanism comprises:

[0137] After preliminary training, comparing the training loss, precision and recall rate under different parameter combinations to select the optimal parameter combination;

[0138] Based on the optimal parameter combination, incrementally training by gradually adding additional training samples and fine-tuning the network structure until the model reaches a preset recognition accuracy threshold.

[0139] Optionally, target type features, spatial position information and behavior trajectory data are extracted from the target information, including:

[0140] Based on the target category of the target information, the key target category with potential risk properties is determined through risk attribute analysis processing, and the corresponding basic risk weight parameters are preset for different types to generate target type features;

[0141] Based on the position information of the target information, the distance of the target from the boundary of the predefined safety area and the relative position in the monitoring interface are determined through spatial distance and regional attribution analysis processing to generate spatial position information;

[0142] Based on the position information of the target information in consecutive frames, the appearance frequency and duration of the target are determined through time series analysis processing, combined with motion speed, direction change and dwell time features to identify abnormal behavior patterns, and behavior trajectory data is generated.

[0143] Optionally, a visual stimulus signal with an independent frequency and a preset phase interval is assigned to each target in the key target list, and a periodic brightness change is superimposed on the corresponding target position in the monitoring interface, including:

[0144] A visual stimulus signal with an independent frequency in the range of 8Hz-15Hz is assigned to each target in the key target list, and a preset phase interval is maintained between each frequency to ensure that the signals do not interfere with each other;

[0145] Adjust the contrast of the visual stimulus according to the size of the target, including using low contrast for targets larger than a threshold and high contrast for targets smaller than a threshold;

[0146] With the adjusted visual stimulus contrast, a periodic brightness change is superimposed on the corresponding target position in the monitoring interface, and the visual stimulus is kept synchronized with the brain electrical signal acquisition frame level through hardware triggering or software timestamp.

[0147] Optionally, the time domain features, frequency domain features and spatial domain features of the brain electrical signals of the monitored person are collected and extracted in real time, and the SSVEP response frequency is decoded through a typical correlation analysis and task-related component analysis fusion algorithm to determine the attention target number of the monitored person, including:

[0148] Real-time acquisition of brain electrical signals of the monitored person, respectively using peak valley analysis, power spectral density analysis, and electroencephalogram distribution topographic map analysis to extract time domain features, frequency domain features and spatial domain features of the brain electrical signals;

[0149] The first SSVEP frequency recognition result is generated by calculating a canonical correlation coefficient between the multi-channel electroencephalogram signal and the reference stimulation signal through canonical correlation analysis, and the second SSVEP frequency recognition result is determined by extracting a cross-trial stable neural response feature through task-related component analysis;

[0150] The first SSVEP frequency recognition result and the second SSVEP frequency recognition result are combined with a confidence level for decision-level fusion determination, and the SSVEP response frequency is decoded to determine the attention target number of the monitored person.

[0151] Optionally, after the electroencephalogram signal of the monitored person is collected in real time, the method further includes:

[0152] A corresponding notch frequency is set based on the environmental power frequency interference, and the power frequency interference periodic noise in the electroencephalogram signal is filtered out through notch filtering;

[0153] A band-pass filtering passband range is set based on the frequency characteristics of the SSVEP signal, and the effective electroencephalogram signal related to the SSVEP in the frequency band of the electroencephalogram signal is obtained through band-pass filtering;

[0154] The effective electroencephalogram signal after the notch filtering and the band-pass filtering is decomposed into a plurality of independent components, and the artifact component is identified and separated to obtain the electroencephalogram signal after removing the artifact.

[0155] The attention monitoring device provided by the embodiments of the application can be applied to an attention monitoring device. Figure 4 A hardware structure block diagram of the attention monitoring device is shown, referring to Figure 4 The hardware structure of the attention monitoring device can include at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4.

[0156] In the embodiments of the application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4.

[0157] The processor 1 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the application, etc.

[0158] The memory 3 can include a high-speed RAM memory, and can also include a non-volatile memory, etc., such as at least one disk memory.

[0159] The memory stores a program, and the processor can call the program stored in the memory, and the program is used for:

[0160] detecting a target in an image or a video in a monitoring interface through a target detection convolutional network, outputting target information containing a target category, position information and a confidence, and marking the target in the interface;

[0161] extracting target type features, spatial position information and behavior trajectory data from the target information, calculating a risk score and sorting through a multi-factor risk assessment model, and generating a key target list;

[0162] allocating a visual stimulation signal with an independent frequency and maintaining a preset phase interval to each target in the key target list, and superimposing a periodic brightness change on the corresponding target position in the monitoring interface;

[0163] real-time acquisition and extraction of time domain features, frequency domain features and space domain features of the brain electrical signals of the monitored person, decoding of the SSVEP response frequency through a canonical correlation analysis and a task-related component analysis fusion algorithm, and determination of the attention target number of the monitored person;

[0164] comparing the key target list with the attention target number, and if there is a key target that is not paid attention to, generating an attention deviation signal to activate an attention prompt guidance strategy.

[0165] Optionally, the detailed functions and extended functions of the program can refer to the description above.

[0166] The embodiments of the application also provide a readable storage medium which can store a program suitable for processor execution, and the program is used for:

[0167] detecting a target in an image or a video in a monitoring interface through a target detection convolutional network, outputting target information containing a target category, position information and a confidence, and marking the target in the interface;

[0168] extracting target type features, spatial position information and behavior trajectory data from the target information, calculating a risk score and sorting through a multi-factor risk assessment model, and generating a key target list;

[0169] allocating a visual stimulation signal with an independent frequency and maintaining a preset phase interval to each target in the key target list, and superimposing a periodic brightness change on the corresponding target position in the monitoring interface;

[0170] real-time acquisition and extraction of time domain features, frequency domain features and space domain features of the brain electrical signals of the monitored person, decoding of the SSVEP response frequency through a canonical correlation analysis and a task-related component analysis fusion algorithm, and determination of the attention target number of the monitored person;

[0171] The key target list is compared with the attention target number, and if there is a key target that is not paid attention to, an attention deviation signal is generated to activate the attention prompt guide strategy.

[0172] Optionally, the refinement function and the extension function of the program can refer to the description above.

[0173] The embodiment of the application further provides a computer program product, comprising a computer program, which, when executed by a processor, performs the method.

[0174] The target detection convolutional network is used to detect targets in images or videos in a monitoring interface, output target information containing target categories, location information and confidence, and mark the targets in the interface;

[0175] Target type features, spatial location information and behavior trajectory data are extracted from the target information, a risk score is calculated and sorted by a multi-factor risk assessment model, and a key target list is generated;

[0176] A visual stimulation signal with an independent frequency and a preset phase interval is assigned to each target in the key target list, and a periodic brightness change is superimposed on the corresponding target position in the monitoring interface;

[0177] The time domain features, frequency domain features and space domain features of the brain electrical signals of the monitored person are collected and extracted in real time, the SSVEP response frequency is decoded by a canonical correlation analysis and task-related component analysis fusion algorithm, and the attention target number of the monitored person is determined;

[0178] The key target list is compared with the attention target number, and if there is a key target that is not paid attention to, an attention deviation signal is generated to activate the attention prompt guide strategy.

[0179] Optionally, the refinement function and the extension function of the program can refer to the description above.

[0180] Finally, it should be noted that in this document, relational terms such as first and second and the like can only be used to distinguish one entity or action from another entity or action, without necessarily requiring or implying that these entities or actions are in any way mutually exclusive or in any way in a required sequence. Moreover, the terms "comprise", "comprises", or "comprising" or any other variant thereof are intended to cover non-exclusive inclusions, such that processes, methods, articles or apparatuses that comprise a list of elements do not necessarily include only those elements in the list, but can include other elements not expressly listed or inherent to such processes, methods, articles or apparatuses. Without more limitations, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0181] The various embodiments described in this specification are presented by way of example, and each embodiment is not necessarily composed of all features described with respect to other embodiments. Each embodiment described in this specification can be combined with one or more other embodiments to produce new embodiments that are not explicitly described in this specification.

[0182] The above description of disclosed embodiments is intended to be illustrative, and not restrictive. Many embodiments will be apparent to those of skill in the art upon reviewing this disclosure. The scope of embodiments includes all of the familiar advantages and features that have been mentioned herein, along with one or more of the following. The language used in this specification should not be used to argue that any claim is dependent on another claim except for purposes of compliance with the Patent Statute and Rules, and the scope of embodiments will include all such alternatives, modifications and variations without departing from the spirit and scope of the disclosure.

Claims

1. An attention monitoring method characterized by, The method comprises the following steps: target detection is performed on images or videos in a monitoring interface by a target detection convolutional network, target information containing target categories, position information and confidence is output, and the target is marked in the interface; target type features, spatial position information and behavior trajectory data are extracted from the target information, a risk score is calculated by a multi-factor risk assessment model, and a key target list is generated; a visual stimulation signal with an independent frequency and a preset phase interval is assigned to each target in the key target list, and a periodic brightness change is superimposed on the corresponding target position in the monitoring interface; time domain features, frequency domain features and space domain features of brain electrical signals of a monitoring person are collected and extracted in real time, a SSVEP response frequency is decoded by a canonical correlation analysis and task-related component analysis fusion algorithm, and a target number of interest of the monitoring person is determined; the key target list and the target number of interest are compared, and if there is a key target that is not focused on, an attention deviation signal is generated to activate an attention prompt guidance strategy.

2. The method of claim 1, wherein, The target detection convolutional network adopts a single parameter training strategy and an incremental training mechanism; The single parameter training strategy comprises: data enhancement operations are performed on the training set images, and image mixing technology is used to fuse multiple images of different scenes in proportion; the input image size is dynamically adjusted for multi-scale training, the learning rate is dynamically adjusted by a cosine annealing strategy, and the number of frozen layers and the number of learning rate warm-up rounds in the initial training stage are controlled by a hot start mechanism; The incremental training mechanism comprises: After preliminary training, compare the training loss, precision and recall rate under different parameter combinations to select the optimal parameter combination; Based on the optimal parameter combination, additional training samples are gradually added and the network structure is fine-tuned for incremental training until the model reaches a preset recognition accuracy threshold.

3. The method of claim 1, wherein, Extracting target type features, spatial position information and behavior trajectory data from the target information comprises: Based on the target categories of the target information, determine the key target categories with potential risk properties by risk attribute analysis processing, and preset corresponding basic risk weight parameters for different types to generate target type features; Based on the position information of the target information, determine the distance of the target from the pre-defined safety area boundary and the relative position in the monitoring interface by spatial distance and region attribution analysis processing to generate spatial position information; Based on the position information of the continuous frames in the target information, determine the appearance frequency and duration of the target by time series analysis processing, and identify abnormal behavior patterns by combining motion speed, direction change and dwell time features to generate behavior trajectory data.

4. The method of claim 1, wherein, Assigning a visual stimulation signal with an independent frequency and a preset phase interval to each target in the key target list and superimposing a periodic brightness change on the corresponding target position in the monitoring interface comprises: Assign a visual stimulation signal with an independent frequency in the range of 8Hz-15Hz to each target in the key target list, and maintain a preset phase interval between each frequency to ensure that the signals do not interfere with each other; Adjust the contrast of the visual stimulation according to the target size, including using low contrast for targets larger than a threshold and high contrast for targets smaller than a threshold. The periodic brightness change is superimposed on the corresponding target position in the monitoring interface, and the visual stimulus is kept synchronous with the brain electrical signal acquisition frame level through hardware triggering or software time stamping.

5. The method of claim 1, wherein, The time domain feature, the frequency domain feature and the space domain feature of the brain electrical signal of the monitored person are collected and extracted in real time, the SSVEP response frequency is decoded through a typical correlation analysis and a task-related component analysis fusion algorithm, and the target number of attention of the monitored person is determined. The time domain feature, the frequency domain feature and the space domain feature of the brain electrical signal are extracted through peak-valley analysis, power spectral density analysis and electroencephalogram distribution topographic map analysis. The first SSVEP frequency recognition result is generated by calculating the typical correlation coefficient between the multi-channel brain electrical signal and the reference stimulus signal through the typical correlation analysis, and the second SSVEP frequency recognition result is determined by extracting the neural response feature stable across trials through the task-related component analysis. The first SSVEP frequency recognition result and the second SSVEP frequency recognition result are combined with the confidence for decision-level fusion judgment, the SSVEP response frequency is decoded, and the target number of attention of the monitored person is determined.

6. The method of claim 5, wherein, After the brain electrical signal of the monitored person is collected in real time, the following steps are further included: A corresponding notch frequency is set based on the environmental power frequency interference, and the power frequency interference periodic noise in the brain electrical signal is filtered out through notch filtering; A band-pass filtering passband range is set based on the frequency characteristics of the SSVEP signal, and the effective brain electrical signal related to the SSVEP in the frequency band of the brain electrical signal is obtained through band-pass filtering; The effective brain electrical signal after the notch filtering and the band-pass filtering is decomposed into a plurality of independent components, and the artifact component is identified and separated, so as to obtain the brain electrical signal after removing the artifact.

7. An attention monitoring apparatus characterized by comprising: It includes: A target recognition module is configured to perform target detection on images or videos in a monitoring interface through a target detection convolutional network, output target information containing target categories, location information and confidence, and mark targets in the interface; A risk level evaluation module is configured to extract target type features, spatial location information and behavior trajectory data from the target information, calculate risk scores and sort them through a multi-factor risk evaluation model, and generate a key target list; A target visual coding module is configured to assign independent frequency and preset phase interval visual stimulus signals to each target in the key target list, and superimpose periodic brightness changes on the corresponding target positions in the monitoring interface; An attention analysis module is configured to collect and extract the time domain feature, the frequency domain feature and the space domain feature of the brain electrical signal of the monitored person in real time, decode the SSVEP response frequency through a typical correlation analysis and a task-related component analysis fusion algorithm, and determine the target number of attention of the monitored person; An attention transfer module is configured to compare the key target list with the target number of attention, and if there is a key target that is not paid attention to, an attention deviation signal is generated to activate an attention prompt guide strategy.

8. An attention monitoring device, characterized in that It includes a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the attention monitoring method according to any one of claims 1-6.

9. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements each step of the attention monitoring method according to any one of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements each step of the attention monitoring method according to any one of claims 1-6.

Citation Information

Cited By

  • Multi-modal processing method and device based on electroencephalogram signals and storage medium

    CN121774535A

  • Multimodal processing methods, devices, and storage media based on electroencephalogram (EEG) signals

    CN121774535B