Vision-physiology multi-mode pilot workload detection method and device

Through the visual-physiological multimodal detection method, combined with the deep fusion of multiple data sources, the problems of inaccurate and insufficient real-time evaluation of pilot workloads in the existing technology are solved, and a comprehensive, accurate and real-time evaluation of pilot workloads is achieved.

CN120408263APending Publication Date: 2025-08-01TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510443883.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing pilot workload detection methods mainly rely on single-modal data fusion, and cannot comprehensively evaluate the pilot's multiple physiological states and behavioral responses in complex flight environments, and lack adaptability to individual differences, resulting in inaccurate assessment and insufficient real-time performance.

Method used

The visual-physiological multimodal detection method is adopted, combining facial image data, EEG signals, eye movement data and task performance data, and standardized processing of entropy weight method, feature extraction based on ResNet34 and LSTM, and deep fusion of Transformer, a multimodal detection model is constructed to monitor the changes in pilot workload in real time.

Benefits of technology

It realizes a comprehensive and accurate assessment of pilot workloads, improves the accuracy and robustness of detection, can capture instantaneous load changes in real time, and adapts to individual differences between different pilots, improving the applicability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408263A_ABST
    Figure CN120408263A_ABST
Patent Text Reader

Abstract

The invention provides a visual-physiological multi-mode pilot workload detection method and device. The method comprises the following steps: acquiring facial image data, physiological sequence data, subjective questionnaire data and objective task performance data of a pilot in a flight simulation environment; and carrying out standardization processing and comprehensive analysis on subjective and objective data through an entropy weight method, and calculating the comprehensive workload level of the pilot. An image encoder is adopted to perform feature extraction on face image data, an LSTM time sequence encoder is combined to extract time sequence features, multi-modal feature fusion and deep learning training are performed through a Transform model, a workload detection model is constructed, workloads of pilots under different task difficulties can be accurately predicted, and the workload detection efficiency is improved. And the accuracy and the real-time performance of pilot workload detection are improved. The instantaneous workload change of the pilot can be effectively captured, and the limitation of a single data source in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of aviation flight, and particularly relates to a method and device for detecting the workload of pilots in a visual-physiological multimodal manner. Background Art

[0002] Human factors are the main factors leading to aviation accidents. According to statistics, fatal crashes caused by pilot operation errors account for 58% of all crashes; among them, improper pilot workload is the main cause of human errors.

[0003] Flight workload refers to the cognitive resources required for pilots to concentrate, perceive the situation, make reasonable decisions and take actions, that is, the total workload and energy required to process information borne by pilots per unit time. Whether the pilot workload is too high or too low is not conducive to the safe completion of flight tasks. When the workload is too high, the pilot's ability reaches the limit level, and it is almost impossible to handle any unexpected situation; when the workload is too low, the pilot will be lax about the task and easily ignore flight information. Therefore, timely and accurate assessment of flight workload is of great significance for aviation safety and efficiency.

[0004] The measurement of pilot workload mainly includes three main methods: subjective measurement method, task performance measurement method, and physiological measurement method. At present, the detection accuracy and reliability of single-modal recognition are low, and there is less application of the fusion of multiple physiological characteristics and the combination of subjective and objective detection methods.

[0005] In the prior art, Chinese Patent CN119523486A discloses a method for monitoring and warning the cognitive load of pilots based on electroencephalogram and eye movement data, including the steps of: 1) designing an experimental paradigm for inducing different cognitive load states of pilots in a simulated flight task scenario to obtain multi-modal physiological signals and task performance of the subjects during the experiment; the multi-modal physiological signals include electroencephalogram signals and eye movement signals; 2) preprocessing the multi-modal physiological signals and task performance and extracting corresponding characteristic indicators, fusing the characteristic indicators and selecting to construct an optimal characteristic subset; 3) constructing a cognitive state classification and discrimination model based on the optimal characteristic subset to monitor and warn the cognitive load of pilots. The present invention can realize the monitoring and evaluation of the cognitive load level of pilots, especially warning the state mutation point of cognitive load overload to improve flight safety.

[0006] However, the existing workload monitoring methods mainly focus on single-modal data fusion, especially the processing of electroencephalogram (EEG) and eye movement data, lacking comprehensive integration of pilots' visual, cognitive, and physiological states. Although the existing methods can evaluate the cognitive workload level of pilots to a certain extent, they cannot comprehensively consider the multiple physiological states and behavioral responses of pilots in complex flight environments. The multi-modal data fusion of the existing methods is still not mature enough to effectively capture the workload changes of pilots under different task difficulties. Secondly, the existing cognitive workload assessment methods rely heavily on complex feature selection and signal processing processes, such as multi-stage feature selection strategies and extraction of EEG signal features in specific frequency bands. Although these methods can improve the accuracy under certain conditions, due to the cumbersome feature selection process and susceptibility to data noise and collinearity problems, the computational efficiency is low and they are sensitive to abnormal data. In addition, these methods lack sufficient adaptability to individual differences among different pilots, which may affect the generalization ability of the model. Moreover, the existing methods are insufficient in dealing with the temporal dynamic changes of pilots' cognitive workload. Although EEG and eye movement data provide certain temporal information, due to the failure to effectively integrate the temporal features in image data and physiological data, the existing methods fail to capture the changes in pilots' instantaneous workload during flight. This defect makes the existing technology unable to evaluate the comprehensive workload of pilots in real time and accurately in the complex scenarios of flight tasks. Summary of the Invention

[0007] The purpose of the present invention is to provide a visual-physiological multi-modal pilot workload detection method and device to overcome the above-mentioned defects existing in the prior art.

[0008] The purpose of the present invention can be achieved through the following technical solutions:

[0009] On the one hand, the present invention provides a visual-physiological multi-modal pilot workload detection method, including the following steps:

[0010] Step S1: Collect multi-modal data of a pilot performing tasks with different difficulty coefficients in a flight simulation environment, including facial image data, physiological sequence data, subjective questionnaire data, and objective task performance data;

[0011] Step S2: Standardize, calculate ratios, and analyze information entropy for the collected subjective questionnaire data and objective task performance data through the entropy weight method to obtain the corresponding comprehensive workload levels of the pilot when performing tasks with different difficulty coefficients;

[0012] Step S3: Preprocess and extract features from the collected facial image data and physiological sequence data, and perform time synchronization alignment to obtain synchronized and aligned multi-modal data;

[0013] Step S4: Use the synchronized and aligned multimodal data as input and the comprehensive workload level as the label to train a multimodal detection model for pilot workload, and obtain a trained multimodal detection model for pilot workload. The multimodal detection model for pilot workload includes an image encoder, a temporal encoder, a fusion module, and a Transformer-based encoder;

[0014] Step S5: Detect the workload of the pilot to be detected based on the trained multimodal detection model for pilot workload.

[0015] Further, the collection of multimodal data of the pilot when performing tasks with different difficulty coefficients in the flight simulation environment specifically includes:

[0016] Real-time collect the facial image data of the pilot in the flight simulation environment by setting up a camera, and the arrangement position of the camera does not affect the flight state of the pilot;

[0017] Collect the electroencephalogram data and eye movement data of the pilot through the electroencephalograph and eye tracker worn by the pilot, and the electroencephalogram data and eye movement data form physiological sequence data;

[0018] The pilot fills out the NASA-TLX questionnaire after completing tasks with different difficulty coefficients to obtain subjective questionnaire data about the pilot. The subjective questionnaire data includes the cognitive load, physical load, time pressure, performance level, effort level, and frustration level perceived by the pilot during the task;

[0019] The pilot conducts a psychomotor vigilance test after completing the task, and obtains objective task performance data through the test. The objective task performance data includes the reaction time and the number of reaction delays of the pilot.

[0020] Further, the specific steps of Step S2 include:

[0021] Collect the scores and weights of 6 dimensions of cognitive load, physical load, time pressure, performance level, effort level, and frustration level through the NASA-TLX questionnaire scale, and calculate the subjective quantification value of workload. The calculation formula is as follows:

[0022] y i =5×p i ×m i (i=1,2,…,6)

[0023] Where y i is the subjective quantification value of workload for the i-th dimension, p i is the weight for the i-th dimension, and m i is the score for the i-th dimension;

[0024] Objective task performance data is obtained through a psychomotor vigilance test, specifically the difference between the key-pressing time of the pilot for the salient signal and the actual occurrence time of the salient signal, so as to calculate the mean reaction time R of the pilot t and the number of reaction delays R d , as an objective quantification value of workload;

[0025] After obtaining the subjective quantification value data y i of workload and the objective task performance data R t and R d , the entropy weight method is used to standardize, calculate ratios, analyze information entropy, and determine weights for the subjective quantification value data y i and the objective task performance data R t and R d , and the comprehensive workload level is calculated.

[0026] Furthermore, the entropy weight method is used to standardize, calculate ratios, analyze information entropy, and determine weights for the subjective quantification value data y and the objective task performance data R t and R d , and the comprehensive workload level is calculated, specifically including:

[0027] Standardize the subjective quantification value data y and the objective task performance data R t and R d , including:

[0028] For positive indicators, including cognitive load, physical load, time pressure, and mean reaction time, standardize through the following formula:

[0029] (Positive indicator)

[0030] where, where, Y ij is the standardized value of the jth positive indicator of the ith sample, X ij is the original value of the jth positive indicator of the ith sample, min(X j ) and max(X j ) are the minimum and maximum values of the jth positive indicator respectively;

[0031] For negative indicators, including performance level, effort level, frustration level, and number of reaction delays, standardize through the following formula:

[0032] (Negative indicator)

[0033] where, where, Y ij is the standardized value of the jth negative indicator of the ith sample, X ijis the original value of the jth negative indicator of the i-th sample, min(X j ) and max(X j ) are the minimum and maximum values of the j-th negative indicator respectively;

[0034] Calculate the ratio p of each indicator ij , that is, the ratio of the normalized index value to the sum of all sample values of the index:

[0035]

[0036] Among them, p ij represents the normalized ratio of the jth indicator of the i-th sample, n is the number of samples, m is the number of indicators for each sample, m = 8;

[0037] According to the normalized index ratio p ij , calculate the information entropy E of each indicator j :

[0038]

[0039] Among them, E j is the information entropy of the j-th indicator;

[0040] According to the information entropy value E j , calculate the weight w of each indicator j :

[0041]

[0042] Among them, w j is the weight of the jth indicator;

[0043] According to the weight w j , and obtain the pilot's comprehensive workload level s when performing tasks with different difficulty coefficients i :

[0044]

[0045] Among them, s i is the comprehensive workload level of the i-th sample.

[0046] Furthermore, the step S3 specifically includes:

[0047] The collected facial image data is decomposed into single-frame images at a frame rate that matches the physiological sequence data, and converted into a grayscale image format;

[0048] Compare the proportion of pixels below the first preset threshold h1 in the grayscale image to the total pixels, and set this proportion as p dark , when the proportion p darkWhen it is less than the second preset threshold h2, the low-light enhancement module is enabled to optimize the image. Among them, the low-light enhancement module uses the Retinexformer model based on the deep fusion of retinal theory and Transformer algorithm to output the processed image data;

[0049] Perform electrode localization, remove useless electrodes, re-reference, filter, segment, and baseline correction preprocessing steps on the electroencephalogram data in the physiological sequence data;

[0050] Perform short-time Fourier transform STFT on the preprocessed electroencephalogram data to obtain the electroencephalogram time-series feature data of the electroencephalogram data, including the spectral data of different frequency components, including δ, θ, α, and β waves. The short-time Fourier transform formula is:

[0051]

[0052] Among them, ω represents the angular frequency, t is the time variable, ω(t) is the window function, * represents the conjugate, and e -jωτ is the complex exponential of the Fourier transform; by sliding the window function ω(τ - t) on the signal, STFT performs local Fourier transform on the signal at each time point t to obtain the amplitude and phase information of the frequency component ω;

[0053] Perform validity tests on the eye movement sequence data in the physiological sequence data to exclude outliers and noise; fill in the missing values through interpolation technology, and use filtering to remove high-frequency noise, and output the eye movement feature data, including pupil diameter, saccade features, fixation features, and blink features;

[0054] Perform linear interpolation on the facial image data, electroencephalogram data, and eye movement sequence data with different sampling frequencies to ensure the consistency of all modal data on the time axis, and output the synchronized and aligned multi-modal data.

[0055] Further, using the synchronized and aligned multi-modal data as the input and the comprehensive workload level as the label, training the pilot workload multi-modal detection model specifically includes:

[0056] Use an image encoder based on the ResNet34 network to extract features from the facial image data in the multi-modal data to obtain the feature representation of the image. The ResNet34 network includes multiple residual modules, and effectively extracts local and global features in the image through the residual learning framework, and generates the corresponding image feature vector;

[0057] Use a time-series encoder based on the long short-term memory network LSTM to extract time-series features from the electroencephalogram time-series feature data and eye movement feature data in the multi-modal data, capture the long-term dependencies in the time-series data, and output the physiological time-series feature vector at each time step;

[0058] Project and fuse the image feature vector and the physiological time series feature vector, map the image features and time series features to the same dimensional space, and obtain a unified multi-modal feature representation;

[0059] Input the fused multi-modal feature representation into the Transformer-based encoder, deeply fuse the features of different modalities through the multi-head attention mechanism and the feed-forward network FFN, and accurately capture the interactions between modalities to form a comprehensive feature representation;

[0060] Input the output comprehensive feature representation into the classifier, predict the work load level of the pilot through the classifier, output the predicted comprehensive work load level, and update the parameters of the pilot work load multi-modal detection model according to the predicted comprehensive work load level, the actual comprehensive work load level and the loss function.

[0061] Furthermore, the image encoder based on the ResNet34 network is used to extract features from the facial image data in the multi-modal data, specifically including:

[0062] The facial image data in the multi-modal data is input into the ResNet34 network, where H is the height of the image, W is the width of the image, and C is the number of color channels;

[0063] Use a 7×7 convolutional kernel with a stride of 2 for initial feature extraction, and use a max pooling layer to reduce the spatial dimension of the feature map;

[0064] Extract features through multiple residual blocks, each residual block contains three convolutional layers: 1×1, 3×3 and 1×1 convolutional layers, and the residual block adds the input feature map and the output feature map through a skip connection, and the formula is expressed as:

[0065]

[0066] where x is the input feature map, Conv represents the convolution operation, BN represents batch normalization, and ReLU is the activation function

[0067] Apply a global average pooling layer at the end of the ResNet34 network to compress the spatial dimension of the feature map to 1 and obtain a fixed-length feature vector;

[0068] Map the features after global average pooling to a specific task space through a fully connected layer to generate the final image feature vector.

[0069] Furthermore, the time series encoder based on the long short-term memory network LSTM is used to extract time series features from the electroencephalogram time series feature data and eye movement feature data in the multi-modal data, specifically including:

[0070] Input the EEG time series feature data and eye movement feature data into the LSTM network for time series feature extraction. For each time step t, concatenate the hidden state h of the previous moment t-1 and the input x at the current moment t and then calculate through different weight matrices and biases:

[0071] Calculate the forget gate f t :

[0072] f t = σ(W f · [h t-1 , x t + b f )

[0073] where σ represents the sigmoid activation function, W f is the weight matrix of the forget gate, and b f is the bias term;

[0074] Calculate the input gate i t :

[0075] i t = σ(W i · [h t-1 , x t + b i )

[0076] where W i is the weight matrix of the input gate, and b i is the bias term;

[0077] Calculate the candidate cell state

[0078]

[0079] where W C is the weight matrix of the candidate cell state, and b C is the bias term;

[0080] Update the cell state C t :

[0081]

[0082] where C t-1 is the cell state of the previous moment;

[0083] Calculate the output gate o t :

[0084] o t = σ(W o · [ht-1 , x t + b o )

[0085] Among them, W o is the weight matrix of the output gate, and b o is the bias term;

[0086] Calculate the hidden state h at the current moment t :

[0087] h t = o t ·tanh(C t )

[0088] Output the hidden state h t as the physiological time series feature vector.

[0089] Furthermore, the loss function is:

[0090]

[0091] Among them, is the loss function, is the predicted comprehensive workload level of the i-th sample, s i is the actual comprehensive workload level of the i-th sample, and n is the number of samples.

[0092] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the above is implemented.

[0093] Compared with the prior art, the present invention has the following advantages:

[0094] (1) By integrating the pilot's facial images, electroencephalogram signals, eye movement data, and task performance data, the present invention overcomes the problem that single-modal data in the prior art cannot comprehensively evaluate the pilot's workload. Through multi-angle comprehensive analysis of visual and physiological signals, the present invention can more comprehensively and accurately reflect the cognitive load changes of pilots in complex flight environments, thereby improving the accuracy and robustness of monitoring.

[0095] (2) The present invention combines subjective and objective workload assessment methods, comprehensively considers the pilot's subjective mental state and objective operation performance, obtains data through the NASA-TLX questionnaire and the psychomotor vigilance test respectively, and uses the entropy weight method to standardize, calculate ratios, analyze information entropy, and determine weights for data with different dimensions. This method can effectively eliminate the bias between data sources, optimize the calculation of the comprehensive workload level, provide more accurate and reliable evaluation labels, and significantly improve the accuracy and practicality of the workload detection model.

[0096] (3) The present invention uses an image encoder based on the ResNet34 network to extract the facial image data features, and combines an LSTM temporal encoder to process the electroencephalogram temporal features and eye movement feature data. This fusion method can effectively capture the complex temporal dependence relationship between the images and physiological data, and comprehensively improve the monitoring accuracy of the pilot's workload.

[0097] (4) In the feature fusion stage, the present invention uses an encoder based on Transformer, and performs deep fusion through the multi-head attention mechanism and the feed-forward network FFN, accurately capturing the interaction between the image, temporal, and physiological features, and improving the effect of multi-modal data fusion. This deep fusion method can fully consider the interaction relationship between different data modalities, thereby providing more accurate workload prediction results.

[0098] (5) By extracting the temporal features from the electroencephalogram and eye movement sequence data, the present invention can monitor the instantaneous workload changes of the pilot during the flight in real time. Compared with the prior art, the present invention can better capture the dynamic workload of the pilot in complex flight tasks, enhancing the real-time performance and accuracy of the evaluation.

[0099] (6) In view of the actual environment with large changes in the light in the pilot's cockpit, the present invention proposes an image processing and recognition technology based on low-light enhancement, and uses the Retinexformer model that combines the retina theory and the Transformer algorithm to enhance and optimize the low-light images. This technology can effectively overcome the interference of the cockpit light fluctuations on the image quality, ensuring the acquisition of high-quality facial image features under low-light conditions, providing stable and reliable visual data support for the pilot workload detection, and improving the robustness and adaptability of the detection system.

[0100] (7) By integrating multiple data sources, the present invention overcomes the problem of insufficient adaptability of traditional methods to individual differences. The workload evaluation of the pilot not only depends on the physiological data, but also takes into account the subjective perception differences of individuals, making the present invention have stronger personalized adaptation ability and can be more widely applied to different pilots. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 is the main step flow chart of the detection method provided by the embodiment of the present invention;

[0102] Figure 2 is the detailed flow chart of the label module of the present invention;

[0103] Figure 3 is the detailed flow chart of the data preprocessing module of the present invention;

[0104] Figure 4This is a detailed flowchart of the model construction module of the present invention. Specific embodiments

[0105] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0106] This embodiment provides a visual-physiological multimodal pilot workload detection method, as Figure 1 shown, including the following steps:

[0107] Step S1: Collect multimodal data of the pilot when performing tasks with different difficulty coefficients in a flight simulation environment, including facial image data, physiological sequence data, subjective questionnaire data, and objective task performance data;

[0108] Step S2: Standardize, calculate ratios, and analyze information entropy for the collected subjective questionnaire data and objective task performance data through the entropy weight method to obtain the corresponding comprehensive workload levels of the pilot when performing tasks with different difficulty coefficients;

[0109] Step S3: Preprocess and extract features from the collected facial image data and physiological sequence data, and perform time synchronization alignment to obtain synchronously aligned multimodal data;

[0110] Step S4: Use the synchronously aligned multimodal data as input and the comprehensive workload level as a label to train a pilot workload multimodal detection model to obtain a trained pilot workload multimodal detection model. The pilot workload multimodal detection model includes an image encoder, a temporal encoder, a fusion module, and a Transformer-based encoder;

[0111] Step S5: Perform workload detection on the pilot to be detected based on the trained pilot workload multimodal detection model.

[0112] Collect multimodal data of the pilot when performing tasks with different difficulty coefficients in a flight simulation environment, specifically including:

[0113] Real-time collect the pilot's facial image data in the flight simulation environment by setting up a camera, and the arrangement position of the camera does not affect the pilot's flight state;

[0114] Collect the pilot's electroencephalogram data and eye movement data through the electroencephalograph and eye tracker worn by the pilot. The electroencephalogram data and eye movement data constitute physiological sequence data;

[0115] After completing tasks with different difficulty coefficients, the pilot fills out the NASA-TLX questionnaire to obtain subjective questionnaire data of the pilot. The subjective questionnaire data includes the cognitive load, physical load, time pressure, performance level, effort level, and frustration level perceived by the pilot during the task;

[0116] After completing the task, the pilot conducts a Psychomotor Vigilance Test to obtain objective task performance data. The objective task performance data includes the reaction time and the number of reaction delays of the pilot.

[0117] The experiment will be completed on the flight simulation platform in the pilot's simulator cockpit. The experimenters are pilots with rich flight experience. The main acquisition devices are cameras, electroencephalographs, eye trackers, etc.; among them, the camera needs to be arranged at a position that does not affect the pilot's flight state to collect complete facial image data of the pilot; the electroencephalograph and eye tracker are worn by the pilot being tested to ensure accurate and complete electroencephalogram and eye movement data are collected. The experiment needs to adjust environmental factors such as flight light, wind speed, and noise, set different workload modes, induce different workload states of the pilot, and collect balanced sample data; after the pilot completes tasks with different loads, the NASA-TLX questionnaire needs to be filled out in a timely manner to obtain subjective questionnaire scale data, and a Psychomotor Vigilance Test (PVT) needs to be conducted to obtain objective task performance data, completing the acquisition of subjective and objective workload labels.

[0118] As Figure 2 shown, step S2 specifically includes:

[0119] The scores and weights of 6 dimensions, namely cognitive load, physical load, time pressure, performance level, effort level, and frustration level, are collected through the NASA-TLX questionnaire scale, and the subjective quantification value of the workload is calculated. The calculation formula is as follows:

[0120] y i = 5 × p i × m i (i = 1, 2, …, 6)

[0121] Among them, y i is the subjective quantification value of the workload of the i-th dimension, p i is the weight of the i-th dimension, and m i is the score of the i-th dimension;

[0122] Objective task performance data is obtained through the Psychomotor Vigilance Test. Specifically, it is the difference between the key-pressing time of the pilot for the salient signal and the actual appearance time of the salient signal, so as to calculate the average reaction time R t and the number of reaction delays R d of the pilot, which are used as the objective quantification value of the workload;

[0123] After obtaining the subjective quantization value data y of the workload i and the objective task performance data R t and R d then, the entropy weight method is used to standardize, calculate the ratio, analyze the information entropy, and determine the weight of the subjective quantization value data y i and the objective task performance data R t and R d to calculate the comprehensive workload level

[0124] Using the entropy weight method to standardize, calculate the ratio, analyze the information entropy, and determine the weight of the subjective quantization value data y and the objective task performance data R t and R d to calculate the comprehensive workload level, specifically including

[0125] Standardize the subjective quantization value data y and the objective task performance data R t and R d including

[0126] For positive indicators, including cognitive load, physical load, time pressure, and mean response time, standardize through the following formula

[0127] (Positive indicator)

[0128] where, where, Y ij is the standardized value of the jth positive indicator of the ith sample, X ij is the original value of the jth positive indicator of the ith sample, min(X j ) and max(X j ) are the minimum and maximum values of the jth positive indicator respectively

[0129] For negative indicators, including performance level, effort level, frustration level, and number of response delays, standardize through the following formula

[0130] (Negative indicator)

[0131] where, where, Y ij is the standardized value of the jth negative indicator of the ith sample, X ij is the original value of the jth negative indicator of the ith sample, min(X j ) and max(X j ) are the minimum and maximum values of the jth negative indicator respectively

[0132] Calculate the ratio p of each indicator ij, that is, the ratio of the normalized index value to the sum of all sample values of this index:

[0133]

[0134] where p ij represents the normalized ratio of the j-th index of the i-th sample. n is the number of samples, m is the number of indexes for each sample, and m = 8;

[0135] According to the normalized index ratio p ij , calculate the information entropy E j of each index:

[0136]

[0137] where E j is the information entropy of the j-th index;

[0138] According to the information entropy value E j , calculate the weight w j of each index:

[0139]

[0140] where w j is the weight of the j-th index;

[0141] According to the weight w j , obtain the comprehensive workload level s i of the pilot when performing tasks with different difficulty coefficients:

[0142]

[0143] where s i is the comprehensive workload level of the i-th sample.

[0144] As Figure 3 shown, step S3 specifically includes:

[0145] S301: Image data preprocessing module. Considering the large variation in the light environment of the pilot's cockpit, when the light environment is relatively dark, low-light enhancement technology is enabled to perform image enhancement processing on facial data. The specific process is as follows: The facial video data collected in S1 is decomposed into single-frame images at a frame rate matching the physiological signals, and the facial image data is converted into grayscale format. By comparing the proportion of the number of pixels with gray values lower than a preset threshold h1 in the total number of pixels in the picture, this is used as the key basis for judging the brightness of the picture. Once this ratio is less than a specific threshold h2, it means that the image brightness is at a low level, and the low-light enhancement module needs to be immediately enabled to optimize and improve the image quality; among them, the low-light enhancement module uses the Retinexformer model based on the deep fusion of retinal theory and Transformer algorithm, which can quickly output high-definition and bright high-quality pictures.

[0146] S302: Physiological data preprocessing module. It is necessary to perform data preprocessing and synchronization alignment on physiological data such as electroencephalogram sequences and eye movement sequences respectively. Among them, the electroencephalogram sequence needs to go through preprocessing work such as electrode positioning, removing useless electrodes, rereferencing, filtering, segmentation and baseline correction, interpolating bad leads and removing bad segments, independent component analysis, and removing artifact components; through the short-time Fourier transform method, the electroencephalogram signal is decomposed into components of different frequencies such as δ, θ, α, β, etc.

[0147] Furthermore, the short-time Fourier transform divides a long signal into many short signals for Fourier transform processing, and has been widely used in the analysis of non-stationary signals. Suppose there is a continuous signal x(t), and a window function ω(t) with a very narrow time width is given, which slides along the time axis. Then the short-time Fourier transform (STFT) of the signal x(t) is defined as:

[0148]

[0149] where ω represents the angular frequency, t is the time variable, ω(t) is the window function, * represents the conjugate, and e -jωτ is the complex exponential of the Fourier transform; by sliding the window function ω(τ - t) on the signal, STFT performs local Fourier transform on the signal at each time point t to obtain the amplitude and phase information of the frequency component ω;

[0150] In the aspect of eye movement sequence processing, it is necessary to go through validity tests to exclude outliers and noise; fill in the missing gaps through interpolation technology, use filtering technology to remove high-frequency noise in the signal, and maintain the continuity and integrity of the data; extract pupil diameter, saccade characteristics, fixation characteristics, blink characteristics, etc. to construct an eye movement feature set.

[0151] S303: Data Synchronization and Alignment Module. The data synchronization and alignment module needs to keep different modalities of data consistent in frequency. In view of the fact that the data volumes of different modalities of data per unit time are different due to different sampling frequencies of sensor devices, a linear interpolation method needs to be adopted to resample some of the modalities of data. Taking the example that the sampling frequency of a certain physiological data is lower than that of image data, through linear interpolation, based on the existing physiological data points, the missing data points within the corresponding time interval are calculated according to the linear relationship, so that the number of data points of the physiological data per unit time increases, and finally matches the time resolution of other modalities of data such as image data. During specific operation, a linear function relationship is constructed between two adjacent known data points, and according to the time interval corresponding to the set target time resolution, the values of the data points to be inserted are calculated, so as to realize the resampling of the physiological data and make its performance on the time scale consistent with other modalities.

[0152] In view of the problem of different time resolutions caused by Fourier transform, it is necessary to perform time regularization on the spectrogram according to the timestamps of the original physiological data and the time resolution of the transformed spectrogram. By analyzing the order and interval characteristics of the timestamps of the original physiological data, combined with the existing time resolution of the spectrogram, by means of appropriate interpolation or adjustment of time indexes, etc., the data on the time axis of the spectrogram are rearranged so that its timestamps can accurately correspond to the timestamps of the image data, ensuring that different modalities of data are synchronized and aligned in time and meeting the requirements of multi-modal fusion processing.

[0153] As Figure 4 shown, step S4 specifically includes:

[0154] S401: Image Encoder. ResNet34 is used to extract visual features from image data, capture local and global features in the image, and provide rich visual information for subsequent modal fusion. ResNet34 is a popular convolutional neural network (CNN) architecture that solves the degradation problem in the training of deep networks by introducing a residual learning framework.

[0155] The original image is input into the network, and a 7×7 convolutional kernel with a stride of 2 is applied to extract initial features, and max pooling is used to further reduce the spatial dimension of the feature map.

[0156] Furthermore, residual modules are constructed. ResNet34 consists of multiple residual blocks, and each residual block contains 3 convolutional layers: 1×1, 3×3, and 1×1. These convolutional layers are responsible for dimensionality reduction, feature extraction, and dimensionality increase respectively. The core of the residual block is a skip connection that adds the input directly to the output of the block, which is expressed by the formula:

[0157]

[0158] Among them, x is the input feature map, Conv represents the convolution operation, BN represents batch normalization, and ReLU is the activation function.

[0159] Furthermore, batch normalization is applied after each convolutional layer to accelerate the training process, and the ReLU activation function is used to introduce non-linearity. By stacking multiple residual blocks, the network can learn hierarchical feature representations from low-level to high-level. Each residual block learns a residual mapping, allowing gradients to flow directly through the skip connections, thus alleviating the vanishing gradient problem. At the end of the network, a global average pooling layer is used to compress the spatial dimension of the feature map to 1, obtaining a fixed-length feature vector.

[0160] Furthermore, the result of global average pooling is mapped to a specific task space through a fully connected layer, which can effectively extract rich feature representations from the input image.

[0161] S402: The temporal encoder uses a long short-term memory network (LSTM) to extract latent feature representations from temporal physiological signal data, capture long-term dependencies in time series data, and provide key temporal information. LSTM is a special recurrent neural network (RNN) structure designed to address the vanishing gradient and exploding gradient problems that occur in traditional RNNs when dealing with long sequence data. It contains a cell state and three gate structures, namely the input gate, forget gate, and output gate.

[0162] Among them, the forget gate concatenates the hidden state h at the previous moment t-1 and the input x at the current moment t , then passes through a weight matrix W f and a bias b f , and then through the sigmoid function to obtain the value f of the forget gate t , and the formula is

[0163] f t = σ(W f · [h t-1 , x t + b f )

[0164] Furthermore, the input gate also concatenates h t-1 and x t , passes through different weight matrices W i and a bias b i , and obtains the value i of the input gate through the sigmoid function t . At the same time, it passes through another weight matrix W C and a bias b C , and obtains the candidate cell state through the tanh function , and the formula is

[0165] i t = σ(W i · [h t-1 , x t + b i )

[0166]

[0167] Furthermore, update the cell state: Multiply the cell state C at the previous moment t-1 element-wise with the value f of the forget gate t , and then add the product of the value i of the input gate t and the candidate cell state to obtain the new cell state C t . The formula is

[0168]

[0169] Furthermore, the output gate concatenates h t-1 and x t , passes through the weight matrix W o and the bias b o through the sigmoid function to obtain the value o of the output gate t . Then, the new cell state C t is activated by the tanh function and multiplied by the value o of the output gate t to obtain the hidden state h at the current moment t . The formula is

[0170] o t = σ(W o · [h t-1 , x t + b o )

[0171] h t = o t · tanh(C t )

[0172] After a series of LSTM calculations at multiple time steps, the hidden state h at each time step t contains the temporal signal feature information from the start of the sequence to that time step and can be used for subsequent multi-modal data fusion.

[0173] S403: With the help of the projection layer, map the image features and physiological temporal features to the same dimensional space. Then, connect the two types of features after projection to construct a unified feature representation form, and input the fused features into the Transformer-based encoder. This encoder consists of multiple layers, and each layer is equipped with a self-attention module and a feed-forward network FFN.

[0174] Furthermore, in the self-attention module, the multi-head attention mechanism is applied to process the features in different subspaces in parallel, enabling the model to understand the features of different modalities more deeply. The specific implementation process is as follows: First, the input features are linearly transformed into three different representation spaces, namely query Q, key K, and value V. At the same time, multiple heads are set, and the model synchronously focuses on the feature interactions in the input sequence from multiple different subspace perspectives. For each head, first calculate the similarity score between the query and the key, and after normalization, use it as a weight to perform weighted summation on the values, thereby obtaining the output result of this head. Finally, a series of operations such as concatenating and linearly transforming the outputs of multiple heads are performed in sequence to merge them into a final output. This output comprehensively integrates the feature interaction information mined by each head from different dimensions, achieving a deep fusion of the features of different modalities at the semantic level, and effectively enhancing the comprehensive expression efficiency of the fused features for multi-modal information.

[0175] Finally, after successfully fusing the features of different modalities and accurately capturing the interaction relationships between them, the multi-modal Transformer encoder feeds its output to a linear classifier, which predicts the final label, that is, predicts the workload level of the pilot.

[0176] The loss function is as follows:

[0177]

[0178] Where, is the loss function, is the predicted comprehensive workload level of the i-th sample, s i is the actual comprehensive workload level of the i-th sample, and n is the number of samples.

[0179] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0180] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A visual-physiological multimodal pilot workload detection method, characterized in that, It includes the following steps: Step S1: Collect multimodal data of a pilot performing tasks with different difficulty coefficients in a flight simulation environment, including facial image data, physiological sequence data, subjective questionnaire data, and objective task performance data; Step S2: Standardize, calculate ratios, and analyze information entropy for the collected subjective questionnaire data and objective task performance data through the entropy weight method to obtain the corresponding comprehensive workload levels of the pilot when performing tasks with different difficulty coefficients; Step S3: Preprocess and extract features from the collected facial image data and physiological sequence data, and perform time synchronization alignment to obtain synchronized multimodal data; Step S4: Use the synchronized multimodal data as input and the comprehensive workload level as a label to train a multimodal detection model for pilot workload, and obtain a trained multimodal detection model for pilot workload. The multimodal detection model for pilot workload includes an image encoder, a temporal encoder, a fusion module, and a Transformer-based encoder; Step S5: Detect the workload of the pilot to be detected based on the trained multimodal detection model for pilot workload.

2. The visual-physiological multi-modal pilot workload detection method according to claim 1, wherein The collection of multimodal data of a pilot performing tasks with different difficulty coefficients in a flight simulation environment specifically includes: Collect the facial image data of the pilot in real time in the flight simulation environment by setting up a camera, and the arrangement position of the camera does not affect the flight state of the pilot; Collect the electroencephalogram data and eye movement data of the pilot through an electroencephalograph and an eye tracker worn by the pilot, and the electroencephalogram data and eye movement data form physiological sequence data; The pilot fills out the NASA-TLX questionnaire after completing tasks with different difficulty coefficients to obtain subjective questionnaire data about the pilot. The subjective questionnaire data includes the cognitive load, physical load, time pressure, performance level, effort level, and frustration level perceived by the pilot during the task; The pilot conducts a psychomotor vigilance test after completing the task, and obtains objective task performance data through the test. The objective task performance data includes the reaction time and the number of reaction delays of the pilot.

3. A visual - physiological multimodal pilot workload detection method according to claim 1 or 2, characterized in that, The specific content of step S2 includes: Collect the scores and weights of 6 dimensions, namely cognitive load, physical load, time pressure, performance level, effort level, and frustration level, through the NASA-TLX questionnaire scale, and calculate the subjective quantification value of the workload. The calculation formula is as follows: y i = 5 × p i × m i (i = 1, 2, …, 6) Among them, y i is the subjective quantification value of the workload of the i-th dimension, p i is the weight of the i-th dimension, m i is the score of the i-th dimension; Objective task performance data is obtained through a psychomotor vigilance test, specifically the difference between the time when the pilot presses the key for the salient signal and the actual occurrence time of the salient signal, so as to calculate the mean reaction time R of the pilot t and the number of reaction delays R d , which serves as an objective quantification value of workload; After obtaining the subjective quantification value data y of the workload i and the objective task performance data R t and R d After that, the entropy weight method is used to perform standardization processing, ratio calculation, information entropy analysis, and weight determination on the subjective quantification value data y i and the objective task performance data R t and R d to calculate the comprehensive workload level.

4. The visual-physiological multi-modal pilot workload detection method according to claim 3, characterized in that, The subjective quantitative value data y and the objective task performance data R are processed by the entropy weight method t and R d for standardization processing, ratio calculation, information entropy analysis and weight determination, and the comprehensive workload level is calculated, specifically including: Subjective quantitative value data y and objective task performance data R t and R d Perform normalization processing, including: For positive indicators, including cognitive load, physical load, time pressure, and mean reaction time, standardize through the following formula: Among them, Y ij is the standardized value of the j-th positive index of the i-th sample, and X ij is the original value of the j-th positive index of the i-th sample. min(X j ) and max(X j ) are the minimum and maximum values of the j-th positive index respectively; For negative indicators, including performance level, effort level, frustration level, and number of reaction delays, standardize through the following formula: Among them, Y ij is the standardized value of the j-th negative index of the i-th sample, and X ij is the original value of the j-th negative index of the i-th sample, min(X j ) and max(X j ) are the minimum and maximum values of the j-th negative index respectively; Calculate the ratio p of each metric ij , that is, the ratio of the normalized metric value to the sum of all sample values of the metric: where p ij represents the normalized ratio of the j-th index of the i-th sample, n is the number of samples, m is the number of indices for each sample, and m = 8; According to the normalized index ratio p ij , calculate the information entropy E of each index j : Among them, E j is the information entropy of the j-th index; According to the information entropy value E j , calculate the weight w of each index j : where, w j is the weight of the j-th index; According to the weight w j , the comprehensive workload level s of the pilot when performing tasks with different difficulty coefficients is obtained i : Among them, s i is the comprehensive workload level of the i-th sample.

5. A visual - physiological multimodal pilot workload detection method according to claim 1, characterized in that, The specific content of step S3 includes: Decompose the collected facial image data into single-frame images according to the frame rate matching the physiological sequence data, and convert it into a grayscale image format; Compare the proportion of pixels below the first preset threshold h1 in the grayscale image, and let this proportion be p dark , when the proportion p dark is less than the second preset threshold h2, enable the low-light enhancement module to optimize the image. Among them, the low-light enhancement module uses the Retinexformer model based on the deep fusion of retinal theory and Transformer algorithm to output the processed image data; Perform electrode localization, remove useless electrodes, re-reference, filter, segment, and baseline correction preprocessing steps on the electroencephalogram data in the physiological sequence data; Perform short-time Fourier transform (STFT) on the preprocessed electroencephalogram (EEG) data to obtain the EEG temporal feature data of the EEG data, including the spectral data of different frequency components, including δ, θ, α, and β waves. The short-time Fourier transform formula is as follows: where ω represents the angular frequency, t is the time variable, ω(t) is the window function, * represents the conjugate, and e -jωτ is the complex exponential of the Fourier transform; by sliding the window function ω(τ - t) over the signal, the STFT performs a local Fourier transform on the signal at each time point t to obtain the amplitude and phase information of the frequency component ω; Perform validity tests on the eye movement sequence data in the physiological sequence data to exclude outliers and noise; fill in the missing values through interpolation techniques, and use filtering to remove high-frequency noise, and output the eye movement feature data, including pupil diameter, saccade characteristics, fixation characteristics, and blink characteristics; Perform linear interpolation on the facial image data, EEG data, and eye movement sequence data with different sampling frequencies to ensure the consistency of all modal data on the time axis, and output the synchronized and aligned multi-modal data.

6. A visual - physiological multimodal pilot workload detection method according to claim 1 or 5, characterized in that, Use the synchronized and aligned multi-modal data as the input and the comprehensive workload level as the label to train the multi-modal detection model for pilot workload, which specifically includes: Use an image encoder based on the ResNet34 network to extract features from the facial image data in the multi-modal data to obtain the feature representation of the image. The ResNet34 network includes multiple residual modules, and effectively extracts local and global features in the image through the residual learning framework, and generates corresponding image feature vectors; Use a temporal encoder based on the long short-term memory network (LSTM) to extract temporal features from the EEG temporal feature data and eye movement feature data in the multi-modal data, capture the long-term dependencies in the temporal data, and output the physiological temporal feature vectors at each time step; Perform projection and fusion on the image feature vectors and physiological temporal feature vectors, map the image features and temporal features to the same dimensional space, and obtain a unified multi-modal feature representation; Input the fused multi-modal feature representation into an encoder based on Transformer, deeply fuse the features of different modalities through the multi-head attention mechanism and the feed-forward network (FFN), and accurately capture the interactions between modalities to form a comprehensive feature representation; Input the output comprehensive feature representation into a classifier, predict the workload level of the pilot through the classifier, output the predicted comprehensive workload level, and update the parameters of the multi-modal detection model for pilot workload according to the predicted comprehensive workload level, the actual comprehensive workload level, and the loss function.

7. A visual-physiological multi-modal pilot workload detection method according to claim 6, characterized in that, The specific steps of using an image encoder based on the ResNet34 network to extract features from the facial image data in the multi-modal data include: Input the facial image data in the multi-modal data into the ResNet34 network, where H is the height of the image, W is the width of the image, and C is the number of color channels; Use a 7×7 convolutional kernel with a stride of 2 for initial feature extraction, and use a max-pooling layer to reduce the spatial dimension of the feature map; Extract features through multiple residual blocks. Each residual block contains three convolutional layers: 1×1, 3×3, and 1×1 convolutional layers. The residual block adds the input feature map and the output feature map through a skip connection. The formula is expressed as: where x is the input feature map, Conv represents the convolutional operation, BN represents batch normalization, and ReLU is the activation function Apply a global average pooling layer at the end of the ResNet34 network to compress the spatial dimension of the feature map to 1 and obtain a fixed-length feature vector; The features after global average pooling are mapped to a specific task space through a fully connected layer to generate the final image feature vector.

8. A visual-physiological multi-modal pilot workload detection method according to claim 6, characterized in that The time series encoder based on the long short-term memory network (LSTM) is used to extract the time series features of the electroencephalogram (EEG) time series feature data and the eye movement feature data in the multi-modal data, specifically including: Input the EEG time series feature data and eye movement feature data into the LSTM network for time series feature extraction. For each time step t, concatenate the hidden state h of the previous moment t-1 and the input x of the current moment t and then perform calculations through different weight matrices and biases: Calculate the forget gate f t : f t = σ(W f · [h t-1 , x t + b f ) where, σ represents the sigmoid activation function, W f is the weight matrix of the forget gate, b f is the bias term; Calculate the input gate i t : i t = σ(W i · [h t-1 , x t + b i ) Among them, W i is the weight matrix of the input gate, and b i is the bias term; Calculating candidate cell states Among them, W C is the weight matrix of the candidate cell state, and b C is the bias term; Update cell state C t : Among them, C t-1 is the cell state at the previous moment; Calculate the output gate o t : o t = σ(W o · [h t-1 , x t + b o ) Among them, W o is the weight matrix of the output gate, and b o is the bias term; Calculate the hidden state h at the current moment t : h t = o t ·tanh(C t ) Output the hidden state h t as a physiological timing feature vector.

9. A visual-physiological multimodal pilot workload detection method according to claim 6, characterized in that, The loss function is: Among them, is the loss function, is the predicted comprehensive workload level of the i-th sample, s i is the actual comprehensive workload level of the i-th sample, and n is the number of samples.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Pilot cognitive load monitoring and early warning method based on electroencephalogram and eye movement data

    CN119523486A