A multi-modal fusion physiological fatigue recognition method and device
By constructing standardized speech datasets and physiological information datasets, and integrating physiological indicators, biochemical indicators, and voiceprint features, a fatigue recognition model was trained and optimized. This solved the problems of insufficient objectivity and low accuracy in the existing technology for physiological fatigue detection and recognition, and achieved fatigue recognition with higher accuracy and better robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ACADEMY OF MILITARY MEDICAL SCIENCES
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-05
AI Technical Summary
Existing physiological fatigue detection and recognition technologies suffer from insufficient objectivity, significant individual variability, expensive and complex equipment, and difficulty in meeting the continuous and accurate monitoring needs of fatigue states in complex work scenarios. In particular, fatigue recognition methods based on voiceprint features lack unified fatigue level standards and public databases, resulting in low recognition accuracy.
A standardized speech dataset and physiological information dataset are constructed, and physiological indicators, biochemical indicators and speech print features are integrated. The fatigue recognition model is trained and optimized through multi-dimensional physiological fatigue feature information to achieve multi-modal fusion physiological fatigue recognition.
It improves the accuracy of physiological fatigue identification and the robustness and generalization ability of the model, making it suitable for real-time monitoring of complex work scenarios.
Smart Images

Figure CN122140188A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of physiological state detection and speech recognition technology, and in particular to a multimodal fusion physiological fatigue recognition method and device. Background Technology
[0002] Physiological fatigue refers to the physiological degeneration of the human body, especially the musculoskeletal, nervous, cardiopulmonary, and respiratory systems, resulting in decreased function, sluggish reactions, and weakened coordination under prolonged high-intensity work or training. It is often accompanied by psychological manifestations such as decreased attention, reduced alertness, and mood swings. Physiological fatigue not only significantly reduces work efficiency and safety, but long-term persistent fatigue can also induce chronic bodily damage, decreased immunity, metabolic disorders, and other potential health threats. With the accelerating pace of social life and work, physiological fatigue detection and monitoring technologies have developed rapidly in recent years, particularly in key fields such as industrial manufacturing, sports training, clinical rehabilitation, transportation, and military aerospace. The demand for real-time, objective, and accurate assessment of worker fatigue is increasing, prompting the emergence of a series of new fatigue recognition technologies based on physiological signal sensing, artificial intelligence recognition, and wearable device integration. These technologies play a crucial role in improving work system safety, optimizing human factors engineering design, and enhancing work efficiency.
[0003] Existing technologies for detecting and identifying physiological fatigue mainly focus on: subjective assessment through questionnaires, behavioral feature monitoring, physiological signal monitoring, and multimodal fusion methods. Among these, questionnaires and subjective assessment methods are widely used in the initial screening and self-assessment stages of fatigue due to their simplicity, low cost, and rapid implementation. However, they are highly dependent on individual subjective feelings, greatly influenced by psychological state and environmental interference, and lack objectivity, making it difficult to meet the needs of continuous and accurate monitoring of fatigue states in complex work scenarios. Behavioral feature monitoring methods infer fatigue states by modeling indicators such as eye movement trajectories, blink frequency, fixation point drift, motion capture and gait analysis, and facial micro-expression recognition. These methods have advantages such as non-invasiveness and ease of deployment, but they suffer from significant individual variability, weak model generalization ability, and a lack of unified standard datasets and labeling systems. Physiological signals used in fatigue monitoring and assessment mainly include electroencephalography (EEG), heart rate variability, skin conductance and body temperature, electromyography (EMG), and near-infrared brain imaging. Although physiological signal monitoring provides relatively objective and continuous data support, it relies on expensive and complex specialized instruments, has high requirements for experimental environment and task design, and requires an experimental paradigm highly matched to the target task to induce an effective fatigue response. The entire process is time-consuming and burdensome for the user experience, limiting its practical deployment. Multimodal fusion fatigue monitoring methods attempt to improve the accuracy and robustness of fatigue detection by integrating multiple signal sources. This direction is still in a rapid evolutionary stage and faces technical challenges such as data fusion strategies, signal synchronization, and label consistency.
[0004] Fatigue detection technology based on voiceprint features falls under the category of fatigue monitoring methods based on behavioral characteristics, and has gradually become an important direction in non-invasive fatigue recognition research in recent years. This method mainly involves collecting voice signals from volunteers under different fatigue states and extracting their voiceprint features for modeling and classification. Voice, as a medium for natural human interaction, is highly dependent on the coordinated control of the respiratory system, vocal system, and central nervous system. When an individual is fatigued, this coordination undergoes subtle changes, resulting in physiological manifestations such as decreased pitch, slower speech rate, weakened phoneme stability, and changes in spectral energy. Compared with other behavioral or physiological features, voice recognition technology has advantages such as non-invasiveness, simple equipment, real-time performance, and low deployment costs, making it suitable for various scenarios such as remote monitoring, mobile terminals, or in-vehicle systems. However, traditional label acquisition relies on subjective questionnaires, lacking unified fatigue level standards and public voice fatigue databases, leading to low accuracy in fatigue recognition and a lack of objective comparison options. There is an urgent need to construct an objective assessment dataset and develop accurate fatigue level recognition algorithms. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multimodal fusion physiological fatigue recognition method and device, which integrates subjective judgment, physiological indicators, biochemical indicators and voiceprint features to improve the accuracy of physiological fatigue recognition.
[0006] To address the aforementioned technical problems, a first aspect of this invention discloses a multimodal fusion method for physiological fatigue recognition, the method comprising: S1, construct standardized speech dataset and physiological information dataset; The physiological information dataset includes physiological indicators and biochemical indicators; the physiological indicators include heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicators include saliva information, urine information, and fingertip blood sample information. S2, Process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; S3, using the multidimensional physiological fatigue feature information, train the preset fatigue recognition model to obtain an optimized fatigue recognition model; S4. Using the optimized fatigue recognition model, the multidimensional physiological fatigue feature information to be processed is processed to obtain the multimodal fusion physiological fatigue recognition result.
[0007] As an optional implementation, in the first aspect of the present invention, the construction of the standardized speech dataset includes: S11, acquire voice data information; S12, perform frequency domain analysis on the speech data information to obtain speech audio domain information; The speech domain information expression is: in For audio domain information, For voice data information, For frequency domain variables, For time-domain variables; S13, process the audio domain information to obtain the first audio domain information; The expression for the first audio domain information is: S14, process the first audio domain information to obtain sampled speech information; The expression for the sampled speech information is: in, To sample voice information, For the largest spectral variable, , , , The number of sampling points. For the first speech audio domain information, for The minimum value; S15, the sampled speech information is processed to obtain a standardized speech dataset.
[0008] As an optional implementation, in the first aspect of the present invention, processing the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information includes: S21, Process the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information; S22, Process the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information; S23, integrate the multidimensional voiceprint spectrum feature vector information, the physiological indicator feature information and the biochemical indicator feature information to obtain multidimensional physiological fatigue feature information.
[0009] As an optional implementation, in the first aspect of the present invention, processing the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information includes: S211, Perform first feature extraction on the standardized speech dataset to obtain Mel frequency cepstral coefficients; S212, perform second feature extraction on the standardized speech dataset to obtain the Mel frequency energy map; S213, perform third feature extraction on the standardized speech dataset to obtain spectral contrast information; S214, the Mel frequency cepstral coefficients, the Mel frequency energy spectrum, and the spectral contrast information are fused to obtain multidimensional acoustic signature spectral feature vector information.
[0010] As an optional implementation, in the first aspect of the present invention, processing the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information includes: S221, Perform feature extraction on the physiological indicator information to obtain physiological indicator feature information; S222, feature extraction is performed on the biochemical indicator information to obtain biochemical indicator feature information.
[0011] As an optional implementation, in the first aspect of the present invention, the step of extracting features from the physiological indicator information to obtain physiological indicator feature information includes: S2211, Perform feature extraction on the heart rate information to obtain heart rate feature information; The expression for the heart rate feature information is: in, , For heart rate characteristic information, For heart rate information, for Length, , for The first in Sample points ; S2212, Perform feature extraction on the heart rate variability information to obtain heart rate variability feature information; The expression for the heart rate variability feature information is: In the formula, Represents the imaginary unit. express The instantaneous phase, This is information on heart rate variability. For heart rate variability information The first in Sample points ; S2213, Perform feature extraction on the electrodermal information to obtain electrodermal feature information; The expression for the skin electrodermal feature information is: in For skin electrophysiological characteristics, For skin electrical information, for Length, ; S2214, the heart rate feature information, the heart rate variability feature information and the skin conductance feature information are fused to obtain the first physiological index feature information; S2215, Perform feature extraction on the eye movement information to obtain eye movement feature information; S2216, Perform feature extraction on the blink frequency information to obtain blink frequency feature information; S2217, Process the eye movement feature information and the blink frequency feature information to obtain the second physiological index feature information; S2218, Process the first physiological indicator feature information and the second physiological indicator feature information to obtain physiological indicator feature information.
[0012] As an optional implementation, in the first aspect of the present invention, the step of training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model includes: S31, acquire physiological fatigue recognition scene information; S32, Perform scene feature modeling on the physiological fatigue recognition scene information to obtain key dimension information for fatigue recognition; S33, process the fatigue identification key dimension information and the multidimensional physiological fatigue feature information to obtain scene physiological fatigue feature information; S34, using the physiological fatigue feature information of the scene, train the preset fatigue recognition model to obtain an optimized fatigue recognition model.
[0013] A second aspect of this invention discloses a multimodal fusion physiological fatigue recognition device, the device comprising: The dataset building module is used to build standardized speech datasets and physiological information datasets; The physiological information dataset includes physiological indicator information and biochemical indicator information; the physiological indicator information includes heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicator information includes saliva information, urine information, and fingertip blood sample information; The feature extraction module is used to process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; The training module is used to train the preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model. The recognition module is used to process the multidimensional physiological fatigue feature information to be processed using the optimized fatigue recognition model, and obtain the multimodal fusion physiological fatigue recognition result.
[0014] As an optional implementation, in the second aspect of the present invention, the construction of the standardized speech dataset includes: S11, acquire voice data information; S12, perform frequency domain analysis on the speech data information to obtain speech audio domain information; The speech domain information expression is: in For audio domain information, For voice data information, For frequency domain variables, For time-domain variables; S13, process the audio domain information to obtain the first audio domain information; The expression for the first audio domain information is: S14, process the first audio domain information to obtain sampled speech information; The expression for the sampled speech information is: in, To sample voice information, For the largest spectral variable, , , , The number of sampling points. For the first speech audio domain information, for The minimum value; S15, the sampled speech information is processed to obtain a standardized speech dataset.
[0015] As an optional implementation, in the second aspect of the present invention, the processing of the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information includes: S21, Process the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information; S22, Process the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information; S23, integrate the multidimensional voiceprint spectrum feature vector information, the physiological indicator feature information and the biochemical indicator feature information to obtain multidimensional physiological fatigue feature information.
[0016] As an optional implementation, in the second aspect of the present invention, processing the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information includes: S211, Perform first feature extraction on the standardized speech dataset to obtain Mel frequency cepstral coefficients; S212, perform second feature extraction on the standardized speech dataset to obtain the Mel frequency energy map; S213, perform third feature extraction on the standardized speech dataset to obtain spectral contrast information; S214, the Mel frequency cepstral coefficients, the Mel frequency energy spectrum, and the spectral contrast information are fused to obtain multidimensional acoustic signature spectral feature vector information.
[0017] As an optional implementation, in the second aspect of the present invention, processing the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information includes: S221, Perform feature extraction on the physiological indicator information to obtain physiological indicator feature information; S222, feature extraction is performed on the biochemical indicator information to obtain biochemical indicator feature information.
[0018] As an optional implementation, in the second aspect of the present invention, the step of extracting features from the physiological indicator information to obtain physiological indicator feature information includes: S2211, Perform feature extraction on the heart rate information to obtain heart rate feature information; The expression for the heart rate feature information is: in, , For heart rate characteristic information, For heart rate information, for Length, , for The first in Sample points ; S2212, Perform feature extraction on the heart rate variability information to obtain heart rate variability feature information; The expression for the heart rate variability feature information is: In the formula, Represents the imaginary unit. express The instantaneous phase, This is information on heart rate variability. For heart rate variability information The first in Sample points ; S2213, Perform feature extraction on the electrodermal information to obtain electrodermal feature information; The expression for the skin electrodermal feature information is: in For skin electrophysiological characteristics, For skin electrical information, for Length, ; S2214, the heart rate feature information, the heart rate variability feature information and the skin conductance feature information are fused to obtain the first physiological index feature information; S2215, Perform feature extraction on the eye movement information to obtain eye movement feature information; S2216, Perform feature extraction on the blink frequency information to obtain blink frequency feature information; S2217, Process the eye movement feature information and the blink frequency feature information to obtain the second physiological index feature information; S2218, Process the first physiological indicator feature information and the second physiological indicator feature information to obtain physiological indicator feature information.
[0019] As an optional implementation, in the second aspect of the present invention, the step of training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model includes: S31, acquire physiological fatigue recognition scene information; S32, Perform scene feature modeling on the physiological fatigue recognition scene information to obtain key dimension information for fatigue recognition; S33, process the fatigue identification key dimension information and the multidimensional physiological fatigue feature information to obtain scene physiological fatigue feature information; S34, using the physiological fatigue feature information of the scene, train the preset fatigue recognition model to obtain an optimized fatigue recognition model.
[0020] A third aspect of the present invention discloses another multimodal fusion physiological fatigue recognition device, the device comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps in the multimodal fusion physiological fatigue recognition method disclosed in the first aspect of the present invention.
[0021] The fourth aspect of the present invention discloses a computer-storable medium storing computer instructions, which, when invoked, are used to execute some or all of the steps in the multimodal fusion physiological fatigue recognition method disclosed in the first aspect of the present invention.
[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: This invention discloses a multimodal fusion method for physiological fatigue recognition. The method includes constructing a standardized speech dataset and a physiological information dataset; the physiological information dataset includes physiological and biochemical indicators; processing the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model; and using the optimized fatigue recognition model to process the multidimensional physiological fatigue feature information to obtain a multimodal fusion physiological fatigue recognition result. The multimodal fusion physiological fatigue recognition method constructed in this invention has higher accuracy and better robustness and generalization of model transfer. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating a multimodal fusion physiological fatigue recognition method disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multimodal fusion physiological fatigue recognition device disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of another multimodal fusion physiological fatigue recognition device disclosed in an embodiment of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] This invention discloses a multimodal fusion physiological fatigue recognition method and apparatus. The method includes: constructing a standardized speech dataset and a physiological information dataset; the physiological information dataset includes physiological indicator information and biochemical indicator information; the physiological indicator information includes heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicator information includes saliva information, urine information, and fingertip blood sample information; processing the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model; and using the optimized fatigue recognition model to process the multidimensional physiological fatigue feature information to obtain a multimodal fusion physiological fatigue recognition result. The multimodal fusion physiological fatigue recognition method constructed by this invention has higher accuracy and better robustness and generalization of model transfer. Detailed explanations follow.
[0029] Example 1 Please see Figure 1 , Figure 1This is a flowchart illustrating a multimodal fusion-based physiological fatigue recognition method disclosed in an embodiment of the present invention. Figure 1 The described multimodal fusion physiological fatigue recognition method is applied in the fields of physiological state detection and speech recognition technology, and the embodiments of this invention are not limited thereto. Figure 1 As shown, the multimodal fusion physiological fatigue recognition method may include the following operations: S1, construct standardized speech dataset and physiological information dataset; The physiological information dataset includes physiological indicators and biochemical indicators; the physiological indicators include heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicators include saliva information, urine information, and fingertip blood sample information. S2, Process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; S3, using the multidimensional physiological fatigue feature information, train the preset fatigue recognition model to obtain an optimized fatigue recognition model; S4. Using the optimized fatigue recognition model, the multidimensional physiological fatigue feature information to be processed is processed to obtain the multimodal fusion physiological fatigue recognition result.
[0030] Optionally, constructing the standardized speech dataset includes: S11, acquire voice data information; S12, perform frequency domain analysis on the speech data information to obtain speech audio domain information; The speech domain information expression is: in For audio domain information, For voice data information, For frequency domain variables, For time-domain variables; S13, process the audio domain information to obtain the first audio domain information; The expression for the first audio domain information is: S14, process the first audio domain information to obtain sampled speech information; The expression for the sampled speech information is: in, To sample voice information, For the largest spectral variable, , , , The number of sampling points. For the first speech audio domain information, for The minimum value; S15, the sampled speech information is processed to obtain a standardized speech dataset.
[0031] The sampled speech information is processed into frames to obtain a frame sequence; The frame sequence is embedded to obtain a high-dimensional phase space matrix; Singular value decomposition is performed on the high-dimensional phase space matrix to obtain a set of singular values ordered by importance; Based on the singular values, the decomposed feature components are divided into two groups: one group is the effective signal group (corresponding to large singular values, containing the core laws of the signal), and the other group is the noise component group (corresponding to small singular values).
[0032] By utilizing the singular values and singular vectors corresponding to the effective signal groups, a reconstruction matrix containing only the information of the effective signals is reconstructed. Then, through an inverse embedding operation, the reconstruction matrix is mapped back to a one-dimensional time series, resulting in a standardized speech dataset.
[0033] Optionally, the processing of the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information includes: S21, Process the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information; Mel frequency cepstral coefficients, Chroma features, Mel frequency energy map, spectral contrast, Tonnetz features, and zero crossover rate (ZCR). The feature vectors are then fused to obtain multidimensional acoustic signature feature vector information; The fusion method is canonical correlation analysis, and this embodiment does not impose any limitations on it; S22, Process the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information; Biochemical indicators include: collecting saliva samples from volunteers 30 minutes after mouth cleaning in the morning and evening, and measuring the average cortisol concentration; Urine samples were collected from volunteers in a dark environment at night to measure melatonin concentration. Finger-prick blood samples were collected from volunteers 2 hours after exercise, and the concentrations of lactate and creatine kinase were detected using a portable biochemical analyzer or ELISA. S23, integrate the multidimensional voiceprint spectrum feature vector information, the physiological indicator feature information and the biochemical indicator feature information to obtain multidimensional physiological fatigue feature information.
[0034] Optionally, the processing of the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information includes: S211, Perform first feature extraction on the standardized speech dataset to obtain Mel frequency cepstral coefficients; S212, perform second feature extraction on the standardized speech dataset to obtain the Mel frequency energy map; S213, perform third feature extraction on the standardized speech dataset to obtain spectral contrast information; S214, the Mel frequency cepstral coefficients, the Mel frequency energy spectrum, and the spectral contrast information are fused to obtain multidimensional acoustic signature spectral feature vector information.
[0035] Optionally, the processing of the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information includes: S221, Perform feature extraction on the physiological indicator information to obtain physiological indicator feature information; S222, feature extraction is performed on the biochemical indicator information to obtain biochemical indicator feature information.
[0036] Optionally, the step of extracting features from the physiological indicator information to obtain physiological indicator feature information includes: S2211, Perform feature extraction on the heart rate information to obtain heart rate feature information; The expression for the heart rate feature information is: in, , For heart rate characteristic information, For heart rate information, for Length, , for The first in Sample points ; S2212, Perform feature extraction on the heart rate variability information to obtain heart rate variability feature information; The expression for the heart rate variability feature information is: In the formula, Represents the imaginary unit. express The instantaneous phase, This is information on heart rate variability. For heart rate variability information The first in Sample points ; S2213, Perform feature extraction on the electrodermal information to obtain electrodermal feature information; The expression for the skin electrodermal feature information is: in For skin electrophysiological characteristics, For skin electrical information, for Length, ; S2214, the heart rate feature information, the heart rate variability feature information and the skin conductance feature information are fused to obtain the first physiological index feature information; The specific fusion method expression is as follows: Norm represents standardization. , express Length; in, express function, and These are preset, learnable model parameters. Initializing with a normal distribution ensures stable variance of the initial output, preventing gradient vanishing or exploding. For bias information, tanh is the hyperbolic tangent function; The expression for the first physiological indicator characteristic information is: S2215, Perform feature extraction on the eye movement information to obtain eye movement feature information; Slow eye closing state: duration of eye closure ≥2.5s, denoted as state B; Blinking status: Eye closure time <1s, denoted as state C; Transitional closing state: Eye closure time ≤ 1s <2.5s, denoted as state D; 1) Filter out isolated false states. If there is no similar state within 5 seconds before and after a single occurrence of slow eye closing / blinking, it is judged as detection noise and is not included in the valid statistics. Transition state processing: The transitional closed state (state D) lasting 1~2.5 seconds is not included in the calculation, but is only used as an abnormal eye movement marker to assist in subsequent feature fusion; The start time of the detected state is averaged at three neighborhood points to eliminate millisecond-level time errors in the detection equipment.
[0037] 2) , This indicates the number of valid instances of the "slowly closing eyes" state within the statistics window. This is the precise start time of the first slow, closed-eye state. For the first i The precise start time of the slow, closed-eye state. For the first The exact start time of the slow, closed-eye state; 3) , For the first i The actual duration of eye closure during the slow eye-closing state; 4) , To count the effective number of blinks within the window, Standard deviation; The length of the statistical window is set experimentally, with an ideal value of 1 to 3 minutes.
[0038] Eye movement feature information The expression is: S2216, Perform feature extraction on the blink frequency information to obtain blink frequency feature information; 1) , This indicates the valid number of blinks within the statistics window. This is the exact start time of the first blink. For the first i The exact start time of the blinking state. For the first The exact start time of the blinking state; 2) Average blink duration , For the first i The actual duration of eye closure during a blink; 3) To calculate the standard deviation of the blink frequency of sub-windows within the statistical window every 10 seconds; Blink frequency characteristic information The expression is: S2217, Process the eye movement feature information and the blink frequency feature information to obtain the second physiological index feature information; The specific method is shown in S2214, and this embodiment does not impose any limitations; S2218, Process the first physiological indicator feature information and the second physiological indicator feature information to obtain physiological indicator feature information.
[0039] The specific method is shown in S2214, and this embodiment does not impose any limitations; Optionally, the step of training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model includes: S31, acquire physiological fatigue recognition scene information; Based on actual application needs, we first divide the physiological fatigue recognition scenario information into fatigue recognition scenarios, and then define exclusive information dimensions for each scenario. Physiological fatigue recognition scenario information includes in-vehicle driving, indoor office work, student learning, etc.; S32, Perform scene feature modeling on the physiological fatigue recognition scene information to obtain key dimension information for fatigue recognition; Key dimensions of fatigue identification information include basic features and derived features; Basic feature extraction: Perform statistical feature calculations on numerical fields in the physiological fatigue recognition scenario information, such as: Driving scenario: mean / variance of vehicle speed (reflecting vehicle speed stability), continuous driving duration (cumulative value); Office setting: cumulative working hours, fluctuation range of temperature and humidity.
[0040] Derivative Feature Construction: Constructing derivative scenario features strongly correlated with fatigue based on domain knowledge. Example: Driving scenario-derived features: Night driving indicator (night = 1, day = 0) × driving duration (reflecting the fatigue risk of long-term night driving); Learning scenario-derived characteristics: Task difficulty (high / medium / low) × focus duration (reflecting the rate of fatigue accumulation under high-difficulty tasks).
[0041] S33, process the fatigue identification key dimension information and the multidimensional physiological fatigue feature information to obtain scene physiological fatigue feature information; The scene's physiological fatigue feature information is obtained by processing it using a feature concatenation method. S34, using the physiological fatigue feature information of the scene, train the preset fatigue recognition model to obtain an optimized fatigue recognition model.
[0042] The fatigue recognition model consists of an input layer, an improved VGG16 convolutional model, an improved bidirectional LSTM model, and an improved fully connected layer, supplemented by a cross-module collaborative optimization layer. It achieves local extraction, temporal capture, deep fusion, and fatigue level classification of scene-specific physiological fatigue features. Specific structure: The input layer takes scene-related physiological fatigue features as input. Through dimensional expansion and standardization, it adapts to the processing needs of subsequent convolutional layers, while ensuring the alignment of timestamps between scene features and physiological features.
[0043] The improved VGG16 convolutional model features a lightweight one-dimensional structure at its core, including a LayerNorm normalization layer for the input layer, four convolutional groups (a total of eight one-dimensional convolutional layers), four one-dimensional max-pooling layers, a feature flattening layer, and a dimension compression layer. All convolutional layers employ the ReLU6 activation function, combined with BatchNorm1d and a regularization strategy that progressively increases the dropout rate. The final output... Local temporal features with 256 dimensions are adapted to subsequent LSTM modules.
[0044] The improved bidirectional LSTM model is an attention-enhanced residual bidirectional LSTM structure, built based on VGG16 output features. The core consists of two bidirectional LSTM layers (the first hidden layer has a dimension of 256, the second has a dimension of 128, and the dimension is reduced layer by layer). A temporal attention layer is added between the two LSTM layers to focus on key time points of fatigue. Residual connections are added inside each LSTM layer to alleviate the gradient vanishing problem of long sequences. At the same time, LayerNorm temporal normalization and recursive Dropout regularization are added. The output is concatenated with the output feature of the last time step through global average pooling to obtain a 512-dimensional temporal fusion feature, which is then passed to the fully connected layer.
[0045] The improved fully connected layer adopts a layered structure design to achieve deep fusion of multi-source features and scene adaptive adaptation; The cross-module collaborative optimization layer serves as the core connecting various modules. It includes an input unification and standardization layer and a feature adaptation layer between modules (dimensionality transformation + LayerNorm). At the same time, it adopts an end-to-end joint training strategy (cosine annealing learning rate) to ensure seamless connection of features between VGG16, bidirectional LSTM, and fully connected layers, thereby improving the model training stability and convergence speed.
[0046] The fatigue recognition model of this invention realizes end-to-end processing of local temporal feature extraction, long-term fatigue dependency capture, multi-source feature deep fusion and scene adaptive classification. The number of parameters is greatly reduced, the noise resistance and cross-scene recognition accuracy are significantly improved, and it is adapted to the real-time fatigue detection requirements.
[0047] Example 2 This embodiment discloses a multimodal fusion method for physiological fatigue recognition, including: Step 1: Construct a standardized speech dataset; Step 2: Extract multidimensional acoustic signature feature vectors; Step 3: Integrate multimodal metrics and label speech samples; Step 4: Train the fatigue state recognition model; Step 5: Model performance verification and evaluation.
[0048] Furthermore, Step 1 includes the following steps: Step 11: Collect voice data from a batch of volunteers; Step 12: Perform pre-emphasis processing on the voice data collected in Step 11; Step 13: Perform a short-time Fourier transform (STFT) on the data processed in Step 12. Step 14: Number the data processed in Step 13 (no.1, no.2, no.3...) and construct a standardized speech dataset.
[0049] Furthermore, Step 2 includes the following steps: Step 21: Extract the following feature parameters from the speech data, including but not limited to: Mel frequency cepstral coefficients, Chroma features, Mel frequency energy spectrum, spectral contrast, Tonnetz features, and zero crossover rate (ZCR). Step 22: Integrate the feature parameters from Step 21 to construct a feature vector matrix with a unified dimension. .
[0050] Furthermore, Step 3 includes the following steps: Step 31: Collect physiological indicators of volunteers simultaneously with Step 11. Heart rate, heart rate variability, skin conductance, eye movement signal, blink rate; Step 32: Collect biochemical indicators from volunteers simultaneously with Step 11. : Saliva, urine, and blood samples from fingertip; Step 33: Simultaneously with Step 11, collect subjective evaluation parameters of volunteers' own fatigue levels. ; Step 34: Preprocess the index parameters from Steps 31, 32, and 33 to construct a weighted fusion fatigue assessment calculation model. (1) Step 35: Calculate the comprehensive fatigue score based on the model in Step 34, and label the speech samples with different fatigue levels according to the numerical range.
[0051] Furthermore, Step 4 includes the following steps: Step 41: Randomly select 80% of the sample data from the S14 speech dataset as the training set; Step 42: Use multiple classifiers to analyze the speaker characteristics (feature vector matrix) of the training set speech. Fatigue identification modeling is performed using fatigue labels; Furthermore, Step 5 includes the following steps: Step 51: Use the remaining 20% of the sample data extracted in Step 41 as the test set input to compare the prediction accuracy of various classifier models. Step 52: Select the classifier model with the highest accuracy as the optimal model; Step 53: Combine cross-validation and grid parameter tuning to optimize the model structure. Detailed implementation method: Step 1: Collect voice recordings from no fewer than 480 volunteers. Each volunteer's recording duration is 6 seconds. Preprocess the voice recordings from Step 1 to establish a standardized voice dataset. Step 2: Extract six spectral feature parameters from the standard dataset in Step 1: Mel frequency cepstral coefficients, Chroma features, Mel frequency energy map, spectral contrast, Tonnetz features, and zero crossover rate (ZCR). Use these parameters to construct a multidimensional speech spectral feature vector. ; Step 3: Simultaneously record the other three types of indicator data of the volunteers in Step 1: physiological indicators, biochemical indicators, and subjective evaluation indicators of their own fatigue level; By integrating the above indicators, a three-category fatigue state classification model was constructed, dividing S1 volunteers into three label levels: "not fatigued", "mildly fatigued", and "moderately fatigued". Randomly select 80% of the volunteer speech data from the Step1 sample, perform three-class label matching on the selected speech, and complete the construction of the speech training set; Step 4: Based on the training set in Step 3, train a fatigue state three-classification model for speech and voiceprint features using multiple machine learning models. Step 5: Use the remaining 20% of the volunteers' voice data from Step 1 as a test set to test the model's classification accuracy, recall, and F1 score.
[0053] Furthermore, Step 1 includes the following steps: Step 11: Establish a corpus of no fewer than 480 short daily conversations (2-8 characters each); Step 12: Volunteers select one passage from Step 11 without repetition, read it aloud, and save the recording as a ".WAV" file. Step 13: Use a first-order high-pass filter to pre-emphasize the speech from Step 12; Step 14: Perform Short Time Fourier Transform (STFT) processing on the speech processed in Step 13: frame division and windowing, frame length selected as 10-30ms, window type selected as Hamming window; Step 15: Based on the processed speech from Step 14, construct a standardized speech dataset.
[0054] Furthermore, Step 3 includes the following steps: Step 31: Collect physiological indicators from volunteers, including but not limited to: heart rate, heart rate variability, skin conductance, eye movement signal, and blink frequency; Step 32: Collect saliva samples from volunteers 30 minutes after they cleanse their mouths in the morning and evening, and measure the average cortisol concentration. Step 33: Collect urine samples from volunteers at night in a dark environment to test melatonin concentration. Step 34: Collect fingertip blood samples from volunteers 2 hours after exercise, and use a portable biochemical analyzer or ELISA method to detect lactate and creatine kinase concentrations. Step 35: Use the standardized fatigue scale KSS to subjectively evaluate and record the volunteers' own fatigue level. Volunteers should score their own fatigue level on a scale of 1-9 (not tired - tired) and classify it according to the following standards: 1-3 points for not tired, 4-6 for mild fatigue, and 7-9 for moderate fatigue. Furthermore, Step 4 includes the following steps: Step 36: Filter and normalize the physiological indicators described in Step 31: Heart rate, heart rate variability, eye movement and blink data were filtered using a high-pass filter with a cutoff frequency of 0.5 Hz and a low-pass filter with a cutoff frequency of 45 Hz. The skin conduction signal data was filtered using a high-pass filter with a cutoff frequency of 0.05 Hz and a low-pass filter with a cutoff frequency of 5 Hz. The data above was normalized using the z-score method.
[0055] Step 37: Preprocess the biochemical indicators described in Steps 32-34: Normalization was performed using the z-score method, and then [the following was removed]: External outliers; Use the Shapiro-Wilk test to check for normality; if the data does not follow a normal distribution, perform a log transformation. Step 38: Record the three types of indicator data as follows: , , Based on the fatigue level mapping to specific numerical values: no fatigue = 0, mild fatigue = 1, moderate fatigue = 2; a multi-indicator weighted evaluation model is constructed, where the initial weighting coefficients are... = =0.4, =0.2, the optimized weighting coefficients can be obtained by weighting based on expert experience or by using optimization algorithms based on PCA or GridSearch.
[0056] Step 39: Based on the weighted calculation value F, divide the selected 80% of volunteers into three fatigue level labels: F < 0.5 (no fatigue), 0.5 ≤ F ≤ 1.5 (mild fatigue), and F > 1.5 (moderate fatigue). According to the individual volunteer's F value and corresponding category label, match their voice data with the corresponding category training samples and input them into the classification model. Furthermore, Step 4 includes the following steps: Step 41: Select multiple classifier models to train the speech samples, including but not limited to: Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), Multilayer Perceptron (MLP), and Gradient Boosting Tree (XGBoost). Step 42: Use the speech data of the remaining 20% of volunteers as a test set to verify the classification accuracy of the model; compare the classification accuracy of various classifier models and select the model with the highest accuracy as the optimal model. Step 43: After 5-fold cross-validation and parameter tuning, calculate the average accuracy, F1-score, and recall for each model to prevent performance fluctuations due to different training set partitions, and determine the final selected model architecture and parameters.
[0057] Example 3 Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a multimodal fusion physiological fatigue recognition device disclosed in an embodiment of the present invention. Figure 2 The described multimodal fusion physiological fatigue recognition device is applied in the fields of physiological state detection and speech recognition technology, and the embodiments of the present invention are not limited thereto. Figure 2 As shown, the multimodal fusion physiological fatigue recognition device may include the following operations: S301, Dataset Construction Module, is used to construct standardized speech datasets and physiological information datasets; The physiological information dataset includes physiological indicator information and biochemical indicator information; the physiological indicator information includes heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicator information includes saliva information, urine information, and fingertip blood sample information; S302, Feature extraction module, used to process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; S303, Training module, used to train the preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model; S304, the recognition module, is used to process the multi-dimensional physiological fatigue feature information to be processed using the optimized fatigue recognition model, and obtain the multi-modal fusion physiological fatigue recognition result.
[0058] Example 4 Please see Figure 3 , Figure 3 This is a schematic diagram of another multimodal fusion physiological fatigue recognition device disclosed in an embodiment of the present invention. Figure 3 The described multimodal fusion physiological fatigue recognition device is applied in the fields of physiological state detection and speech recognition technology, and the embodiments of the present invention are not limited thereto. Figure 3 As shown, the multimodal fusion physiological fatigue recognition device may include the following operations: Memory 401 storing executable program code; Processor 402 coupled to memory 401; The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the multimodal fusion physiological fatigue recognition method described in Embodiment 1 and Embodiment 2.
[0059] Example 5 This invention discloses a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program enables a computer to perform the steps in the multimodal fusion physiological fatigue recognition method described in Embodiments 1 and 2.
[0060] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0061] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0062] Finally, it should be noted that the multimodal fusion physiological fatigue recognition method and device disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal fusion method for physiological fatigue recognition, characterized in that, The method includes: S1, construct standardized speech dataset and physiological information dataset; The physiological information dataset includes physiological indicators and biochemical indicators; the physiological indicators include heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicators include saliva information, urine information, and fingertip blood sample information. S2, Process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; S3, using the multidimensional physiological fatigue feature information, train the preset fatigue recognition model to obtain an optimized fatigue recognition model; S4. Using the optimized fatigue recognition model, the multidimensional physiological fatigue feature information to be processed is processed to obtain the multimodal fusion physiological fatigue recognition result.
2. The multimodal fusion physiological fatigue recognition method according to claim 1, characterized in that, The construction of the standardized speech dataset includes: S11, acquire voice data information; S12, perform frequency domain analysis on the speech data information to obtain speech audio domain information; The speech domain information expression is: in For audio domain information, For voice data information, For frequency domain variables, For time-domain variables; S13, process the audio domain information to obtain the first audio domain information; The expression for the first audio domain information is: S14, process the first audio domain information to obtain sampled speech information; The expression for the sampled speech information is: in, To sample voice information, For the largest spectral variable, , , , The number of sampling points. For the first speech audio domain information, for The minimum value; S15, the sampled speech information is processed to obtain a standardized speech dataset.
3. The multimodal fusion physiological fatigue recognition method according to claim 1, characterized in that, The process of processing the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information includes: S21, Process the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information; S22, Process the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information; S23, integrate the multidimensional voiceprint spectrum feature vector information, the physiological indicator feature information and the biochemical indicator feature information to obtain multidimensional physiological fatigue feature information.
4. The multimodal fusion physiological fatigue recognition method according to claim 3, characterized in that, The process of processing the standardized speech dataset to obtain multidimensional speaker spectrum feature vector information includes: S211, Perform first feature extraction on the standardized speech dataset to obtain Mel frequency cepstral coefficients; S212, perform second feature extraction on the standardized speech dataset to obtain the Mel frequency energy map; S213, perform third feature extraction on the standardized speech dataset to obtain spectral contrast information; S214, the Mel frequency cepstral coefficients, the Mel frequency energy spectrum, and the spectral contrast information are fused to obtain multidimensional acoustic signature spectral feature vector information.
5. The multimodal fusion physiological fatigue recognition method according to claim 3, characterized in that, The process of processing the physiological information dataset to obtain physiological indicator feature information and biochemical indicator feature information includes: S221, Perform feature extraction on the physiological indicator information to obtain physiological indicator feature information; S222, feature extraction is performed on the biochemical indicator information to obtain biochemical indicator feature information.
6. The multimodal fusion physiological fatigue recognition method according to claim 5, characterized in that, The step of extracting features from the physiological indicator information to obtain physiological indicator feature information includes: S2211, Perform feature extraction on the heart rate information to obtain heart rate feature information; The expression for the heart rate feature information is: in, , For heart rate characteristic information, For heart rate information, for Length, , for The first in Sample points ; S2212, Perform feature extraction on the heart rate variability information to obtain heart rate variability feature information; The expression for the heart rate variability feature information is: In the formula, Represents the imaginary unit. express The instantaneous phase, This is information on heart rate variability. For heart rate variability information The first in Sample points ; S2213, Perform feature extraction on the electrodermal information to obtain electrodermal feature information; The expression for the skin electrodermal feature information is: in For skin electrophysiological characteristics, For skin electrical information, for Length, ; S2214, the heart rate feature information, the heart rate variability feature information and the skin conductance feature information are fused to obtain the first physiological index feature information; S2215, Perform feature extraction on the eye movement information to obtain eye movement feature information; S2216, Perform feature extraction on the blink frequency information to obtain blink frequency feature information; S2217, Process the eye movement feature information and the blink frequency feature information to obtain the second physiological index feature information; S2218, Process the first physiological indicator feature information and the second physiological indicator feature information to obtain physiological indicator feature information.
7. The multimodal fusion physiological fatigue recognition method according to claim 1, characterized in that, The step of training a preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model includes: S31, acquire physiological fatigue recognition scene information; S32, Perform scene feature modeling on the physiological fatigue recognition scene information to obtain key dimension information for fatigue recognition; S33, process the fatigue identification key dimension information and the multidimensional physiological fatigue feature information to obtain scene physiological fatigue feature information; S34, using the physiological fatigue feature information of the scene, train the preset fatigue recognition model to obtain an optimized fatigue recognition model.
8. A multimodal fusion physiological fatigue recognition device, characterized in that, The device includes: The dataset building module is used to build standardized speech datasets and physiological information datasets; The physiological information dataset includes physiological indicator information and biochemical indicator information; the physiological indicator information includes heart rate information, heart rate variability information, skin conductance information, eye movement information, and blink frequency information; the biochemical indicator information includes saliva information, urine information, and fingertip blood sample information; The feature extraction module is used to process the standardized speech dataset and the physiological information dataset to obtain multidimensional physiological fatigue feature information; The training module is used to train the preset fatigue recognition model using the multidimensional physiological fatigue feature information to obtain an optimized fatigue recognition model. The recognition module is used to process the multidimensional physiological fatigue feature information to be processed using the optimized fatigue recognition model, and obtain the multimodal fusion physiological fatigue recognition result.
9. A multimodal fusion physiological fatigue recognition device, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the multimodal fusion physiological fatigue recognition method as described in any one of claims 1-7.
10. A computer-storable medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the multimodal fusion physiological fatigue recognition method as described in any one of claims 1-7.