Pilot non-contact physiological status assessment method based on signal fusion
By constructing a multi-source signal fusion network of rPPG signal, sound and expression signal feature extraction model, and combining with a deep learning model, non-contact assessment of pilot physiological status is achieved, and the problem of lack of real-time monitoring tools and inconvenient use of contact equipment in the prior art is solved, and the accuracy and reliability of the evaluation are improved.
Patent Information
- Application Number
- CN202510503111.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The prior art lacks real-time monitoring tools in the evaluation of pilot physiological status, and existing contact equipment is inconvenient to use, is easily lost, and affects measurement accuracy and convenience.
Using a signal fusion-based method, a rPPG signal feature extraction model, a sound and expression fusion signal feature extraction model is constructed, and a multi-source signal fusion network of a deep learning model is combined to achieve contactless assessment of pilot physiological status.
It improves the accuracy and reliability of physiological status assessment, provides more comprehensive and accurate information, and is suitable for integration into aircraft on-board equipment to achieve real-time supervision and recording of pilots.
Smart Images

Figure CN120021955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multiple technical fields such as computer vision, image processing and physiological signal detection, and in particular to a pilot non-contact physiological state assessment method based on signal fusion. Background Art
[0002] Remote Photoplethysmography (rPPG) is a technology that uses sensors such as cameras to capture the periodic changes in skin color caused by the cardiac cycle. By capturing these tiny color changes, the blood volume pulse signal (BVP signal) can be extracted to measure physiological indicators related to the cardiac cycle, such as heart rate (HR), respiratory rate (RR), and heart rate variability (HRV). rPPG technology plays a very important role in scenarios that require non-contact physiological parameter monitoring. It can be used in the medical field, home health monitoring, fatigue driving prevention, human-computer interaction and other scenarios.
[0003] Heart rate variability is an important indicator of autonomic nervous system activity. By analyzing the time and frequency domain characteristics of rPPG signals, the trend of heart rate variability can be evaluated to understand the individual's stress response and health status. Audio signals such as pitch, volume, and speech speed can reflect an individual's emotional state, fatigue level, and psychological state. For example, a low voice or a decrease in volume may indicate that the individual is in a state of fatigue or tension. Video images can also provide rich facial expression information. Facial expression features, such as smiling, frowning, and blinking, can determine an individual's emotional state. Therefore, combining the three modalities of rPPG signals, audio signals, and expression signals for physiological state assessment can obtain more comprehensive and accurate information.
[0004] Currently, airlines and flight training institutions lack tools to monitor pilots' heart rates in real time, and mainly rely on regular physical examinations and pre-flight psychological analysis at aviation hospitals to assess pilots' mental health. Most of the existing physiological status analysis equipment is portable and wearable, and needs to be worn on the earlobes, fingers or wrists, etc., which are rich in blood perfusion. It can accurately reflect the changes in the heart beat cycle and calculate the trend of heart rate variability (HRV). However, these contact devices have problems such as inconvenience in use, easy loss and psychological resistance, which affect the measurement accuracy and convenience. Therefore, heart rate monitoring equipment is gradually changing from wearable to remote sensing. Remote sensing equipment has the characteristics of non-contact, easy installation and strong concealment. It is very suitable for integration into aircraft onboard equipment to achieve real-time supervision and recording of pilots throughout the flight. As a result, remote sensing non-contact heart rate monitoring equipment has shown a high application demand and broad development potential in the civil aviation field. Summary of the invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a pilot non-contact physiological state assessment method based on signal fusion to improve the accuracy and reliability of physiological state assessment.
[0006] The objective of the present invention is achieved through the following technical solutions: A non-contact physiological status assessment method for a pilot based on signal fusion, comprising: Step 1: construct an rPPG signal feature extraction model, and use contrastive learning as a self-supervised learning scheme to train the rPPG signal feature extraction model; use the trained rPPG signal feature extraction model to extract rPPG signal features from the video; Step 2: construct a sound and expression fusion signal feature extraction model and perform pre-training, and use the trained sound and expression fusion signal feature extraction model to extract sound and expression signal features from the video; Step 3: construct a multi-source signal fusion network based on a deep learning model, take the rPPG signal features and the sound and expression signal features as input, and perform continuous evaluation of the physiological state of the human body.
[0007] Preferably, the step 1 specifically includes: The videos in the training set , M represents the total number of video samples, and sparse time enhancement and random horizontal flipping are used to enhance the data to generate two enhanced clips ; The two enhanced clips are input into the video encoder VVT to generate rPPG features corresponding to different videos. and ; rPPG features and Send it to an MLP projection head and get the projection vectors and ; Contrastive learning is performed on the projection vector. For a pair of positive samples in the feature learning stage , using the following contrastive loss: ;in is an indicator function, if Then the indicator function It means 1, otherwise it means 0. k represents the index of the sample, is the temperature hyperparameter, Represents the calculation of the projection vector and the projection vector The cosine similarity between ; In the testing phase, the rPPG features are obtained by passing the test set video through the video encoder VVT. Directly pass it to the rPPG estimator to obtain the rPPG signal features r , and its calculation formula is ; where rPPGExtractor represents the rPPG estimator.
[0008] Preferably, the process of obtaining the rPPG signal features through the video encoder VVT specifically includes: Video As the input of the video encoder VVT, T is the number of frames, W is the width of each video frame, H is the height of each video frame, and C is the number of channels; Use pipeline embedding to embed videos V Convert to sub-video block sequence ,in Represents a pipeline, t represents the time length of the pipeline, h Indicates the height of the pipe, w Indicates the width of the pipe. c Indicates the number of channels of the pipeline, N is the total number of pipelines, Represents space; the total number of pipes N The calculation method is: ,in and They represent the number of blocks divided along the time dimension, height, width and channel dimension respectively, and the corresponding calculation formulas are and ; Then the pipeline is connected through three-dimensional convolution Rasterize into feature blocks and linearly project all feature blocks into space Then add the position embedding vector including the position inductive bias , get the feature block embedding sequence , which is calculated as ,in E represents a three-dimensional convolution, LP represents the linear projection operation, which embeds the feature block into the sequence through the self-attention layer Change to self-attention enhancement feature , self-attention enhanced features through transformer neural network Perform feature extraction and use average pooling and flattening to obtain rPPG features .
[0009] Preferably, the step 2 specifically includes: Split the video into windows of size aand a fragment of step size s: given a window size a and step length s, ,have n The video frames will be divided into fragments, of which Video clips Contains frames, specifically a continuous sequence from the start frame to the end frame , where F represents the video frame; Extract video segments using a pre-trained audio feature extractor Audio features , extract video clips using a pre-trained expression feature extractor facial features ; The audio features are and facial features Input into the temporal convolutional network TCN for temporal encoding, the formula is as follows: ; ;in, represents the spatiotemporal audio features, Represents spatiotemporal expression features; connects spatiotemporal expression features and spatiotemporal audio features to obtain spatiotemporal features ; Use Transformer encoder to learn the association between sound and image in the video and generate emotional features of the characters , expressed as follows: , and then the emotional characteristics of the characters in all the clips Splice into a sequence to get the emotional signal characteristics e .
[0010] Preferably, the step 3 specifically includes: The rPPG signal characteristics r and voice and emotional signal characteristics e As input, the spatial features are extracted by convolutional neural network (CNN), and then the temporal information of the features is extracted by long short-term memory network (LSTM) to obtain the spatiotemporal signal. and ; Design two identical encoders and To compare the time and space signals on the two branches and To encode: ; ; and The two features are extracted from the spatiotemporal signals of the two branches, and the two features are fused by superposition, and finally passed through the decoder Decode to get the fused features: ;in, is the final fusion signal, represents the superposition of features, L Indicates the length of the fused signal; the final fused signal The activity level of a person's physiological state is predicted through a multi-layer perceptron and input into a regression model.
[0011] The beneficial effects of the present invention are: The present invention analyzes video clips containing human faces, extracts rPPG signal features, sound and expression signal features, and then realizes the feature fusion of these three different modal information. Combining the three modalities of rPPG signal, audio signal and expression signal to evaluate physiological state can obtain more comprehensive and accurate information. The multimodal fusion analysis method not only improves the accuracy and reliability of physiological state evaluation, but also provides strong support for subsequent health management, disease diagnosis, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is the architecture diagram of the rPPG signal feature extraction model based on video vision Transformer; Figure 2 Schematic diagram of the self-supervised learning framework; Figure 3 The architecture diagram of the model for extracting features from voice and expression fusion signals; Figure 4 This is a diagram of the multi-source signal fusion network architecture based on the deep learning model. DETAILED DESCRIPTION
[0013] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0014] See also Figure 1-Figure 4 , the present invention provides a technical solution: A non-contact physiological status assessment method for a pilot based on signal fusion, comprising: Step 1: Build an rPPG signal feature extraction model, use contrastive learning as a self-supervised learning scheme to train the rPPG signal feature extraction model; use the trained rPPG signal feature extraction model to extract rPPG signal features from the video. The schematic diagram of the self-supervised learning framework is shown in the figure. Figure 2 shown.
[0015] In this embodiment, the step 1 specifically includes: Step 11: Put the videos in the training set , M represents the total number of video samples, and sparse time enhancement and random horizontal flipping are used to enhance the data to generate two enhanced clips .
[0016] Step 12: Input the two enhanced clips into the video encoder VVT to generate rPPG features corresponding to different videos and .
[0017] Step 13: rPPG features and Send it to an MLP projection head and get the projection vectors and .
[0018] Step 14: Perform contrastive learning on the projection vector. For a pair of positive samples in the feature learning stage , using the following contrastive loss: ;in is an indicator function, if Then the indicator function It means 1, otherwise it means 0. k represents the index of the sample, is the temperature hyperparameter, Represents the calculation of the projection vector and the projection vector ; this loss brings positive pairs (i.e., augmented images from the same input) close together in feature space and moves negative pairs away from each other.
[0019] Step 15: During the testing phase, the rPPG features are obtained by passing the video in the test set through the video encoder VVT. Directly pass it to the rPPG estimator to obtain the rPPG signal features r : ; Among them, rPPGExtractor represents the rPPG estimator.
[0020] like Figure 1 As shown, the rPPG characteristics , d’ Represents rPPG characteristics The process of obtaining the rPPG signal feature through the video encoder VVT specifically includes: Video As the input of the video encoder VVT, T is the number of frames, W is the width of each video frame, H is the height of each video frame, and C is the number of channels; Use pipeline embedding to embed videos V Convert to sub-video block sequence ,in Represents a pipeline, t represents the time length of the pipeline, h Indicates the height of the pipe, w Indicates the width of the pipe. c Indicates the number of channels of the pipeline, N is the total number of pipelines, Represents space; the total number of pipes N The calculation method is: ,in and They represent the number of blocks split along the time dimension, height, width and channel dimension (rounded down), and the corresponding calculation formulas are and ; The space for a single pipe is ; Then the pipeline is connected through three-dimensional convolution Rasterize into feature blocks and linearly project all feature blocks into space Then add the position embedding vector including the position inductive bias , get the feature block embedding sequence , which is calculated as ,in E represents a three-dimensional convolution, LP represents the linear projection operation, which embeds the feature block into the sequence through the self-attention layer Change to self-attention enhancement feature , self-attention enhanced features through transformer neural network Perform feature extraction and use average pooling and flattening to obtain rPPG features rPPG Features Effectively represents the spatiotemporal information of RGB video.
[0021] In the task of extracting rPPG signals from videos, the video sequence should be considered as a signal sequence problem of context clues. The present invention uses an end-to-end video vision Transformer (VVT) structure to extract remote local and global spatiotemporal features from the video, and finally obtains rPPG signal features, which can better reflect the rPPG information of the characters in the video. When training the rPPG extraction model, contrastive learning is used as a self-supervised learning (SSL) solution to solve the problem of scarce data sets and difficulty in capturing higher-level semantic information in the rPPG estimation task.
[0022] Step 2: construct a sound and expression fusion signal feature extraction model and perform pre-training, and use the trained sound and expression fusion signal feature extraction model to extract sound and expression signal features from the video.
[0023] In this embodiment, the step 2 specifically includes: Step 21: Split the video into windows of different sizes a and a fragment of step size s: given a window size a and step length s, ,have n The video frames will be divided into fragments, of which Video clips Contains frames, specifically a continuous sequence from the start frame to the end frame , where F represents the video frame; Step 22: Extract video clips using pre-trained audio feature extractor Audio features , extract video clips using a pre-trained expression feature extractor facial features ; Step 23: The audio features and facial features Input into the temporal convolutional network TCN for temporal encoding, the formula is as follows: ; ;in, represents the spatiotemporal audio features, Represents spatiotemporal expression characteristics; Step 24: Connect the spatiotemporal expression features and the spatiotemporal audio features to obtain the spatiotemporal features ; Step 25: Use the Transformer encoder to learn the association between sound and image in the video and generate emotional features of the characters , expressed as follows: , and then the emotional characteristics of the characters in all the clips Splice into a sequence to get the emotional signal characteristics e .
[0024] The architecture diagram of the sound and expression fusion signal feature extraction model is as follows Figure 3 As shown in the figure, the video contains facial expressions and voice information, which carry the emotional characteristics of the characters. Combining the information of these two modalities to extract the characteristics of sound and expression signals can improve the robustness of the model for emotion perception.
[0025] Step 3: construct a multi-source signal fusion network based on a deep learning model, take the rPPG signal features and the sound and expression signal features as input, and perform continuous evaluation of the physiological state of the human body.
[0026] In this embodiment, step 3 specifically includes: Step 31: The rPPG signal feature r and voice and emotional signal characteristics e As input, the spatial features are extracted by convolutional neural network (CNN), and then the temporal information of the features is extracted by long short-term memory network (LSTM) to obtain the spatiotemporal signal. and ; Step 32: Design two identical encoders and To compare the time and space signals on the two branches and To encode: ; ; and These are the spatiotemporal signals extracted by the two branches; Step 33: The two features are superimposed to fuse the features, and finally the decoder is used to Decode to get the fused features: ;in, is the final fusion signal, represents the superposition of features, L Indicates the length of the fusion signal; Step 34: Final fusion signal The activity level of a person's physiological state is predicted through a multi-layer perceptron and input into a regression model.
[0027] like Figure 4 As shown, the present invention designs a multi-source signal fusion network based on a deep learning model to fuse multi-source information and continuously evaluate the physiological state of the human body.
[0028] The embodiment of the present invention provides a non-contact human physiological state assessment method combining rPPG signals, sound and expression signals. The method analyzes video clips containing human faces, extracts rPPG signal features, sound and expression signal features, and then realizes feature fusion of these three different modal information. In order to further improve the fusion efficiency and effect, a deep regression model based on multi-source signal fusion is designed, which can continuously and real-timely assess the physiological state of pilots.
[0029] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. A non-contact physiological status assessment method for pilots based on signal fusion, characterized in that: include: Step 1: construct an rPPG signal feature extraction model, and use contrastive learning as a self-supervised learning scheme to train the rPPG signal feature extraction model; use the trained rPPG signal feature extraction model to extract rPPG signal features from the video; Step 2: construct a sound and expression fusion signal feature extraction model and perform pre-training, and use the trained sound and expression fusion signal feature extraction model to extract sound and expression signal features from the video; Step 3: construct a multi-source signal fusion network based on a deep learning model, take the rPPG signal features and the sound and expression signal features as input, and perform continuous evaluation of the physiological state of the human body.
2. The method for non-contact physiological status assessment of a pilot based on signal fusion according to claim 1, characterized in that: The step 1 specifically includes: The videos in the training set (j ), M represents the total number of video samples, and sparse time enhancement and random horizontal flipping are used to enhance the data to generate two enhanced clips ; The two enhanced clips are input into the video encoder VVT to generate rPPG features corresponding to different videos. and ; rPPG features and Send it to an MLP projection head and get the projection vectors and ; Contrastive learning is performed on the projection vector. For a pair of positive samples in the feature learning stage ( ), using the following contrastive loss: ;in is an indicator function, if The indicator function It means 1, otherwise it means 0. k represents the index of the sample, is the temperature hyperparameter, Represents the calculation of the projection vector and the projection vector The cosine similarity between ; In the testing phase, the rPPG features are obtained by passing the test set video through the video encoder VVT. Directly pass it to the rPPG estimator to obtain the rPPG signal features , and its calculation formula is ; where rPPGExtractor represents the rPPG estimator.
3. The pilot non-contact physiological status assessment method based on signal fusion according to claim 2 is characterized by: The process of obtaining the rPPG signal features through the video encoder VVT specifically includes: Video As the input of the video encoder VVT, T is the number of frames, W is the width of each video frame, H is the height of each video frame, and C is the number of channels; Use pipeline embedding to embed videos Convert to sub-video block sequence ,in Represents a pipeline, t represents the time length of the pipeline, h Indicates the height of the pipe, w Indicates the width of the pipe. c Indicates the number of channels of the pipeline, N is the total number of pipelines, Represents space; the total number of pipes N The calculation method is: ,in , , and They represent the number of blocks split along the time dimension, height, width and channel dimension (rounded down), and the corresponding calculation formulas are , , and ; Then the pipeline is connected through three-dimensional convolution Rasterize into feature blocks and linearly project all feature blocks into space Then add the position embedding vector including the position inductive bias , get the feature block embedding sequence , which is calculated as ,in E represents a three-dimensional convolution, LP represents the linear projection operation, which embeds the feature block into the sequence through the self-attention layer Change to self-attention enhancement feature , self-attention enhanced features through transformer neural network Perform feature extraction and use average pooling and flattening to obtain rPPG features .
4. The method for non-contact pilot physiological status assessment based on signal fusion according to claim 3 is characterized by: The step 2 specifically includes: Split the video into windows of size a and step length A snippet of: Given a window size a and step length , ,have The video frames will be divided into fragments, of which Video clips Contains frames, specifically a continuous sequence from the start frame to the end frame , F represents the video frame; Extract video segments using a pre-trained audio feature extractor Audio features , extract video clips using a pre-trained expression feature extractor facial features ; The audio features are and facial features Input into the temporal convolutional network TCN for temporal encoding, the formula is as follows: ; ;in, represents the spatiotemporal audio features, Represents spatiotemporal expression features; connects spatiotemporal expression features and spatiotemporal audio features to obtain spatiotemporal features ; Use Transformer encoder to learn the association between sound and image in the video and generate emotional features of the characters , expressed as follows: , and then the emotional characteristics of the characters in all the clips Splice into a sequence to get the emotional signal characteristics e .
5. The method for non-contact pilot physiological status assessment based on signal fusion according to claim 4, characterized in that: The step 3 specifically includes: The rPPG signal characteristics and voice and emotional signal characteristics e As input, the spatial features are extracted by convolutional neural network (CNN), and then the temporal information of the features is extracted by long short-term memory network (LSTM) to obtain the spatiotemporal signal. and ; Design two identical encoders , To compare the time and space signals on the two branches and To encode: ; ; and The two features are extracted from the spatiotemporal signals of the two branches, and the two features are fused by superposition, and finally passed through the decoder Decode to get the fused features: ;in, is the final fusion signal, represents the superposition of features, L Indicates the length of the fused signal; the final fused signal The activity level of a person's physiological state is predicted through a multi-layer perceptron and input into a regression model.
Citation Information
Patent Citations
Voice-and-facial-expression-based identification method and system for dual-modal emotion fusion
CN105976809A
Non-contact heart rate measurement method based on space-time attention network and input optimization
CN113343821A
Sound event detection method and device, equipment and storage medium
CN114882911A
Non-contact multi-modal physiological signal detection method based on self-supervision and lifelong learning
CN115497143A
Emotion detection method and device based on multimedia signals
CN115844403A
Cited By
Non-contact athlete physiological status assessment method and system based on feature fusion and medium
CN120616480A