Non-contact Psychological Stress Detection Method and System Based on Visual Perception

By obtaining the gaze angle and head posture data in the facial visible light video, combined with the VSDNet model of the two-stage learning strategy, the problem of insufficient accuracy of single-dimensional data detection is solved, and more accurate psychological stress detection is achieved.

CN118743551BActive Publication Date: 2025-07-08HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410918045.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-07-08
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

In the prior art, psychological stress detection is only relied on single-dimensional data, resulting in extremely limited stress information obtained and insufficient classification accuracy of the detection model.

Method used

By obtaining the user's facial visible light video, extracting the gaze angle timing data and head posture timing data, performing standardization and gradient calculations, and combining the remote photoelectric volume pulse wave signal data, a VSDNet model is constructed using a two-stage learning strategy for psychological stress detection.

Benefits of technology

The fusion analysis of physiological and behavioral data is realized, providing more comprehensive and accurate information on individual psychological stress status, and improving the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118743551B_ABST
    Figure CN118743551B_ABST
Patent Text Reader

Abstract

The present invention provides a non-contact psychological stress detection method, system, storage medium and electronic device based on visual perception, which relates to the field of non-contact psychological stress detection. In the present invention, by combining rPPG data, eye gaze angle gradient data and head movement posture gradient data, the fusion analysis of physiological and behavioral data is realized, and more comprehensive and accurate individual state information can be provided, thereby improving the accuracy and reliability of stress detection. In addition, a VSDNet model is constructed by combining a two-stage learning strategy. In the first stage, the ability of the model to obtain general stress features is trained by judging the similarity of input sample pairs. In the second stage, the pre-training result of the first stage is fine-tuned and the model inference layer is extended to improve the ability of the model to obtain high-level stress features; compared with single-stage detection, the classification performance of the psychological stress detection model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of non-contact psychological stress detection, and particularly to a non-contact psychological stress detection method, system, storage medium and electronic device based on visual perception. Background Art

[0002] Non-contact psychological stress detection includes non-contact collection of detection signals through relevant devices for psychological stress analysis.

[0003] In related technologies, the common problem is to rely only on single-dimensional data for detection. For example, Patent CN112597949A discloses a method and system for measuring psychological stress based on video. Among them, the method for measuring psychological stress based on video includes the following steps: receiving a video to be measured; using a face detection algorithm to detect faces in the video to be measured, and after detecting a face, selecting face feature points; dynamically selecting the position of the region of interest according to the face feature points; calculating the brightness values of each channel of pixels within the position of the region of interest to obtain a PPG waveform; preprocessing the PPG waveform to obtain a processed PPG waveform; processing the processed PPG waveform to extract multiple features of the processed PPG waveform; inputting the multiple features of the processed PPG waveform into a pre-trained machine learning model for cognitive load degree classification, and giving a classification result to achieve cognitive load measurement.

[0004] However, the method of relying only on single-dimensional data for detection as described above can obtain extremely limited stress information, resulting in insufficient classification accuracy of the detection model. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a non-contact psychological stress detection method, system, storage medium and electronic device based on visual perception, and solves the technical problem that the stress information obtained by the method of relying only on single-dimensional data for detection is extremely limited.

[0007] (2) Technical Solutions

[0008] To achieve the above object, the present invention is realized through the following technical solutions:

[0009] A non-contact psychological stress detection method based on visual perception includes:

[0010] Obtaining a visible light video of the user's face;

[0011] Respectively extracting gaze angle time series data and head pose time series data from the visible light video of the face, and performing normalization and gradient calculation processing;

[0012] Cut the face image sequence based on the facial visible light video to extract the initial rppg signal data, and perform detrending, filtering, and normalization processing;

[0013] Use the normalized rppg data, gaze angle gradient data, and head pose gradient data as the input of the VSDNet model constructed based on the two-stage learning strategy to detect whether the user is stressed or not.

[0014] Preferably, the process of constructing the VSDNet model based on the two-stage learning strategy includes:

[0015] Based on the historical facial visible light videos of the user, obtain the historical normalized rppg data, gaze angle gradient data, and head pose gradient data respectively;

[0016] Slice the historical facial visible light videos based on a preset sampling interval, construct a single sample with the historical normalized rppg data, gaze angle gradient data, head pose gradient data, and corresponding stress labels within each preset sampling interval, and match different samples pairwise to form sample pairs;

[0017] Set similar labels respectively based on the similarities and differences of the stress labels within each sample pair; where the similar labels include similar and dissimilar;

[0018] Train the first-stage VSDNet model based on the sample pairs with similar labels; where the core network of the first-stage VSDNet model is ExpLearnNet;

[0019] Fine-tune the ExpLearnNet of the first-stage VSDNet model, use the samples with stress labels as the input, and train to obtain the second-stage VSDNet model.

[0020] Preferably, the first-stage VSDNet model includes a Siamese network ExpLearnNet, an L1 distance calculation layer, and a sigmoid function output layer; where the ExpLearnNet is a one-dimensional CNN and BiLSTM Siamese network with shared parameters for two groups, and all CNN layers apply the LeakyReLU activation function network;

[0021] The training of the first-stage VSDNet model based on the sample pairs with similar labels includes:

[0022] For any sample in the sample pair, use its corresponding historical standardized rppg data, gaze angle gradient data, and head pose gradient data as the input of a one-dimensional CNN. Perform two one-dimensional convolutions and max pooling for downsampling respectively, and perform multi-channel feature fusion to obtain fused features; use the fused features as the input of a BiLSTM siamese network to obtain the final output that captures long-term dependencies in the data;

[0023] Use the two final outputs of the sample pair as the input of the L1 distance calculation layer to obtain the L1 distance between the two samples, and then use it as the input of the sigmoid function output layer to obtain the similarity prediction result;

[0024] Based on the similarity label and its similarity prediction result of each sample pair, construct a first loss function; minimize the first loss function, and after convergence, obtain the VSDNet model in the first stage.

[0025] Preferably, for the ExpLearnNet that fine-tunes the VSDNet model in the first stage, use the sample with a stress label as the input and train to obtain the VSDNet model in the second stage; including:

[0026] For the ExpLearnNet of the VSDNet model obtained after convergence in the first stage, sequentially add a first fully connected layer, a Dropout layer, and a second fully connected layer after it, and select softmax regression as the activation function of the output layer after the second fully connected layer; where the LeakyReLU activation function is applied in the first fully connected layer;

[0027] Use the sample with a stress label as the input to obtain the stress prediction result;

[0028] Based on the stress label and its stress prediction result of each sample, construct a second loss function; minimize the second loss function, and after convergence, obtain the VSDNet model in the second stage.

[0029] Preferably, the gaze angle time series data is any one or any combination of the gaze direction vectors of the leftmost eye along the x-axis, the gaze direction vectors of the leftmost eye along the y-axis, the gaze direction vectors of the leftmost eye along the z-axis, the gaze direction vectors of the rightmost eye along the x-axis, the gaze direction vectors of the rightmost eye along the y-axis, and the gaze direction vectors of the rightmost eye along the z-axis in the image; where the x-axis, y-axis, and z-axis respectively include the three coordinate axes in the three-dimensional coordinate system constructed by the spatial Cartesian coordinates.

[0030] Preferably, the head pose time series data is any one or a combination of any several of the deflection of the head along the x-axis, the deflection of the head along the y-axis, and the deflection of the head along the z-axis; where the x-axis, y-axis, and z-axis respectively include three coordinate axes in a three-dimensional coordinate system constructed by spatial Cartesian coordinates.

[0031] Preferably, the extraction process of the initial rPPG signal data includes:

[0032] Using a skin color adaptive segmentation method based on the YCbCr color space to perform skin segmentation on the facial image sequence of the person;

[0033] Using a projection plane orthogonal skin algorithm to extract the initial rPPG signal data from the facial image sequence after skin segmentation.

[0034] A non-contact psychological stress detection system based on visual perception includes:

[0035] A video acquisition module for acquiring the visible light video of the user's face;

[0036] A data acquisition module for respectively extracting the gaze angle time series data and the head pose time series data from the visible light video of the face, and performing normalization and gradient calculation processing;

[0037] And for cutting the facial image sequence based on the visible light video of the face to extract the initial rPPG signal data, and performing detrending, filtering, and normalization processing;

[0038] A stress detection module for using the normalized rPPG data, gaze angle gradient data, and head pose gradient data as the input of a VSDNet model constructed based on a two-stage learning strategy to detect whether the user is stressed or not

[0039] A storage medium stores a computer program for non-contact psychological stress detection based on visual perception, wherein the computer program causes a computer to execute the non-contact psychological stress detection method as described above.

[0040] An electronic device includes:

[0041] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the non-contact psychological stress detection method as described above.

[0042] (III) Beneficial effects

[0043] The present invention provides a non-contact psychological stress detection method, system, storage medium and electronic device based on visual perception. Compared with the prior art, the present invention has the following beneficial effects:

[0044] In the present invention, first, a visible light video of the user's face is acquired; secondly, gaze angle time series data and head pose time series data are respectively extracted from the visible light video of the face, and are subjected to normalization and gradient calculation processing; at the same time, a face image sequence is cut based on the visible light video of the face to extract initial rppg signal data, and detrending, filtering and normalization processing are performed; finally, the normalized rppg data, gaze angle gradient data and head pose gradient data are used as the inputs of the VSDNet model constructed based on a two-stage learning strategy to detect whether the user is stressed or not. By combining rPPG data, eye gaze angle gradient data and head movement posture gradient data, the fusion analysis of physiological and behavioral data is realized, and more comprehensive and accurate individual state information can be provided, thereby improving the accuracy and reliability of stress detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0046] Figure 1 It is a block diagram of a non-contact psychological stress detection method based on visual perception provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic structural diagram of the VSDNet model in the first stage provided by an embodiment of the present invention;

[0048] Figure 3 It is a schematic structural diagram of a network structure of ExpLearnNet provided by an embodiment of the present invention;

[0049] Figure 4 It is a schematic structural diagram of the VSDNet model in the second stage provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0051] In an embodiment of the present application, by providing a non-contact psychological stress detection method, system, storage medium, and electronic device based on visual perception, the technical problem that the stress information that can be obtained by a method relying only on single-dimensional data is extremely limited is solved.

[0052] The overall idea of the technical solution in the embodiment of the present application to solve the above technical problem is as follows:

[0053] An embodiment of the present invention proposes a non-contact psychological stress detection method based on visual perception, which fuses physiological behavior data obtained from facial videos, including remote photoplethysmograph (rppg) data, eye gaze angle gradient data, and head movement posture gradient data, to identify the psychological stress state of an individual.

[0054] In addition, the stress detection in the embodiment of the present invention is divided into two stages:

[0055] In the first stage, a VSDNet (Video-based Stress Detection Network) model is designed to determine whether the input data sample pairs are similar, and the model is trained to obtain the ability to acquire general physiological behavior stress characteristics.

[0056] In the second stage, the VSDNet model is inherited and fine-tuned for the stress state detection task. On the basis of the first stage, further learning and reasoning of the model are realized, so that the model has the ability to learn and extract higher-level stress characteristics, thereby more accurately identifying the psychological stress state of an individual, which has a certain significance for the long-term intelligent detection of psychological stress and disease prevention, and assisting managers in personnel health management and decision-making.

[0057] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0058] Embodiment 1:

[0059] As Figure 1 shown, an embodiment of the present invention provides a non-contact psychological stress detection method based on visual perception, including:

[0060] S1. Obtain the visible light video of the user's face;

[0061] S21. Extract the gaze angle time series data and head pose time series data from the visible light video of the face respectively, and perform standardization and gradient calculation processing;

[0062] S22. Cut the face image sequence based on the facial visible light video to extract the initial rPPG signal data, and perform detrending, filtering, and normalization processing;

[0063] S3. Use the normalized rPPG data, gaze angle gradient data, and head pose gradient data as the input of the VSDNet model constructed based on the two-stage learning strategy to detect whether the user is stressed or not.

[0064] Through the embodiments of the present invention, by combining rPPG data, eye gaze angle gradient data, and head movement pose gradient data, the fusion analysis of physiological and behavioral data is realized, and more comprehensive and accurate individual state information can be provided, thereby improving the accuracy and reliability of stress detection.

[0065] Next, each step of the above solution will be introduced in detail:

[0066] First, in step S1, obtain the facial visible light video of the user.

[0067] Based on the facial visible light video, obtain physiological and behavioral data, specifically as follows:

[0068] In step S21, extract the gaze angle time series data and head pose time series data from the facial visible light video respectively, and perform normalization and gradient calculation processing.

[0069] Use the OpenFace2.2.0 open source toolkit on GitHub to extract the gaze angle time series data gaze_0_x, gaze_0_y, gaze_0_z, gaze_1_x, gaze_1_y, gaze_1_z and head pose time series data pose_Rx, pose_Ry, pose_Rz from the facial visible light video. The specific meanings of the relevant parameters are shown in Table 1.

[0070] Table 1 Gaze angle time series data and head pose time series data

[0071] Index Name Meaning gaze_0_x Gaze direction vector of the leftmost eye in the image along the x-axis gaze_0_y Gaze direction vector of the leftmost eye in the image along the y-axis gaze_0_z Gaze direction vector of the leftmost eye in the image along the z-axis gaze_1_x Gaze direction vector of the rightmost eye in the image along the x-axis gaze_1_y Gaze direction vector of the rightmost eye in the image along the y-axis gaze_1_z Gaze direction vector of the rightmost eye in the image along the z-axis pose_Rx Deflection of the head along the x-axis pose_Ry Deflection of the head along the y-axis pose_Rz Deflection of the head along the z-axis

[0072] Among them, the x-axis, y-axis, and z-axis respectively include the three coordinate axes in the three-dimensional coordinate system constructed by the spatial Cartesian coordinates.

[0073] After that, for the extracted gaze angle time series data and head pose time series data, perform normalization and gradient calculation processing to obtain the gaze angle gradient data Head pose gradient data where L is the length of the time series data, D is the number of features of the time series data, D G is the feature dimension of the gaze angle gradient data, D Pis the feature dimension of the head pose gradient data.

[0074] Next, in step S22, the face facial image sequence is cut based on the facial visible light video to extract the initial rppg signal data, and detrending, filtering, and normalization processing are performed.

[0075] The extraction process of the initial rppg signal data includes:

[0076] (1). Adopt a skin color adaptive segmentation method based on the YCbCr space to segment the skin of the face facial image sequence to reduce the influence of movements of eyebrows, eyes, and mouth on the extraction of the facial rppg signal.

[0077] It should be supplementary explained that YCbCr is a conversion standard used to encode the RGB color space, and it has been widely used in digital video compression systems to improve the efficiency and compression ratio of video processing. The specific process is as follows:

[0078] Convert the color space of any obtained face facial image from the RGB three channels to the Y, Cb, Cr three channels. By setting appropriate Y, Cb, Cr channel thresholds, judge whether each pixel in the current face facial image is skin. If this pixel meets the Y, Cb, and Cr channel thresholds, it belongs to skin, and set the mask at the corresponding position to 1. Otherwise, it does not belong to skin, and set the mask at the corresponding position to 0. Specifically, as shown in the formula:

[0079]

[0080] After that, the skin image with mask = 1 at the corresponding position of the original RGB image can be obtained through the addition operation function of OpenCV.

[0081] (2). Adopt the projection plane orthogonal skin algorithm to extract the initial rppg signal data from the face facial image sequence after skin segmentation.

[0082] Use the Plane-Orthogonal-to-Skin (POS) algorithm to extract the initial rppg signal from the face facial image sequence after skin segmentation. For each frame image in the face facial image sequence, obtain the average value of its R, G, B three channels respectively, and stack them to obtain the R, G, B sequence average value data Through the following calculation to Perform fusion processing on the data to obtain the initial rppg signal data S.

[0083]

[0084] where std is the standard deviation function.

[0085] For the extracted initial rppg signal data, detrending and third-order Butterworth band-pass filtering in the range of 0.7 - 2.5 Hz are performed to reduce the influence of clutter and noise in the signal on subsequent pressure detection, and normalization processing is carried out to obtain normalized rppg data where D R is the feature dimension of the normalized rppg data.

[0086] Finally, in step S3, the normalized rppg data, gaze angle gradient data, and head pose gradient data are used as the inputs of the VSDNet model constructed based on the two-stage learning strategy to detect whether the user is stressed or not.

[0087] It should be noted that through deep learning technology, the embodiments of the present invention can automatically learn and extract features in the data, and the powerful ability of deep learning enables the model to effectively fuse multi-dimensional data and process complex non-linear relationships, combined with the two-stage learning strategy, thereby constructing a model that can more accurately predict the user's stress state.

[0088] Among them, the process of constructing the VSDNet model based on the two-stage learning strategy includes:

[0089] (1) Data preparation

[0090] Based on the user's historical facial visible light videos, historical normalized rppg data, gaze angle gradient data, and head pose gradient data are respectively obtained;

[0091] (2) Constructing samples and sample pairs

[0092] Based on a preset sampling interval, data slicing is performed on the historical facial visible light videos, and the historical normalized rppg data, gaze angle gradient data, head pose gradient data, and corresponding stress labels within each preset sampling interval are used to construct a single sample, and different samples are paired pairwise to form sample pairs.

[0093] Based on the similarities and differences of the stress labels within each sample pair, similar labels are respectively set; where the similar labels include similar and dissimilar. That is, assuming that the stress labels of the samples within the sample pair are k1 and k2 respectively, if k1 and k2 are the same, the sample pair is considered similar and the similar label is set to 1, and if k1 and k2 are different, the sample pair is considered dissimilar and the similar label is set to 0.

[0094] (3) The first stage

[0095] Based on sample pairs with similar tags, train the VSDNet model in the first stage; the core network of the VSDNet model in the first stage is ExpLearnNet.

[0096] Use the sample pairs as the input for the first stage of the non-contact psychological stress detection framework. First, a neural network VSDNet model needs to be trained to identify whether the input sample pairs are similar or dissimilar, enabling the model to learn certain empirical knowledge. These empirical knowledge can preliminarily judge and extract data stress characteristics, that is, obtain general physiological behavior stress characteristics.

[0097] As Figure 2 shown, the VSDNet model in the first stage includes the twin network ExpLearnNet, the L1 distance calculation layer, and the sigmoid function output layer.

[0098] Among them, as Figure 3 shown, the ExpLearnNet is a one-dimensional CNN and BiLSTM twin network with shared parameters for two groups, and the LeakyReLU activation function is applied to all CNN layers to obtain the output feature map. The model is mainly designed based on one-dimensional CNN and BiLSTM. One-dimensional CNN is suitable for processing one-dimensional sequence data. The CNN layer can learn local features in 3 types of data. BiLSTM is bidirectional, and it can consider both forward and backward information in the feature sequence, so as to more comprehensively understand the overall structure of the sequence.

[0099] Then the training process of the VSDNet model in the first stage is as follows:

[0100] S101. For any sample in the sample pair, use its corresponding historical normalized rppg data, gaze angle gradient data, and head pose gradient data as the input of the one-dimensional CNN. Perform one-dimensional convolution and max pooling twice for downsampling, and perform multi-channel feature fusion to obtain the fusion feature; use the fusion feature as the input of the BiLSTM twin network to obtain the final output that captures the long-term dependencies in the data.

[0101] S102. Use the two final outputs of the sample pair as the input of the L1 distance calculation layer to obtain the L1 distance between the two samples, and then use it as the input of the sigmoid function output layer to obtain the similarity prediction result.

[0102] S103. Based on the similarity label and its similarity prediction result of each sample pair, construct the first loss function; minimize the first loss function, and after convergence, obtain the VSDNet model in the first stage.

[0103] Exemplarily, considering that the model output result is a single value and it is a binary classification task, binary cross-entropy loss can be used as the first loss function.

[0104] (4), The second stage

[0105] Fine-tune the ExpLearnNet of the VSDNet model in the first stage, use the samples with pressure labels as input, and train to obtain the VSDNet model in the second stage

[0106] After the model training and saving in the first stage, the core ExpLearnNet network of the VSDNet model has to a certain extent mastered the ability to obtain general physiological behavior stress characteristics, but this is not enough for identifying specific stress states. Therefore, the model structure in the first stage is fine-tuned to the model shown in the second stage, and the pre-trained ExpLearnNet is continued to be trained to obtain the deep physiological behavior stress characteristics of the stress state.

[0107] As Figure 4 shown, for the ExpLearnNe of the VSDNet model obtained after convergence in the first stage, a first fully connected layer, a Dropout layer, and a second fully connected layer are sequentially added after it, and softmax regression is selected as the activation function of the output layer after the second fully connected layer.

[0108] Among them, the LeakyReLU activation function is applied in the first fully connected layer to obtain the output feature map; since the amount of data itself is small, in the design of the neural network, to prevent overfitting, a Dropout(0.5) layer is also added to the two fully connected layers to reduce the dependence between features; and considering that the goal of the target learning stage is a binary classification task, softmax regression is selected as the activation function of the output layer.

[0109] Then the training process of the VSDNet model in the second stage is specifically as follows:

[0110] S201. Use the samples with pressure labels as input to obtain the pressure prediction result;

[0111] S202. Based on the pressure label and its pressure prediction result of each sample, construct a second loss function; minimize the second loss function, and after convergence, obtain the VSDNet model in the second stage.

[0112] Exemplarily, considering that the goal of the target learning stage is a binary classification task, cross-entropy loss is used as the second loss function.

[0113] So far, the VSDNet model constructed based on the two-stage learning strategy is obtained.

[0114] On this basis, this step respectively includes:

[0115] Since the feature dimensions D R 、D G 、D P of the standardized rppg data, gaze angle gradient data, and head pose gradient data are all different, the standardized rppg data, gaze angle gradient data, and head pose gradient data are uniformly represented as

[0116] First, perform two one-dimensional convolutions and max pooling on X R 、X P 、X G respectively for downsampling. The pooling layer can reduce the data dimension while retaining the most significant features, obtaining the feature representation At this time, their lengths are all L1, and the number of features are respectively

[0117] Then, perform multi-channel feature fusion, and splice the feature representations obtained from the three parts together to get X cat , forming a more abundant feature representation.

[0118] Subsequently, input X cat into the BiLSTM network. The role of the BiLSTM network is to process the feature sequence extracted by the convolutional layer and the pooling layer, and further learn the long-term dependence relationship between these features. The BiLSTM consists of two independent LSTM layers. One processes the input in chronological order (from left to right), and the other processes the input in reverse chronological order (from right to left). These two LSTM layers respectively capture the forward and backward information in the input sequence. In a single LSTM layer, for each time step t ∈ [1, L1], the following are the main calculation steps and formulas

[0119] f t = σ(W f · [h t-1 , x t + b f )

[0120] i t = σ(W i · [h t-1 , x t + b i )

[0121]

[0122] O t = σ(W O · [h t-1 ,x t+b O )

[0123] h t = O t ☉tanh(C t )

[0124] Among them, W f , W i , W C , v O are the weight matrices of the forget gate, input gate, candidate state, and output gate respectively, h t-1 is the hidden state of the previous time step, x t is the input of the current time step, b f , b i , b C , b O are the corresponding bias terms, σ is the sigmoid function, and tanh is the hyperbolic tangent function. The output of the BiLSTM is composed of the hidden states of the forward LSTM and the backward LSTM concatenated at each time step.

[0125] For the last time step T, the output of the BiLSTM is

[0126]

[0127] Among them, is the hidden state of the forward LSTM at the last time step, is the hidden state of the backward LSTM at the last time step. Take y T as the final output output of the VSDNet model.

[0128] Finally, pass the final output output through the first fully connected layer, Dropout layer, second fully connected layer, and softmax regression layer added in the second stage in sequence to obtain the detection result of whether the user is stressed or not.

[0129] Example 2:

[0130] The embodiment of the present invention provides a non-contact psychological stress detection system based on visual perception, including:

[0131] A video acquisition module for acquiring the visible light video of the user's face;

[0132] A data acquisition module for respectively extracting the gaze angle time series data and head pose time series data from the visible light video of the face, and performing normalization and gradient calculation processing;

[0133] and for cutting a sequence of human face images based on the facial visible light video to extract initial rppg signal data, and performing detrending, filtering, and normalization processing;

[0134] A pressure detection module, configured to use the normalized rppg data, gaze angle gradient data, and head pose gradient data as inputs to a VSDNet model constructed based on a two-stage learning strategy to detect whether the user is stressed or not stressed

[0135] Example 3:

[0136] An embodiment of the present invention provides a storage medium storing a computer program for non-contact psychological stress detection based on visual perception, wherein the computer program causes a computer to execute the non-contact psychological stress detection method as described in Example 1.

[0137] Example 4:

[0138] An embodiment of the present invention provides an electronic device, including:

[0139] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the non-contact psychological stress detection method as described in Example 1.

[0140] It can be understood that the non-contact psychological stress detection system, storage medium, and electronic device provided by the embodiments of the present invention correspond to the non-contact psychological stress detection method provided by the embodiments of the present invention. Explanations, examples, beneficial effects, and other parts of the relevant content can refer to the corresponding parts in the non-contact psychological stress detection method, which will not be elaborated here.

[0141] In summary, compared with the prior art, the following beneficial effects are achieved:

[0142] 1. The embodiments of the present invention integrate the physiological and behavioral characteristics of the user's facial visible light video, fully obtain psychological stress information to identify their psychological stress state. Compared with contact-based psychological stress state recognition or general mental health assessment, the present invention can detect the psychological stress state with the advantages of low cost and low interference.

[0143] 2. The embodiments of the present invention combine a two-stage learning strategy. In the first stage, the model is trained to obtain the ability to acquire general stress features by judging the similarity of input sample pairs. In the second stage, the classification performance of the psychological stress detection model can be effectively improved compared with single-stage detection by fine-tuning the pre-training results of the first stage and expanding the model inference layer to improve the model's ability to acquire high-level stress features.

[0144] 3. The embodiment of the present invention adopts a non-contact detection method, which can realize the detection of psychological stress only based on the visible light video of the face. The required equipment is simple, and the user can perform the detection without wearing any equipment.

[0145] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0146] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A non-contact psychological stress detection method based on visual perception, characterized in that, Including: Obtain the user's facial visible light video; Extract the gaze angle time series data and head pose time series data from the facial visible light video respectively, and perform normalization and gradient calculation processing; Based on the facial visible light video, cut the facial image sequence to extract the initial rppg signal data, and perform detrending, filtering, and normalization processing; Use the normalized rppg data, gaze angle gradient data, and head pose gradient data as the input of the VSDNet model constructed based on the two-stage learning strategy to detect whether the user is stressed or not; The process of constructing the VSDNet model based on the two-stage learning strategy; including: Based on the user's historical facial visible light video, obtain the historical normalized rppg data, gaze angle gradient data, and head pose gradient data respectively; Slice the historical facial visible light video based on a preset sampling interval, construct a single sample with the historical normalized rppg data, gaze angle gradient data, head pose gradient data within each preset sampling interval and the corresponding stress label, and match different samples pairwise to form sample pairs; Based on the similarities and differences of the stress labels within each sample pair, set similarity labels respectively; where the similarity labels include similar and dissimilar; Train the first-stage VSDNet model based on the sample pairs with similarity labels; where the core network of the first-stage VSDNet model is ExpLearnNet; Fine-tune the ExpLearnNet of the first-stage VSDNet model, use the samples with stress labels as the input, and train to obtain the second-stage VSDNet model; The VSDNet refers to a video-based stress detection network. The first-stage VSDNet model includes a Siamese network ExpLearnNet, an L1 distance calculation layer, and a sigmoid function output layer; where the ExpLearnNet refers to an exploration learning network, which is a one-dimensional CNN and BiLSTM Siamese network with shared parameters for two groups, and all CNN layers apply the LeakyReLU activation function network.

2. The non-contact psychological stress detection method according to claim 1, wherein The training of the first-stage VSDNet model based on the sample pairs with similarity labels; including: For any sample in the sample pair, use its corresponding historical normalized rppg data, gaze angle gradient data, and head pose gradient data as the input of the one-dimensional CNN, perform two one-dimensional convolutions and max pooling for downsampling respectively, and perform multi-channel feature fusion to obtain the fusion feature; use the fusion feature as the input of the BiLSTM Siamese network to obtain the final output that captures the long-term dependencies in the data; Use the two final outputs of the sample pair as the input of the L1 distance calculation layer to obtain the L1 distance between the two samples, and then use it as the input of the sigmoid function output layer to obtain the similarity prediction result; Based on the similarity label of each sample pair and its similarity prediction result, construct the first loss function; minimize the first loss function, and after convergence, obtain the first-stage VSDNet model.

3. The non-contact psychological stress detection method according to claim 2, wherein The ExpLearnNet that fine-tunes the VSDNet model in the first stage takes samples with pressure labels as inputs and trains to obtain the VSDNet model in the second stage, including: For the ExpLearnNet of the VSDNet model obtained after convergence in the first stage, a first fully connected layer, a Dropout layer, and a second fully connected layer are successively added after it, and softmax regression is selected as the activation function of the output layer after the second fully connected layer; where the LeakyReLU activation function is applied in the first fully connected layer. Taking samples with pressure labels as inputs to obtain pressure prediction results. Based on the pressure label and its pressure prediction result of each sample, a second loss function is constructed; the second loss function is minimized, and after convergence, the VSDNet model in the second stage is obtained.

4. The non-contact psychological stress detection method according to claim 1, characterized in that, The gaze angle time series data is any one or any combination of the gaze direction vectors of the leftmost eye along the x-axis, the gaze direction vectors of the leftmost eye along the y-axis, the gaze direction vectors of the leftmost eye along the z-axis, the gaze direction vectors of the rightmost eye along the x-axis, the gaze direction vectors of the rightmost eye along the y-axis, and the gaze direction vectors of the rightmost eye along the z-axis in the image; where the x-axis, y-axis, and z-axis respectively include three coordinate axes in a three-dimensional coordinate system constructed by spatial Cartesian coordinates.

5. The non-contact psychological stress detection method according to claim 1, characterized in that, The head pose time series data is any one or any combination of the deflections of the head along the x-axis, the deflections of the head along the y-axis, and the deflections of the head along the z-axis; where the x-axis, y-axis, and z-axis respectively include three coordinate axes in a three-dimensional coordinate system constructed by spatial Cartesian coordinates.

6. The non-contact psychological stress detection method according to claim 1, wherein, The extraction process of the initial rppg signal data includes: Using a skin color adaptive segmentation method based on the YCbCr color space to perform skin segmentation on the sequence of face images. Using the projection plane orthogonal skin algorithm to extract the initial rppg signal data from the sequence of face images after skin segmentation.

7. A non-contact psychological stress detection system based on visual perception, characterized in that, For implementing the non-contact psychological stress detection method as claimed in claim 1, including: A video acquisition module for acquiring the visible light video of the user's face. A data acquisition module for respectively extracting the gaze angle time series data and the head pose time series data from the visible light video of the face, and performing normalization and gradient calculation processing. And for cutting the sequence of face images based on the visible light video of the face to extract the initial rppg signal data, and performing detrending, filtering, and normalization processing. A stress detection module for using the normalized rppg data, gaze angle gradient data, and head pose gradient data as inputs to the VSDNet model constructed based on a two-stage learning strategy to detect whether the user is stressed or not.

8. A storage medium, characterized in that, It stores a computer program for non-contact psychological stress detection based on visual perception, where the computer program causes the computer to execute the non-contact psychological stress detection method as claimed in any one of claims 1 to 6.

9. An electronic device, characterized in that, Including: One or more processors; A memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including those for performing the non-contact psychological stress detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video-based psychological pressure measurement method and system

    CN112597949A

  • Multi-modal emotional pressure recognition method and device, computer equipment and storage medium

    CN113057633A