Non-contact intelligent health monitoring method and equipment based on physiological signals

By adopting face video image processing and improved rPPG technology in contactless health monitoring technology, combined with deep learning emotion recognition model, the problems of instability of heart rate monitoring signal and insufficient accuracy of emotion recognition in dynamic environments in the prior art are solved, and high-precision joint monitoring of heart rate and emotions are achieved.

CN120199518APending Publication Date: 2025-06-24CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510275633.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing contactless health monitoring technology has unstable heart rate monitoring signals and insufficient accuracy and robustness of emotion recognition under dynamic environments and complex lighting conditions.

Method used

A contactless intelligent health monitoring method based on physiological signals is adopted, and the division and tracking of face video images is achieved through the improvement of rPPG technology and deep learning-driven emotion recognition model to achieve joint monitoring of heart rate and emotions.

Benefits of technology

It improves the robustness and accuracy of heart rate monitoring, enhances the accuracy and robustness of emotion recognition, and realizes stable health monitoring in a dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199518A_ABST
    Figure CN120199518A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact intelligent health monitoring method and equipment based on a physiological signal, and relates to the technical field of non-contact health monitoring, and the non-contact intelligent health monitoring method based on the physiological signal mainly comprises the steps: carrying out the division according to a face video image, and obtaining a region of interest; obtaining an aligned region of interest by utilizing a facial feature point tracking and alignment technology, and fusing and inputting an improved rPPG technology into the region to obtain an actual heart rate value; according to the original audio signal, performing framing and discrete Fourier transform by using a Hamming window to obtain a spectrogram of each audio clip; mel feature extraction is carried out by using a librosa library to obtain MFCC features, and the three features are fused to obtain a predicted emotion type of the voice. By means of the non-contact intelligent health monitoring method and device based on the physiological signals, the robustness and precision of heart rate monitoring can be improved, the accuracy and robustness of emotion recognition are improved, and combined monitoring of the emotion and the heart rate is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-contact health monitoring, and more particularly, to a non-contact intelligent health monitoring method and device based on physiological signals. Background Art

[0002] In recent years, with the increasing demand for health management, non-contact health monitoring technologies have received extensive attention. According to different monitoring methods, current health monitoring technologies mainly include Remote Photoplethysmography (PPG) and remote photoplethysmography (rPPG) technologies. The PPG technology monitors the heart rate by detecting the change in light reflection caused by blood flow through a photoelectric sensor, but it is easily interfered by motion artifacts, light changes, and skin color differences in a dynamic environment, resulting in unstable signals. The rPPG technology uses a camera to remotely collect minute skin reflection changes to extract pulse signals. Although its advantage lies in the fact that no sensor needs to be worn, it still faces the problem that the signal is easily affected by environmental changes and facial expressions. Especially in a complex or moving environment, the accuracy and stability are greatly reduced. In addition, as an important part of health monitoring, emotion recognition technology currently mostly relies on the analysis of facial expressions, speech, or behavioral characteristics. Although certain progress has been made in speech emotion recognition technology, factors such as background noise, speech quality, and individual differences still affect its accuracy, and the emotion analysis of facial expressions requires the user to remain static, which limits its usability in practical applications. Therefore, there are certain defects in the integration and accuracy of emotion recognition and physiological signal monitoring in the prior art, and further improvement is urgently needed. Summary of the Invention

[0003] The purpose of the present invention is to provide a non-contact intelligent health monitoring method and device based on physiological signals, which can improve the robustness and accuracy of heart rate monitoring, improve the accuracy and robustness of emotion recognition, and realize the joint monitoring of emotion and heart rate.

[0004] The present invention provides a non-contact intelligent health monitoring method based on physiological signals, including the following steps: S1, dividing a face video image to obtain a region of interest, tracking and aligning the region of interest to obtain an aligned region of interest; S2, according to the aligned region of interest, combining with an improved rPPG technology input to obtain an actual heart rate value; S3, performing frame segmentation and discrete Fourier transform on the original audio signal using a Hamming window to obtain a spectrogram of each audio segment; using the librosa library to extract Mel features to obtain MFCC features; S4, performing feature extraction and fusion on the spectrogram, MFCC features, and original audio signal of each audio segment to obtain the predicted emotion type of the speech.

[0005] Further, the above-mentioned region of interest includes the forehead, the lower jaw, and the left and right cheeks.

[0006] Further, step S1 specifically includes: obtaining the region of interest from the face video image, aligning the region of interest to obtain the aligned region of interest, as shown in the formula:

[0007] F = FM(I),

[0008] P t+1 = P t + ΔP,

[0009] T = ETransform(P t , P t+1 ),

[0010] F aligned = T(FM(It)) where T = ETransform(P t , P t+1 ),

[0011] where F is the extracted face feature, I is the input image, FM(·) represents the Fachmesh model; P t+1 represents the tracking point in frame t + 1, P t is the feature point extracted in frame t, ΔP is the optical flow calculated by the Kanade-Lucas-Tomasi algorithm, T represents the transformation matrix, ETransform(·) represents the feature point transformation matrix, and ETransform(P t , P t+1 ) accepts the feature point P t extracted in frame t and the feature point P t+1 tracked in frame (t + 1), and calculates the transformation matrix T required to transform P t to P t+1 , and F aligned represents the aligned image.

[0012] Further, step S2 specifically includes: obtaining the actual heart rate value according to the aligned region of interest, as shown in the formula:

[0013]

[0014] where bpm is the estimated heart rate value of each region, spec is the value converted from the time domain to the frequency domain by Fourier transform, len(spec) is the number of elements returned by the sequence spec, argmax() is the index of the maximum value in the first half of the sequence spec, seq represents the first half of the sequence, and fps represents the video frame rate.

[0015] Further, step S3 specifically includes: according to the original audio signal, using a Hamming window for framing and discrete Fourier transform to obtain the spectrogram of each audio segment, as shown in the formula:

[0016]

[0017] x t (n) = x(n + t×160)·ω(n), 0 ≤ n < 640,

[0018]

[0019] where ω(n) represents the Hamming window function value used to reduce spectral leakage, n represents the sampling point serial number within the window (0 <= n < 640), N represents the window length, x t (n) represents the t-th frame signal after windowing, x(n) represents the discrete sampling value of the original speech signal, k represents the frequency point serial number, and X t (k) represents the amplitude value corresponding to the k-th frequency point of the t-th frame. is the DFT complex basis function.

[0020] Further, step S4 specifically includes: S41. Using the convolutional block of the AlexNet network to extract features from the spectrogram of each audio segment and using the Coordinate Attention mechanism for feature enhancement to obtain the final spectrogram features; S42. Using the BiLSTM network to process the MFCC features to obtain emotional information; S43. Using a pre-trained self-supervised model as an encoder to process the original audio signal to obtain HuBERT features; S44. Using a module based on the KAN fusion attention mechanism to perform feature fusion and discrimination on the final spectrogram features, emotional information, and HuBERT features to obtain the predicted emotion type of the speech.

[0021] Further, step S44 specifically includes: S441. Using a module based on the KAN fusion attention mechanism to perform feature fusion on the final spectrogram features, emotional information, and HuBERT features to obtain fused features; S442. According to the fused features, using a module based on the KAN fusion attention mechanism to obtain the probability values of each emotion category; selecting the emotion label with the highest probability to determine the predicted emotion type of the speech.

[0022] Further, step S441 specifically includes: using a module based on the KAN fusion attention mechanism to perform feature fusion on the final spectrogram features, emotional information, and HuBERT features to obtain fused features, as shown in the formula:

[0023] W = σ(W2·ReLU(W1[F s ′; F m ′] + b1) + b2),

[0024] F' hubert = We F hubert

[0025]

[0026] where W is the dynamic weight matrix, F s ' and F m ' are the spectrogram features and MFCC features respectively, σ is the Sigmoid function, W1 and W2 are trainable parameters; ReLU is the activation function, b1 and b2 represent the bias vectors of the first and second fully connected layers respectively, F fusion is the fused feature matrix, ⊙ represents element-wise multiplication, F hubert represents the HuBERT feature matrix; P(y|F fusion ) represents the probability distribution of the emotion category corresponding to the input feature F fsinon , Softmax() represents the Softmax function, which maps the output to a probability distribution, K represents the number of basis functions, c i represents the learnable coefficient of the i-th basis function, φ i (g) represents the i-th B-spline basis function, W c represents the weight matrix of the classifier; α represents the weight coefficient for adjusting the fusion ratio of the spectrogram and MFCC features, β represents the weight coefficient for adjusting the fusion ratio of the HuBERT features, F spec represents the final spectrogram feature; F mfcc represents the MFCC feature matrix; represents element-wise addition.

[0027] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned non-contact intelligent health monitoring method based on physiological signals are implemented.

[0028] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the above-mentioned non-contact intelligent health monitoring method based on physiological signals are implemented.

[0029] Implementing the non-contact intelligent health monitoring method and device based on physiological signals provided by the present invention has the following beneficial effects:

[0030] The present invention is different from the traditional dlib face detection technology. The present invention optimizes the face ROI extraction technology, uses the Mediapipe software developed by Google, and combines the FaceMesh model for three-dimensional all-round face detection and feature point completion. It accurately extracts the region of interest (ROI) of the face and divides it into four regions with different weights, improving the efficiency and accuracy of face detection, especially the adaptability in the case of head tilt and dynamic environment. The present invention adopts a non-contact heart rate monitoring method, which only collects the face area of the subject through an ordinary camera and uses the rPPG technology to capture the skin reflection signal, avoiding the contact problem brought by the traditional electrode monitoring method. Combining machine learning with traditional signal processing methods, it can non-contact detect physiological indexes such as heart rate and heart rate variability, and achieve an accuracy comparable to that of contact devices. Specifically, for the rPPG signal, signal processing technologies such as Butterworth band-pass filter, Savitzky-Golay smoothing filter and FFT are adopted, and the noise is removed by quadratic polynomial trend fitting to refine the signal, and finally the heart rate and heart rate variability (HRV) are calculated, optimizing the accuracy and efficiency of signal processing and providing more reliable data support for health monitoring. By using the optical flow method and Kalman filter, the problem of heart rate monitoring in the case of motion state or environmental light change is successfully solved, significantly improving the robustness and accuracy.

[0031] The present invention adopts a deep learning-driven emotion recognition model, preprocesses the collected audio signals by using three different methods, and uses a knowledge-aware network (KAN) to fuse the attention mechanism, integrating spectrogram features, Mel-frequency cepstral coefficient (MFCC) features and HuBERT features, so as to obtain more comprehensive acoustic information. Using a deep learning network for emotion classification and combining the attention mechanism to fuse multiple audio features, the accuracy and robustness of emotion recognition are effectively improved.

[0032] The present invention not only monitors the heart rate through the rPPG technology, but also combines the voice-based emotion recognition technology. It can not only monitor physiological indexes such as heart rate, but also evaluate the psychological state - emotion, and conduct multi-signal fusion health monitoring: by real-time collecting voice signals and extracting emotion-related features, and combining deep learning algorithms for emotion discrimination, the joint monitoring of emotion and heart rate is realized, thus providing a more comprehensive monitoring of the subject. This multi-signal fusion method can more comprehensively evaluate the health status of an individual and provide a more accurate health management plan. In addition, the present invention can fuse multiple influencing factors such as heart rate, heart rate variability, emotion, body temperature, and distance to realize intelligent health monitoring of personnel. Its software part can be deployed on an embedded development board and has the characteristics of being simple and easy to use. Description of the Drawings

[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. In the accompanying drawings:

[0034] Figure 1 is a flowchart of a non-contact intelligent health monitoring method based on physiological signals provided by the present invention;

[0035] Figure 2 is a schematic diagram of speech emotion recognition of multi-modal hierarchical acoustic information based on KAN (Knowledge-Aware Network) fusion attention provided by the present invention;

[0036] Figure 3 is a block diagram of the composition of a non-contact and non-invasive health monitoring system provided by the present invention;

[0037] Figure 4 is an algorithm flowchart of a non-contact intelligent health monitoring method based on physiological signals provided by the present invention;

[0038] Figure 5 is an effect diagram of the hardware implementation of a non-contact intelligent health monitoring method based on physiological signals provided by the present invention;

[0039] Figure 6 is a block diagram of the structure of a computer device provided by the present invention. Detailed Embodiments

[0040] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0041] Figure 1 shows a schematic diagram of a non-contact intelligent health monitoring method based on physiological signals in this embodiment. In this embodiment, the non-contact intelligent health monitoring method based on physiological signals includes the following steps:

[0042] S1. Divide the face video image to obtain the region of interest, track and align the region of interest, and obtain the aligned region of interest;

[0043] In an exemplary embodiment, the region of interest includes the forehead, the lower jaw, and the left and right cheeks;

[0044] In an exemplary embodiment, step S1 specifically includes: dividing the face video image to obtain the region of interest, tracking and aligning the region of interest, and obtaining the aligned region of interest, as shown in the formula:

[0045] F = FM(I),

[0046] P t+1 = P t + ΔP,

[0047] T = ETransform(P t , P t+1 ),

[0048] F aligned = T(FM(It)) where T = ETransform(P t , P t+1 ),

[0049] where F is the extracted face feature, I is the input image, FM(·) represents the Fachmesh model; P t+1 represents the tracking points in frame t + 1, P t are the feature points extracted in frame t, ΔP is the optical flow calculated by the Kanade - Lucas - Tomasi algorithm, T represents the transformation matrix, ETransform(·) represents the feature point transformation matrix, ETransform(P t , P t+1 ) accepts the feature points P t extracted in frame t and the feature points P t+1 tracked in frame (t + 1), and calculates the transformation matrix T required to transform P t to P t+1 , and F aligned represents the aligned image;

[0050] As an exemplary embodiment, in step S1, a camera with a higher pixel is used to collect face video images. The region of interest ROI is selected from the collected face video images and divided into four regions: the forehead, the mandible, and the left and right cheeks. When monitoring the heart rate, face detection and tracking are extremely important parts of the entire framework. After converting the BGR format in opencv to the RGB format in mediapipe, the FM (Fachmesh) model and the KLT (Kanade - Lucas - Tomasi) optical flow method are used together for accurate face alignment. First, the FM model is used to extract face features:

[0051] F = FM(I) (1)

[0052] where I is the input image and F is the extracted face feature;

[0053] When using the KLT algorithm to track feature points, the feature points extracted in frame t are set as P t , and the tracking points in frame t + 1 are P t+1 :

[0054] P t+1 = P t + ΔP (2)

[0055] where ΔP is the optical flow calculated by the KLT algorithm; alignment is achieved by calculating the transformation matrix TTT of the feature points:

[0056] T = ETransform(P t , P t+1 ) (3)

[0057] That is, the comprehensive formula is expressed as:

[0058] F aligned = T(FM(It)) where T = ETransform(P t , P t+1 ) (4)

[0059] This step solves the problem of real-time tracking of facial feature points in a dynamic video stream, not only improving the problems of face detection and face tracking extraction, but also ensuring the accurate extraction of rPPG signals;

[0060] S2. According to the aligned region of interest, combined with the input improved rPPG technology, obtain the actual heart rate value;

[0061] In an exemplary embodiment, step S2 specifically includes: according to the aligned region of interest, combined with the input improved rPPG technology, obtain the actual heart rate value, as shown in the formula:

[0062]

[0063] where bpm is the estimated heart rate value of each region, spec is the value converted from the time domain to the frequency domain by Fourier transform, len(spec) is the number of elements returned by the sequence spec, argmax() is the index of the maximum value in the first half of the sequence spec, represents the first half of the sequence spec, and fps represents the video frame rate.

[0064] As an exemplary embodiment, in step S2, obtaining the ROI region image and separating it into three RGB channels includes: R = ROI[:, :, 0], G = ROI[:, :, 1], B = ROI[:, :, 2]; where R, G, and B respectively represent the red channel value, green channel value, and blue channel value of all pixels within the ROI region, ROI[:, :, 0] represents extracting the first color channel of all pixels in the ROI, ROI[:, :, 1] represents extracting the second color channel of all pixels in the ROI, and ROI[:, :, 2] represents extracting the third color channel of all pixels in the ROI; processing each frame of the image, calculating the average value of each channel in the ROI region, and forming the time series of each channel as follows:

[0065] 1. Red channel average value sequence

[0066]

[0067] 2. Green channel average value sequence

[0068]

[0069] 3. Blue channel average value sequence

[0070]

[0071] Among them, R series , G series , B series respectively represent the average value sequences of the red, green, and blue channels; n t represents the total number of pixels in the ROI region at time t; R t,i , G t,i , B t,i represent the values of the red, green, and blue channels of the i-th pixel in the ROI region at time t. T represents the total number of frames, indicating the time length considered during the entire processing.

[0072] In terms of signal processing, a quadratic polynomial is used for trend fitting to remove the linear trend component in the data. A Butterworth band-pass filter and a Savitzky-Golay smoothing filter are used to refine the BVP pulse signal and remove the noise frequencies unrelated to the heart rate; the RGB information of the filtered signal is converted to the feature space to obtain the PPG signal; the PPG signal is respectively converted to the time domain and the frequency domain, and the fast Fourier transform (FFT) is used, and different weights are assigned according to the skin capillary richness and calibrated to obtain the final heart rate value; the heart rate time domain and frequency domain values of each obtained ROI region are substituted into formula (5):

[0073]

[0074] Among them, bpm is the estimated heart rate value of each region, spec is the value converted from the time domain to the frequency domain by Fourier transform, len(spec) is the sequence of spec, returning the number of elements in spec, and argmax() returns the index of the maximum value in the first half of spec; finally, the obtained bpm values are weighted according to the skin capillary richness to obtain the actual heart rate value;

[0075] In the above method, the influence of ambient light on the measurement result is weakened, and emphasis is placed on the division of the region of interest in the face and the application of the optical flow method to enhance the robustness of the measurement in the motion state. Its accuracy is significantly improved compared with the existing heart rate measurement methods. The experimental results of different methods are shown in Table 1:

[0076] Table 1: Comparison table of experimental results of different methods

[0077]

[0078]

[0079] In Table 1, the detection errors of six rPPG methods are listed respectively in the motion and static states. The mean absolute error (MAD), root mean square error (RMSE), standard deviation (SD), and Pearson correlation coefficient (PCC) are used as evaluation indicators to verify the accuracy of the proposed method, which has great reference value.

[0080] S3. According to the original audio signal, use a Hamming window for framing and discrete Fourier transform to obtain the spectrogram of each audio segment; use the librosa library for Mel feature extraction to obtain MFCC features.

[0081] In an exemplary embodiment, step S3 specifically includes: According to the original audio signal, use a Hamming window for framing and discrete Fourier transform to obtain the spectrogram of each audio segment, as shown in the formula:

[0082]

[0083] x t (n) = x(n + t×160)·ω(n), 0 ≤ n < 640,

[0084]

[0085] where ω(n) represents the Hamming window function value, which is used to reduce spectral leakage, n represents the sampling point serial number in the window (0 <= n < 640), N represents the window length, x t (n) represents the t-th frame signal after windowing, x(n) represents the discrete sampling value of the original speech signal, k represents the frequency point serial number, X t (k) represents the amplitude value corresponding to the k-th frequency point of the t-th frame, is the DFT complex basis function.

[0086] As an exemplary embodiment, in step S3, the original audio signals of the IEMOCAP emotional speech database are sampled at a sampling rate of 16 kHz; the audio signals are preprocessed by three different methods to obtain comprehensive acoustic information; firstly, the audio segments are framed using a Hamming window with a step size of 10 ms and a window length of 40 ms, and a discrete Fourier transform of length 800 is applied to each frame, and the first 200 points are selected as input features to obtain a spectrogram of 300*200 for each audio segment; secondly, 40-dimensional HTK-style Mel features are extracted using the librosa library as MFCC features; thirdly, the original audio signals are used as features.

[0087] The Hamming window function is defined as:

[0088]

[0089] The signal of the t-th frame is:

[0090] x t (n) = x(n + t×160)·ω(n), 0 ≤ n < 640

[0091] Spectrum calculation: Apply an 800-point discrete Fourier transform (DFT) to each frame, and take the logarithm of the magnitude of the energy of the first 200 frequency points:

[0092]

[0093] S4. Feature extraction and fusion are performed on the spectrogram, MFCC features, and original audio signals of each audio segment to obtain the predicted emotion type of the speech.

[0094] In an exemplary embodiment, step S4 specifically includes:

[0095] S41. Use the convolutional block of the AlexNet network to extract features from the spectrogram of each audio segment, and use the Coordinate Attention mechanism to enhance the features to obtain the final spectrogram features.

[0096] S42. Use the BiLSTM network to process the MFCC features to obtain emotional information.

[0097] S43. Use the pre-trained self-supervised model as an encoder to process the original audio signals to obtain HuBERT features.

[0098] S44. Use the module based on the KAN fusion attention mechanism to perform feature fusion and discrimination on the final spectrogram features, emotional information, and HuBERT features to obtain the predicted emotion type of the speech.

[0099] In an exemplary embodiment, step S44 specifically includes:

[0100] S441. Use a module based on the KAN fusion attention mechanism to fuse the final spectrogram features, emotional information, and HuBERT features to obtain fused features;

[0101] In an exemplary embodiment, step S441 specifically includes: Use a module based on the KAN fusion attention mechanism to fuse the final spectrogram features, emotional information, and HuBERT features to obtain fused features, as shown in the formula:

[0102] W = σ(W2·ReLU(W1[F s ′; F m ′]+b1)+b2),

[0103] F' hubert = We F hubert

[0104]

[0105] where W is the dynamic weight matrix, F s ' and F m ' are the spectrogram features and MFCC features respectively, σ is the Sigmoid function, W1 and W2 are trainable parameters; ReLU is the activation function, b1 and b2 represent the bias vectors of the first and second fully connected layers respectively, F fusion is the fused feature matrix, ⊙ represents element-wise multiplication, F hubert represents the HuBERT feature matrix; P(y|F fusion ) represents the probability distribution of the emotional category corresponding to the input feature F fsinon , Softmax() represents the Softmax function, which maps the output to a probability distribution, K represents the number of basis functions, c i represents the learnable coefficient of the i-th basis function, φ i (g) represents the i-th B-spline basis function, W c represents the weight matrix of the classifier; α represents the weight coefficient for adjusting the fusion ratio of the spectrogram and MFCC features, β represents the weight coefficient for adjusting the fusion ratio of the HuBERT features, F spec represents the final spectrogram feature; F mfcc represents the MFCC feature matrix; represents element-wise addition.

[0106] S442. According to the fused features, use a module based on the KAN fusion attention mechanism to obtain the probability values of each emotional category; select the emotional label with the maximum probability to determine the predicted emotional type of the speech;

[0107] As an exemplary embodiment, in step S4, three different encoders are used for feature extraction; the convolutional blocks of the AlexNet network are used to extract features from the spectrogram features, and the Coordinate Attention mechanism is used for feature enhancement to obtain the final spectrogram features; the BiLSTM network is used to process the MFCC features to capture the emotional information with context information; the pre-trained self-supervised model is used as an encoder to process the original audio signal to obtain the HuBERT features of the deep speech signal; in order to better fuse multi-level features (spectrogram features, MFCC features, and HuBERT features), this embodiment designs a module based on the KAN (Knowledge-Aware Network) fusion attention mechanism; KAN (Knowledge-Aware Network) is a multi-modal feature fusion network that realizes emotion recognition through dynamic weight modulation and an interpretable classifier; its core is: generating frame-level attention weights based on spectrogram and MFCC features, adaptively enhancing the emotion-sensitive components in the HuBERT features; and using a B-spline basis function combination to replace the traditional neural network layer to explicitly model the non-linear classification boundary; as Figure 2 shown, this module can use the concatenation of spectrogram features and MFCC features as frame-level weights, multiply them with the HuBERT features, and further extract the emotion information in the speech signal;

[0108] Frame-level weight generation: Concatenate F s ' and F m ' along the feature dimension, and generate a dynamic weight matrix through learnable parameters:

[0109] W = σ(W2·ReLU(W1[F s ′; F m ′]+b1)+b2)

[0110] where σ is the Sigmoid function, are trainable parameters.

[0111] Feature weighted fusion: Multiply the weights with the HuBERT features frame by frame to enhance the emotion-related components:

[0112] F' hubert = We F hubert

[0113] Specifically, the fusion process can be represented by the following formula:

[0114]

[0115] Composite function fitting: The KAN classifier uses a piecewise polynomial basis function φi (x) Fitting a non - linear decision boundary:

[0116]

[0117] Next, all acoustic features are input into the KAN - based emotion classifier for discrimination; the classifier utilizes the ability of the KAN network to fit non - linear composite functions. After multiple rounds of fitting iterations, it can output the probability values of each emotion category based on the input features;

[0118] Finally, by selecting the emotion label with the highest probability, the predicted emotion type of the speech is determined.

[0119] The experiments have achieved good results on multiple datasets.

[0120] To evaluate the effectiveness of the model proposed in this embodiment and deeply analyze the impact of different combinations of acoustic features on the model performance, multiple groups of ablation experiments were conducted, and the results are shown in Table 2.

[0121] Table 2: Comparison chart of multiple groups of ablation experiments

[0122] Number Model WA(%) UA(%) ACC(%) 1 Spectrogram 60.06 60.26 60.16 2 MFCC 53.59 55.78 54.69 3 HuBERT 71.15 71.88 71.52 4 Spectrogram(CA) 61.07 60.99 61.03 5 Spectrogram + MFCC 60.01 59.80 59.91 6 Spectrogram(CA) + MFCC 59.87 61.17 60.52 7 Spectrogram + HuBERT 71.96 73.21 72.58 8 Spectrogram(CA) + HuBERT 72.07 73.11 72.59 9 MFCC + HuBERT 71.62 73.14 72.38 10 Spectrogram(CA) + MFCC + HuBERT 72.16 73.56 72.86 11 Spectrogram(CA) + HuBERT + att 72.43 74.11 73.27 12 MFCC + HuBERT + att 72.21 72.86 72.53 13 Spectrogram(CA) + MFCC + HuBERT + att 73.66 74.80 74.23

[0123] As Figure 3 shown is the block diagram of a non - contact and non - invasive health monitoring system. In terms of heart rate monitoring, a method based on the rPPG technology is adopted, combining machine learning and traditional signal processing methods, effectively overcoming the problems of motion artifacts and environmental interference during the detection process. In terms of emotion monitoring, a voice - based emotion recognition unit is integrated, and characteristic parameters capable of characterizing emotion changes are extracted from the real - time collected voice signals. Through the extraction and analysis of effective parameters expressing emotions, and then through deep - learning algorithms for emotion discrimination, the emotion type of the voice is obtained. As Figure 4 shown is the algorithm flowchart of the non - contact intelligent health monitoring method based on physiological signals; through non - contact and non - invasive means, the system realizes the effective monitoring of heart rate and emotion, integrates multiple physiological monitoring indicators, has good accuracy and stability, solves some limitations in traditional monitoring methods, and provides a new means of health monitoring. As Figure 5 shown is the hardware implementation effect diagram of the non - contact intelligent health monitoring method based on physiological signals.

[0124] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned non-contact intelligent health monitoring method based on physiological signals are implemented. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memories.

[0125] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned non-contact intelligent health monitoring method based on physiological signals are implemented.

[0126] As Figure 6 shown, the computer device 120 may include: at least one processor 121, such as a central processing unit (CPU), at least one communication interface 123, a memory 124, and at least one communication bus 122. Among them, the communication bus 122 is used to realize the connection and communication between these components. Among them, the communication interface 123 may include a display screen and a keyboard. Optionally, the communication interface 123 may further include a standard wired interface and a wireless interface. The memory 124 may be a high-speed random access memory (RAM), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 124 may further be at least one storage device located far from the aforementioned processor 121. Among them, an application program is stored in the memory 124, and the processor 121 calls the program code stored in the memory 124 to execute any of the above method steps. Among them, the communication bus 122 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 122 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6It is represented by only one line in the figure, but it does not mean that there is only one bus or one type of bus. Among them, the memory 124 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 124 may also include a combination of the above types of memories. Among them, the processor 121 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. Among them, the processor 121 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Optionally, the memory 124 is also used to store program instructions. The processor 121 may call the program instructions to implement the non-contact intelligent health monitoring method based on physiological signals as described in this embodiment.

[0127] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A non-contact intelligent health monitoring method based on physiological signals, characterized in that: The following steps are involved: S1. Divide the face video image to obtain regions of interest, and track and align the regions of interest to obtain aligned regions of interest; S2. According to the aligned region of interest, combined with inputting improved rPPG technology, an actual heart rate value is obtained; S3, according to the original audio signal, use the Hamming window to perform frame division and discrete Fourier transform to obtain the spectrum diagram of each audio clip; Use librosa library to extract Mel features and obtain MFCC features; S4. Extract and fuse the spectrogram, MFCC features and original audio signal of each audio clip to obtain the predicted emotion type of the speech.

2. The non-contact intelligent health monitoring method based on physiological signals according to claim 1 is characterized in that: The regions of interest include the forehead, the jaw, and the left and right cheeks.

3. The non-contact intelligent health monitoring method based on physiological signals according to claim 1 is characterized in that: Step S1 specifically includes: obtaining a region of interest according to the face video image, aligning the region of interest, and obtaining an aligned region of interest, such as formula: F=FM(I), P t+1 =P t +ΔP, T=ETransform(P t ,P t+1 ), F alined =T(FM(It))where T=ETransform(P t ,P t+1 ), Where F is the extracted facial features, I is the input image, FM(·) represents the Fachmesh model; P t+1 represents the tracking point in frame t+1, P t The feature points extracted in frame t, ΔP, are the optical flows calculated by the Kanade-Lucas-Tomasi algorithm, T represents the transformation matrix, ETransform(·) represents the feature point transformation matrix, and ETransform(P t ,P t+1 ) Accept the feature point P extracted in frame t t and the feature point P tracked in frame (t+1) t+1 , and calculate P t Transform to P t+1 The required transformation matrix T, F aligned Represents the aligned image.

4. The non-contact intelligent health monitoring method based on physiological signals according to claim 1, characterized in that: Step S2 specifically includes: obtaining the actual heart rate value according to the aligned region of interest, such as the formula: Among them, bpm is the estimated heart rate value of each area, spec is the value converted from the time domain to the frequency domain by Fourier transform, len(spec) is the number of elements in the sequence spec returned by spec, and argmax() is the index of the maximum value of the first half of the sequence spec. It indicates the first half of the sequence spec, and fps indicates the video frame rate.

5. The non-contact intelligent health monitoring method based on physiological signals according to claim 1, characterized in that: Step S3 specifically includes: according to the original audio signal, using the Hamming window to perform frame division and discrete Fourier transform to obtain the frequency spectrum of each audio segment, such as formula: x t (n)=x(n+t×160)·ω(n),0≤n\640, Where ω(n) represents the Hamming window function value, which is used to reduce spectrum leakage, n represents the sampling point number in the window (0 <= n < 640), N represents the window length, and x t (n) represents the t-th frame signal after windowing, x(n) represents the discrete sampling value of the original speech signal, k represents the frequency point number, X t (k) represents the corresponding amplitude value of the kth frequency point in the tth frame, are the DFT complex basis functions.

6. The non-contact intelligent health monitoring method based on physiological signals according to claim 1, characterized in that: Step S4 specifically includes: S41, using the convolution block of the AlexNet network to extract features from the spectrogram of each audio clip, and using the Coordinate Attention mechanism to enhance the features, to obtain the final spectrogram features; S42, using a BiLSTM network to process the MFCC features to obtain emotional information; S43, using the pre-trained self-supervised model as an encoder to process the original audio signal to obtain HuBERT features; S44. Using a module based on the KAN fusion attention mechanism, the final spectrogram features, sentiment information and HuBERT features are subjected to feature fusion and discrimination to obtain the predicted emotion type of the speech.

7. The non-contact intelligent health monitoring method based on physiological signals according to claim 6 is characterized in that: Step S44 specifically includes: S441, using a module based on the KAN fusion attention mechanism, the final spectrogram features, sentiment information and HuBERT features are fused to obtain fused features; S442. According to the fusion features, a module based on the KAN fusion attention mechanism is used to obtain the probability value of each emotion category; the emotion label with the maximum probability is selected to determine the predicted emotion type of the speech.

8. The non-contact intelligent health monitoring method based on physiological signals according to claim 6, characterized in that: Step S441 specifically includes: using a module based on the KAN fusion attention mechanism to fuse the final spectrogram features, sentiment information and HuBERT features to obtain fusion features, such as the formula: W=σ(W2·ReLU(W1[Fs′;Fm′]+b1)+b2), F'hubert=We Fhubert Among them, W is the dynamic weight matrix, F s 'With F m ' are the spectrogram features and MFCC features, σ is the Sigmoid function, W1 and W2 are trainable parameters; ReLU is the activation function, b1 and b2 represent the bias vectors of the first and second fully connected layers, respectively, fusion is the fused feature matrix, ⊙ represents element-by-element multiplication, F hubert represents the HuBERT feature matrix; P(y|F fusion ) represents the input feature F fsinon The corresponding probability distribution of emotion categories, Softmax() represents the Softmax function, which maps the output to a probability distribution, K represents the number of basis functions, and c i represents the learnable coefficient of the i-th basis function, φ i (g) represents the i-th B-spline basis function, W c represents the weight matrix of the classifier; α represents the weight coefficient, which is used to adjust the fusion ratio of the spectrogram and MFCC features; β represents the weight coefficient, which is used to adjust the fusion ratio of HuBERT features; F spec Represents the final spectrum feature; F mfcc Represents the MFCC feature matrix; Represents element-by-element addition.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the non-contact intelligent health monitoring method based on physiological signals as described in any one of claims 1 to 8 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the non-contact intelligent health monitoring method based on physiological signals as described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Method and system for identifying, monitoring and early warning abnormal state of patient

    CN120878285A