Safety condition monitoring method and device for old people

Through feature extraction and fusion of facial videos and infrared videos, the accuracy problem of non-contact heart rate measurement methods under external interference is solved, and the accuracy and efficiency of heart rate monitoring are improved, which is suitable for real-time monitoring of the health status of the elderly.

CN120678407APending Publication Date: 2025-09-23广州新华学院
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510849488.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Non-contact heart rate measurement methods are susceptible to external interference. During exercise or when lighting changes significantly, measurements are inaccurate and inefficient, resulting in poor monitoring accuracy.

Method used

By obtaining the target user's facial video and infrared video, RGB features and NIR features are extracted respectively, and feature fusion is performed in the spatiotemporal feature fusion module to utilize multimodal information to improve the accuracy and efficiency of heart rate monitoring.

Benefits of technology

The accuracy and efficiency of non-contact heart rate monitoring are improved, and the user's health status can be monitored and evaluated in real time, which improves the user's comfort and monitoring accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120678407A_ABST
    Figure CN120678407A_ABST
Patent Text Reader

Abstract

The invention discloses a safety condition monitoring method and device for old people, and relates to the technical field of heart rate monitoring. Acquiring a face video and an infrared video of a target user to obtain a first final space-time diagram and a second final space-time diagram; respectively inputting the first final space-time diagram and the second final space-time diagram into a space-time feature extraction module for feature extraction to obtain RGB features and NIR features, and fusing the RGB features and the NIR features to obtain fused features; the heart rate of the target user is determined according to the fusion features, and state monitoring is conducted. Acquiring a face and an infrared video of a target user, and extracting spatio-temporal features of the face video according to a preset rule to obtain a first spatio-temporal diagram; the infrared video features are extracted and amplified frame by frame to obtain a second space-time diagram, RGB and NIR features are extracted through a space-time feature extraction module and fused into fused features, non-contact heart rate monitoring is achieved, the accuracy and efficiency of heart rate monitoring are improved, the heart rate is determined according to the fused features, and the health condition of a user is monitored and evaluated in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of heart rate monitoring, and in particular relates to a method and device for monitoring the safety status of an elderly person. Background Art

[0002] Heart rate is a key parameter for assessing cardiac physiological activity. It not only reflects the frequency of the heartbeat but is also directly linked to the heart's pumping function and blood circulation. By monitoring and analyzing heart rate, doctors can gain insight into heart health and promptly identify potential cardiovascular problems, providing crucial evidence for prevention and treatment.

[0003] Existing heart rate measurement methods primarily involve contact-based methods such as electrocardiography (ECG) and photoplethysmography (PPG). ECG uses electrode patches to capture the heart's electrical activity, providing detailed information about the heart's condition. PPG, on the other hand, uses light variations to sense blood flow in blood vessels to measure heart rate. However, these contact-based methods are relatively complex and difficult for the elderly to perform independently. Therefore, non-contact signal measurement methods address this difficulty for the elderly.

[0004] The above technologies still have the following problems, for example: the non-contact heart rate measurement method is easily affected by external interference, and the measurement is inaccurate and inefficient during exercise or when there are large changes in lighting, resulting in poor monitoring accuracy. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem that non-contact heart rate measurement methods are easily affected by external interference, and the measurement is inaccurate and inefficient during exercise or when the lighting changes greatly, resulting in poor monitoring accuracy. A method and device for monitoring the safety status of the elderly are proposed.

[0006] In a first aspect of the present invention, a method for monitoring the safety of an elderly person is first proposed, the method comprising: Obtaining a facial video and an infrared video of a target user, and extracting facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; Extract facial features from the infrared video according to preset rules to obtain an initial spatiotemporal feature atlas, and amplify the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; Inputting the first final spatiotemporal graph and the second final spatiotemporal graph into a spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, and inputting the RGB features and the NIR features into a spatiotemporal feature fusion module for feature fusion to obtain fused features; The heart rate of the target user is determined according to the fusion feature, and the state of the target user is monitored according to the heart rate.

[0007] Optionally, facial features are extracted from the facial video according to preset rules to obtain a first final spatiotemporal graph, wherein the preset rules include: Extracting frames from a facial video at a preset frequency to obtain a facial image set, and locating facial features on the facial image to obtain a facial feature area; the facial image set includes a plurality of facial images; The facial feature area is divided into i×j sub-areas, and the channels in each sub-area are average pooled to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The facial image set is processed frame by frame to obtain a space-time graph, each space-time graph is connected according to a preset frequency to obtain a space-time continuous graph, and the space-time continuous graph is transposed to obtain a first final space-time graph.

[0008] Optionally, the spatiotemporal feature extraction module is composed of a convolution detector and an attention screening mechanism, and the first final spatiotemporal graph and the second final spatiotemporal graph are input into the spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, including: Inputting the first final spatiotemporal graph and the second final spatiotemporal graph into a spatiotemporal feature extraction module respectively, performing a channel shuffling operation on the first final spatiotemporal graph and the second final spatiotemporal graph to obtain S1, S2, and S3; Perform dilated convolution on S1, S2, and S3 through parallel branches to obtain feature maps X1, X2, and X3 respectively. Connect the feature maps X1, X2, and X3, and restore the channels to the preset number to obtain the final feature map x. The first aggregated feature map and the second aggregated feature map are obtained by respectively encoding the horizontal coordinates and the vertical coordinates of the channels of the feature map x through the pooling kernel, the first aggregated feature map and the second aggregated feature map are connected to obtain a third aggregated feature map, and the third aggregated feature map is transformed by 1×1 convolution to obtain the final aggregated feature map; The final aggregated feature map is divided into a preset number of tensors along the spatial dimension, each tensor is converted into a tensor with the same number of channels as the input feature map, and the attention weight of the attention screening mechanism is determined according to the tensor to obtain RGB features and NIR features.

[0009] Optionally, the RGB feature and the NIR feature are input into a spatiotemporal feature fusion module for feature fusion to obtain a fused feature, including: The RGB features and the NIR features are respectively subjected to intra-modal fusion and inter-modal fusion through a spatiotemporal context module, and the RGB features and the NIR features are respectively subjected to intra-modal fusion to obtain first cross-modal information and second cross-modal information; Performing intermodal fusion on the first cross-modal information and the second cross-modal information to obtain Z1 and Z2 respectively; the first cross-modal information includes Q1, K1 and V1, and the second cross-modal information includes Q2, K2 and V2; The Z1 and the Z2 are aggregated through a convolutional layer to obtain a deep feature, and the deep feature is input into a spatiotemporal feature fusion module for feature fusion to obtain a fusion feature.

[0010] Optionally, the target user's status is monitored based on their heart rate, including: If the first preset threshold < the target user's heart rate ≤ the second preset threshold, increase the monitoring frequency; If the target user's heart rate is greater than a second preset threshold, an alarm is issued.

[0011] In a second aspect of the present invention, a device for monitoring the safety of an elderly person is provided, comprising: a video acquisition module, a feature extraction module, a feature fusion module, and a status monitoring module. The video acquisition module is used to acquire a facial video and an infrared video of a target user, and extract facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; The feature extraction module is used to extract facial features from the infrared video according to preset rules to obtain an initial spatiotemporal feature atlas, and to amplify the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; The feature fusion module is used to input the first final spatiotemporal graph and the second final spatiotemporal graph into the spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, and input the RGB features and the NIR features into the spatiotemporal feature fusion module for feature fusion to obtain fused features; The state monitoring module is used to determine the heart rate of the target user according to the fusion feature, and perform state monitoring on the target user according to the heart rate.

[0012] Optionally, the feature extraction module includes: a region division module, a sub-region processing module and a spatiotemporal graph connection module: The region division module is used to extract frames from the facial video at a preset frequency to obtain a facial image set, and locate facial features on the facial image to obtain facial feature regions; the facial image set includes multiple facial images; The sub-region processing module is used to divide the facial feature region into i×j sub-regions, and perform average pooling on the channels in each sub-region to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The spatiotemporal graph connection module is used to process the facial image set frame by frame to obtain a spatiotemporal graph, connect the spatiotemporal graphs according to a preset frequency to obtain a spatiotemporal continuous graph, and transpose the spatiotemporal continuous graph to obtain a first final spatiotemporal graph.

[0013] Optionally, the feature fusion module includes: a channel shuffling module, a dilated convolution module, a feature map aggregation module and an aggregated feature extraction module: The channel shuffling module is configured to input the first final spatiotemporal graph and the second final spatiotemporal graph into the spatiotemporal feature extraction module, and perform a channel shuffling operation on the first final spatiotemporal graph and the second final spatiotemporal graph to obtain S1, S2, and S3; The dilated convolution module is used to perform dilated convolution transfer on S1, S2 and S3 through parallel branches to obtain feature maps X1, X2 and X3 respectively, connect the feature maps X1, X2 and X3, and restore the channels to a preset number to obtain the final feature map x; The feature map aggregation module is used to perform horizontal coordinate encoding and vertical coordinate encoding on the channels of the feature map x through a pooling kernel to obtain a first aggregated feature map and a second aggregated feature map, concatenate the first aggregated feature map and the second aggregated feature map to obtain a third aggregated feature map, and perform a 1×1 convolution transformation on the third aggregated feature map to obtain a final aggregated feature map; The aggregate feature extraction module is used to split the final aggregate feature map into a preset number of tensors along the spatial dimension, convert each tensor into a tensor with the same number of channels as the input feature map, determine the attention weight of the attention screening mechanism based on the tensor, and obtain RGB features and NIR features.

[0014] Optionally, the feature fusion module further includes: an intra-modality fusion module, an inter-modality fusion module and a final feature fusion module: The intra-modality fusion module is used to perform intra-modality fusion and inter-modality fusion on the RGB features and the NIR features through a spatiotemporal context module, and perform intra-modality fusion on the RGB features and the NIR features to obtain first cross-modality information and second cross-modality information; The inter-modal fusion module is configured to perform inter-modal fusion on the first cross-modal information and the second cross-modal information to obtain Z1 and Z2 respectively; the first cross-modal information includes Q1, K1 and V1, and the second cross-modal information includes Q2, K2 and V2; The final feature fusion module is used to aggregate the Z1 and the Z2 through a convolutional layer to obtain a deep feature, and input the deep feature into the spatiotemporal feature fusion module for feature fusion to obtain a fused feature.

[0015] Optionally, the status monitoring module includes: a first alarm module and a second alarm module: The first alarm module is configured to increase the monitoring frequency if the first preset threshold is less than the heart rate of the target user and less than or equal to the second preset threshold; The second alarm module is configured to issue an alarm if the heart rate of the target user is greater than a second preset threshold.

[0016] Beneficial effects of the present invention: The present invention proposes a method for monitoring the safety status of the elderly. The method obtains a facial video and an infrared video of a target user to obtain a first final spatiotemporal graph and a second final spatiotemporal graph. The first final spatiotemporal graph and the second final spatiotemporal graph are respectively input into a spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, which are then fused to obtain fused features. The heart rate of the target user is determined based on the fused features, and the status of the target user is monitored based on the heart rate. The target user's facial and infrared videos are obtained, and spatiotemporal features are extracted from the facial videos according to preset rules to obtain a first spatiotemporal graph. After feature extraction, the infrared video is amplified frame by frame to obtain a second spatiotemporal graph. The RGB and NIR features are extracted by the spatiotemporal feature extraction module and fused into fused features to achieve non-contact heart rate monitoring, improve the accuracy and efficiency of heart rate monitoring, determine the heart rate based on the fused features, and monitor and evaluate the user's health status in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings.

[0018] Figure 1 A flowchart of a method for monitoring the safety status of an elderly person is provided in accordance with an embodiment of the present invention; Figure 2 Schematic diagram of a method for monitoring the safety status of an elderly person provided by an embodiment of the present invention Figure 3 A schematic structural diagram of a device for monitoring the safety status of an elderly person is provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments represent only a portion of the embodiments of the present invention, not all of them. The term "and / or" herein simply describes an association relationship between associated objects, indicating that three possible relationships exist. For example, "A" and "B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, references to "first," "second," and so on in the present invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of these features. Furthermore, the technical solutions of the various embodiments may be combined, but only if they are achievable by a person of ordinary skill in the art. If a combination of technical solutions contradicts or is unachievable, such combination shall be deemed non-existent and outside the scope of protection claimed by the present invention.

[0020] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.

[0021] The embodiment of the present invention provides a method for monitoring the safety status of an elderly person. Figure 1 , Figure 1 This is a flow chart of a method for monitoring the safety status of an elderly person provided by an embodiment of the present invention. The method includes the following steps: S101, obtaining a facial video and an infrared video of a target user, and extracting facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; S102, extracting facial features from the infrared video according to a preset rule to obtain an initial spatiotemporal feature atlas, and amplifying the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; S103: Input the first final spatiotemporal graph and the second final spatiotemporal graph into a spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, and input the RGB features and NIR features into a spatiotemporal feature fusion module for feature fusion to obtain fused features; S104: Determine the heart rate of the target user based on the fused features, and perform status monitoring on the target user based on the heart rate.

[0022] A method for monitoring the safety status of an elderly person provided by an embodiment of the present invention obtains the face and infrared video of a target user, extracts spatiotemporal features from the facial video according to preset rules, and obtains a first spatiotemporal graph. After feature extraction, the infrared video is amplified frame by frame to obtain a second spatiotemporal graph. RGB and NIR features are extracted by a spatiotemporal feature extraction module and fused into fused features to achieve non-contact heart rate monitoring, improve the accuracy and efficiency of heart rate monitoring, determine the heart rate based on the fused features, and monitor and evaluate the user's health status in real time.

[0023] In one implementation, facial videos correspond to RGB features (color features), and infrared videos correspond to NIR features (infrared features). The facial and infrared videos are acquired using a normal camera and an infrared camera, which are installed indoors to monitor the heart rate of the elderly. Heart rate prediction is performed in a non-contact manner, combining RGB and NIR features as input data for a neural network model. Multimodal data training makes heart rate prediction more accurate.

[0024] In one implementation, heart rate can be extracted from infrared video based on the reflection or transmission of infrared light by blood in human blood vessels. Near-infrared light can penetrate the skin, travel through blood, and be reflected and scattered by tissue before returning to the camera. Heartbeats cause blood to pulse, and hemoglobin in blood absorbs and scatters near-infrared light. Therefore, subtle variations in light intensity can be observed in the video as the heart beats. These variations correspond to the rhythm of the heart rate, allowing the heart rate signal to be extracted using signal processing and computer vision algorithms. Each heartbeat generates blood flow, producing a change in light intensity in the infrared video. Therefore, the heart rate value can be derived by calculating the frequency of these light intensity variations. This mapping relationship is stable and has certain commonalities across individuals. For example, suppose 10 distinct pulse signals related to heartbeats are extracted from a signal, with time intervals of 0.8, 0.82, 0.79, 0.81, 0.8, 0.83, 0.78, 0.8, 0.81, and 0.82 seconds, respectively. Based on the extracted pulse signal periods, the following calculation can be performed: Heart rate (bpm) = 60 / average period. In this example, the average period is approximately 0.806 seconds, resulting in a heart rate of approximately 74.4 bpm. This example is for illustrative purposes only.

[0025] In one implementation, the heart rate of the target user is determined based on the fused features, that is, the fused features are used as the input of the neural network model to obtain the heart rate of the target user; historical heart rate data and historical fused features are obtained (historical heart rate data and historical fused features are one-to-one corresponding), and the historical heart rate data and historical fused features are input into the neural network model (convolutional neural network, recurrent neural network and convolution + recurrent neural network) for training. After the above model update is completed, the neural network model will learn the mapping relationship from RGB features and NIR features to heart rate, and then determine the heart rate of the target user.

[0026] In one implementation, facial video and infrared video of the target user are acquired. Spatiotemporal features are extracted from the facial video according to preset rules to produce a first final spatiotemporal map containing spatial and temporal information of facial features. Feature extraction is then performed on the infrared video to form an initial spatiotemporal feature atlas, which is then zoomed in frame by frame over time to obtain a more refined second final spatiotemporal map. The spatiotemporal feature extraction module extracts RGB features and NIR features from these two types of spatiotemporal maps, respectively, and then fuses them into fused features in the spatiotemporal feature fusion module. This avoids the discomfort and limitations associated with traditional contact heart rate monitoring devices, improves user comfort and acceptance, and significantly enhances the accuracy of heart rate monitoring by fusing multimodal information. Based on the fused features, the target user's heart rate can be accurately determined, enabling real-time monitoring and assessment of their health status, improving the accuracy and robustness of heart rate monitoring.

[0027] In one implementation, the initial spatiotemporal feature atlas is magnified frame by frame in a time series to obtain a second final spatiotemporal image: each frame of the initial spatiotemporal feature atlas is decomposed, usually using a Laplacian pyramid decomposition for each frame, and magnified along the time axis t. On the decomposed image, a specific algorithm is used to detect small motion changes (motion changes are usually manifested as small fluctuations in pixel intensity, for example: in time Motion changes are detected at pixel locations (a, b) where motion is present. A specific motion frequency is selected for amplification, for example, 0.4 and 4 Hz, a range that covers typical human heart rates, including high rates such as supraventricular tachycardia. The amplification factor (e.g., α = 120) determines the degree of amplification. The amplified image components are reassembled into a complete image frame. This step aims to restore image integrity and coherence. To optimize computational efficiency, motion amplification is applied specifically to the extracted forehead region rather than the entire face.

[0028] In one embodiment, facial features are extracted from the facial video according to preset rules to obtain a first final spatiotemporal graph, wherein the preset rules include: Extract frames from the facial video at a preset frequency to obtain a facial image set, and locate facial features in the facial image to obtain facial feature areas; the facial image set includes multiple facial images; The facial feature area is divided into i×j sub-areas, and the channels in each sub-area are average pooled to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The facial image set is processed frame by frame to obtain a space-time graph, each space-time graph is connected according to a preset frequency to obtain a space-time continuous graph, and the space-time continuous graph is transposed to obtain a first final space-time graph.

[0029] In one implementation, the preset rule is a method for locating facial regions in videos. By extracting frames from facial videos at a preset frequency, the amount of data processed is reduced while preserving key facial information (reducing interference from non-skin areas, such as hair and background), significantly reducing noise interference and improving signal quality. This helps save computing resources and time when processing large amounts of facial video data. The facial feature region is divided into i×j subregions and average pooling is performed on each subregion, which captures facial features in greater detail, facilitating the identification and analysis of more specific facial details, and improving the accuracy and robustness of facial recognition.

[0030] In one implementation, the technology effectively fuses spatiotemporal information by processing facial regions in a video sequence frame by frame and concatenating the processed results in chronological order into a spatiotemporal graph. This fusion helps understand the dynamic changes in facial features and improves the accuracy of physiological signal extraction. The spatiotemporal graph is then transposed, with each row representing the average pixel value of a facial feature region evolving over time, and each column representing the average pixel value of a different facial feature region at a given time point. This transposition optimizes the data layout, facilitating subsequent analysis and processing.

[0031] In one implementation, the concatenation operation is to concatenate multiple feature vectors or matrices into a larger vector or matrix. In the above steps, the concatenation operation is mainly used in two aspects: to concatenate the feature vectors after processing the facial feature area. (n∈[1,N]) are connected to obtain feature representation This step is to merge the spatially distributed feature vectors into an overall representation for easy subsequent processing. (t∈[1,T]) are connected in time order to obtain the spatiotemporal graph M. This step is to merge the temporally distributed feature representations into a spatiotemporal graph, realizing the fusion of spatiotemporal information.

[0032] In one implementation, the transposition operation swaps the rows and columns of the space-time graph matrix. In the transposed spatiotemporal graph M, T represents the number of video frames, N represents the number of sub-regions in the facial feature region, and 3 represents the number of RGB channels of each block (i.e., three RGB channels). In the transposed spatiotemporal graph M, each row corresponds to a column of the original spatiotemporal graph, i.e., it represents the average pixel value of a facial feature region at different time points; each column corresponds to a row of the original spatiotemporal graph, i.e., it represents the average pixel value of the facial feature region at a certain time point. The transposition operation makes the data layout more in line with the needs of subsequent analysis, facilitating the extraction and analysis of the temporal evolution information of specific facial feature regions.

[0033] In one embodiment, the spatiotemporal feature extraction module is composed of a convolution detector and an attention screening mechanism. The first final spatiotemporal graph and the second final spatiotemporal graph are input into the spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, including: Inputting the first final spatiotemporal graph and the second final spatiotemporal graph into the spatiotemporal feature extraction module respectively, performing a channel shuffling operation on the first final spatiotemporal graph and the second final spatiotemporal graph to obtain S1, S2 and S3; Perform dilated convolution on S1, S2, and S3 through parallel branches to obtain feature maps X1, X2, and X3 respectively. Connect the feature maps X1, X2, and X3, and restore the channels to the preset number to obtain the final feature map x. The first aggregated feature map and the second aggregated feature map are obtained by respectively encoding the horizontal coordinates and the vertical coordinates of the channels of the feature map x through the pooling kernel, and the first aggregated feature map and the second aggregated feature map are connected to obtain the third aggregated feature map, and the third aggregated feature map is transformed by 1×1 convolution to obtain the final aggregated feature map; The final aggregated feature map is divided into a preset number of tensors along the spatial dimension, and each tensor is converted into a tensor with the same number of channels as the input feature map. The attention weight of the attention screening mechanism is determined according to the tensor to obtain RGB features and NIR features.

[0034] In one implementation, the contraction of the pulse (heart rate) causes slight changes in facial skin color over a continuous period of time, indicating that skin color changes between predicted frames are time-dependent. By using skin color changes over a continuous period to characterize the heart rate signal, we can obtain heart rate information in both temporal and spatial dimensions, thereby detecting subtle changes in skin color. Conversely, heart rate can be determined from skin color (RGB and NIR features). The spatiotemporal feature extraction module consists of a convolutional detector and an attention filtering mechanism. The convolutional detector uses dilated convolution to merge signals from different receptive fields into a signal pool. The attention filtering mechanism uses a coordinated attention mechanism to filter these signals, thereby selecting (filtering) heart rate signal features within the spatiotemporal feature map.

[0035] In one implementation, dilated convolution: S1, S2, and S3 are processed through dilated convolutions with dilation rates of 1, 3, and 5, respectively. Dilated convolution is a special convolution operation that allows the convolution kernel to expand the receptive field without increasing the number of parameters, thereby capturing a wider range of contextual information: Feature map concatenation: The output feature maps x1, x2, and x3 of the three branches after dilated convolution are concatenated, restoring the number of channels to the original c to obtain the final feature map x; Pooling: Two spatial range pooling kernels are used to encode each channel along the horizontal and vertical coordinates respectively. This can obtain the average value of each channel in the horizontal and vertical directions, thereby capturing spatial information; Coordinate Attention: The spatial information is encoded by performing average pooling in the horizontal and vertical directions, followed by a transformation operation, and finally the spatial information is fused by weighting the channels. This method can capture long-distance dependencies while retaining precise location information; Spatial Dimension Splitting: The aggregated feature map is split into two independent tensors along the spatial dimension, and then these tensors are converted into tensors with the same number of channels as the input feature map through a 1x1 convolution transformation. Sigmoid activation function: In coordinate attention, the sigmoid activation function is used to obtain attention weights, which are then used to weight the original feature map to enhance the features at specific locations.

[0036] In one implementation, S1, S2, and S3 are three parts of the feature map obtained after the channel shuffling operation. The channel shuffling operation is to disrupt the channel order of the input feature map and divide it into three equal parts. The purpose of this is to ensure that each part can contain part of the information of the original feature map in subsequent processing, so that features can be extracted from different angles in subsequent grouping and dilated convolution operations. By inputting the first final spatiotemporal map and the second final spatiotemporal map into the spatiotemporal feature extraction module and performing the channel shuffling operation, the expressive ability of the feature map can be enhanced, so that the network can learn richer and more diverse spatiotemporal features. The dilated convolution is performed on S1, S2 and S3 through parallel branches to obtain feature maps X1, X2 and X3, and they are connected. This step not only improves the efficiency of feature extraction, but also can fuse the feature information of different branches, further enhancing the robustness of the feature map. The connected feature map channels are restored to the preset number (the feature map channels are restored to the preset number, for example: the feature map channels are restored to c; c refers to the cth channel in the feature map, which is a one-dimensional array representing the response intensity of the feature at all positions in the channel), and the final feature map x is obtained. This process ensures the consistency of the feature map in subsequent processing. At the same time, the horizontal and vertical coordinates of the channels of the feature map are encoded through the pooling kernel to obtain an aggregated feature map. This process helps to capture the spatial distribution information of the feature map; the final aggregated feature map is divided into multiple tensors along the spatial dimension, and the attention weights of the attention screening mechanism are determined based on these tensors. This step can adaptively screen out feature information for the final task (heart rate monitoring), thereby improving the accuracy and robustness of the model; by converting the features after attention screening into a tensor with the same number of channels as the input feature map, and obtaining RGB features and NIR features respectively, this process realizes the fusion and representation of multi-feature information, which helps the model better understand and utilize the diversity of input data.

[0037] In one embodiment, the RGB features and the NIR features are input into the spatiotemporal feature fusion module for feature fusion to obtain fused features, including: The RGB features and NIR features are fused intramodally and intermodally through the spatiotemporal context module, and the first cross-modal information and the second cross-modal information are obtained by intra-modal fusion of the RGB features and the NIR features. Perform intermodal fusion on the first cross-modal information and the second cross-modal information to obtain Z1 and Z2 respectively; the first cross-modal information includes Q1, K1 and V1, and the second cross-modal information includes Q2, K2 and V2; Z1 and Z2 are aggregated through the convolution layer to obtain deep features, and the deep features are input into the spatiotemporal feature fusion module for feature fusion to obtain fused features.

[0038] In one implementation, see Figure 2 , Figure 2 This is a schematic diagram of a method for monitoring the safety of an elderly person, provided by an embodiment of the present invention. Modules 1 and 3 form the first stage, and modules 2 and 4 form the second stage. Modules 1 and 2 are shown above, and modules 3 and 4 are shown below. The RGB feature X1 is the input to the spatiotemporal context module (above), and the NIR feature X2 is the input to the spatiotemporal context module (below). Each input data type is processed through independent PatchEmbedding layers, which divide the image into small patches and map them into a high-dimensional embedding space. Phase 1: RGB branch (above): Q1, K1, and V1 are the query vector, key vector, and value vector generated from the RGB data, respectively. A multi-head self-attention mechanism (MSA) uses these vectors to capture the internal relationships of the RGB data. The output passes through Layer Normalization (LN) and a fully connected layer (MLP), with residual connections (RGB residual) used to preserve the input information. NIR branch (below): Q2, K2, and V2 are the query, key, and value generated from the NIR data. Similarly, MSA captures the feature relationships within the NIR data, and the output passes through Layer Nomm and Multilayer Perceptron (MLP). Residual connections (NIR Residual) preserve the original information. The second stage: A multi-modal cross-attention mechanism (MCA) is introduced to integrate RGB and NIR features. For RGB data, Q1 from RGB interacts with K2 and V2 from NIR. The cross-modal attention mechanism combines the features of RGB data with those of NIR data. For NIR data, Q2 from NIR interacts with K1 and V1 from RGB, fusing the RGB information. The output features are further processed by Layer Nomm (LN) and Multilayer Perceptron (MLP). Residual connections (RGB Residual and NIR Residual) preserve the information after intermodal fusion. Overall Process: In the first stage, the model independently extracts features from RGB and NIR images. In the second stage, the information from the two modalities is fused through a cross-modal attention mechanism. Multiple Laver Nomm, MLP, and Residual Connections help maintain stability and prevent gradient vanishing. Multi-head Self-Attention (MSA) captures global relationships within a single modality. Cross-modal Attention (MCA) enables information exchange and feature fusion between RGB and NIR modalities. Residual Connections preserve input information, reduce information loss, and improve training efficiency.

[0039] In one implementation, in the actual measurement process, the non-contact measurement effect is not good, so it is necessary to supplement the RGB and NIR modal information to improve the robustness of the heart rate measurement; in order to achieve comprehensive global information aggregation and prevent the loss of the contextual relationship of information in continuous time periods and adjacent spatial ranges, feature fusion is performed through the spatiotemporal feature fusion module (composed of three Refine modules, each task branch includes three Refine modules, and the Refine module refines the fused features at each level to extract task-oriented features and improve the heart rate estimation performance); the Refine module has two inputs: deep features and . represents the features of the nth task learned by the m−1th Refine module; the output of the Refine module can be expressed as: ConvBlock1 consists of a 3×3 DSConv (distributed shift convolution), a batch normalization layer, and a sigmoid layer; ConvBlock2 consists of a 3×3 convolution, a batch normalization layer, a ReLU layer, and a pooling layer. ⊕ represents element-wise addition. DAMBlock (⋅) includes both channel attention and spatial attention.

[0040] In one implementation, this method leverages the internal information of each modality by fusing RGB and NIR features separately, enhancing the expressiveness and robustness of each feature. Intra-modal fusion helps capture the inherent connections and subtle differences between data within the same modality, thereby improving feature effectiveness. After determining the first cross-modal information (Q1, K1, V1) and the second cross-modal information (Q2, K2, V2), inter-modal fusion is performed to obtain Z1 (Q1, K2, V2) and Z2 (Q2, K1, V1). This process promotes information interaction and complementarity between the RGB and NIR modalities, so that the fused features can simultaneously contain key information from different modalities, improving the comprehensiveness and accuracy of the features; through the cross-modal fusion (inter-modal fusion) strategy and combined with the convolutional layer to aggregate deep features, this method effectively improves the efficiency of feature fusion; finally, the deep features are input into the spatiotemporal feature fusion module for feature fusion to obtain fused features. This step not only considers the spatial distribution of the features, but also incorporates information in the time dimension, thereby achieving comprehensive extraction and efficient fusion of spatiotemporal features.

[0041] In one embodiment, monitoring the target user's status based on the heart rate includes: If the first preset threshold < the target user's heart rate ≤ the second preset threshold, increase the monitoring frequency; If the target user's heart rate is greater than a second preset threshold, an alarm is issued.

[0042] In one implementation, if the first preset threshold is less than the target user's heart rate and less than or equal to the second preset threshold, the system will first increase the target user's monitoring frequency. This is intended to more closely observe the user's heart rate trends and promptly identify potential health risks. By increasing the frequency of monitoring, the system can capture more details about the user's heart rate status, providing more comprehensive and accurate data support for subsequent health analysis and decision-making.

[0043] In one implementation, when the target user's heart rate exceeds (is greater than) the second preset threshold, that is, enters a higher-risk range, the system will immediately activate a more urgent alarm mechanism. This alarm method not only attracts the user's attention through simple sound or visual prompts, but also adopts a strategy that combines multiple alarm methods based on the user's specific situation (for example, age, health status, past medical history, etc.) and the urgency of the current heart rate. For example: for elderly users or users with poor health, the system may choose a more intuitive and easy-to-understand graphical alarm, and at the same time cooperate with voice prompts to ensure that users can quickly understand the current emergency situation; or, the system will automatically trigger a linkage mechanism with medical institutions to transmit the user's health data to medical professionals in real time, so that medical staff can provide remote guidance or arrange emergency rescue as soon as possible. The above alarm method effectively improves the efficiency and accuracy of responding to emergencies.

[0044] Based on the same inventive concept, the present invention also provides a device for monitoring the safety of the elderly. Figure 3 , Figure 3 A schematic diagram of the structure of a device for monitoring the safety status of an elderly person provided by an embodiment of the present invention includes: a video acquisition module, a feature extraction module, a feature fusion module, and a status monitoring module: A video acquisition module is used to acquire a facial video and an infrared video of a target user, and extract facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; A feature extraction module is used to extract facial features from the infrared video according to preset rules to obtain an initial spatiotemporal feature atlas, and to amplify the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; A feature fusion module is used to input the first final space-time map and the second final space-time map into the space-time feature extraction module for feature extraction to obtain RGB features and NIR features, and input the RGB features and NIR features into the space-time feature fusion module for feature fusion to obtain fused features; The state monitoring module is used to determine the heart rate of the target user based on the fusion features and perform state monitoring on the target user based on the heart rate.

[0045] A device for monitoring the safety of an elderly person provided by an embodiment of the present invention obtains a target user's face and infrared video, extracts spatiotemporal features from the facial video according to preset rules, and obtains a first spatiotemporal graph. After feature extraction, the infrared video is amplified frame by frame to obtain a second spatiotemporal graph. A spatiotemporal feature extraction module extracts RGB and NIR features and fuses them into fused features to achieve non-contact heart rate monitoring, improve the accuracy and efficiency of heart rate monitoring, determine the heart rate based on the fused features, and monitor and evaluate the user's health status in real time.

[0046] In one embodiment, the feature extraction module includes: a region division module, a sub-region processing module, and a spatiotemporal graph connection module: A region division module is used to extract frames from a facial video at a preset frequency to obtain a facial image set, and locate facial features in the facial image to obtain facial feature regions; the facial image set includes multiple facial images; The sub-region processing module is used to divide the facial feature area into i×j sub-regions and perform average pooling on the channels in each sub-region to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The spatiotemporal graph connection module is used to process the facial image set frame by frame to obtain a spatiotemporal graph, connect the spatiotemporal graphs according to a preset frequency to obtain a spatiotemporal continuous graph, and transpose the spatiotemporal continuous graph to obtain a first final spatiotemporal graph.

[0047] In one embodiment, the feature fusion module includes: a channel shuffling module, a dilated convolution module, a feature map aggregation module, and an aggregated feature extraction module: a channel shuffling module, configured to input the first final space-time graph and the second final space-time graph into the space-time feature extraction module respectively, and perform a channel shuffling operation on the first final space-time graph and the second final space-time graph to obtain S1, S2, and S3; The dilated convolution module is used to perform dilated convolution on S1, S2, and S3 through parallel branches to obtain feature maps X1, X2, and X3 respectively, connect the feature maps X1, X2, and X3, and restore the channels to the preset number to obtain the final feature map x; A feature map aggregation module is used to encode the horizontal coordinates and vertical coordinates of the channels of the feature map x using a pooling kernel to obtain a first aggregated feature map and a second aggregated feature map, concatenate the first aggregated feature map and the second aggregated feature map to obtain a third aggregated feature map, and perform a 1×1 convolution transformation on the third aggregated feature map to obtain a final aggregated feature map; The aggregate feature extraction module is used to split the final aggregate feature map into a preset number of tensors along the spatial dimension, convert each tensor into a tensor with the same number of channels as the input feature map, determine the attention weight of the attention screening mechanism based on the tensor, and obtain RGB features and NIR features.

[0048] In one embodiment, the feature fusion module further includes: an intra-modality fusion module, an inter-modality fusion module, and a final feature fusion module: The intra-modal fusion module is used to perform intra-modal fusion and inter-modal fusion on RGB features and NIR features through the spatiotemporal context module, respectively, to obtain the first cross-modal information and the second cross-modal information by performing intra-modal fusion on RGB features and NIR features respectively. An intermodal fusion module is configured to intermodally fuse the first cross-modal information and the second cross-modal information to obtain Z1 and Z2, respectively; the first cross-modal information includes Q1, K1, and V1, and the second cross-modal information includes Q2, K2, and V2; The final feature fusion module is used to aggregate Z1 and Z2 through the convolution layer to obtain deep features, and input the deep features into the spatiotemporal feature fusion module for feature fusion to obtain fused features.

[0049] In one embodiment, the status monitoring module includes: a first alarm module and a second alarm module: A first alarm module is configured to increase the monitoring frequency if the first preset threshold is less than the heart rate of the target user and less than or equal to the second preset threshold; The second alarm module is configured to issue an alarm if the heart rate of the target user exceeds a second preset threshold.

[0050] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A method for monitoring the safety status of an elderly person, characterized in that: The method comprises: Obtaining a facial video and an infrared video of a target user, and extracting facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; Extract facial features from the infrared video according to preset rules to obtain an initial spatiotemporal feature atlas, and amplify the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; Inputting the first final spatiotemporal graph and the second final spatiotemporal graph into a spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, and inputting the RGB features and the NIR features into a spatiotemporal feature fusion module for feature fusion to obtain fused features; The heart rate of the target user is determined according to the fusion feature, and the state of the target user is monitored according to the heart rate.

2. The method for monitoring the safety of an elderly person according to claim 1, wherein: The facial features of the facial video are extracted according to preset rules to obtain a first final spatiotemporal graph, wherein the preset rules include: Extracting frames from a facial video at a preset frequency to obtain a facial image set, and locating facial features on the facial image to obtain a facial feature area; the facial image set includes a plurality of facial images; The facial feature area is divided into i×j sub-areas, and the channels in each sub-area are average pooled to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The facial image set is processed frame by frame to obtain a space-time graph, each space-time graph is connected according to a preset frequency to obtain a space-time continuous graph, and the space-time continuous graph is transposed to obtain a first final space-time graph.

3. The method for monitoring the safety status of an elderly person according to claim 1, wherein: The spatiotemporal feature extraction module is composed of a convolution detector and an attention screening mechanism. The first final spatiotemporal graph and the second final spatiotemporal graph are input into the spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, including: Inputting the first final spatiotemporal graph and the second final spatiotemporal graph into a spatiotemporal feature extraction module respectively, performing a channel shuffling operation on the first final spatiotemporal graph and the second final spatiotemporal graph to obtain S1, S2, and S3; Perform dilated convolution on S1, S2, and S3 through parallel branches to obtain feature maps X1, X2, and X3 respectively. Connect the feature maps X1, X2, and X3, and restore the channels to the preset number to obtain the final feature map x. The first aggregated feature map and the second aggregated feature map are obtained by respectively encoding the horizontal coordinates and the vertical coordinates of the channels of the feature map x through the pooling kernel, the first aggregated feature map and the second aggregated feature map are connected to obtain a third aggregated feature map, and the third aggregated feature map is transformed by 1×1 convolution to obtain the final aggregated feature map; The final aggregated feature map is divided into a preset number of tensors along the spatial dimension, each tensor is converted into a tensor with the same number of channels as the input feature map, and the attention weight of the attention screening mechanism is determined according to the tensor to obtain RGB features and NIR features.

4. The method for monitoring the safety of an elderly person according to claim 1, wherein: Inputting the RGB features and the NIR features into a spatiotemporal feature fusion module for feature fusion to obtain fused features, including: The RGB features and the NIR features are respectively subjected to intra-modal fusion and inter-modal fusion through a spatiotemporal context module, and the RGB features and the NIR features are respectively subjected to intra-modal fusion to obtain first cross-modal information and second cross-modal information; Performing intermodal fusion on the first cross-modal information and the second cross-modal information to obtain Z1 and Z2 respectively; the first cross-modal information includes Q1, K1 and V1, and the second cross-modal information includes Q2, K2 and V2; The Z1 and the Z2 are aggregated through a convolutional layer to obtain a deep feature, and the deep feature is input into a spatiotemporal feature fusion module for feature fusion to obtain a fusion feature.

5. The method for monitoring the safety status of an elderly person according to claim 1, characterized in that: Monitor the target user's status based on their heart rate, including: If the first preset threshold < the target user's heart rate ≤ the second preset threshold, increase the monitoring frequency; If the target user's heart rate is greater than a second preset threshold, an alarm is issued.

6. A safety status monitoring device for the elderly, characterized in that: The device includes: a video acquisition module, a feature extraction module, a feature fusion module and a status monitoring module: The video acquisition module is used to acquire a facial video and an infrared video of a target user, and extract facial features from the facial video according to preset rules to obtain a first final spatiotemporal graph; The feature extraction module is used to extract facial features from the infrared video according to preset rules to obtain an initial spatiotemporal feature atlas, and to amplify the initial spatiotemporal feature atlas frame by frame in a time sequence to obtain a second final spatiotemporal map; The feature fusion module is used to input the first final spatiotemporal graph and the second final spatiotemporal graph into the spatiotemporal feature extraction module for feature extraction to obtain RGB features and NIR features, and input the RGB features and the NIR features into the spatiotemporal feature fusion module for feature fusion to obtain fused features; The state monitoring module is used to determine the heart rate of the target user according to the fusion feature, and perform state monitoring on the target user according to the heart rate.

7. The device for monitoring the safety of an elderly person according to claim 6, characterized in that: The feature extraction module includes: a region division module, a sub-region processing module and a spatiotemporal graph connection module: The region division module is used to extract frames from the facial video at a preset frequency to obtain a facial image set, and locate facial features on the facial image to obtain facial feature regions; the facial image set includes multiple facial images; The sub-region processing module is used to divide the facial feature region into i×j sub-regions, and perform average pooling on the channels in each sub-region to obtain , the sub-regions Connect to get ; represents the first region of N regions, which has three channels; The spatiotemporal graph connection module is used to process the facial image set frame by frame to obtain a spatiotemporal graph, connect the spatiotemporal graphs according to a preset frequency to obtain a spatiotemporal continuous graph, and transpose the spatiotemporal continuous graph to obtain a first final spatiotemporal graph.

8. The elderly safety monitoring device according to claim 6, characterized in that: The feature fusion module includes: a channel shuffling module, a dilated convolution module, a feature map aggregation module and an aggregated feature extraction module: The channel shuffling module is configured to input the first final spatiotemporal graph and the second final spatiotemporal graph into the spatiotemporal feature extraction module, and perform a channel shuffling operation on the first final spatiotemporal graph and the second final spatiotemporal graph to obtain S1, S2, and S3; The dilated convolution module is used to perform dilated convolution transfer on S1, S2 and S3 through parallel branches to obtain feature maps X1, X2 and X3 respectively, connect the feature maps X1, X2 and X3, and restore the channels to a preset number to obtain the final feature map x; The feature map aggregation module is used to perform horizontal coordinate encoding and vertical coordinate encoding on the channels of the feature map x through a pooling kernel to obtain a first aggregated feature map and a second aggregated feature map, concatenate the first aggregated feature map and the second aggregated feature map to obtain a third aggregated feature map, and perform a 1×1 convolution transformation on the third aggregated feature map to obtain a final aggregated feature map; The aggregate feature extraction module is used to split the final aggregate feature map into a preset number of tensors along the spatial dimension, convert each tensor into a tensor with the same number of channels as the input feature map, determine the attention weight of the attention screening mechanism based on the tensor, and obtain RGB features and NIR features.

9. The elderly safety monitoring device according to claim 6, characterized in that: The feature fusion module also includes: an intra-modality fusion module, an inter-modality fusion module and a final feature fusion module: The intra-modality fusion module is used to perform intra-modality fusion and inter-modality fusion on the RGB features and the NIR features through a spatiotemporal context module, and perform intra-modality fusion on the RGB features and the NIR features to obtain first cross-modality information and second cross-modality information; The inter-modal fusion module is configured to perform inter-modal fusion on the first cross-modal information and the second cross-modal information to obtain Z1 and Z2 respectively; the first cross-modal information includes Q1, K1 and V1, and the second cross-modal information includes Q2, K2 and V2; The final feature fusion module is used to aggregate the Z1 and the Z2 through a convolutional layer to obtain a deep feature, and input the deep feature into the spatiotemporal feature fusion module for feature fusion to obtain a fused feature.

10. The elderly safety monitoring device according to claim 6, characterized in that: The status monitoring module includes: a first alarm module and a second alarm module: The first alarm module is configured to increase the monitoring frequency if the first preset threshold is less than the heart rate of the target user and less than or equal to the second preset threshold; The second alarm module is configured to issue an alarm if the heart rate of the target user is greater than a second preset threshold.

Citation Information

Cited By

  • Physiological parameter detection method and device based on RGB-NIR combination, terminal and storage medium

    CN121370108A

  • Physiological parameter detection method, device, terminal and storage medium based on RGB-NIR combination

    CN121370108B