Physiological signal extraction method based on self-adaptive ROI and ViM dual paths

By combining adaptive ROI and ViM dual-path physiological signal extraction method with adaptive region of interest selection and low-rank tensor completion, and utilizing the Vision Mamba dual-path network structure, the robustness and computational complexity issues of physiological signal extraction in existing technologies are solved, and efficient and accurate physiological signal extraction is achieved in different scenarios.

CN122045784APending Publication Date: 2026-05-15BEIJING SPORT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SPORT UNIV
Filing Date
2025-12-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing non-contact physiological signal extraction technologies struggle to simultaneously capture the long-term temporal dependence of heart rate signals and remove spatiotemporal redundancy. They also suffer from high computational complexity, and robustness is affected by facial occlusion, motion artifacts, and changes in illumination. Furthermore, a single fixed model cannot adequately address the differentiated needs of both static and dynamic scenarios.

Method used

A physiological signal extraction method using adaptive ROI and ViM dual-path is adopted. A quality-aware spatiotemporal graph is constructed through an adaptive region of interest selection mechanism and a low-rank tensor completion method. Then, the time-frequency domain features are extracted by combining the time path and the frequency path using the Vision Mamba dual-path network structure.

Benefits of technology

It improves the robustness, accuracy and efficiency of physiological signal extraction, solves the regional quality fluctuations caused by facial occlusion, motion artifacts and lighting changes, and meets the needs of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045784A_ABST
    Figure CN122045784A_ABST
Patent Text Reader

Abstract

The invention discloses a physiological signal extraction method based on self-adaptive ROI and ViM dual paths, which comprises the following steps: constructing a first quality perception space-time diagram of a first user and a second quality perception space-time diagram of a second user based on a self-adaptive region-of-interest selection mechanism and a low-rank tensor completion method; extracting a first time sequence prediction signal of the first quality perception space-time diagram and a second time sequence prediction signal of the second quality perception space-time diagram through a time path of the ViM dual-path network structure, and extracting a first frequency domain prediction signal of the first quality perception space-time diagram and a second frequency domain prediction signal of the second quality perception space-time diagram through a frequency path; and fusing the first time sequence prediction signal and the first frequency domain prediction signal to obtain a first physiological signal of the first user, and fusing the second time sequence prediction signal and the second frequency domain prediction signal to obtain a second physiological prediction signal of the second user. According to the method, the robustness, precision and efficiency of physiological signal extraction are improved, and different scene requirements can be considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and remote physiological signal monitoring technology, and in particular to a physiological signal extraction method based on adaptive ROI and ViM dual paths. Background Technology

[0002] Current non-contact physiological signal extraction technologies face multiple challenges. On the one hand, traditional deep learning methods struggle to simultaneously capture the long-term temporal dependencies of heart rate signals and remove spatiotemporal redundancy. On the other hand, mainstream spatiotemporal modeling techniques, such as 3D CNNs or Transformers, suffer from excessively high computational complexity, limiting their deployment on resource-constrained devices. Furthermore, in real-world scenarios, facial occlusion, motion artifacts, and lighting variations can easily lead to fluctuations in the quality of regions of interest (ROIs), severely impacting the robustness of signal extraction. Moreover, a single, fixed model cannot adequately address the differentiated needs of both static and dynamic scenes, making it difficult to achieve an effective balance between accuracy and efficiency.

[0003] Therefore, there is an urgent need to provide a technical solution to address the above problems. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a method, system, electronic device, and storage medium for extracting physiological signals based on adaptive ROI and ViM dual paths.

[0005] Firstly, this invention provides a physiological signal extraction method based on adaptive ROI and ViM dual paths, the technical solution of which is as follows: Based on the adaptive region of interest selection mechanism and low-rank tensor completion method, the first quality perception spatiotemporal graph of the first user and the second quality perception spatiotemporal graph of the second user are constructed. The first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map are input into a Vision Mamba dual-path network structure. The first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map are extracted through the time path of the Vision Mamba dual-path network structure. The first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map are extracted through the frequency path of the Vision Mamba dual-path network structure. The first time-series prediction signal and the first frequency-domain prediction signal are fused to obtain the first physiological prediction signal of the first user, and the second time-series prediction signal and the second frequency-domain prediction signal are fused to obtain the second physiological prediction signal of the second user. Output the first physiological prediction signal and the second physiological prediction signal.

[0006] The beneficial effects of the physiological signal extraction method based on adaptive ROI and ViM dual paths of the present invention are as follows: The method of this invention constructs a quality-aware spatiotemporal graph through an adaptive region of interest selection mechanism and low-rank tensor completion, and simultaneously extracts time-frequency domain features by combining the Vision Mamba dual-path network. This solves the problems of regional quality fluctuations caused by facial occlusion, motion artifacts, and changes in illumination, as well as the high computational complexity and insufficient long-period dependency capture of traditional methods. It improves the robustness, accuracy and efficiency of physiological signal extraction, and takes into account the needs of different scenarios.

[0007] Based on the above scheme, the physiological signal extraction method based on adaptive ROI and ViM dual paths of the present invention can be further improved as follows.

[0008] In one alternative approach, the step of constructing the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user based on the adaptive region of interest selection mechanism and low-rank tensor completion method includes: Extract a first facial region from multiple consecutive frames of the first user's facial video data, and extract a second facial region from multiple consecutive frames of the second user's facial video data. Through the adaptive region of interest selection mechanism, the green channel intensity and signal-to-noise ratio of each first preset ROI region in the first facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K first preset ROI regions are determined as the first ROI regions. Similarly, the green channel intensity and signal-to-noise ratio of each second preset ROI region in the second facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K second preset ROI regions are determined as the second ROI regions. The RGB signals corresponding to each first ROI region are spliced ​​together to obtain the first original spatiotemporal map of the first user, and the RGB signals corresponding to each second ROI region are spliced ​​together to obtain the second original spatiotemporal map of the second user. The low-rank tensor completion method is used to repair the first original spatiotemporal graph to obtain the first quality-aware spatiotemporal graph, and the second original spatiotemporal graph is repaired to obtain the second quality-aware spatiotemporal graph.

[0009] In one alternative approach, the time path includes: a physiological encoder and a non-physiological encoder; the step of extracting the first temporal prediction signal corresponding to the first quality-aware spatiotemporal map and the second temporal prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure includes: The physiological encoder obtains the first original physiological features of the first quality-perceived spatiotemporal map and the second original physiological features of the second quality-perceived spatiotemporal map, and the non-physiological encoder obtains the first original non-physiological features of the first quality-perceived spatiotemporal map and the second original non-physiological features of the second quality-perceived spatiotemporal map. The first original physiological feature and the first original non-physiological feature are reconstructed to obtain a first reconstructed spatiotemporal map; the first original physiological feature and the second original non-physiological feature are reconstructed to obtain a second reconstructed spatiotemporal map; the first original non-physiological feature and the second original physiological feature are reconstructed to obtain a third reconstructed spatiotemporal map; and the second original physiological feature and the second original non-physiological feature are reconstructed to obtain a fourth reconstructed spatiotemporal map. The physiological encoder obtains the first reconstructed physiological features of the second reconstructed spatiotemporal map and the second reconstructed physiological features of the third reconstructed spatiotemporal map, and the non-physiological encoder obtains the first reconstructed non-physiological features of the second reconstructed spatiotemporal map and the second reconstructed non-physiological features of the third reconstructed spatiotemporal map. The first original physiological feature is spliced ​​with the first reconstructed physiological feature to obtain the first spliced ​​feature, which is then input into the first regression layer to obtain the first time-series prediction signal. The second original physiological feature is spliced ​​with the second reconstructed physiological feature to obtain the second spliced ​​feature, which is then input into the second regression layer to obtain the second time-series prediction signal.

[0010] In one alternative approach, the step of extracting the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure includes: Fourier transforms are performed on the first quality-sensing spatiotemporal map and the second quality-sensing spatiotemporal map respectively to obtain the first frequency domain signal of the first quality-sensing spatiotemporal map and the second frequency domain signal of the second quality-sensing spatiotemporal map; The first frequency domain signal is encoded by a frequency domain encoder to obtain a first frequency domain feature, and the second frequency domain signal is encoded to obtain a second frequency domain feature. The first frequency domain feature and the second frequency domain feature are passed sequentially through a first linear layer, a normalization layer, an activation layer, and a second linear layer to obtain the first frequency domain prediction signal corresponding to the first frequency domain feature and the second frequency domain prediction signal corresponding to the second frequency domain feature.

[0011] In one optional approach, the steps of fusing the first time-series prediction signal and the first frequency-domain prediction signal to obtain the first physiological prediction signal of the first user, and fusing the second time-series prediction signal and the second frequency-domain prediction signal to obtain the second physiological prediction signal of the second user, include: The first time-series prediction signal and the first frequency-domain prediction signal are fused through a fusion layer containing a multilayer perceptron to obtain the first physiological prediction signal of the first user. The second time-series prediction signal and the second frequency-domain prediction signal are fused through the fusion layer to obtain the second physiological prediction signal of the second user.

[0012] In one alternative approach, the prediction process of physiological signals employs a two-stage training strategy; the first stage uses the L1 loss function for training, and the second stage uses the L1+a·Lp loss function for training; where p in the Lp loss function is a preset hyperparameter, and 0 <p<1。

[0013] In one alternative approach, the L1 loss function is used to measure the differences between spatiotemporal maps, the differences between features, and the differences between predicted physiological signals and actual physiological signals. The differences between the spatiotemporal maps are determined by calculating the difference between the first quality-perceived spatiotemporal map and the first reconstructed spatiotemporal map, and by calculating the difference between the second quality-perceived spatiotemporal map and the fourth reconstructed spatiotemporal map; the differences between the features are determined by calculating the differences between the first original physiological feature and the first reconstructed physiological feature, the differences between the first original non-physiological feature and the first reconstructed non-physiological feature, the differences between the second original physiological feature and the second reconstructed physiological feature, and the differences between the second original non-physiological feature and the second reconstructed non-physiological feature; the differences between the physiological signals are determined by calculating the difference between the first physiological prediction signal and the first physiological real signal, and by calculating the difference between the second physiological prediction signal and the second physiological real signal.

[0014] Secondly, this invention provides a physiological signal extraction system based on adaptive ROI and ViM dual paths, the technical solution of which is as follows: The module is constructed based on an adaptive region of interest selection mechanism and a low-rank tensor completion method to build the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user. The extraction module is used to input the first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map into a Vision Mamba dual-path network structure, extract the first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure, and extract the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure. The fusion module is used to fuse the first time-series prediction signal and the first frequency-domain prediction signal to obtain the first physiological prediction signal of the first user, and to fuse the second time-series prediction signal and the second frequency-domain prediction signal to obtain the second physiological prediction signal of the second user. The output module is used to output the first physiological prediction signal and the second physiological prediction signal.

[0015] The beneficial effects of the physiological signal extraction system based on adaptive ROI and ViM dual paths of the present invention are as follows: The system of this invention constructs a quality-aware spatiotemporal graph through an adaptive region of interest selection mechanism and low-rank tensor completion, and simultaneously extracts time-frequency domain features by combining the Vision Mamba dual-path network. This solves the problems of regional quality fluctuations caused by facial occlusion, motion artifacts, and changes in illumination, as well as the high computational complexity and insufficient long-period dependency capture of traditional methods. It improves the robustness, accuracy, and efficiency of physiological signal extraction, and takes into account the needs of different scenarios.

[0016] Thirdly, the technical solution of an electronic device according to the present invention is as follows: It includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the physiological signal extraction method based on adaptive ROI and ViM dual paths as described in this invention.

[0017] Fourthly, the technical solution of a computer-readable storage medium provided by the present invention is as follows: The computer-readable storage medium stores instructions that, when read, cause the computer-readable storage medium to perform the steps of the physiological signal extraction method based on adaptive ROI and ViM dual paths of the present invention.

[0018] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0019] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating an embodiment of a physiological signal extraction method based on adaptive ROI and ViM dual paths according to the present invention. Figure 2 Schematic diagram of the process for generating a quality-aware spatiotemporal map (RQS-STMap) Figure 3 This is a schematic diagram of the overall framework; Figure 4 A schematic diagram of the general adaptive rPPG framework; Figure 5 A schematic diagram of the low-rank tensor completion process; Figure 6 This is a schematic diagram illustrating the change in training loss during two-stage training. Figure 7 This is a schematic diagram illustrating the change in test loss during the two-stage training loss process. Figure 8 This is a schematic diagram of an embodiment of a physiological signal extraction system based on adaptive ROI and ViM dual paths according to the present invention. Figure 9 This is a schematic diagram of an embodiment of an electronic device according to the present invention. Detailed Implementation

[0020] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0021] Figure 1This diagram illustrates a flowchart of an embodiment of a physiological signal extraction method based on an adaptive ROI and ViM dual-path approach provided by the present invention. This method can be executed by electronic devices such as terminal devices or servers. The terminal device can be any fixed or mobile terminal, such as user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the physiological signal extraction method based on the adaptive ROI and ViM dual-path approach by having its processor call computer-readable instructions stored in its memory. Figure 1 As shown, it includes the following steps: S1. Based on the adaptive region of interest selection mechanism and low-rank tensor completion method, construct the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user.

[0022] The adaptive region of interest (ROI) selection mechanism refers to an algorithm that dynamically selects high-quality regions from facial videos for signal extraction. It evaluates and selects regions in real-time based on the signal-to-noise ratio (SNR) of the green channel signal. For example, when processing user A's video, the algorithm calculates the SNR of sub-regions such as the forehead and cheeks frame by frame, dynamically selecting the region with the highest SNR in the current frame as the ROI, rather than a fixed region. The low-rank tensor completion method is a mathematical method based on the low-rank assumption to repair missing data, used to restore parts of the spatiotemporal map damaged by occlusion. For example, if some data in user A's original spatiotemporal map is missing due to blinking, this method uses the inherent structure of the data to fill the gaps and generate a complete spatiotemporal map. The first user refers to one specific object from which physiological signals are extracted. For example, in a dual-channel video processing scenario, user A is defined as the first user. The first quality-perceived spatiotemporal map refers to a two-dimensional matrix constructed for the first user, where rows represent time frames, columns represent adaptively selected facial spatial points, and element values ​​are the corresponding green channel intensity and SNR information, after data restoration. For example, the matrix constructed for user A contains fused information from 300 time points and 50 spatial points. The second user refers to another specific object that is processed simultaneously with the first user; for example, user B who inputs video at the same time as user A. The second quality-aware spatiotemporal graph refers to a two-dimensional matrix constructed for the second user with the same structure as the first quality-aware spatiotemporal graph; for example, a matrix constructed for user B containing 300 time points and 50 spatial points.

[0023] S2. Input the first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map into the Vision Mamba dual-path network structure. Extract the first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure. Extract the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure.

[0024] The Vision Mamba dual-path network structure refers to a parallel dual-branch neural network with a core Vision Mamba module, used to extract time-domain and frequency-domain signals from the spatiotemporal graph. For example, this network simultaneously inputs the spatiotemporal graphs of user A and user B, performing time-domain and frequency-domain analyses in parallel. The time path refers to the branch in the Vision Mamba dual-path network structure specifically handling temporal features, containing physiological and non-physiological encoders. For example, this path analyzes user A's spatiotemporal graph, extracting temporal patterns related to heartbeat. The first temporal prediction signal refers to the preliminary temporal physiological waveform predicted by the time path based on the first user's spatiotemporal graph; for example, outputting a 300-point sequence reflecting user A's estimated heartbeat rhythm. The second temporal prediction signal refers to the preliminary temporal physiological waveform predicted by the time path based on the second user's spatiotemporal graph; for example, outputting a 300-point sequence reflecting user B's estimated heartbeat rhythm. The frequency path refers to the branch in the Vision Mamba dual-path network structure that specifically handles frequency domain features. It first converts the spatiotemporal graph to the frequency domain and then extracts the feature signals. For example, this path performs a Fourier transform on user A's spatiotemporal graph and then analyzes its frequency components. The first frequency domain predicted signal refers to the preliminary frequency domain physiological signal predicted by the frequency path based on the first user's spatiotemporal graph. For example, it outputs a vector whose peak value corresponds to the estimated heart rate frequency of user A. The second frequency domain predicted signal refers to the preliminary frequency domain physiological signal predicted by the frequency path based on the second user's spatiotemporal graph. For example, it outputs a vector whose peak value corresponds to the estimated heart rate frequency of user B.

[0025] S3. The first time-series prediction signal and the first frequency-domain prediction signal are fused to obtain the first physiological prediction signal of the first user, and the second time-series prediction signal and the second frequency-domain prediction signal are fused to obtain the second physiological prediction signal of the second user.

[0026] The first physiological prediction signal refers to the final physiological signal estimate obtained by fusing the time-series and frequency-domain prediction signals of the first user; for example, fusing the two prediction signals of user A to generate the final heart rate waveform. The second physiological prediction signal refers to the final physiological signal estimate obtained by fusing the time-series and frequency-domain prediction signals of the second user; for example, fusing the two prediction signals of user B to generate the final heart rate waveform.

[0027] S4. Output the first physiological prediction signal and the second physiological prediction signal.

[0028] The technical solution of this embodiment constructs a quality-aware spatiotemporal graph through an adaptive region of interest selection mechanism and low-rank tensor completion, and combines it with the Vision Mamba dual-path network to extract time-frequency domain features synchronously. This solves the problems of regional quality fluctuations caused by facial occlusion, motion artifacts, and changes in illumination, as well as the high computational complexity and insufficient long-period dependency capture of traditional methods. It improves the robustness, accuracy and efficiency of physiological signal extraction, and takes into account the needs of different scenarios.

[0029] In one alternative approach, S1 specifically includes: From the facial video data of the first user, extract the first facial region from multiple consecutive frames; from the facial video data of the second user, extract the second facial region from multiple consecutive frames.

[0030] Here, facial video data refers to: a video sequence containing user facial images; for example, a 30-second color video of user A's face and a 30-second color video of user B's face. A first facial region consisting of multiple consecutive frames refers to: a standardized sequence of facial images extracted frame-by-frame from the first user video; for example, 900 consecutive facial region images extracted from user A's video. A second facial region consisting of multiple consecutive frames refers to: a standardized sequence of facial images extracted frame-by-frame from the second user video; for example, 900 consecutive facial region images extracted from user B's video.

[0031] The adaptive region of interest (ROI) selection mechanism calculates the green channel intensity and signal-to-noise ratio of each first preset ROI region in the first facial region of the consecutive frames and sorts them in descending order. The first K first preset ROI regions are determined as the first ROI regions. Similarly, the green channel intensity and signal-to-noise ratio of each second preset ROI region in the second facial region of the consecutive frames are calculated and sorted in descending order. The first K second preset ROI regions are determined as the second ROI regions.

[0032] The default number of first preset ROI regions is 64. Multiple first ROI regions refer to one or more high-quality sub-regions dynamically selected from the facial regions of a first user in each frame; for example, two small areas selected from the forehead and left cheek in a frame of user A's face. Multiple second ROI regions refer to one or more high-quality sub-regions dynamically selected from the facial regions of a second user in each frame; for example, two small areas selected from the forehead and right cheek in a frame of user B's face.

[0033] The RGB signals corresponding to each first ROI region are spliced ​​together to obtain the first original spatiotemporal map of the first user, and the RGB signals corresponding to each second ROI region are spliced ​​together to obtain the second original spatiotemporal map of the second user.

[0034] The green channel intensity refers to the average green channel pixel value calculated from the ROI region image; for example, calculating the average green component of all pixels in a small ROI patch of user A. The signal-to-noise ratio (SNR) is a metric for evaluating the signal quality of the ROI region, representing the ratio of physiological signal energy to noise energy; for example, calculating the SNR of the green channel signal in a ROI region of user A using the POS algorithm. The first original spatiotemporal map refers to the initial matrix formed by stitching together the RGB signals of all ROI regions of the first user, which may contain missing data; for example, arranging the ROI information of each frame of user A into a 900-row, 4-column matrix. The second original spatiotemporal map refers to the initial matrix formed by stitching together the green intensity and SNR of all ROIs of the second user; for example, arranging the ROI information of each frame of user B into a 900-row, 4-column matrix.

[0035] The low-rank tensor completion method is used to repair the first original spatiotemporal graph to obtain the first quality-aware spatiotemporal graph, and the second original spatiotemporal graph is repaired to obtain the second quality-aware spatiotemporal graph.

[0036] Regarding the construction process of the original spatiotemporal graph, it is necessary to explain that: like Figure 2 As shown, taking the facial video data of the first user as an example, the input video frame undergoes facial key point detection, and the facial region in each frame is spatially aligned based on the detected key points to eliminate spatial position deviations caused by head movement. The aligned facial region is divided into 64 regularly gridded regions of interest (first preset ROI regions). For each first preset ROI region, the average value of its green channel pixel intensity is calculated, and the signal-to-noise ratio of the signal in that region is estimated based on the POS algorithm.

[0037] For each frame, the 64 regions of interest are sorted in descending order based on the calculated green channel intensity value. The pooled RGB values ​​corresponding to the top K regions are selected and stitched together according to their spatial position. This data from all video frames is then stacked along the time dimension to form a spatiotemporal map based on green channel intensity filtering.

[0038] For the same frame, the 64 regions of interest are sorted in descending order according to the calculated signal-to-noise ratio (SNR) values. The pooled RGB values ​​corresponding to the top K regions are selected and stitched together according to their spatial position. This data from all video frames is then stacked along the time dimension to form a spatiotemporal map based on SNR intensity filtering.

[0039] The generated spatiotemporal map based on green channel intensity filtering is stitched together with the spatiotemporal map based on signal-to-noise ratio intensity filtering to form the first original spatiotemporal map of the first user. This first original spatiotemporal map serves as the input for the subsequent low-rank tensor completion method, and after repair, the first quality-aware spatiotemporal map is finally obtained.

[0040] Among the above-mentioned optional methods, a high-quality region is further selected from the facial video through an adaptive ROI selection mechanism to generate the original spatiotemporal map, and then the missing data is repaired by using low-rank tensors to improve the quality and robustness of the spatiotemporal map.

[0041] In one alternative approach, the time path includes: a physiological encoder and a non-physiological encoder; the step of extracting the first temporal prediction signal corresponding to the first quality-aware spatiotemporal map and the second temporal prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure includes: The physiological encoder obtains the first original physiological features of the first quality-perceived spatiotemporal map and the second original physiological features of the second quality-perceived spatiotemporal map, and the non-physiological encoder obtains the first original non-physiological features of the first quality-perceived spatiotemporal map and the second original non-physiological features of the second quality-perceived spatiotemporal map.

[0042] The physiological encoder refers to a neural network module that extracts physiologically relevant features from the time path; for example, a module that encodes user A's spatiotemporal graph into a 512-dimensional feature vector. The non-physiological encoder refers to a neural network module that extracts motion and other interference features from the time path; for example, a module that encodes user A's spatiotemporal graph into a 256-dimensional interference feature vector. The first original physiological feature refers to the feature obtained by the physiological encoder from encoding the first user's spatiotemporal graph; for example, the 512-dimensional physiological feature vector corresponding to user A's spatiotemporal graph. The second original physiological feature refers to the feature obtained by the physiological encoder from encoding the second user's spatiotemporal graph; for example, the 512-dimensional physiological feature vector corresponding to user B's spatiotemporal graph. The first original non-physiological feature refers to the feature obtained by the non-physiological encoder from encoding the first user's spatiotemporal graph; for example, the 256-dimensional non-physiological feature vector corresponding to user A's spatiotemporal graph. The second original non-physiological feature refers to the feature obtained by the non-physiological encoder from encoding the second user's spatiotemporal graph; for example, the 256-dimensional non-physiological feature vector corresponding to user B's spatiotemporal graph.

[0043] The first original physiological feature and the first original non-physiological feature are reconstructed to obtain a first reconstructed spatiotemporal map. The first original physiological feature and the second original non-physiological feature are reconstructed to obtain a second reconstructed spatiotemporal map. The first original non-physiological feature and the second original physiological feature are reconstructed to obtain a third reconstructed spatiotemporal map. The second original physiological feature and the second original non-physiological feature are reconstructed to obtain a fourth reconstructed spatiotemporal map.

[0044] The first reconstructed spatiotemporal map refers to a spatiotemporal map reconstructed using the physiological and non-physiological characteristics of the first user; for example, reconstructing a spatiotemporal map using the original characteristics of user A. The second reconstructed spatiotemporal map refers to a spatiotemporal map reconstructed using the physiological characteristics of the first user and the non-physiological characteristics of the second user; for example, reconstructing a spatiotemporal map using the physiological characteristics of user A and the non-physiological characteristics of user B. The third reconstructed spatiotemporal map refers to a spatiotemporal map reconstructed using the non-physiological characteristics of the first user and the physiological characteristics of the second user; for example, reconstructing a spatiotemporal map using the non-physiological characteristics of user A and the physiological characteristics of user B. The fourth reconstructed spatiotemporal map refers to a spatiotemporal map reconstructed using the physiological and non-physiological characteristics of the second user; for example, reconstructing a spatiotemporal map using the original characteristics of user B.

[0045] The physiological encoder obtains the first reconstructed physiological features of the second reconstructed spatiotemporal map and the second reconstructed physiological features of the third reconstructed spatiotemporal map, and the non-physiological encoder obtains the first reconstructed non-physiological features of the second reconstructed spatiotemporal map and the second reconstructed non-physiological features of the third reconstructed spatiotemporal map.

[0046] The first reconstructed physiological feature refers to the feature obtained by passing the second reconstructed spatiotemporal map through a physiological encoder again; for example, inputting the second reconstructed spatiotemporal map into the physiological encoder to obtain a new 512-dimensional feature. The second reconstructed physiological feature refers to the feature obtained by passing the third reconstructed spatiotemporal map through a physiological encoder again; for example, inputting the third reconstructed spatiotemporal map into the physiological encoder to obtain a new 512-dimensional feature. The first reconstructed non-physiological feature refers to the feature obtained by passing the second reconstructed spatiotemporal map through a non-physiological encoder again; for example, inputting the second reconstructed spatiotemporal map into the non-physiological encoder to obtain a new 256-dimensional feature. The second reconstructed non-physiological feature refers to the feature obtained by passing the third reconstructed spatiotemporal map through a non-physiological encoder again; for example, inputting the third reconstructed spatiotemporal map into the non-physiological encoder to obtain a new 256-dimensional feature.

[0047] The first original physiological feature is spliced ​​with the first reconstructed physiological feature to obtain the first spliced ​​feature, which is then input into the first regression layer to obtain the first time-series prediction signal. The second original physiological feature is spliced ​​with the second reconstructed physiological feature to obtain the second spliced ​​feature, which is then input into the second regression layer to obtain the second time-series prediction signal.

[0048] The first splicing feature refers to a combined feature formed by connecting the first original physiological feature and the first reconstructed physiological feature; for example, splicing two 512-dimensional features of user A into a 1024-dimensional vector. The first regression layer refers to a fully connected layer that maps the spliced ​​features to a time-series prediction signal; for example, mapping the 1024-dimensional first spliced ​​feature to a 300-point first time-series prediction signal. The second splicing feature refers to a combined feature formed by connecting the second original physiological feature and the second reconstructed physiological feature; for example, splicing two 512-dimensional features of user B into a 1024-dimensional vector. The second regression layer refers to a fully connected layer that maps the second spliced ​​feature to a time-series prediction signal; for example, mapping the 1024-dimensional second spliced ​​feature to a 300-point second time-series prediction signal.

[0049] In the above-mentioned optional methods, features are further extracted by separating physiological and non-physiological encoders, and combined with cross-user feature reconstruction and contrastive learning mechanisms to enhance the ability to distinguish physiological signals. The accuracy of temporal prediction is improved by splicing original and reconstructed physiological features.

[0050] In one alternative approach, the step of extracting the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure includes: Perform Fourier transforms on the first quality-sensing spatiotemporal map and the second quality-sensing spatiotemporal map respectively to obtain the first frequency domain signal of the first quality-sensing spatiotemporal map and the second frequency domain signal of the second quality-sensing spatiotemporal map.

[0051] The Fourier transform refers to the mathematical transformation that converts a time-domain signal into a frequency-domain representation. For example, performing an FFT along the time dimension on the spatiotemporal graph of user A yields a frequency-domain signal. The first frequency-domain signal refers to the result of the Fourier transform of the first user's spatiotemporal graph; for example, the complex matrix corresponding to the spatiotemporal graph of user A. The second frequency-domain signal refers to the result of the Fourier transform of the second user's spatiotemporal graph; for example, the complex matrix corresponding to the spatiotemporal graph of user B.

[0052] The first frequency domain signal is encoded by a frequency domain encoder to obtain a first frequency domain feature, and the second frequency domain signal is encoded to obtain a second frequency domain feature.

[0053] In this context, a frequency domain encoder refers to a neural network module that encodes frequency domain signals in the frequency path; for example, a Vision Mamba module used to process frequency domain signals. The first frequency domain feature refers to the feature obtained by the frequency domain encoder encoding the first frequency domain signal; for example, the 256-dimensional feature vector corresponding to user A's frequency domain signal. The second frequency domain feature refers to the feature obtained by the frequency domain encoder encoding the second frequency domain signal; for example, the 256-dimensional feature vector corresponding to user B's frequency domain signal.

[0054] The first frequency domain feature and the second frequency domain feature are passed sequentially through a first linear layer, a normalization layer, an activation layer, and a second linear layer to obtain the first frequency domain prediction signal corresponding to the first frequency domain feature and the second frequency domain prediction signal corresponding to the second frequency domain feature.

[0055] The first linear layer refers to the fully connected layer in the frequency path that performs the first linear transformation on the frequency domain features; for example, a linear layer that maps 256-dimensional features to 128-dimensional features. The normalization layer refers to the network layer that standardizes the data; for example, layer normalization is applied to the output of the linear layer to stabilize the data distribution. The activation layer refers to the network layer that provides nonlinear transformations; for example, applying the ReLU function to the normalized data. The second linear layer refers to the last fully connected layer in the frequency path that generates the final frequency domain prediction signal; for example, a linear layer that maps a 128-dimensional vector to a 10-dimensional frequency domain prediction signal.

[0056] In the above-mentioned optional methods, the quality-aware spatiotemporal map is further subjected to Fourier transform to obtain the frequency domain signal, the frequency domain features are extracted by the frequency domain encoder, and the frequency domain prediction signal is obtained by linear layer processing, so as to realize the in-depth mining and utilization of frequency domain information.

[0057] In one alternative approach, S3 specifically includes: The first physiological prediction signal of the first user is obtained by fusing the first time-series prediction signal and the first frequency-domain prediction signal through a fusion layer including a multi-layer perceptron.

[0058] Among them, the multi-layer perceptron refers to a neural network structure stacked by multiple fully-connected layers and non-linear functions; for example, a simple network including two fully-connected layers. The fusion layer refers to a network layer that fuses time-series and frequency-domain prediction signals into the final physiological signal, usually a multi-layer perceptron; for example, a multi-layer perceptron that receives 310-dimensional input and outputs a 300-dimensional final waveform.

[0059] The second physiological prediction signal of the second user is obtained by fusing the second time-series prediction signal and the second frequency-domain prediction signal through the fusion layer.

[0060] In the above optional method, an adaptive weighted fusion of the time-series prediction signal and the frequency-domain prediction signal is further performed using a multi-layer perceptron fusion layer, fully combining the dynamic changes in the time domain and the periodic characteristics in the frequency domain, and improving the accuracy and stability of physiological signal prediction.

[0061] In an optional method, a two-stage training strategy is adopted for the prediction process of physiological signals; in the first stage, the L1 loss function is used for training, and in the second stage, the L1 + a·Lp loss function is used for training; where p in the Lp loss function is a preset hyperparameter, and 0 < p < 1.

[0062] Among them, the two-stage training strategy refers to a training method using different loss functions in two stages; for example, first using the L1 loss function to train for the 1st - 20th rounds, and then using the L1 + a·Lp loss function to train for the 21st - 32nd rounds. The L1 loss function refers to a loss function that calculates the average of the absolute differences between the predicted value and the true value; for example, it is used to measure the absolute error between the predicted signal and the true signal. The Lp loss function refers to a loss function that calculates the average of the p-th power of the difference between the predicted value and the true value, 0 < p < 1; for example, taking p = 0.5, it is used to penalize the differences between features to enhance sparsity.

[0063] It should be noted that the value of a is determined by the experimental test results.

[0064] In the above optional method, a two-stage training strategy is further adopted. First, the L1 loss is used for coarse-grained optimization, and then the L1 + a·Lp loss is used for fine-grained constraint, gradually improving the model convergence speed and prediction accuracy, and enhancing the training stability.

[0065] In an optional method, the L1 loss function is used to measure the differences between spatio-temporal graphs, the differences between features, and the differences between the predicted physiological signal and the true physiological signal; The differences between the spatiotemporal maps are determined by calculating the difference between the first quality-perceived spatiotemporal map and the first reconstructed spatiotemporal map, and by calculating the difference between the second quality-perceived spatiotemporal map and the fourth reconstructed spatiotemporal map; the differences between the features are determined by calculating the differences between the first original physiological feature and the first reconstructed physiological feature, the differences between the first original non-physiological feature and the first reconstructed non-physiological feature, the differences between the second original physiological feature and the second reconstructed physiological feature, and the differences between the second original non-physiological feature and the second reconstructed non-physiological feature; the differences between the physiological signals are determined by calculating the difference between the first physiological prediction signal and the first physiological real signal, and by calculating the difference between the second physiological prediction signal and the second physiological real signal.

[0066] Among the above-mentioned optional methods, the model's reconstruction capability, feature discrimination, and prediction accuracy can be further enhanced by using differences in spatiotemporal graphs, features, and physiological signals.

[0067] Figure 3 The overall process of this embodiment is illustrated, specifically: 1) Facial video data of the first user Facial video data of the second user As input to the entire process, these data first undergo preprocessing and feature construction. At this stage, an adaptive region of interest selection mechanism and a low-rank tensor completion method are applied to construct a first-level quality-aware spatiotemporal graph for the first user. And construct a second quality-perception spatiotemporal map for the second user. .

[0068] 2) First quality perception spatiotemporal diagram With the second quality perception spatiotemporal map The data is fed in parallel into a Vision Mamba two-path network structure. The Vision Mamba two-path network structure contains two processing paths. In the temporal path, the first quality-aware spatiotemporal graph... With the second quality perception spatiotemporal map The first time-series prediction signal was extracted after processing by both physiological encoders and non-physiological encoders. With the second time-series prediction signal In the frequency path, the first quality-sensing spatiotemporal diagram With the second quality perception spatiotemporal map First, the signal undergoes Fourier transform, then is processed by a frequency domain encoder and subsequent linear, normalization, and activation layers to extract the first frequency domain prediction signal. With the second frequency domain prediction signal .

[0069] 3) Subsequently, the first time-series prediction signal from the time path With the first frequency domain prediction signal from the frequency path It is fed to a fusion layer. The fusion layer consists of multiple sensing layers, which process the first time-series prediction signal. With the first frequency domain prediction signal The data is fused to output the first physiological prediction signal for the first user. Symmetrically, the second time-series prediction signal from the time path With the second frequency domain prediction signal from the frequency path They are fed into the same fusion layer for fusion, outputting the second physiological prediction signal for the second user. .

[0070] 4) In this embodiment, the L1 loss function is used to measure the differences between spatiotemporal graphs and the differences between features. The differences between spatiotemporal graphs are calculated by measuring the first quality-perceived spatiotemporal graph. With the first reconstructed spacetime map The differences between them, and the calculation of the second quality perception spatiotemporal map With the fourth reconstruction of the spacetime map The differences between them are determined. Differences between characteristics are calculated by using the first original physiological characteristic. With the first reconstruction of physiological characteristics Differences between them, primary non-physiological characteristics With the first reconstruction of non-physiological characteristics Differences between them, secondary primitive physiological characteristics With the second reconstruction of physiological characteristics Differences between them and the second primitive non-physiological characteristics With the second reconstruction of non-physiological characteristics The difference between them is used to determine this.

[0071] Figure 4 A workflow diagram of a general adaptive remote photoplethysmography (rPPG) framework is shown. The framework's input is raw video data containing user faces. This video data is fed into a scene classifier for processing.

[0072] The scene classifier analyzes the content of the input video, primarily evaluating factors such as the amplitude of facial movements, the drastic changes in lighting, and the presence of significant occlusions, thereby determining the overall stability and complexity of the video scene. Based on this analysis, the scene classifier makes a binary path selection decision.

[0073] When the scene classifier determines that the current video scene is relatively stable, such as when the user is stationary and the lighting is uniform, the framework will activate the traditional algorithm path. In the traditional algorithm path, the system will call mature traditional rPPG algorithms such as POS and CHROM to quickly estimate preliminary physiological signals such as heart rate directly from the raw video data. This path focuses on processing efficiency and low computational overhead.

[0074] Conversely, when the scene classifier determines that the current video scene is complex, with obvious motion, lighting fluctuations, or partial occlusion, the framework will activate the deep learning path. In the deep learning path, the raw video data will enter the complete processing flow of the method in this embodiment. This flow includes constructing a quality-aware spatiotemporal map, extracting temporal and frequency domain features through the Vision Mamba dual-path network structure, and fusing them to generate the final physiological prediction signal. This path focuses on ensuring the accuracy and robustness of signal extraction in high-interference environments.

[0075] Ultimately, regardless of the processing path, the framework outputs the corresponding physiological signal estimation results for the user, such as heart rate waveforms or heart rate values. This general adaptive framework achieves an effective balance between efficiency and high accuracy requirements on resource-constrained devices by dynamically selecting the processing path most suitable for the current scenario.

[0076] Figure 5 This paper demonstrates the processing flow of a low-rank tensor completion method for repairing original spatiotemporal maps. Specifically, the method receives a first original spatiotemporal map as input. Due to facial occlusion, motion, or illumination variations, the original spatiotemporal map often contains incomplete or damaged data points, represented as missing or outlier values ​​in the matrix. The low-rank tensor completion method is based on a core assumption that high-quality, clean physiological signal spatiotemporal maps possess an inherent low-rank structure. This method establishes an optimization objective that minimizes the kernel norm of the spatiotemporal map matrix while satisfying data observation constraints, thereby forcing it to approximate a low-rank space. By iteratively solving this optimization problem, the algorithm can intelligently estimate and fill in the missing values ​​by utilizing the data correlations of the undamaged parts of the spatiotemporal map, while smoothing out anomalous fluctuations. Finally, the repaired and completed matrix is ​​output, resulting in a more complete, continuous, and noise-suppressed first quality-perceived spatiotemporal map. This process effectively improves the quality of the input data, laying a reliable foundation for subsequent deep feature extraction.

[0077] Figure 6This diagram illustrates the changes in training loss during the two-stage training strategy employed for physiological signal prediction in this embodiment. Specifically, it shows the variation of the line-following loss (i.e., training loss) with the number of training epochs. The first stage is epoch 1-epoch 20, and the second stage is epoch 21-epoch 32. The horizontal axis represents the number of training epochs, and the vertical axis represents the loss value. Loss curves for different hyperparameter α values ​​(α=0, 0.6, 0.7, 0.8, 0.9) are plotted for comparative analysis. The first stage uses only the L1 loss function, i.e., α=0, while the second stage uses L1+α·Lp. Figure 6 The diagram illustrates the impact of different α values ​​on the second-stage training loss, with a magnified view of the detailed second-stage loss in the upper right corner. The training loss reaches its optimum when α = 0.7.

[0078] Figure 7 This diagram illustrates the loss variation during the two-stage training strategy of this embodiment, specifically showing the change in test loss with the number of training epochs. It shows the model's loss on the test set (i.e., test loss) with the number of training epochs for different hyperparameter α values ​​(α=0, 0.6, 0.7, 0.8, 0.9). The horizontal axis represents the number of training epochs, and the vertical axis represents the test loss value. In the second stage, the model is trained using a joint loss function of L1 + α·Lp. Observing the curves, it can be seen that the test loss reaches its optimum when α==0.7.

[0079] Figure 8 This diagram illustrates a structural schematic of an embodiment of a physiological signal extraction system 200 based on adaptive ROI and ViM dual-path provided by the present invention. Figure 8 As shown, the physiological signal extraction system 200 based on adaptive ROI and ViM dual paths includes: Module 201 is constructed based on an adaptive region of interest selection mechanism and a low-rank tensor completion method to construct the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user. Extraction module 202 is used to input the first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map into a Vision Mamba dual-path network structure, extract the first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure, and extract the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure. The fusion module 203 is used to fuse the first time-series prediction signal and the first frequency-domain prediction signal to obtain the first physiological prediction signal of the first user, and to fuse the second time-series prediction signal and the second frequency-domain prediction signal to obtain the second physiological prediction signal of the second user. The output module 204 is used to output the first physiological prediction signal and the second physiological prediction signal.

[0080] In an alternative embodiment, the building module 201 is specifically used for: Extract a first facial region from multiple consecutive frames of the first user's facial video data, and extract a second facial region from multiple consecutive frames of the second user's facial video data. Through the adaptive region of interest selection mechanism, the green channel intensity and signal-to-noise ratio of each first preset ROI region in the first facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K first preset ROI regions are determined as the first ROI regions. Similarly, the green channel intensity and signal-to-noise ratio of each second preset ROI region in the second facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K second preset ROI regions are determined as the second ROI regions. The RGB signals corresponding to each first ROI region are spliced ​​together to obtain the first original spatiotemporal map of the first user, and the RGB signals corresponding to each second ROI region are spliced ​​together to obtain the second original spatiotemporal map of the second user. The low-rank tensor completion method is used to repair the first original spatiotemporal graph to obtain the first quality-aware spatiotemporal graph, and the second original spatiotemporal graph is repaired to obtain the second quality-aware spatiotemporal graph.

[0081] In one alternative approach, the time path includes: a physiological encoder and a non-physiological encoder; the extraction module 202 is specifically used for: The physiological encoder obtains the first original physiological features of the first quality-perceived spatiotemporal map and the second original physiological features of the second quality-perceived spatiotemporal map, and the non-physiological encoder obtains the first original non-physiological features of the first quality-perceived spatiotemporal map and the second original non-physiological features of the second quality-perceived spatiotemporal map. The first original physiological feature and the first original non-physiological feature are reconstructed to obtain a first reconstructed spatiotemporal map; the first original physiological feature and the second original non-physiological feature are reconstructed to obtain a second reconstructed spatiotemporal map; the first original non-physiological feature and the second original physiological feature are reconstructed to obtain a third reconstructed spatiotemporal map; and the second original physiological feature and the second original non-physiological feature are reconstructed to obtain a fourth reconstructed spatiotemporal map. The physiological encoder obtains the first reconstructed physiological features of the second reconstructed spatiotemporal map and the second reconstructed physiological features of the third reconstructed spatiotemporal map, and the non-physiological encoder obtains the first reconstructed non-physiological features of the second reconstructed spatiotemporal map and the second reconstructed non-physiological features of the third reconstructed spatiotemporal map. The first original physiological feature is spliced ​​with the first reconstructed physiological feature to obtain the first spliced ​​feature, which is then input into the first regression layer to obtain the first time-series prediction signal. The second original physiological feature is spliced ​​with the second reconstructed physiological feature to obtain the second spliced ​​feature, which is then input into the second regression layer to obtain the second time-series prediction signal.

[0082] In an alternative embodiment, the extraction module 202 is specifically used for: Fourier transforms are performed on the first quality-sensing spatiotemporal map and the second quality-sensing spatiotemporal map respectively to obtain the first frequency domain signal of the first quality-sensing spatiotemporal map and the second frequency domain signal of the second quality-sensing spatiotemporal map; The first frequency domain signal is encoded by a frequency domain encoder to obtain a first frequency domain feature, and the second frequency domain signal is encoded to obtain a second frequency domain feature. The first frequency domain feature and the second frequency domain feature are passed sequentially through a first linear layer, a normalization layer, an activation layer, and a second linear layer to obtain the first frequency domain prediction signal corresponding to the first frequency domain feature and the second frequency domain prediction signal corresponding to the second frequency domain feature.

[0083] In one alternative embodiment, the fusion module 203 is specifically used for: The first time-series prediction signal and the first frequency-domain prediction signal are fused through a fusion layer containing a multilayer perceptron to obtain the first physiological prediction signal of the first user. The second time-series prediction signal and the second frequency-domain prediction signal are fused through the fusion layer to obtain the second physiological prediction signal of the second user.

[0084] In one alternative approach, the physiological signal prediction process employs a two-stage training strategy; the first stage uses the L1 loss function for training, and the second stage uses the L1+a·Lp loss function for training; where p in the Lp loss function is a preset hyperparameter, and 0 < 1. <p<1。

[0085] In one alternative approach, the L1 loss function is used to measure the differences between spatiotemporal maps, the differences between features, and the differences between predicted physiological signals and actual physiological signals. The differences between the spatiotemporal maps are determined by calculating the difference between the first quality-perceived spatiotemporal map and the first reconstructed spatiotemporal map, and by calculating the difference between the second quality-perceived spatiotemporal map and the fourth reconstructed spatiotemporal map; the differences between the features are determined by calculating the differences between the first original physiological feature and the first reconstructed physiological feature, the differences between the first original non-physiological feature and the first reconstructed non-physiological feature, the differences between the second original physiological feature and the second reconstructed physiological feature, and the differences between the second original non-physiological feature and the second reconstructed non-physiological feature; the differences between the physiological signals are determined by calculating the difference between the first physiological prediction signal and the first physiological real signal, and by calculating the difference between the second physiological prediction signal and the second physiological real signal.

[0086] It should be noted that the beneficial effects of the physiological signal extraction system 200 based on adaptive ROI and ViM dual paths provided in the above embodiments are the same as those of the physiological signal extraction method based on adaptive ROI and ViM dual paths, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0087] The physiological signal extraction system 200 based on adaptive ROI and ViM dual paths of the present invention can be a computer program (including program code) running on a computer device. For example, the physiological signal extraction system 200 based on adaptive ROI and ViM dual paths of the present invention is an application software that can be used to execute the corresponding steps in the physiological signal extraction method based on adaptive ROI and ViM dual paths of the present invention.

[0088] In some embodiments, the physiological signal extraction system 200 based on adaptive ROI and ViM dual paths of the present invention can be implemented in a combination of hardware and software. As an example, the physiological signal extraction system 200 based on adaptive ROI and ViM dual paths of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the physiological signal extraction method based on adaptive ROI and ViM dual paths of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0089] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0090] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned physiological signal extraction methods based on adaptive ROI and ViM dual paths. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the physiological signal extraction method based on adaptive ROI and ViM dual paths shown in any embodiment of the present invention by calling the computer program.

[0091] In one alternative embodiment, an electronic device is provided, such as Figure 9 As shown, Figure 9 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0092] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0093] Bus 4002 may include a path for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus 4002 is represented by only one thick line, but this does not mean that there is only one bus or one type of bus.

[0094] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0095] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0096] Among them, electronic devices can also be terminal devices. A terminal device can be any terminal device that can install applications and access web pages through applications, including at least one of smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, smart TVs, and smart in-vehicle devices.

[0097] It should be noted that, Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0098] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned physiological signal extraction methods based on adaptive ROI and ViM dual paths.

[0099] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0100] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the aforementioned physiological signal extraction method based on adaptive ROI and ViM dual paths.

[0101] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The computer-readable storage medium provided in this invention can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0104] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0105] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0106] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0107] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0108] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A physiological signal extraction method based on adaptive ROI and ViM dual-path, characterized in that, include: Based on the adaptive region of interest selection mechanism and low-rank tensor completion method, the first quality perception spatiotemporal graph of the first user and the second quality perception spatiotemporal graph of the second user are constructed. The first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map are input into a Vision Mamba dual-path network structure. The first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map are extracted through the time path of the Vision Mamba dual-path network structure. The first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map are extracted through the frequency path of the Vision Mamba dual-path network structure. The first time-series prediction signal and the first frequency-domain prediction signal are fused to obtain the first physiological prediction signal of the first user, and the second time-series prediction signal and the second frequency-domain prediction signal are fused to obtain the second physiological prediction signal of the second user. Output the first physiological prediction signal and the second physiological prediction signal.

2. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 1, characterized in that, The steps for constructing the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user based on the adaptive region of interest selection mechanism and low-rank tensor completion method include: Extract a first facial region from multiple consecutive frames of the first user's facial video data, and extract a second facial region from multiple consecutive frames of the second user's facial video data. Through the adaptive region of interest selection mechanism, the green channel intensity and signal-to-noise ratio of each first preset ROI region in the first facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K first preset ROI regions are determined as the first ROI regions. Similarly, the green channel intensity and signal-to-noise ratio of each second preset ROI region in the second facial region of the consecutive multiple frames are calculated and sorted in descending order. The first K second preset ROI regions are determined as the second ROI regions. The RGB signals corresponding to each first ROI region are spliced ​​together to obtain the first original spatiotemporal map of the first user, and the RGB signals corresponding to each second ROI region are spliced ​​together to obtain the second original spatiotemporal map of the second user. The low-rank tensor completion method is used to repair the first original spatiotemporal graph to obtain the first quality-aware spatiotemporal graph, and the second original spatiotemporal graph is repaired to obtain the second quality-aware spatiotemporal graph.

3. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 2, characterized in that, The time path includes: a physiological encoder and a non-physiological encoder; the step of extracting the first temporal prediction signal corresponding to the first quality-aware spatiotemporal map and the second temporal prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the Vision Mamba dual-path network structure includes: The physiological encoder obtains the first original physiological features of the first quality-perceived spatiotemporal map and the second original physiological features of the second quality-perceived spatiotemporal map, and the non-physiological encoder obtains the first original non-physiological features of the first quality-perceived spatiotemporal map and the second original non-physiological features of the second quality-perceived spatiotemporal map. The first original physiological feature and the first original non-physiological feature are reconstructed to obtain a first reconstructed spatiotemporal map; the first original physiological feature and the second original non-physiological feature are reconstructed to obtain a second reconstructed spatiotemporal map; the first original non-physiological feature and the second original physiological feature are reconstructed to obtain a third reconstructed spatiotemporal map; and the second original physiological feature and the second original non-physiological feature are reconstructed to obtain a fourth reconstructed spatiotemporal map. The physiological encoder obtains the first reconstructed physiological features of the second reconstructed spatiotemporal map and the second reconstructed physiological features of the third reconstructed spatiotemporal map, and the non-physiological encoder obtains the first reconstructed non-physiological features of the second reconstructed spatiotemporal map and the second reconstructed non-physiological features of the third reconstructed spatiotemporal map. The first original physiological feature is spliced ​​with the first reconstructed physiological feature to obtain the first spliced ​​feature, which is then input into the first regression layer to obtain the first time-series prediction signal. The second original physiological feature is spliced ​​with the second reconstructed physiological feature to obtain the second spliced ​​feature, which is then input into the second regression layer to obtain the second time-series prediction signal.

4. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 3, characterized in that, The step of extracting the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the Vision Mamba dual-path network structure includes: Fourier transforms are performed on the first quality-sensing spatiotemporal map and the second quality-sensing spatiotemporal map respectively to obtain the first frequency domain signal of the first quality-sensing spatiotemporal map and the second frequency domain signal of the second quality-sensing spatiotemporal map; The first frequency domain signal is encoded by a frequency domain encoder to obtain a first frequency domain feature, and the second frequency domain signal is encoded to obtain a second frequency domain feature. The first frequency domain feature and the second frequency domain feature are passed sequentially through a first linear layer, a normalization layer, an activation layer, and a second linear layer to obtain the first frequency domain prediction signal corresponding to the first frequency domain feature and the second frequency domain prediction signal corresponding to the second frequency domain feature.

5. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 4, characterized in that, The steps of fusing the first time-series prediction signal and the first frequency-domain prediction signal to obtain the first physiological prediction signal of the first user, and fusing the second time-series prediction signal and the second frequency-domain prediction signal to obtain the second physiological prediction signal of the second user, include: The first time-series prediction signal and the first frequency-domain prediction signal are fused through a fusion layer containing a multilayer perceptron to obtain the first physiological prediction signal of the first user. The second time-series prediction signal and the second frequency-domain prediction signal are fused through the fusion layer to obtain the second physiological prediction signal of the second user.

6. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 3, characterized in that, The prediction process for physiological signals employs a two-stage training strategy; the first stage uses the L1 loss function for training, and the second stage uses the L1+a·Lp loss function for training; where p in the Lp loss function is a preset hyperparameter, and 0 < 1. <p<1。 7. The physiological signal extraction method based on adaptive ROI and ViM dual paths according to claim 6, characterized in that, The L1 loss function is used to measure the differences between spatiotemporal maps, the differences between features, and the differences between predicted physiological signals and actual physiological signals. The differences between the spatiotemporal maps are determined by calculating the difference between the first quality-perceived spatiotemporal map and the first reconstructed spatiotemporal map, and by calculating the difference between the second quality-perceived spatiotemporal map and the fourth reconstructed spatiotemporal map; the differences between the features are determined by calculating the differences between the first original physiological feature and the first reconstructed physiological feature, the differences between the first original non-physiological feature and the first reconstructed non-physiological feature, the differences between the second original physiological feature and the second reconstructed physiological feature, and the differences between the second original non-physiological feature and the second reconstructed non-physiological feature; the differences between the physiological signals are determined by calculating the difference between the first physiological prediction signal and the first physiological real signal, and by calculating the difference between the second physiological prediction signal and the second physiological real signal.

8. A physiological signal extraction system based on adaptive ROI and ViM dual paths, characterized in that, include: The module is constructed based on an adaptive region of interest selection mechanism and a low-rank tensor completion method to build the first quality-perceived spatiotemporal graph of the first user and the second quality-perceived spatiotemporal graph of the second user. The extraction module is used to input the first quality-aware spatiotemporal map and the second quality-aware spatiotemporal map into a VisionMamba dual-path network structure, extract the first time-series prediction signal corresponding to the first quality-aware spatiotemporal map and the second time-series prediction signal corresponding to the second quality-aware spatiotemporal map through the time path of the VisionMamba dual-path network structure, and extract the first frequency domain prediction signal corresponding to the first quality-aware spatiotemporal map and the second frequency domain prediction signal corresponding to the second quality-aware spatiotemporal map through the frequency path of the VisionMamba dual-path network structure. The fusion module is used to fuse the first time-series prediction signal and the first frequency-domain prediction signal to obtain the first physiological prediction signal of the first user, and to fuse the second time-series prediction signal and the second frequency-domain prediction signal to obtain the second physiological prediction signal of the second user. The output module is used to output the first physiological prediction signal and the second physiological prediction signal.

9. An electronic device, characterized in that, The electronic device includes a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to implement the physiological signal extraction method based on adaptive ROI and ViM dual paths as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which, when executed by a processor, implements the physiological signal extraction method based on adaptive ROI and ViM dual paths as described in any one of claims 1 to 7.