Non-contact physiological signal processing method with low-latency interference identification and dynamic filtering
By constructing an interference prediction model and a dynamic filtering method using multi-scale dilated convolutional layers, the problem of insufficient dynamic interference processing capability in traditional methods is solved, and low-latency and high-precision physiological parameter extraction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2026-02-04
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional non-contact physiological signal processing methods lack real-time assessment and adaptive capabilities for dynamic interference, resulting in poor signal robustness, wasted computational resources, and increased latency.
An interference prediction model is adopted, which uses a one-dimensional depthwise separable convolutional structure, a bottleneck mapping path, and dilated convolutional branches. Combined with multi-scale dilated convolutional layers, dynamic filtering and physiological parameter decoding are achieved. Soft switching is performed by calculating a depth adjustment factor to optimize the feature extraction and filtering process.
It enables real-time, robust, and high-precision extraction of physiological parameters under non-contact conditions, reducing latency and improving the adaptability and accuracy of signal processing.
Smart Images

Figure CN122135927A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biological signal processing, specifically relating to a non-contact physiological signal processing method for low-latency interference recognition and dynamic filtering. Background Technology
[0002] Traditional non-contact physiological signal processing methods typically follow a fixed linear processing chain. First, the raw sensor signal is bandpass filtered to isolate a preset physiological frequency band (e.g., respiration 0.1-0.5Hz, heart rate 0.8-2.5Hz), and baseline correction is performed using moving average or median filtering. Subsequently, time-frequency analysis (e.g., short-time Fourier transform, wavelet transform) or peak detection algorithms are directly applied to the preprocessed signal to estimate physiological parameters. To address motion interference, common methods include using multi-antenna beamforming for spatial filtering in hardware, or introducing reference noise cancellation techniques based on auxiliary sensors such as accelerometers in the algorithm. The entire processing flow's model and parameters are usually statically preset, lacking real-time, fine-grained evaluation and response mechanisms for signal quality and interference types.
[0003] Static processing chains cannot handle dynamic interference. Fixed filters and analysis windows, when faced with sudden, non-periodic bodily motion interference, may fail to effectively filter out the signal, leading to signal overload; or they may over-smooth the signal, losing high-frequency heartbeat details, directly compromising robustness. Secondly, there is a lack of intelligent front-end interference prediction. Traditional methods often still perform a full set of calculations when encountering highly interfering segments, ultimately outputting unreliable results or requiring complex post-verification, which wastes computational resources and introduces unnecessary latency. Finally, the model and strategy lack adaptability. A uniform processing intensity cannot achieve an optimal balance between "quiet" and "interference" scenarios: it may be computationally redundant in quiet conditions, while insufficient in interfering conditions. Therefore, traditional methods struggle to simultaneously meet the requirements of low latency and high accuracy in dynamic real-world scenarios; their fundamental flaw lies in the rigidity and passive response characteristics of the processing flow. Summary of the Invention
[0004] To address the aforementioned problems in the existing technology, this invention provides a non-contact physiological signal processing method with low-latency interference identification and dynamic filtering.
[0005] The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, the present invention provides a non-contact physiological signal processing method for low-latency interference identification and dynamic filtering, the method comprising: Human body time-series signals are acquired through sensors; wherein, the sensors are imaging sensors or non-imaging sensors; when the sensor is an imaging sensor, the human body time-series signal is a sequence of human skin images; when the sensor is a non-imaging sensor, the human body time-series signal is a micro-motion signal on the human body surface. The human time series signal is input into a pre-constructed interference prediction model to obtain the interference probability and corresponding interference type label of the human time series signal; wherein, the interference prediction model includes a one-dimensional depthwise separable convolutional structure, a bottleneck mapping path, a dilated convolutional branch and a pruned dilated convolutional branch connected in sequence. The interference prediction model is soft-switched to the dilated convolution branch or the pruned dilated convolution branch according to the computational depth adjustment factor, and the human temporal signal is input into the corresponding branch to obtain the enhanced human features; the computational depth adjustment factor is obtained according to the interference probability of the human temporal signal. The enhanced human features are input into a multi-scale dilated convolutional layer to obtain rhythmic fusion features; wherein, the multi-scale dilated convolutional layer includes three parallel dilated convolutional sub-modules with different dilation rates and a convolutional module; After obtaining the spectral information by performing a fast Fourier transform on the fusion features of the rhythm, the main peak of the physiological rhythm is located based on the energy distribution according to the spectral information to complete the decoding of physiological parameters.
[0006] Optionally, acquiring human body time-series signals via sensors includes: When the human body time-series signal is acquired through the imaging sensor, an image sequence of the human skin region is acquired, and spatial localization and tracking of the region of interest of the human skin are performed on the image sequence. Then, the pixel intensity or multispectral channel information of the region of interest of the human skin is spatially aggregated to obtain the initial human body time-series signal. When the human body time-series signal is acquired through the non-imaging sensor, the human body surface micro-motion signal is acquired, and after uniform resampling of the human body surface micro-motion signal, it is divided into continuous time-series segments. Then, normalization and baseline drift correction processing are sequentially performed on the continuous time-series segments to obtain the initial human body time-series signal. When the initial human body time series signal is a multimodal signal, the modal signals in the initial human body time series signal are aligned on the time axis and combined into multi-channel one-dimensional time series data to obtain the human body time series signal. When the initial human body time series signal is not a multimodal signal, the initial human body time series signal is used as the human body time series signal.
[0007] Optionally, the pruning process of the dilated convolution branch includes: Calculate the importance score of each bottleneck layer in the dilated convolution branch; The importance scores of each bottleneck layer are sorted in descending order, and each bottleneck layer is pruned according to a preset ratio to obtain the pruned dilated convolutional branch.
[0008] Optionally, the importance score of each bottleneck layer is calculated as follows: ; in, Indicates the first The importance score of each bottleneck layer Indicates the first The number of channels for the output features of each bottleneck layer. Indicates the time step of a time sequence segment. For the first The bottleneck layer, the first The first channel in the Output feature response values at each time step.
[0009] Optionally, the interference probability of the human body time-series signal is expressed as follows: ; in, This represents the interference probability of the human body time-series signal. Indicates the first A discriminant characteristic response associated with aperiodic perturbations, express Quantity, Indicates the first A characteristic response associated with periodic physiological rhythms, express The quantity.
[0010] Optionally, the calculation depth adjustment factor is expressed as follows: ; in, This represents the calculation depth adjustment factor. Represents an exponential function. Indicates control factor. This indicates the preset interference threshold.
[0011] Optionally, the step of inputting the human temporal signal into the corresponding branch to obtain the enhanced human features includes: The human body time-series signal is input into the corresponding branch to extract the initial features; The initial features are input into the temporal attention module, global average pooling is performed, the basic weights are adjusted according to the consistency index, and the pooled initial features are enhanced to obtain the enhanced human features.
[0012] Optionally, the consistency index is expressed as follows: ; in, This indicates the consistency index. This indicates the initial features after pooling at time step Instantaneous response on the main frequency channel The mean response is the instantaneous response within a time segment of time step. For time step.
[0013] Optionally, the step of inputting the enhanced human features into a multi-scale dilated convolutional layer to obtain rhythmic fusion features includes: The enhanced human body features are input into three parallel dilated convolution sub-modules with different dilation rates to obtain the corresponding three dilated human body features. Calculate the dynamic fusion weights of three parallel dilated convolutional sub-modules with different dilation rates and normalize them; Based on the dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates after normalization and the corresponding three dilated human features, feature fusion is performed to obtain the fused features. The fused features are reweighted according to the rhythm consistency constraint and input into the convolution module to obtain the rhythm fusion features.
[0014] Optionally, the dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates are represented as follows: ; in, Indicates the expansion rate Dynamic fusion weights of the dilated convolutional submodules This represents the interference probability of the human body time-series signal. To be related to the expansion rate The relevant sensitivity coefficient, This represents an exponential function.
[0015] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: In the above technical solution, this invention improves the MobileNet-V2 model, reducing the number of parameters in the interference prediction model to 42% ± 5% of that in the MobileNet-V2 model. While ensuring low latency, it retains high-frequency physiological rhythm information. With the help of the interference prediction model, it can stably output interference probability and interference type labels. Soft switching is achieved by adjusting the computational depth factor, avoiding frequent path jitter in the medium interference range. The enhanced human features output show significant improvements in temporal continuity and frequency band consistency, providing high-quality input for multi-scale fusion of subsequent multi-scale dilated convolutional layers. Over-amplification or suppression of features under fixed-scale configuration is avoided through three parallel dilated convolutional sub-modules with different dilation rates. The overall effect of real-time, robust, and high-precision extraction of human temporal signals as a physiological parameter under non-contact conditions is achieved.
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0017] Figure 1 This is a flowchart of a non-contact physiological signal processing method for low-latency interference identification and dynamic filtering provided in an embodiment of the present invention; Figure 2 This is a multi-component waveform feature map containing interference provided in an embodiment of the present invention. Detailed Implementation
[0018] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0019] Figure 1 This is a flowchart of a non-contact physiological signal processing method for low-latency interference identification and dynamic filtering provided by an embodiment of the present invention, as shown below. Figure 1 As shown, the method may include the following steps: S101. Obtain human body time-series signals through a sensor; wherein the sensor is an imaging sensor or a non-imaging sensor; when the sensor is an imaging sensor, the human body time-series signal is a sequence of human skin images; when the sensor is a non-imaging sensor, the human body time-series signal is a micro-motion signal on the human body surface.
[0020] Optionally, human body time-series signals are acquired via sensors, including: When acquiring human temporal signals through an imaging sensor, an image sequence of the human skin region is obtained, and spatial localization and tracking of the region of interest of the human skin are performed on the image sequence. Then, the pixel intensity or multispectral channel information of the region of interest of the human skin is spatially aggregated to obtain the initial human temporal signal. When acquiring human body time-series signals through non-imaging sensors, the micro-motion signals on the human body surface are acquired, and after uniform resampling of the micro-motion signals on the human body surface, they are divided into continuous time-series segments. Then, normalization and baseline drift correction are sequentially performed on the continuous time-series segments to obtain the initial human body time-series signal. When the initial human time series signal is a multimodal signal, the modal signals in the initial human time series signal are aligned on the time axis and combined into multi-channel one-dimensional time series data to obtain the human time series signal. When the initial human body time series signal is not a multimodal signal, the initial human body time series signal is used as the human body time series signal.
[0021] For example, image sequences of human skin regions are acquired using visible light or near-infrared imaging sensors, or micro-motion signals of the human surface are acquired using millimeter-wave radar sensors. When using an imaging sensor, spatial localization and tracking of the region of interest (ROI) on the human skin are performed on the image sequence, and spatial aggregation processing is performed on the pixel intensity or multispectral channel information within the ROI to project the two-dimensional spatial information into a corresponding one-dimensional initial human temporal signal. When using a non-imaging sensor, the corresponding one-dimensional micro-motion signal of the human surface is directly acquired. The acquired one-dimensional micro-motion signal of the human surface is then processed at an equivalent sampling rate of not less than 100Hz. A unified resampling process is performed, and the data is divided into continuous time segments with a fixed duration of 100ms. Normalization and baseline drift correction are sequentially applied to each time segment to obtain the initial human body time-series signal. If the initial human body time-series signal contains multimodal sources, the signals of each mode are aligned on the time axis and combined into multi-channel one-dimensional time-series data to obtain the human body time-series signal. The number of channels is dynamically determined based on the type of sensor and the number of modes used, thus providing a unified data input format for subsequent interference prediction and dynamic calculation path selection. When the initial human body time-series signal is not a multimodal signal, it is used as the human body time-series signal.
[0022] Figure 2 This is a multi-component waveform feature map containing interference provided in an embodiment of the present invention, such as... Figure 2 As shown, after S101 completes multimodal signal acquisition and time-series segmentation, the acquired one-dimensional human time-series signal may still contain various typical interferences, specifically manifested as the time-series characteristics of physiological signals with a duration of 1 second: the green curve represents a periodic interference component of 0.8Hz (energy percentage 30%), exhibiting high-frequency regular fluctuations; the blue curve is a mixed signal superimposed with physiological rhythms such as heart rate and respiration and random noise, with slow fluctuations at the baseline; the red peaks are non-periodic sudden interferences occurring at time points such as 0.2s and 0.6s, such as limb micro-movements. These interferences will directly affect the accuracy of subsequent physiological parameter decoding. Therefore, it is necessary to accurately determine the interference status of the current human time-series signal segment through the interference prediction model of S102.
[0023] S102. Input the human time series signal into the pre-constructed interference prediction model to obtain the interference probability of the human time series signal and the corresponding interference type label; wherein, the interference prediction model includes a one-dimensional depthwise separable convolutional structure, a bottleneck mapping path, a dilated convolutional branch and a pruned dilated convolutional branch connected in sequence.
[0024] Understandably, an interference prediction model suitable for one-dimensional human temporal signals is constructed. Its input layer receives a 100ms one-dimensional human temporal signal obtained from S101 via spatial projection of the imaging region or non-imaging methods. To address the temporal characteristics of the one-dimensional signal, the two-dimensional convolutional structure in the original MobileNet-V2 model is replaced with a one-dimensional depthwise separable convolutional structure. The kernel width is set to 3, the stride to 1, and the pooling layer is removed to avoid weakening high-frequency physiological rhythm information. In the main structure of the model, only the bottleneck layer in the inverse residual structure is retained, and structured pruning is performed on the dilated convolutional branches to compress the model size, bringing the number of model parameters to 42% ± 5% of the original model. The pruned model is jointly fine-tuned using a physiological interference dataset containing both optical imaging iPPG signals and non-imaging, non-contact signals, ensuring that the model can stably output the interference probability value and corresponding interference type label of the human temporal signal under different signal acquisition mechanisms, thus obtaining the pre-constructed interference prediction model. It is worth mentioning that the interference prediction model, as the core decision-making unit connecting the front-end time-series segmentation and the back-end dynamic computation path selection, is designed with constraints on low latency and interference discriminability in both its structure and output form. This applies to single-channel or multi-channel one-dimensional human body time-series signals output by the S101. ,in Consistent with the sensor modal number For the number of sampling points within 100ms, the first layer of the model uses one-dimensional depthwise separable convolution to model the temporal local perturbation, and its calculation form is defined as: ; This method is used to model local perturbations in a 100ms one-dimensional time segment. Its computational logic consists of two concatenated stages: depthwise convolution and pointwise convolution. This retains the ability to respond to high-frequency physiological micro-vibrations while significantly reducing redundant computational load across channels. Indicates the first Each channel at time step Convolutional output features at the location; This is the depthwise convolution calculation part. These are the depthwise convolution weights that unfold along the time axis (corresponding to a kernel width of 3). For the first Each channel at time step The input time-series signal at the point is used for feature extraction of local time-series information within a single channel, capturing physiological rhythms and perturbation details in the time dimension. This is the part involving pointwise convolution calculation. These are the pointwise convolution weights between channels. The number of input signal channels is consistent with the number of sensor modes. For the first Each channel at time step The input signal at this stage is responsible for fusing feature information from different channels, enabling cross-channel feature interaction and dimensional adjustment, and ultimately outputting the signal. This provides a stable feature distribution basis for the inverted residual structure of the subsequent model body.
[0025] The interference probability of human body time-series signals is obtained by normalization function after a one-dimensional fully connected mapping, as follows: ; in, It represents the probability of interference with human temporal signals, directly reflecting the proportion of interference energy relative to physiological energy. The total energy of the non-periodic perturbation. Indicates the first A discriminant characteristic response associated with aperiodic perturbations, express Quantity, Indicates the first A characteristic response associated with periodic physiological rhythms, express The formula calculates the probability value between 0 and 1 by using the total energy of aperiodic perturbations as the numerator and the sum of the total energy of aperiodic perturbations and the total energy of periodic physiological rhythms as the denominator. This result is directly used to trigger the subsequent model deep switching mechanism.
[0026] Optionally, the pruning process for the dilated convolutional branch includes: Calculate the importance score of each bottleneck layer in the dilated convolution branch; The importance scores of each bottleneck layer are sorted in descending order, and each bottleneck layer is pruned according to a preset ratio to obtain pruned dilated convolutional branches.
[0027] Understandably, in the main network part, only the bottleneck mapping path in the inverted residual structure of the MobileNet-V2 model is retained, and structured pruning is performed on the dilated convolutional layers. The importance score of each bottleneck layer is ranked by channel importance score, and the calculation method of the importance score of each bottleneck layer is expressed as follows: ; in, Indicates the first The importance score of each bottleneck layer Indicates the first The number of channels for the output features of each bottleneck layer. Indicates the time step of a time sequence segment. For the first The bottleneck layer, the first The first channel in the The formula is the core function in S102 used to score the importance of bottleneck layer channels in the model. It aims to provide a quantitative basis for structured pruning. Its calculation logic revolves around the global statistics of the energy intensity of the feature map. The importance scoring formula first sums the absolute values of the feature responses of all channels and all time steps of a single bottleneck layer, and then divides them by the product of the number of channels and the time step length to obtain the average value. This quantifies the overall feature expression intensity of each bottleneck layer. The higher the score, the more important the bottleneck layer is for extracting physiological rhythm features or discriminating interference. The importance scores of each bottleneck layer are sorted in descending order. According to a preset ratio (e.g., the bottom 50% of bottleneck layers), each bottleneck layer after sorting is pruned to obtain the pruned dilated convolutional branch. While compressing the model parameters to 42% ± 5% of the original model, it ensures the consistent expression scale of the temporal backbone information. The model output adopts a joint discriminant form. S103. Based on the calculated depth adjustment factor, the interference prediction model is softly switched to the dilated convolution branch or the pruned dilated convolution branch, and the human temporal signal is input into the corresponding branch to obtain the enhanced human features; the calculated depth adjustment factor is obtained based on the interference probability of the human temporal signal.
[0028] Understandably, the model calculation path is dynamically adjusted based on the interference probability value output in step two: when When the value is ≤0.3, the full model branch is enabled to perform high-precision feature extraction; when When the value is >0.3, switch to the pruned lightweight branch; embed a temporal attention module at the end of the two branches. This module compresses the channel dimension through a global average pooling layer, generates the weight coefficients of each channel, assigns a weight of 2.5±0.8 times to periodic physiological features, and assigns a suppression weight of 0.2±0.1 times to non-periodic motion interference components. Furthermore, model depth switching can be extended from single-threshold decision to continuous mapping control. By introducing a computational depth adjustment factor driven by the disturbance probability, the model can achieve soft switching rather than hard switching between two branches. The mapping relationship is defined as follows: The depth adjustment factor is calculated as follows: ; in, This represents the depth adjustment factor, whose value ranges from 0 to 1. This represents an exponential function used to distribute computational weights between two model branches. This represents a control factor used to allocate computational weights between the full dilated convolutional branch and the pruned dilated convolutional branch. This avoids frequent path jitter in the moderate interference range and improves feature continuity while maintaining low latency. The core logic is to use the sigmoid function to transform the interference probability into continuously adjustable weight coefficients, achieving a smooth transition between the full dilated convolutional branch and the pruned dilated convolutional branch rather than an abrupt switch. This represents a preset interference threshold, serving as a critical reference point for switching. Preferably, it can be set to 0.3. The interference probability is the core input basis for the adjustment factor. This formula uses the Sigmoid function. Its nonlinear mapping characteristics make the interference probability lower than hour When the value approaches 0 and exceeds the preset high interference threshold, Approaching 1, a smooth transition is formed in the medium interference range, avoiding frequent path jitter, while taking into account both low-latency targets and feature continuity, providing a stable computational basis for subsequent feature enhancement. At the feature enhancement level, although the current weighting coefficients for periodic and non-periodic components are numerically reasonable and have clear physical meanings, they still belong to the fixed interval form. The optimization direction is not to expand the weight range, but to make the weight directly related to the rhythmic stability of the current segment.
[0029] Optionally, the human temporal signal is input into the corresponding branch to obtain enhanced human features, including: Input the human body time-series signal into the corresponding branch and extract the initial features; The initial features are input into the temporal attention module, global average pooling is performed, the basic weights are adjusted according to the consistency index, and the pooled initial features are enhanced to obtain the enhanced human features.
[0030] Understandably, attention weights can be fine-tuned by introducing a rhythm consistency metric, which is defined as: ; in, This indicates a consistency index; a higher index indicates a more stable periodic physiological rhythm. This indicates the initial features after pooling at time step The instantaneous response on the main frequency channel (i.e., the characteristic response related to periodic physiological rhythms such as heart rate and respiration)). The mean response is the instantaneous response within a time segment of time step. This is the time step, consistent with the number of sampling points at a 100Hz sampling rate. This is achieved by first processing each time step... The absolute value of the difference between the instantaneous response and the mean response at a given time step is taken, and then the reciprocal is calculated. This summates the results over all time steps, and finally, the result is divided by the time step size. An averaged rhythm consistency index is obtained, which acts as a weight scaling factor on the original temporal attention weight range. This allows the intensity of feature enhancement and suppression to be dynamically adjusted according to the rhythmic stability of the current segment, improving the temporal continuity and frequency band consistency of the features. Through further optimization, S103 evolves from a static rule-driven mechanism to a continuously adjustable dynamic scheduling mechanism. Its output features are more stable in terms of temporal continuity and frequency band consistency, providing smoother and separable enhanced human features for the parallel modeling of different physiological rhythms by multi-scale dilated convolution in S104.
[0031] S104. Input the enhanced human body features into the multi-scale dilated convolutional layer to obtain rhythmic fusion features; wherein, the multi-scale dilated convolutional layer includes 3 parallel dilated convolutional sub-modules with different dilation rates and a convolutional module.
[0032] Optionally, S104 may include: The enhanced human body features are input into three parallel dilated convolution sub-modules with different dilation rates to obtain the corresponding three dilated human body features. Calculate the dynamic fusion weights of three parallel dilated convolutional sub-modules with different dilation rates and normalize them; Based on the dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates after normalization and the corresponding three dilated human features, feature fusion is performed to obtain the fused features. The fused features are reweighted based on rhythm consistency constraints and then input into the convolution module to obtain rhythm fusion features.
[0033] Understandably, the enhanced human features obtained from S103 are input into a multi-scale dilated convolutional layer. This layer contains three parallel dilated convolutional sub-modules with dilation rates of 1, 2, and 4, respectively, and a kernel width of 5 for each. The convolution with a dilation rate of 1 captures local pulsation details, corresponding to the heart rate frequency band of 2-5Hz; the convolution with a dilation rate of 2 covers the respiratory fundamental frequency; and the convolution with a dilation rate of 4 captures long-period respiratory harmonics. The features output from the three parallel dilated convolutional sub-modules are concatenated along the channel dimension and compressed to a uniform dimension by a 1×1 convolution to generate interference-resistant rhythmic fusion features. Specifically, while existing multi-scale dilated convolutional structures are valid and effective in terms of frequency band coverage and parallel modeling, their scale relationships remain fixed. The optimization space mainly lies in creating an intrinsic coupling between the scale response and the stability of the features output from step three, as well as the interference state, rather than simply increasing the number of convolutional layers or adjusting the dilation rate. Specifically, for the enhanced human features output by S103, while keeping the discrete set of dilation rates (1, 2, 4) constant, a scale-adaptive response coefficient can be introduced. This allows the contribution of different dilated branches to the current segment to continuously change with the interference state, thus avoiding the problem of over-amplification or over-suppression of a certain scale feature within the moderate interference range. Therefore, based on the interference probability... In addition to the rhythmic fusion feature distribution, the dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates are defined as follows: ; in, Indicates the expansion rate The dynamic fusion weights of the dilated convolutional submodule, The values are 1, 2, and 4, corresponding to the heart rate frequency band, respiratory fundamental frequency, and respiratory harmonic frequency band, respectively. This represents the probability of interference with human temporal signals. To be related to the expansion rate The relevant sensitivity coefficient, This represents an exponential function. The formula is a weight calculation expression used for dynamic fusion of multi-scale dilated convolution branches. Its core logic is to adaptively allocate the contributions of each scale branch based on the disturbance state and rhythm stability, avoiding the problem of over-amplification or suppression of features under a fixed scale configuration. This is used to normalize the weights, ensuring that the sum of all branch weights is 1. During calculation, the interference probability and sensitivity coefficient are first normalized using an exponential function. The negative correlation combinations are transformed into non-negative values, and then the relative contribution weights of each branch are obtained through the normalization operation of the denominator. This transforms multi-scale feature fusion from static splicing to continuous fusion of probability modulation, allowing the contribution of each scale branch to dynamically adapt to the interference state, thereby improving the anti-interference capability and frequency band adaptability of the fused features. Based on this, to enhance the separability of features at different scales before channel splicing, a rhythm consistency constraint consistent with the attention result in step three can be introduced into each dilated convolution branch. This is achieved by defining an intra-scale energy concentration index: ; in, Indicates the expansion rate The dilated convolutional submodule in time The response at the location, The mean response of this branch is used to reweight the channels before 1×1 convolution compression, so that scale features that exhibit stronger periodic consistency within the current segment occupy a higher proportion of channel expression.
[0034] S105. After obtaining the spectral information by performing a fast Fourier transform on the rhythm fusion features, locate the main peak of the physiological rhythm based on the energy distribution according to the spectral information to complete the decoding of physiological parameters.
[0035] Understandably, a Fast Fourier Transform (FFT) is performed on the anti-interference rhythm fusion features obtained from S104 to obtain the corresponding spectral information. In the frequency domain, the main peak of the physiological rhythm is located based on energy distribution. First, within the respiratory frequency band (0.1–0.5 Hz), rhythm fusion features from multi-scale dilated convolution branches are combined to determine the candidate interval for the respiratory main frequency, and the peak with the highest energy is selected as the initial value of the respiratory rate. Subsequently, within the heart rate frequency band (1.0–3.0 Hz), the candidate heart rate peaks are screened using the initial value of the respiratory rate and its harmonic relationships as constraints to eliminate harmonic interference caused by respiration and its harmonics, thereby obtaining… The system obtains the main heart rate peak and outputs the corresponding respiratory rate and heart rate parameters. Simultaneously, it determines the reliability of the current segment decoding result based on the interference probability value. When the interference probability exceeds the preset high interference threshold, the physiological parameters of the segment are directly discarded and the sensor is triggered to re-acquire. When the interference probability is higher than the preset interference threshold but lower than the preset high interference threshold, the decoding result is smoothed by combining the rhythm continuity of the preceding and following segments to ensure the stability and reliability of the physiological parameter output in continuous monitoring scenarios. This completes the adaptive physiological parameter decoding process consistent with the aforementioned dynamic model scheduling and multi-scale rhythm fusion.
[0036] This invention constructs a backbone processing pipeline with a lightweight interference prediction model as the decision-making core and dynamic computational path switching capability. This pipeline sequentially executes time-series segmentation, interference probability and interference type discrimination, probability-based model depth soft switching and feature enhancement, and multi-scale physiological rhythm fusion, ultimately completing physiological parameter decoding. Its core architectural feature is that the interference prediction model provides interference probabilities for subsequent processing stages, and uses these probabilities to drive the computational depth, attention weights, and multi-scale fusion strategies of the feature extraction model to adaptively adjust. This achieves closed-loop collaboration between interference identification, path switching, and dynamic filtering, ensuring low-latency processing while optimizing the robustness and accuracy of physiological signal extraction in complex interference environments.
[0037] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.
[0038] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0039] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.
[0040] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A non-contact physiological signal processing method for low-latency interference identification and dynamic filtering, characterized in that, The method includes: Human body time-series signals are acquired through sensors; wherein, the sensors are imaging sensors or non-imaging sensors; when the sensor is an imaging sensor, the human body time-series signal is a sequence of human skin images; when the sensor is a non-imaging sensor, the human body time-series signal is a micro-motion signal on the human body surface. The human time series signal is input into a pre-constructed interference prediction model to obtain the interference probability and corresponding interference type label of the human time series signal; wherein, the interference prediction model includes a one-dimensional depthwise separable convolutional structure, a bottleneck mapping path, a dilated convolutional branch and a pruned dilated convolutional branch connected in sequence. The interference prediction model is soft-switched to the dilated convolution branch or the pruned dilated convolution branch according to the computational depth adjustment factor, and the human temporal signal is input into the corresponding branch to obtain the enhanced human features; the computational depth adjustment factor is obtained according to the interference probability of the human temporal signal. The enhanced human features are input into a multi-scale dilated convolutional layer to obtain rhythmic fusion features; wherein, the multi-scale dilated convolutional layer includes three parallel dilated convolutional sub-modules with different dilation rates and a convolutional module; After obtaining the spectral information by performing a fast Fourier transform on the fusion features of the rhythm, the main peak of the physiological rhythm is located based on the energy distribution according to the spectral information to complete the decoding of physiological parameters.
2. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The acquisition of human body time-series signals through sensors includes: When the human body time-series signal is acquired through the imaging sensor, an image sequence of the human skin region is acquired, and spatial localization and tracking of the region of interest of the human skin are performed on the image sequence. Then, the pixel intensity or multispectral channel information of the region of interest of the human skin is spatially aggregated to obtain the initial human body time-series signal. When the human body time-series signal is acquired through the non-imaging sensor, the human body surface micro-motion signal is acquired, and after uniform resampling of the human body surface micro-motion signal, it is divided into continuous time-series segments. Then, normalization and baseline drift correction processing are sequentially performed on the continuous time-series segments to obtain the initial human body time-series signal. When the initial human body time series signal is a multimodal signal, the modal signals in the initial human body time series signal are aligned on the time axis and combined into multi-channel one-dimensional time series data to obtain the human body time series signal. When the initial human body time series signal is not a multimodal signal, the initial human body time series signal is used as the human body time series signal.
3. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The pruning process of the dilated convolution branch includes: Calculate the importance score of each bottleneck layer in the dilated convolution branch; The importance scores of each bottleneck layer are sorted in descending order, and each bottleneck layer is pruned according to a preset ratio to obtain the pruned dilated convolutional branch.
4. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 3, characterized in that, The importance score for each bottleneck layer is calculated as follows: ; in, Indicates the first The importance score of each bottleneck layer Indicates the first The number of channels for the output features of each bottleneck layer. Indicates the time step of a time sequence segment. For the first The bottleneck layer, the first The first channel in the Output feature response values at each time step.
5. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The interference probability of the human body time-series signal is expressed as follows: ; in, This represents the interference probability of the human body time-series signal. Indicates the first A discriminant feature response associated with aperiodic perturbations, express Quantity, Indicates the first A characteristic response associated with periodic physiological rhythms, express The quantity.
6. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 5, characterized in that, The calculation depth adjustment factor is expressed as follows: ; in, This represents the calculation depth adjustment factor. Represents an exponential function. Indicates control factor. This indicates the preset interference threshold.
7. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The step of inputting the human temporal signal into the corresponding branch to obtain enhanced human features includes: The human body time-series signal is input into the corresponding branch to extract the initial features; The initial features are input into the temporal attention module, global average pooling is performed, the basic weights are adjusted according to the consistency index, and the pooled initial features are enhanced to obtain the enhanced human features.
8. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The consistency index is expressed as follows: ; in, This indicates the consistency index. This indicates the initial features after pooling at time step Instantaneous response on the main frequency channel The mean response is the instantaneous response within a time segment of time step. For time step.
9. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 1, characterized in that, The step of inputting the enhanced human features into a multi-scale dilated convolutional layer to obtain rhythmic fusion features includes: The enhanced human body features are input into three parallel dilated convolution sub-modules with different dilation rates to obtain the corresponding three dilated human body features. Calculate the dynamic fusion weights of three parallel dilated convolutional sub-modules with different dilation rates and perform normalization processing; Based on the dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates after normalization and the corresponding three dilated human features, feature fusion is performed to obtain the fused features. The fused features are reweighted according to the rhythm consistency constraint and then input into the convolution module to obtain the rhythm fusion features.
10. The non-contact physiological signal processing method for low-latency interference identification and dynamic filtering according to claim 9, characterized in that, The dynamic fusion weights of the three parallel dilated convolutional sub-modules with different dilation rates are represented as follows: ; in, Indicates the expansion rate Dynamic fusion weights of the dilated convolutional submodules This represents the interference probability of the human body time-series signal. To be related to the expansion rate The relevant sensitivity coefficient, This represents an exponential function.