A Multi-channel PPG Wide-Range Blood Oxygen Saturation Estimation Method for High-Altitude Environments
By employing a quality-perception feature extraction and attention-based cross-site fusion strategy, the challenges of modeling the heterogeneity of multi-channel PPG signal quality and inter-channel relationships in high-altitude environments were addressed. This enabled stable and accurate blood oxygen saturation estimation for wrist-worn devices, making them suitable for continuous monitoring in high-altitude environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
In high-altitude environments, traditional methods for estimating blood oxygen saturation are difficult to achieve continuous and reliable long-term monitoring. Furthermore, the heterogeneity of signal quality and the difficulty in modeling the relationships between channels in multi-channel PPG signals result in insufficient estimation accuracy and robustness.
We designed a quality-aware feature extraction module and an attention-based cross-site fusion strategy. By evaluating signal reliability and modulating features in the time dimension, we suppressed low-quality signal interference. Furthermore, we integrated complementary vascular features of PPG signals from multiple sites through learnable global representations to construct an estimation model for multi-channel PPG signals.
It achieves stable and accurate blood oxygen saturation estimation within a wide range of 70%-100%, improving the accuracy and robustness of the estimation, and is suitable for daily health monitoring and high-altitude hypoxia risk early warning using wrist-worn devices.
Smart Images

Figure CN122074977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of physiological signal processing and dynamic blood oxygen monitoring, and more specifically, to a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments. Background Technology
[0002] In high-altitude environments, low oxygen and low pressure significantly reduce human arterial blood oxygen saturation. High altitude exposure can trigger a range of altitude-related illnesses, such as acute altitude sickness and high-altitude pulmonary edema, posing a direct threat to individual health and safety. Against this backdrop, continuous and reliable blood oxygen saturation estimation is crucial for risk warning and health management for workers, mountaineers, and those undergoing training at high altitudes. However, the changes in vascular status at high altitudes, unlike those in plains environments, make photoplethysmography (PPG)-based methods less effective. It is estimated that even greater challenges will be faced. Traditional pulse oximeters mostly rely on static, short-term measurements, which are difficult to meet the needs of long-term, continuous monitoring in high-altitude environments. Therefore, developing a continuous pulse oximeter for high-altitude hypoxic environments is crucial. The estimation method is particularly crucial.
[0003] For blood oxygen saturation estimation, the traditional method is represented by the ratio method (RoR), which is based on the Beer-Lambert law and calculates the ratio of the pulsating (AC) and non-pulsating (DC) components of the red and infrared light signals. However, this method assumes that the optical path length is fixed and the AC / DC ratio is... The relationship is linear; therefore, due to tissue scattering and the nonlinear absorption of hemoglobin under low saturation, the effect occurs over a relatively wide range of conditions. Within this range, reliable estimation is difficult to achieve.
[0004] To overcome this limitation, recent studies have employed various deep learning methods to capture nonlinear hemodynamics. Notably, these studies primarily focus on video-based remote PPG for... It is estimated that current methods are typically limited to controlled interaction scenarios and struggle to meet the needs of continuous daily monitoring. In contrast, wrist-worn devices offer the advantage of contact-based measurement without environmental constraints, making them a more suitable option for monitoring blood oxygenation during daily activities. However, existing research primarily relies on PPG signals acquired by a single photodiode, which only reflects local vascular dynamics. In this case, the obtained physiological information is less comprehensive than that provided by multi-site PPG measurements, significantly limiting the scope of research. The availability of various estimated features is limited. Furthermore, the dependence of this estimation on individual measurement sites makes it highly susceptible to local interference, and the unstable contact between the sensor and the skin leads to a significant performance degradation in practical applications.
[0005] In comparison, multi-site PPG measurements offer physical redundancy and diverse signal characteristics, enhancing robustness to local signal degradation. However, multi-site-based... The estimation faces two major challenges. First, the signal quality differences between sensors complicate the integration of multiple PPG channels. Due to the temporal heterogeneity of different channels, applying a uniform feature extraction process is unsuitable, as contaminated PPG channel signals may be misclassified as valid physiological features, thus reducing estimation accuracy. Second, modeling the relationships between channels to obtain a coherent physiological representation is difficult. PPG signals from different sensors are spatially correlated views of the same underlying physiological process and are not independent of each other. These signals contain both redundant information reflecting a common cardiac cycle and complementary details from local vascular sites. Effectively utilizing these signals to construct the optimal feature representation is crucial for accurate and robust estimation. Estimation is crucial.
[0006] Therefore, there is an urgent need for a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments, which can dynamically assess PPG signal quality, efficiently fuse features from multiple sites for wide-range blood oxygen estimation, and promote the practical application of blood oxygen monitoring technology for wearable devices. Summary of the Invention
[0007] This invention provides a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments. Its purpose is to solve key problems such as wide-range adaptability, signal quality heterogeneity, and effective fusion of information from multiple sites in blood oxygen saturation estimation under wrist-worn device scenarios.
[0008] To achieve the above objectives, the technical solution of this invention first designs a quality-aware feature extraction module, which suppresses interference from low-quality signal components through signal reliability assessment and feature modulation in the time dimension; subsequently, it proposes an attention-based cross-site fusion strategy, which integrates complementary vascular features of PPG signals from multiple sites through learnable global representations to achieve precise... Inference. This framework relies solely on multi-channel PPG signals from the wrist to achieve a width of 70%-100%. It achieves stable and accurate continuous estimation within a certain range, while also possessing good subject adaptability and clinical applicability.
[0009] This invention provides a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments, which specifically includes the following steps: S1. Acquire multi-channel PPG signals and corresponding references. The measured values were obtained, and more data were preprocessed to obtain standardized multi-channel PPG signals and corresponding standardized reference labels. S2. Based on the standardized multi-channel PPG signal and the corresponding standardized reference label, construct and train... The estimation model, the The estimation model includes a quality-perceived feature extraction module, an attention cross-region fusion module, and a regression head; among which: S21. The quality-aware feature extraction module is used to extract features, generate reliability masks, and optimize modulation of multi-channel PPG signals to output high-quality features. S22, The attention cross-part fusion module is used to concatenate high-quality features with built-in learnable global nodes and then obtain global fused features through multi-head self-attention processing and linear projection; S23. The regression head is used to infer blood oxygen saturation values based on global fusion features; S24, Training In the process of estimating the model, a composite loss function combining mean squared error and Pearson correlation coefficient is used, and the model parameters are iteratively optimized using standardized reference labels as the true values. S3. Input the newly acquired and preprocessed multi-channel PPG signal into the training completed signal. The estimation model enables accurate and continuous estimation of blood oxygen saturation over a wide range of 70%-100%.
[0010] Preferably, in S1, during the data preprocessing stage, because the testing equipment may automatically terminate recording due to its safety protection mechanism in a simulated hypoxic environment, and subjects may temporarily remove their masks to alleviate discomfort in some cases, these factors can lead to waveform distortion. Therefore, preprocessing is required before modeling to ensure the quality of the multi-channel PPG signal input to the model, laying the foundation for wide-range blood oxygen saturation estimation. The specific content of the preprocessing includes: S11. Compare the acquired multi-channel PPG signal with the corresponding reference. The measured values are time-aligned to ensure the timing consistency between the signal and the tag; S12. Remove invalid records and references from multi-channel PPG signals that have waveform distortion or are missing. Abnormal data with a measured value of zero or exceeding the physiologically reasonable range; S13. Divide the aligned PPG signal into signal segments according to a fixed length and a set overlap, and use the average value within each segment as the basis for the signal segment. The value serves as a reference label for this segment; S14. The baseline drift of the PPG signal is mitigated by using a moving average filter, and then a low-pass filter is used to suppress high-frequency interference noise and improve signal quality. S15. Adjust the sampling rate of PPG signals from different source datasets to ensure that all input data have a uniform format, and obtain standardized multi-channel PPG signals and corresponding standardized reference labels.
[0011] Preferably, in S2, the core objective of the quality-aware feature extraction module is to suppress the influence of unreliable signal regions while extracting effective features. This module includes two key components: a feature encoder and a quality mask generator. The specific implementation process is as follows: S211. The standardized multi-channel PPG signal is extracted using a convolutional neural network encoder to capture physiological correlation information between channels and output an initial feature matrix. S212. A multi-scale convolution kernel (corresponding to 1s, 1.5s, and 3s time scales) is used to process each PPG channel separately. A reliability mask in the time dimension is generated based on the similarity of the signal morphology. The mask value is used to characterize the reliability of the signal at the corresponding time position. The higher the mask value, the lower the reliability of the signal at the corresponding time position. S213, To avoid damage The key amplitude information is estimated, and the reliability mask is aligned with the initial feature matrix in the time dimension and then modulated element by element. At the same time, a learnable positional code is introduced to supplement the temporal information, optimize the feature representation, and output high-quality features.
[0012] Preferably, the initial feature matrix The expression is: ,in The number of feature channels, The length of the feature's time dimension. The reliability mask The expression is: ; The high quality feature The calculation process includes: in, For convolutional neural network encoders, For quality mask generator, Represents element-wise product. It is a learnable positional encoding.
[0013] The feature encoder and the quality mask generator employ different convolutional architectures. The former uses a narrower kernel and greater depth to capture higher-order integrated representations, while the latter uses a multi-scale kernel to capture morphological and contextual dependencies at different time scales, thus decoupling "feature effectiveness" from "signal reliability".
[0014] Preferably, in S2, in the attention cross-site fusion module, a learnable global node is introduced to effectively integrate the complementary information of PPG signals from multiple sites. (D represents the embedding dimension) As the aggregation center, the specific process is as follows: S221. Reshape the high-quality features output by the quality-perceived feature extraction module according to channel pairs to obtain the potential feature representation set corresponding to each part. S222. Construct and initialize learnable global nodes, and concatenate the global nodes with the latent feature representation set to form the input sequence of the fusion module; S223. The input sequence of the fusion module is passed into the multi-head self-attention layer, and a query, key, and value matrix is generated by linear projection. The similarity between features is calculated based on scaled dot product attention, and an attention-weighted representation is generated. S224. Concatenate the weighted representations of all attention heads, perform linear projection to obtain the global attention representation, discard channel-specific representations, extract the updated global node latent representations as global fusion features, and input them into the regression head to obtain the final [attention representation]. Estimated value.
[0015] Preferably, in S222, the number of PPG channel pairs is set to... C represents the total number of PPG signal channels; The high quality feature The expression is: ,in , for the first The potential representation of each channel pair Represents transpose; The fusion module input sequence The expression is: ; in Random initialization is performed before training, and multi-channel global dependencies are adaptively learned during the training process.
[0016] In S223, each attention head h is the number of attention heads, which project the input sequence into a query matrix. Key matrix Sum matrix : ; in, The attention-weighted representation is the input sequence to the fusion module. The expression is: ; in, For each dimension of attention head, , , The first i The query, key, and value projection matrix of each attention head; In S224, the global attention representation The expression is: ; in, The output projection matrix; the global fusion feature The expression is: ; in, = , to select vectors for features.
[0017] Furthermore, this invention designs a composite training loss function. To simultaneously ensure estimation accuracy and linear consistency, a composite loss function combining mean squared error (MSE) and Pearson correlation coefficient (PCC) is adopted. MSE loss is used to quantify the numerical difference between the estimated and true values; PCC loss is used to strengthen the linear correlation between the estimated and true values and suppress systematic bias. The final total training loss function integrates MSE and PCC losses, balancing the contribution ratios of the two losses through weighting coefficients, achieving dual optimization of accuracy and consistency.
[0018] Furthermore, in The estimation model construction and training process employs a cross-validation strategy with retained subjects. In each tradeoff, approximately 80% of the subjects are used for training, and the remaining 20% for testing. 20% of the training data is randomly selected as the validation set for model selection. The testing phase includes two scenarios: no fine-tuning and fine-tuning. In the fine-tuning scenario, 20% of the test set data is used to update the model. All pre-training and test data are non-overlapping to avoid data leakage.
[0019] Therefore, the present invention employs the above-mentioned multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments, which has the following advantages compared with the prior art: (1) In the quality-sensing feature extraction module, the interference of low-quality signals and noise is effectively suppressed by reliability assessment and feature modulation in the time dimension, thus solving the problem of heterogeneous quality of multi-channel PPG signals; (2) The attention cross-site fusion strategy aggregates complementary information from multiple sites through learnable global nodes, making full use of the differences and correlations in vascular features of different sites, and improving the wide range of attention. The accuracy of the estimate; (3) This invention relies solely on multi-channel PPG signals collected by a wrist-worn device, without the need for additional sensors. It balances the convenience of wearing and clinical applicability, and is suitable for daily health monitoring and hypoxia risk warning scenarios in special environments such as high altitude.
[0020] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the process architecture of a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to the present invention. Figure 2 This is a dataset from an embodiment of a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to the present invention. Reference value distribution histogram, where (a) represents the distribution of representative subjects. Distribution, (b) is the overall distribution of all subjects. distributed; Figure 3 This is a visualization of the overall performance of the proposed method in an embodiment of a multi-channel PPG wide-range blood oxygen saturation estimation method for plateau environments; where (a) is a scatter plot and fitting line of predicted values and true values, and (b) is a Bland-Altman analysis plot. Figure 4 This paper presents a comparison of the performance distribution of different methods on various subjects in an embodiment of a multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to the present invention; wherein, (a) is the estimation result of the population-level model, and (b) is the estimation result of the fine-tuned model. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0023] Example 1 This embodiment first describes the dataset and experimental setup used, then verifies the effectiveness of the core module through ablation experiments, and finally analyzes the performance differences of different channel pairs and methods, confirming that the proposed framework is effective across a wide range of applications. Accuracy and robustness in estimation. Figure 1 The proposed method architecture is visualized.
[0024] I. Dataset The datasets used in this embodiment include a self-developed clinical hypoxia training dataset and the publicly available DaLiA dataset.
[0025] The self-tested clinical hypoxia training dataset was collected through a clinical hypoxia training process, including 8 subjects. Each subject sequentially experienced four gradient hypoxia stages, recording for 10 minutes at oxygen concentrations of 15%, 14%, 13%, and 12%, corresponding to simulated altitudes of 2700m, 3200m, 3800m, and 4400m. The hypoxic environment was generated using a GO2Altitude hypoxia generator (Biomedtech, Australia).
[0026] During data acquisition, a commercial finger-clip pulse oximeter was used to record reference data at a sampling rate of 1 Hz. The device simultaneously acquires multi-channel PPG signals at a sampling rate of 100Hz using a self-developed wristband device. This wristband device is equipped with four photodiodes, each containing both red and infrared light sources, providing a total of eight PPG channels.
[0027] The DaLiA dataset contains physiological signals from 15 subjects. PPG data were acquired using wrist and chest wearable devices at a sampling rate of 64Hz, during which subjects performed daily activity protocols. Artifact segmentation labels for this dataset were obtained using a network annotation tool for data quality screening. To maintain consistency with the sampling rate of the self-test dataset, the PPG signals from this dataset were upsampled to 100Hz for pre-training of the quality mask generator.
[0028] II. Experiment Setup The computing platform in this embodiment is a server equipped with two NVIDIA GeForce RTX 4090 GPUs (each with 24GB of video memory), one Intel(R) Core(TM) i9-14900K CPU, and 128GB of RAM. The experiment uses the Python 3.10 programming language and the PyTorch 1.12 deep learning framework.
[0029] In terms of data preprocessing and partitioning, the PPG signal and the reference signal are first... Time-aligned and discarded measurements Invalid records with a value of zero; the aligned PPG signal is divided into 3-second long, 1-second overlapping segments, with the average value within each segment being calculated. The values were used as reference labels; a moving average filter with a window size of 50 data points was used to alleviate baseline drift, and a 5th-order Butterworth low-pass filter (cutoff frequency 3.5Hz) was used to suppress high-frequency noise; finally, signal segments containing at least one effective channel pair within the same time window were retained, resulting in a total of 7246 samples for model training and evaluation. Figure 2 (b) Display all subjects' The overall distribution ranges from 70% to 100%, covering both normal oxygenation and hypoxia. Figure 2 (a) Showing different subjects The distribution shows significant differences, reflecting the heterogeneity of hypoxia tolerance among individuals.
[0030] In terms of model architecture configuration, in the quality-aware feature extraction module, the quality mask generator contains three convolutional layers with kernel sizes of 100, 150 and 300, respectively, to capture signal features at time scales of 1s, 1.5s and 3s; the feature encoder adopts a convolutional architecture with a narrow kernel and a large depth to extract high-order integrated representations.
[0031] In the attention cross-site fusion module, the number of attention heads is configured as needed, and the embedding dimension... satisfy ( (For a single attention head dimension); global nodes are initialized with random values, and multi-channel global dependencies are learned adaptively through training.
[0032] For comparative verification, CNN, LSTM, GRU, and Transformer were selected as baseline models. The traditional ratio method (RoR) was also included in the comparison. The coefficients A and B of RoR were calibrated based on the training dataset and then fine-tuned on a subset of the test set.
[0033] Regarding the training protocol, all models were trained using common deep learning optimizers, with a batch size of 64. The learning rate during the pre-training phase was 5×10⁻⁶. -4 The maximum training epoch is 1000 epochs, employing an early stopping strategy: if the validation loss does not improve for 50 consecutive epochs, training stops, and the model with the lowest loss is retained. The learning rate during the fine-tuning phase is set to 1×10⁻⁶. -4 The training period is 200 epochs, and only the parameters of the final regression head are updated to ensure fairness.
[0034] The model proposed in this invention adopts a composite loss function combining MSE and PCC, and balances the contribution ratio of the two losses by weighting coefficients, thereby enhancing the estimation accuracy and linear consistency.
[0035] III. Results Analysis First, an overall performance comparison was conducted. Table 1 shows the performance comparison between the Proposed method, the baseline model, and the RoR method. In the population-level evaluation, the MAE of the Proposed method was 3.437%, the lowest among all methods, and it met the acceptable error range (<4%) for clinical pulse oximeters. The MAEs of other baseline models ranged from 4.078% (LSTM) to 6.523% (RoR), none of which reached the 4% clinical threshold. In addition, the ME of the Proposed method was -0.611% (closest to zero), and the SD was 4.336% (lowest), indicating that it had the least systematic bias and estimation variability, and had stronger generalization ability in heterogeneous subject populations.
[0036] Table 1. Comparison of blood oxygen saturation estimation performance of different models at the population level and after fine-tuning.
[0037] After fine-tuning, the performance of all methods was significantly improved, reaching clinically acceptable standards. Furthermore, the Proposed method maintained a clear advantage. As shown in Table 1, its PCC significantly increased from 0.489 at the population level to 0.799 after fine-tuning, indicating that the initial representation of the proposed model is superior, and individual differences among subjects can be fully captured with only minor adjustments.
[0038] Figure 3 A comprehensive performance visualization analysis was conducted. Figure 3 In the scatter plot of (a), when fine-tuning for a limited range of subjects (<80%), low The distribution bias has been overestimated. This phenomenon may be due to data imbalance, because... The references are primarily distributed across specific datasets containing atypical measurements. Overall, performance improvements were observed across all data values, and the estimated... The real one in the fine-tuned driving method The proposed method exhibits linear consistency, remaining within the 70%-100% range, indicating that its capabilities outperform other models and demonstrate more accurate estimation power. Furthermore, Figure 3 (b) shows that the mean difference is close to zero and the ME is 0.02, with most data points falling within the 95% range, indicating that the proposed method can provide a wide range of... Stable estimation of the level.
[0039] To further analyze the performance of subject-specificity Figure 4The estimated distributions of different methods across various subjects are presented. In the population-level model, all methods exhibit overall underestimation in some subjects, but the proposed method's estimates are closer to the reference distribution, with the smallest ME. Some baseline models show physiologically abnormal estimates exceeding 100% in some subjects, or the estimated range is more dispersed, leading to higher MAE. In contrast, the proposed method closely matches the true distribution in most subjects, with no significant overestimation or underestimation. After fine-tuning, the estimated distributions of all models are more concentrated near the reference value, but some baseline models still exhibit abnormal estimates. The estimation results of the proposed method remain stable and close to the true value, further confirming its reliability in individual adaptation.
[0040] Next, to verify the effectiveness of the quality-aware extractor and the attention multi-site fusion module, an ablation experiment was conducted in this embodiment. In this process, QMG represents the quality mask generator in the quality-aware feature extraction module, and AMSF represents the attention multi-site fusion module. Table 2 shows the ablation experiment results. Compared with the baseline without these components, incorporating either OMG or AMSF alone reduces MAE. These results indicate that mitigating the impact of low-quality features and achieving cross-channel complementarity contributes to more stable global estimation. Therefore, the model equipped with both OMG and AMSF achieved the lowest MAE in both the population-level and fine-tuning evaluations. Furthermore, a more significant performance improvement was observed in the population-based setting, indicating that the combination of QMG and AMSF can effectively capture generalized features and reduce reliance on personalized data, further validating the rationale of the present invention.
[0041] Table 2 Comparison of Ablation Test Performance
[0042] This embodiment also includes channel pair performance analysis. Table 3 compares the performance of the Proposed and RoR methods on different channel pairs and across all channels. Both methods exhibit instability in the estimation results for single channel pairs. For example, in the C3+C4 channel pair, the RoR method's MAE is as high as 13.627%, and the RMSE is even higher at 446.81%. This is attributed to the poor PPG signal quality of this channel pair, making RoR more sensitive to noise and artifacts. While the Proposed method has the highest MAE (4.735%) in this channel pair, its RMSE is only 6.217%, indicating more stable estimation results. This is attributed to the model's ability to suppress noise and anomalous morphology. Comparing the performance of single channel pairs and across all channels, the RoR method's MAE across all channels is higher than that of the best single channel pair because its static fusion strategy introduces interference from low-quality channels. The proposed method, through dynamic attention weighting, achieves a significantly lower MAE across all channels than any single channel pair, fully demonstrating the effectiveness of multi-part fusion. That is, the model can automatically assign higher weights to high-confidence channels while reducing the influence of low-quality channels, maximizing the utilization of complementary information from multiple parts.
[0043] Table 3. Performance Comparison of the Proposed Method and the Ratio Method for Estimating Blood Oxygen Saturation
[0044] In summary, this invention proposes a quality-aware multi-site PPG fusion framework for wide-range applications in wrist-worn devices. Estimation. Through experimental verification using a clinical hypoxia training dataset and the DaLiA dataset, the MAE of this invention can reach as low as 2.154%, meeting the requirements for clinical application, and within a width of 70%-100%. The system exhibits stable estimation performance within its range. Ablation experiments and channel analysis confirm that the quality-aware feature extraction module effectively suppresses low-quality signal interference, while the attention-based cross-site fusion module fully integrates complementary information from multiple sites. The synergistic effect of these two modules significantly improves estimation accuracy and generalization ability. This framework relies solely on multi-channel PPG signals from the wrist, requiring no additional sensors, providing a practical solution for routine health monitoring and hypoxia risk warning in special environments such as high altitudes.
[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments, characterized in that, Specifically, the following steps are included: S1. Acquire multi-channel PPG signals and corresponding references. Measure the values and acquire the data for preprocessing to obtain standardized multi-channel PPG signals and corresponding standardized reference labels; S2. Based on the standardized multi-channel PPG signal and the corresponding standardized reference label, construct and train... The estimation model, the The estimation model includes a quality-perceived feature extraction module, an attention cross-region fusion module, and a regression head; among which: S21. The quality-aware feature extraction module is used to extract features, generate reliability masks, and optimize modulation of multi-channel PPG signals to output high-quality features. S22, The attention cross-part fusion module is used to concatenate high-quality features with built-in learnable global nodes and then obtain global fused features through multi-head self-attention processing and linear projection; S23. The regression head is used to infer blood oxygen saturation values based on global fusion features; S24, Training In the process of estimating the model, a composite loss function combining mean squared error and Pearson correlation coefficient is used, and the model parameters are iteratively optimized using standardized reference labels as the true values. S3. Input the newly acquired and preprocessed multi-channel PPG signal into the training completed signal. The estimation model enables accurate and continuous estimation of blood oxygen saturation over a wide range of 70%-100%.
2. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 1, characterized in that, In S1, the specific content of the preprocessing includes: S11. Compare the acquired multi-channel PPG signal with the corresponding reference. The measured values are time-aligned to ensure the timing consistency between the signal and the tag; S12. Remove invalid records and references from multi-channel PPG signals that have waveform distortion or are missing. Abnormal data with a measured value of zero or exceeding the physiologically reasonable range; S13. Divide the aligned PPG signal into signal segments according to a fixed length and a set overlap, and use the average value within each segment as the basis for the signal segment. The value serves as a reference label for this segment; S14. The baseline drift of the PPG signal is mitigated by using a moving average filter, and then a low-pass filter is used to suppress high-frequency interference noise and improve signal quality. S15. Adjust the sampling rate of PPG signals from different source datasets to ensure that all input data have a uniform format, and obtain standardized multi-channel PPG signals and corresponding standardized reference labels.
3. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 1, characterized in that, In S2, the specific content of the quality-perceived feature extraction module includes: S211. The standardized multi-channel PPG signal is extracted using a convolutional neural network encoder to capture physiological correlation information between channels and output an initial feature matrix. S212. A multi-scale convolution kernel is used to process each PPG channel separately, and a reliability mask in the time dimension is generated based on the similarity of signal morphology. The mask value is used to characterize the reliability of the signal at the corresponding time position. S213. After aligning the reliability mask with the initial feature matrix in the time dimension, perform element-wise modulation, and introduce learnable positional coding to supplement the temporal information, outputting high-quality features.
4. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 3, characterized in that: The initial feature matrix The expression is: ,in The number of feature channels, The length of the feature's time dimension; The reliability mask The expression is: ; The high quality feature The calculation process includes: in, For convolutional neural network encoders, For quality mask generator, Represents element-wise product. It is a learnable positional encoding.
5. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 1, characterized in that, In S2, the specific content of the attention cross-site fusion module includes: S221. Reshape the high-quality features output by the quality-perceived feature extraction module according to channel pairs to obtain the potential feature representation set corresponding to each part. S222. Construct and initialize learnable global nodes, and concatenate the global nodes with the latent feature representation set to form the input sequence of the fusion module; S223. The input sequence of the fusion module is passed to the multi-head self-attention layer, and a query, key, and value matrix is generated by linear projection. The similarity between features is calculated and an attention-weighted representation is generated. S224. Concatenate the weighted representations of all attention heads, and obtain the global attention representation through linear projection. Extract the updated global node latent representation as the global fusion feature.
6. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 5, characterized in that: In S222, the global node The expression is: Where D is the embedding dimension; Let the number of PPG channel pairs be C represents the total number of PPG signal channels; The high quality feature The expression is: ,in , for the first The potential representation of a pair of channels, where T represents transpose; The fusion module input sequence The expression is: ; In S223, each attention head h is the number of attention heads, which project the input sequence into a query matrix. Key matrix Sum matrix : ; in, The attention-weighted representation is the input sequence to the fusion module. The expression is: ; in, For each dimension of attention head, , , The first i The query, key, and value projection matrix of each attention head; In S224, the global attention representation The expression is: ; in, The output projection matrix; the global fusion feature The expression is: ; in, , to select vectors for features.
7. The multi-channel PPG wide-range blood oxygen saturation estimation method for high-altitude environments according to claim 1, characterized in that, In S24, the numerical difference between the mean squared error loss quantification estimate and the true value in the composite training loss function; The Pearson correlation coefficient loss strengthens the linear correlation between the estimated value and the true value, suppressing systematic bias.