Non-contact aortic pressure measurement method and system, storage medium and equipment
Through the 3DCNN-InvertU-Net model combined with multi-module preprocessing technology, the accuracy problem of contactless aortic pressure measurement is solved, and high-precision aortic pressure estimation is achieved, suitable for telemedicine and intelligent health monitoring.
Patent Information
- Application Number
- CN202510434366.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
The existing non-contact aortic pressure measurement methods have the problem of unsatisfactory accuracy, which is limited by the limitations of deep learning algorithms and external factors, especially lighting conditions and individual differences.
The 3DCNN-InvertU-Net fusion model is adopted, combined with multi-level and multi-module fine design, and the hemodynamic characteristics are extracted through video data preprocessing and deep learning to achieve accurate estimation of aortic pressure.
It significantly improves the accuracy and stability of aortic pressure measurement, is suitable for telemedicine and intelligent health monitoring, reduces information loss and error, and enhances the generalization ability and robustness of the model.
Smart Images

Figure CN120279468A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aortic pressure measurement, and specifically to a non-contact aortic pressure measurement method, system, storage medium and device. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Aortic pressure refers to the pressure exerted by blood on the blood vessel wall when flowing in the aorta (the largest artery in the human body). Among them, aortic systolic pressure determines the afterload during the systolic phase of the left ventricle, while aortic diastolic pressure affects the perfusion of the coronary arteries. Aortic pressure can reflect the blood perfusion status during the resuscitation process. The restoration of spontaneous circulation after cardiopulmonary resuscitation is closely related to maintaining the level of invasive aortic pressure. Aortic pressure is an important physiological index reflecting the heart's pumping function and the state of the arterial system. For example, during cardiopulmonary resuscitation, aortic pressure data can accurately reflect the implementation quality of cardiopulmonary resuscitation.
[0004] Existing aortic pressure measurement methods are mainly divided into two categories: invasive measurement and non-invasive measurement. Invasive blood pressure measurement methods are considered the gold standard for blood pressure measurement and can provide continuous and high-precision blood pressure data. For example, the arterial catheterization method directly inserts a catheter into an artery to monitor the patient's arterial blood pressure and blood gas changes in real time. Due to risks such as nerve damage, bleeding, and thrombosis, invasive measurement methods are generally only used in clinical settings and are not suitable for daily monitoring.
[0005] Non-invasive measurement usually estimates aortic pressure based on sensors or imaging technology, which has higher convenience and comfort and is particularly suitable for telemedicine and long-term monitoring. Among them, Photoplethysmography (PPG) is a typical non-contact technology widely used in wearable devices such as smart watches and health monitoring bracelets.
[0006] PPG indirectly estimates aortic pressure data by monitoring the optical changes caused by blood flow on the skin surface. This technology detects the changes in the pulse wave in the blood vessel through an LED and a photodiode, and estimates blood pressure by combining parameters such as pulse wave propagation time or pulse arrival time. However, the accuracy and stability of the PPG technology are easily affected by factors such as skin pigmentation and body movement.
[0007] Based on PPG, Imaging Photoplethysmography (iPPG) technology uses a camera to record minute color changes in the skin to estimate blood pressure. Different from traditional PPG, iPPG does not require direct contact with the skin, providing a more comfortable experience during blood pressure monitoring. It can not only achieve non-contact blood pressure measurement but also has good privacy protection functions, making it suitable for application scenarios such as health management and telemedicine.
[0008] However, iPPG relies on deep learning methods for measurement and is limited by different deep learning algorithms, resulting in unsatisfactory accuracy. For example:
[0009] The method based on 2DCNN can only extract the spatial features of the image and cannot effectively capture time series information;
[0010] The method based on recurrent neural networks is mainly used for time series modeling but relies on manually extracted PPG signals and is difficult to adapt to individual physiological differences and environmental changes;
[0011] The method based on 3DCNN can extract spatio-temporal information simultaneously, but has a high computational complexity and is prone to overfitting;
[0012] Although the U-Net structure performs well in medical image segmentation, it is not ideal when applied to blood flow signal extraction.
[0013] In summary, the measured aortic pressure results obtained by iPPG are estimated values generated by algorithms. In addition to being affected by various algorithms themselves, they are still affected by factors such as lighting conditions and skin characteristics, resulting in unsatisfactory measurement accuracy. Summary of the Invention
[0014] To solve the technical problems existing in the above background technology, the present invention provides a non-contact aortic pressure measurement method, system, storage medium, and device. A series of preprocessing methods are used to minimize the influence of complex lighting and individual differences on data quality. By using 3DCNN-InvertU-Net in combination, with a fine design of multiple levels and multiple modules, it ensures the effective extraction of hemodynamic features from video data and realizes the accurate estimation of aortic pressure.
[0015] To achieve the above object, the present invention adopts the following technical solutions:
[0016] The first aspect of the present invention provides a non-contact aortic pressure measurement method, including the following steps:
[0017] Obtain video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain facial key point image data;
[0018] Extract the hemodynamic features from the facial key-point image data. Use the invasive aortic pressure signal collected at the corresponding time point as the ground truth label. Through training, enable the aortic pressure estimation model to learn the mapping relationship between the hemodynamic features and the true aortic pressure. During training, use the 3DCNN module to capture the dynamic features of the video data in the temporal and spatial dimensions, and use the InvertU-Net module to fuse the detailed features and global features.
[0019] Use the trained aortic pressure estimation model to obtain the measured value of the aortic pressure of the subject being measured.
[0020] As a further implementation, the preprocessing includes face detection and recognition. Specifically: The acquired facial video data is based on the Swin Transformer structure to extract feature maps of different scales, realize hierarchical modeling of face features, and obtain the facial ROI. During this process, through a multi-level channel attention mechanism, fuse the features through different levels and optimize feature selection.
[0021] As a further implementation, the preprocessing also includes light condition judgment. Specifically: According to the light intensity, judge whether the current ambient light condition is suitable for video acquisition. When the light intensity does not meet the set conditions, call the fill light device.
[0022] As a further implementation, the preprocessing also includes key area detection. Specifically: Based on the Transformer key-point coordinate regression method, use the obtained facial ROI to determine the key points. During this process, adopt an adaptive ROI selection strategy, calculate the brightness mean and contrast of the ROI, and dynamically adjust the position and shape of the ROI to ensure that the quality of the selected ROI meets the requirements. If the quality of a single ROI does not meet the standard, fuse multiple ROIs on the premise of not introducing additional noise.
[0023] As a further implementation, the preprocessing also includes light normalization. Specifically: Obtain the local brightness and contrast of each pixel in the facial key-point image data, and enhance the contrast or suppress the excessive contrast by adjusting the Gamma value corresponding to the pixel.
[0024] As a further implementation, during the training of the aortic pressure estimation model, input the video data with ROI regions and key points, and the invasive aortic pressure data collected at the corresponding time points, and use the 3DCNN module to capture the dynamic features of the video data in the temporal and spatial dimensions.
[0025] As a further implementation, during the training of the aortic pressure estimation model, the obtained dynamic features are used by the InvertU-Net module to fuse the detailed and global features. The output of the InvertU-Net module adopts a multi-scale weighted fusion mechanism to adaptively adjust the weights of features at different scales, and the training is completed in combination with batch normalization, ReLU activation, and the loss function.
[0026] The second aspect of the present invention provides a system required to implement the above method, including:
[0027] A video acquisition and preprocessing module, configured to: obtain video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain facial key point image data;
[0028] An aortic pressure estimation module, configured to: extract hemodynamic features from the facial key point image data, use the invasive aortic pressure signal collected at the corresponding time point as the ground truth label, and through training, enable the aortic pressure estimation model to learn the mapping relationship between the hemodynamic features and the true aortic pressure; during training, use the 3DCNN module to capture the dynamic features of the video data in the time and space dimensions, and use the InvertU-Net module to fuse the detailed features and the global features;
[0029] A result output module, configured to: use the trained aortic pressure estimation model to obtain the aortic pressure measurement value of the measured subject.
[0030] The third aspect of the present invention provides a computer-readable storage medium.
[0031] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the non-contact aortic pressure measurement method as described above.
[0032] The fourth aspect of the present invention provides a computer device.
[0033] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the non-contact aortic pressure measurement method as described above.
[0034] Compared with the prior art, the above one or more technical solutions have the following beneficial effects:
[0035] 1. A camera can be used to capture the minute color changes in the facial skin of a subject caused by blood pulsation. The hemodynamic features in the video data are automatically extracted through a deep learning model and directly mapped to the aortic pressure to achieve non-contact real-time estimation of the aortic pressure. Compared with traditional single optimization schemes, multiple preprocessing methods are used in combination to minimize the impact of complex lighting and individual differences on data quality as much as possible. Through the combined use of a 3DCNN module and an InvertU-Net module, with a fine design of multiple levels and multiple modules, it is ensured that hemodynamic features can be effectively extracted from the video data. Then, taking the invasive aortic pressure signal collected at the corresponding time point as a label, the model is helped to learn the mapping relationship between hemodynamic features and the real aortic pressure, and finally a measurement result with relatively more ideal accuracy is obtained. Since it is a non-contact real-time monitoring of aortic pressure, it is particularly suitable for telemedicine, intelligent health monitoring, and hospital automation monitoring systems, and has good application adaptability and promotion value.
[0036] 2. In terms of technical principles, traditional methods usually rely on indirect physiological signals (such as pulse wave transit time or photoplethysmogram), and estimate aortic pressure after multiple steps of conversion. The information loss and error accumulation are obvious during the process. While this solution directly uses a deep learning model to achieve an end-to-end mapping from video data to aortic pressure, eliminating the intermediate links, greatly reducing information loss and error accumulation, and thus improving the estimation accuracy.
[0037] 3. In terms of preprocessing methods, traditional schemes often only use a single preprocessing technique and are difficult to effectively eliminate signal interference caused by complex lighting environments and individual differences. This solution, however, uses multiple preprocessing methods in combination, including multiple technical means such as image enhancement, dynamic color compensation, adaptive filtering, and personalized standardization, effectively improving data quality and significantly reducing the interference brought by complex lighting conditions and individual differences.
[0038] 4. The 3DCNN module and the InvertU-Net module are integrated for a joint design of multiple scales and multiple modules. Among them, the 3DCNN module realizes efficient modeling of the fine-grained dynamic changes between frames in the video sequence by applying three-dimensional convolution operations in both the spatial and temporal dimensions, and can fully capture the continuous and weakly-amplified spatio-temporal features caused by blood pulsation. The InvertU-Net module conducts a reverse design of the structure based on the traditional U-Net structure. First, it extracts high-resolution local spatial details through successive upsampling, and then integrates deep features and compresses features through successive downsampling. This "upsampling first, then downsampling" structural design enables the model to focus on local features in the image at the shallow stage, enhancing the sensitivity to weak spatial features; the initial upsampling process helps to introduce the overall structural information within a larger receptive field, facilitating the model to capture the global morphology of the image and the periodic patterns of physiological signals, thereby reducing the risk of key physiological features being ignored or lost in the early feature extraction stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0040] Figure 1 is a schematic diagram of the non-contact aortic pressure measurement process provided by one or more embodiments of the present invention;
[0041] Figure 2 is a schematic diagram of the overall process of the non-contact aortic pressure measurement process provided by one or more embodiments of the present invention;
[0042] Figure 3 is a schematic diagram of video acquisition and preprocessing during non-contact aortic pressure measurement provided by one or more embodiments of the present invention;
[0043] Figure 4 is a schematic diagram of the architecture of the 3DCNN-InvertU-Net algorithm during non-contact aortic pressure measurement provided by one or more embodiments of the present invention;
[0044] Figure 5 is a schematic diagram of the architecture of the InvertU-Net algorithm during non-contact aortic pressure measurement provided by one or more embodiments of the present invention;
[0045] Figure 6 is a schematic diagram of a simulation experiment device for verifying the non-contact aortic pressure measurement process provided by one or more embodiments of the present invention.
[0046] In the figure: 1 is an invasive aortic pressure measurement system, 2 is a non-contact aortic pressure measurement system, and 3 is a human simulation model. Detailed implementation manners
[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0048] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0049] It should be noted that the terms used in the following embodiments are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0050] Term explanation:
[0051] 2DCNN, two-dimensional convolutional neural network, which divides the temporal signal (such as the RGB channels of facial video frames) into independent 2D image segments (such as each frame or short-term stack), captures spatial features (such as small color changes in the skin), but ignores temporal continuity.
[0052] 3DCNN, regards the video segment as a spatio-temporal cube (height × width × time depth), and captures both spatial (facial area) and temporal (pulse wave propagation) features
[0053] U-Net structure, the original U-Net is designed for medical image segmentation and is often used for signal enhancement (such as denoising) and feature localization (such as segmenting the facial ROI area) in blood pressure measurement.
[0054] Embodiment 1:
[0055] The first technical problem to be solved in this embodiment is the problem of efficient and accurate video acquisition and video preprocessing.
[0056] In the prior art, video data often contains a large amount of redundant information and noise. If not effectively preprocessed, it is very likely to affect the feature extraction effect of the subsequent deep learning model, resulting in inaccurate aortic pressure estimation results. At the same time, if the ROI area is selected improperly or the illumination is uneven, the extraction quality of physiological signals will be severely reduced.
[0057] Therefore, in the video acquisition part of this embodiment, a face detection method based on SwinFace is adopted, and multi-scale key features are extracted through the SwinTransformer structure to ensure the accurate recognition of the face area and the stable positioning of the ROI.
[0058] To further improve the signal quality, an adaptive ROI selection strategy is used to dynamically adjust the ROI position based on the brightness mean and contrast to ensure the best signal quality in the selected area. In view of the differences in facial features of different individuals, the ROI is uniformly processed through affine transformation to reduce the impact of individual differences and ensure the standardization and consistency of the model input.
[0059] In addition, to address the impact of illumination changes on signal extraction, an adaptive gamma correction method is used to dynamically adjust the image brightness to ensure the stability and consistency of input data under different illumination conditions, thereby effectively improving the stability of blood flow signal extraction and the accuracy of subsequent model learning.
[0060] The second technical problem to be solved by this embodiment is the mapping learning problem between the contactless video signal and the invasive aortic pressure data.
[0061] Existing non-contact blood pressure estimation technology has problems such as insufficient signal extraction accuracy and large differences from real invasive blood pressure waveforms, making it difficult to meet clinical application needs. The main reason is that there are significant differences in data dimensions and feature structures between video signals and invasive aortic pressure data, and the existing model cannot accurately capture the spatiotemporal dynamic change characteristics of blood flow signals, resulting in low estimation accuracy.
[0062] To address this problem, this embodiment proposes a deep learning model that integrates 3DCNN and InvertU-Net. The 3DCNN module in the model can extract the dynamic features of blood flow signals in both time and space dimensions, accurately capturing the tiny changes in blood flow in video data; the InvertU-Net module in the model ensures the integrity of feature recovery through multi-scale feature extraction, jump connection and deconvolution mechanism, avoiding the loss of key detail information during the compression process.
[0063] In addition, invasive aortic pressure data is used as a supervisory signal during model training, and the mean square error loss function is used to guide model optimization, ensuring that the model can accurately learn the mapping relationship between video signals and invasive aortic pressure, significantly improving the accuracy and stability of aortic pressure estimation.
[0064] The third technical problem to be solved in this embodiment is the stability and generalization improvement of the end-to-end model structure. When facing different individuals, complex environments or dynamic scenes, traditional deep learning models are prone to overfitting, unstable estimation or insufficient generalization, which seriously affects the reliability of the model and the feasibility of clinical application.
[0065] Therefore, in this embodiment, a multi-level architecture that integrates 3DCNN and InvertU-Net is adopted in the model structure to ensure the comprehensive extraction of spatio-temporal features and the complete restoration of information. Through a multi-scale weighted fusion mechanism, the weight distribution of features at different scales is dynamically adjusted to ensure that the model can accurately identify complex physiological features and improve the generalization ability and robustness.
[0066] In addition, the Dropout mechanism is introduced during the training process to prevent the model from overfitting. Batch normalization is combined to ensure stable gradient transmission, and the ReLU activation function is used to accelerate the model convergence process. Through the above design, the model can achieve stable and efficient aortic pressure estimation under diverse individuals and dynamic backgrounds, ensuring that the model has good stability and generalization ability and meets the application requirements in multiple scenarios.
[0067] As Figure 1 shown, the non-contact aortic pressure measurement method includes the following steps:
[0068] Obtain the video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain the facial key point image data;
[0069] Extract the hemodynamic features in the facial key point image data, and use the invasive aortic pressure signal collected at the corresponding time point as the ground truth label. Through training, the aortic pressure estimation model learns the mapping relationship between the hemodynamic features and the real aortic pressure; during training, the 3DCNN module is used to capture the dynamic features of the video data in the time and space dimensions, and the InvertU-Net module is used to fuse the detailed features and global features;
[0070] Use the trained aortic pressure estimation model to obtain the aortic pressure measurement value of the subject to be measured.
[0071] The non-contact aortic pressure measurement method proposed in this embodiment includes the following parts: video acquisition, video preprocessing, aortic pressure estimation, and result output and visualization.
[0072] Video acquisition. Obtain the video data of the subject's facial area to provide the original information for subsequent blood flow feature extraction. A high-frame-rate RGB camera can be used to ensure that minute blood flow change information can be captured.
[0073] Video preprocessing. Perform a series of optimizations on the collected video, including ROI extraction and lighting normalization.
[0074] Aortic pressure estimation. An end-to-end aortic pressure estimation is performed using 3DCNN-InvertU-Net. This model utilizes 3DCNN to capture spatio-temporal blood flow information and combines the efficient feature recovery mechanism of InvertU-Net to improve the accuracy of aortic pressure estimation.
[0075] Result output and visualization. The aortic pressure waveform is output and the results are visually displayed to achieve stable and accurate measurement of aortic pressure, and support data storage, remote transmission, and trend analysis.
[0076] The overall process is as Figure 2 shown. After the camera is calibrated, video data is acquired. The face is automatically detected based on the acquired video data. Data processing is performed when the face is detected, the ambient light is suitable, and the data quality is good, and the aortic pressure value is output; if the face is not detected, the previous step is returned to continue detecting the face from the face data; if the ambient light is not suitable, the fill light device is activated; if the data quality does not meet the requirements, the camera is recalibrated and the video data is re-acquired.
[0077] The specific process of video acquisition and preprocessing is as Figure 3 shown. The video information of the subject's face is real-time acquired through a high-sensitivity camera to ensure that high-definition, continuous, and detailed video data is captured. This camera has high frame rate, high resolution, and high sensitivity, and can effectively capture the subtle blood flow change information in the facial area to ensure the accuracy of data acquisition.
[0078] The obtained video data is subjected to face detection and recognition using the SwinFace method. This method is based on the SwinTransformer structure, which can gradually extract key features at different scales, improve the recognition ability of the model, and ensure that the ROI area can accurately align with the facial area. SwinFace uses Swin Transformer as the feature extraction backbone network. First, the input 112×112×3 face image is divided into non-overlapping image patches of size 2×2 and projected into a 96-dimensional feature space through a linear embedding layer. Then, the image patch merging layer and the Swin Transformer block are alternately operated to gradually extract the feature maps at different scales, and finally four feature maps FM1, FM2, FM3, and FM4 of different scales are generated, thus realizing the hierarchical modeling of face features.
[0079] As Figure 3 shown, to improve the recognition accuracy of the model for key regions, SwinFace combines a multi-level channel attention (MLCA) mechanism. Among them, multi-level feature fusion enhances the model's perception ability through the integration of different-level features, and the channel attention mechanism further optimizes feature selection. The optimization objective of the MLCA mechanism can be expressed as: Among them, F i represents feature maps at different levels, and ω i is the weight assigned for channel attention. In addition, to improve the inter-class discrimination, SwinFace adopts the CosFace loss function for face recognition, and its optimization objective is defined as:
[0080]
[0081] Among them, represents the weight matrix of the last fully connected layer, n is the number of identity categories, represents the j-th column of the weight matrix W, represents the deep feature of the i-th sample belonging to the y i category, θ j is the weight W j and the feature x i the included angle between them, the feature x i is L2-normalized and rescaled to s, m is the CosFace cosine margin penalty term (set to s = 64, m = 0.4 in the experiment), and NR is the number of samples with identity labels in each training batch. Since both W j and x i are L2-normalized, and the feature vector x i is further rescaled to s, so the inner product of the two is simplified to cosine similarity, as shown in Equation (2).
[0082] CosFace makes the feature distributions of different classes more compact by imposing an angular margin penalty before Softmax, improving the robustness of face recognition.
[0083] Secondly, the environmental light intensity is monitored in real time to determine whether the current environmental light conditions are suitable for video acquisition. By equipping with a high-precision light sensor, the intensity of the environmental light can be detected in real time and compared with the preset appropriate light range in real time. Once it is detected that the environmental light is insufficient or inappropriate, the system will automatically activate the supplementary light device to perform appropriate light compensation on the environment. The specific implementation method is to control the LED array light source through an embedded controller and accurately adjust the supplementary light brightness and color temperature with an adaptive algorithm, so that the environmental light quickly reaches the optimal state for video acquisition, thereby ensuring that the collected video images have high clarity and stability.
[0084] As Figure 3 shown, during video preprocessing, optimization processes such as ROI extraction and light normalization are performed on the collected video data to reduce the influence of the external environment on the measurement accuracy. After the environmental light is suitable or reaches the suitable state after supplementary light, Transformer keypoint detection is performed.
[0085] As Figure 3As shown in the figure, in this embodiment, a key point coordinate regression method based on Transformer is adopted for key point detection, and this method directly regresses the key point coordinates. At this stage, the model is based on the face ROI region that has been recognized and extracted by SwinFace to further perform key point detection. The Transformer model effectively captures the global dependencies in the ROI region through its built-in self-attention mechanism, accurately recognizes and locates the key point coordinates in the forehead and cheek regions, ensuring the recognition accuracy and stability of the key regions.
[0086] Based on the key point detection results, an adaptive ROI selection strategy is adopted. By calculating the brightness mean and contrast of the ROI, the position and shape of the ROI are dynamically adjusted to ensure that the selected ROI has good signal quality. If the quality of a single ROI does not meet the standard, on the premise of ensuring that no additional noise is introduced, an attempt will be made to fuse multiple ROIs to improve signal stability, and the signal quality will be re-evaluated after fusion to ensure the optimization effect.
[0087] In addition, to reduce the impact of individual physiological differences on signal extraction, an affine transformation is used to standardize the shape of the ROI.
[0088]
[0089] Among them, (x, y) represents the original coordinate point, and (x′, y′) represents the new coordinate after the affine transformation. The matrix is a linear transformation matrix that controls the rotation, scaling, and shearing of the image. Among them, a and d respectively control the scaling in the x-axis and y-axis directions, and b and c control the shearing or rotation effect. The vector is a translation vector, representing the translation distances on the x-axis and y-axis respectively.
[0090] In video preprocessing, illumination normalization is a key step to improve the accuracy of physiological signal extraction. Due to the large differences in illumination conditions in different shooting environments, unnormalized images may cause signal extraction distortion in the ROI region, affecting the recognition and estimation accuracy of the model.
[0091] For this reason, this embodiment proposes an adaptive Gamma correction method based on local brightness and contrast to perform illumination normalization on the ROI region to ensure the stability and consistency of the input data. Adaptive Gamma correction enhances the balance of image brightness and the performance of image details by dynamically adjusting the Gamma value. Its calculation formula is:
[0092] γ(x,y)=1 - c1·(μ(x,y) - 0.5) + c2·(C(x,y) - C ref ) (4)
[0093] Among them, γ(x,y) represents the adaptive Gamma value of the pixel position (x,y), μ(x,y) is the local average luminance centered on this pixel, C(x,y) is the local contrast, and C ref is the reference contrast, and c1 and c2 are the adjustment coefficients for luminance and contrast.
[0094] The above realizes two aspects of adaptive adjustment: when the local luminance is dim (μ(x,y) < 0.5), the Gamma value is less than 1, and the image is automatically brightened; when the luminance is high (μ(x,y) > 0.5), the Gamma value is greater than 1, and the image is automatically darkened. At the same time, when the local contrast is low, the Gamma value will be adjusted accordingly to enhance the contrast. If the contrast is too high, the Gamma value is adjusted to suppress the excessive contrast, ensuring that the image is more natural and balanced. Based on the above adaptive Gamma value, the image correction formula is:
[0095] I′(x,y) = I(x,y) γ(x,y) (5)
[0096] where I(x,y) represents the original pixel value, and I′(x,y) is the corrected pixel value. Each pixel is dynamically adjusted according to its corresponding Gamma value, thereby realizing the optimal luminance and contrast correction of the local area. This method can not only effectively improve the detail performance of the image, avoid the loss of details caused by too high or too low luminance, but also enhance the sense of hierarchy and three-dimensionality of the image. By adjusting the parameters c1 and c2, the adaptive adjustment of luminance and contrast can be flexibly controlled to meet the image processing requirements in different scenarios.
[0097] In this solution, adaptive Gamma correction, as an indispensable part of the process, ensures the consistency and reliability of signal extraction in different lighting environments by precisely adjusting the image details and contrast, further enhancing the adaptability and accuracy of this solution in multiple scenarios.
[0098] This solution organically integrates SwinFace with Transformer keypoint detection, adaptive ROI selection, affine transformation normalization, and adaptive Gamma correction, and designs a multi-link, highly robust dynamic optimization process, significantly improving the stability and accuracy of aortic pressure signal estimation and filling the gap in the existing technology.
[0099] During aortic pressure estimation, this solution proposes a deep learning model that fuses the 3DCNN and InvertU-Net architectures, aiming to achieve high-precision dynamic estimation of aortic pressure based on non-contact video data. Through the fine design of multiple levels and multiple modules, this model ensures the effective extraction of hemodynamic features from video data and realizes the accurate estimation of aortic pressure. The specific architecture of the model is as Figure 4 shown.
[0100] The model consists of an input layer, a feature extraction layer (including feature fusion), a fully connected layer, and an output layer.
[0101] The input layer receives video data of the ROI region and invasive aortic pressure signals. The video data is collected by a high-frame-rate optical sensor, and video frames of regions rich in blood flow information such as the face are selected to ensure the high quality and stability of signal acquisition. At the same time, invasive aortic pressure data at the corresponding time points is synchronously collected as the ground truth label to ensure that the model learns the precise mapping relationship between non-contact video features and real aortic pressure during the training process.
[0102] The feature extraction layer integrates the feature extraction structures of 3DCNN and InvertU-Net, which helps to achieve an accurate non-linear mapping between the input video and the real aortic pressure. This task depends on the joint modeling of temporal dynamic features and spatial detail features in the video sequence. 3DCNN can effectively capture the continuous minute changes between frames caused by physiological rhythms such as blood pulsation by applying three-dimensional convolution operations synchronously in the temporal and spatial dimensions, and has strong temporal modeling capabilities. InvertU-Net, under the structural design of "upsampling first and then downsampling", enhances the early perception of local image details and the integration ability of multi-scale features, and is especially good at extracting weak signals with high resolution from spatial distributions. The combination of the two can form a collaborative modeling mechanism in the spatio-temporal dimension, which not only improves the ability to capture continuous physiological features but also enhances the response to local texture and structural changes, significantly improving the integrity and discriminability of feature expression, and providing a solid foundation for establishing an accurate mapping relationship between the input and target physiological parameters.
[0103] The 3DCNN module consists of 3 convolutional layers, 3 max pooling layers, and 1 fully connected layer, and has the functions of three-dimensional convolution, spatial dimension pooling, and global feature fusion, and can effectively capture the dynamic features of video data in the temporal and spatial dimensions.
[0104] 3DCNN-InvertU-Net constructs a deep neural network framework that integrates spatio-temporal joint modeling and regression functions. By embedding the 3DCNN module into the upsampling and downsampling paths of InvertU-Net, the model has the collaborative ability of spatial structure modeling and time series modeling. Different from the traditional U-Net that adopts the encoding-decoding structure of "downsampling first and then upsampling", InvertU-Net innovatively adopts the structural process of "upsampling first and then downsampling", so that the model focuses on high-resolution image details at the early stage of feature extraction, thus enhancing the perception ability of tiny and local features. At the same time, with the help of the skip connection mechanism and the progressive downsampling path, the model realizes the integration and compressed expression of multi-scale deep features while maintaining spatial details.
[0105] On this basis, 3D convolutional operations are adopted at all stages of the entire network from upsampling to downsampling, modeling simultaneously in the spatial and temporal dimensions to achieve collaborative perception of the inter-frame dynamic changes and spatial structure information. This design not only breaks through the modeling paradigm of traditional 2D convolutional models but also enhances the model's ability to extract continuous and weak spatio-temporal features caused by blood pulsation. 3DCNN and InvertU-Net are highly collaborative in structure and complementary in feature expression, thereby accurately learning the non-linear mapping relationship between hemodynamic features and the true aortic pressure, and finally realizing the end-to-end estimation of aortic pressure through the regression layer.
[0106] To achieve the effective fusion of these two modules, multiple key technical difficulties need to be overcome:
[0107] First, in the face of the high computational load and memory overhead caused by 3DCNN, a streamlined convolutional layer design and channel control strategy are adopted to ensure good trainability while maintaining the model's expressive ability.
[0108] Second, to solve the inconsistency problems in feature dimension, time step, and channel structure between the 3DCNN and InvertU-Net modules, fine-tuning is carried out on the time dimension alignment, channel mapping, and skip connection methods of the network to ensure the effective transmission and fusion of information within the structure.
[0109] In addition, in the design of the regression layer, a linear mapping structure is selected to avoid the bias caused by non-linear transformation, enabling the model to stably output continuous aortic pressure estimation values. During the training process, synchronous invasive aortic pressure signals are introduced as high-precision supervision to guide the model to learn the mapping relationship between video data and true blood pressure, achieving both the accuracy and generalization ability of non-invasive blood pressure estimation.
[0110] As Figure 5 shown, the InvertU-Net module includes 4 expansion path modules and 4 contraction path modules, and through multi-scale feature extraction and skip connections, the effective fusion of detailed and global features is achieved. The expansion path enhances spatial features through bilinear interpolation and convolutional blocks, the contraction path compresses spatial features through max pooling and convolutional blocks, and the bottleneck layer and skip connections ensure the efficient transmission of information, improving the adaptability and robustness of the model to different individuals and environments.
[0111] The feature fusion part receives the output of the InvertU-Net module, adopts a multi-scale weighted fusion mechanism, adaptively adjusts the weights of different-scale features, combines batch normalization and ReLU activation to stabilize the feature distribution, and improves the generalization ability of the model.
[0112] The fully connected layer further integrates and optimizes the fused features, adopts the Dropout mechanism to prevent overfitting, batch normalization to ensure stable gradients, and ReLU activation to accelerate training convergence. The output layer uses a single-output linear regression structure to directly estimate the aortic pressure waveform at the corresponding time point, avoiding the estimation bias caused by non-linear transformation and ensuring the accuracy of aortic pressure estimation.
[0113] During the training phase, the mean squared error is used as the loss function, and the Adam optimizer is used for efficient optimization to ensure the minimization of the estimation error. The invasive aortic pressure data serves as the ground truth label, which is compared with the model estimation value to guide the model in optimizing feature extraction and the regression process, ensuring that the model accurately learns the mapping relationship between video features and aortic pressure.
[0114] After training, the model can independently complete the dynamic estimation of aortic pressure based on video data without relying on invasive aortic pressure. Model evaluation uses metrics such as mean squared error, root mean squared error, mean absolute percentage error, mean error, and correlation coefficient to comprehensively evaluate the estimation accuracy of the model. Referring to the AAMI (American Association for the Advancement of Medical Instrumentation) standard, it is ensured that the model meets the internationally accepted accuracy requirements. The formulas for each metric are as follows:
[0115]
[0116] Among them, \(O\) represents the aortic pressure result estimated from the proposed 3DCNN-InvertU-Net, \(T\) represents the dataset of invasive aortic pressure; \(n\) represents the number of datasets, and \(\overline{O}\) represents the average value of the output set \(O\). According to the above formulas, the lower the MSE, RMSE, MAPE, and ME, the higher the accuracy of the model. In addition, \(R\) represents the correlation, and the stronger the correlation, the closer the correlation coefficient is to ±1.
[0117] The result output and visualization part, as the final link of the system, is responsible for formatting, metric extraction, storage, display, and transmission of the aortic pressure estimation results output by the aortic pressure estimation module, ensuring that the data has good usability and clinical application value. Specifically:
[0118] The model estimation results are standardized to ensure consistent data structure and format, and the data is sorted in chronological order, with additional label information such as timestamps to ensure the traceability of the data. In terms of storage, the module supports local storage and multi-format data export, such as CSV, Excel, JSON, XML, etc., facilitating subsequent analysis and archiving to ensure efficient data management.
[0119] For easy data viewing, the result output and visualization part provides diverse visualization functions, including real-time waveform display, key metric charts, trend comparison charts, and anomaly alarm mechanisms, helping users and doctors quickly grasp the aortic pressure status.
[0120] In addition, it also supports the automatic generation of standardized aortic pressure monitoring reports, covering key indicators, trend analysis, and annotation of abnormal data. The reports support multi-format export for easy printing, archiving, and electronic transmission, and have a result comparison function to assist doctors in diagnosis and decision-making.
[0121] Verification. The method proposed in this solution is verified by building an experimental device, which is implemented using the simulation experimental device shown in Figure 6 . The invasive aortic pressure measurement system 1 therein is a prior art, the non-contact aortic pressure measurement system 2 is this solution, and the subject of aortic pressure measurement is represented by a human simulation model 3 to simulate the actual measurement scenario.
[0122] Compared with the traditional single optimization solution, this solution forms a highly integrated and dynamically optimized signal processing process by organically integrating a variety of advanced technologies. Especially in an environment with complex lighting and obvious individual differences, the adaptive optimization strategy of this solution significantly improves the accuracy and stability of signal extraction. The multi-module deep optimization process is applied to the real-time monitoring of non-contact aortic pressure for the first time, which is particularly suitable for telemedicine, intelligent health monitoring, and hospital automation monitoring systems, and has good application adaptability and promotion value. It not only overcomes the limitations of traditional contact measurement methods but also significantly improves the accuracy and comfort of monitoring.
[0123] Embodiment 2:
[0124] A non-contact aortic pressure measurement system, comprising:
[0125] A video acquisition and preprocessing module, configured to: obtain video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain facial key point image data;
[0126] An aortic pressure estimation module, configured to: extract hemodynamic features from the facial key point image data, use the invasive aortic pressure signal collected at the corresponding time point as the true label, and through training, enable the aortic pressure estimation model to learn the mapping relationship between the hemodynamic features and the true aortic pressure; during training, use the 3DCNN module to capture the dynamic features of the video data in the time and space dimensions, and use the InvertU-Net module to fuse the detailed features and global features;
[0127] A result output module, configured to: use the trained aortic pressure estimation model to obtain the aortic pressure measurement value of the measured subject.
[0128] A series of preprocessing methods are used to minimize the impact of complex lighting and individual differences on data quality as much as possible. By using 3DCNN-InvertU-Net in combination, with a fine design of multiple levels and multiple modules, it is ensured that hemodynamic features can be effectively extracted from video data, and accurate estimation of aortic pressure can be achieved.
[0129] Embodiment 3:
[0130] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the non-contact aortic pressure measurement method described in Embodiment 1 above.
[0131] Embodiment 4:
[0132] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the non-contact aortic pressure measurement method described in Embodiment 1 above.
[0133] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0134] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Non-contact aortic pressure measurement method, characterized in that: It includes the following steps: Obtain the video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain the facial key point image data; Extract the hemodynamic features in the facial key point image data, use the invasive aortic pressure signal collected at the corresponding time point as the ground truth label, and through training, enable the aortic pressure estimation model to learn the mapping relationship between the hemodynamic features and the real aortic pressure; During training, use the 3DCNN module to capture the dynamic features of the video data in the temporal and spatial dimensions, and use the InvertU-Net module to fuse the detailed features and global features; Use the trained aortic pressure estimation model to obtain the aortic pressure measurement value of the measured subject.
2. The non-contact aortic pressure measurement method according to claim 1, wherein The face detection and recognition is specifically as follows: The obtained facial video data is based on the Swin Transformer structure, extracts feature maps of different scales, realizes hierarchical modeling of face features, and obtains the facial ROI. During this period, through a multi-level channel attention mechanism, the features passed through different levels are fused and the feature selection is optimized.
3. The non-contact aortic pressure measurement method according to claim 1, characterized in that, The lighting condition judgment is specifically as follows: According to the light intensity, judge whether the current ambient lighting condition is suitable for video acquisition. When the light intensity does not meet the set conditions, call the fill light device.
4. The non-contact aortic pressure measurement method according to claim 1, wherein, The key area detection is specifically as follows: Based on the Transformer key point coordinate regression method, use the obtained facial ROI to determine the key points. During this period, adopt an adaptive ROI selection strategy, and dynamically adjust the position and shape of the ROI by calculating the brightness mean and contrast of the ROI to ensure that the quality of the selected ROI meets the requirements; If the quality of a single ROI does not meet the standard, fuse multiple ROIs without introducing additional noise.
5. The non-contact aortic pressure measurement method according to claim 1, characterized in that, The lighting normalization processing is specifically as follows: Obtain the local brightness and contrast of each pixel in the facial key point image data, and enhance the contrast or suppress the overstrong contrast by adjusting the Gamma value corresponding to the pixel.
6. The non-contact aortic pressure measurement method according to claim 1, wherein During the training of the aortic pressure estimation model, input the video data with ROI regions and key points, and the invasive aortic pressure data collected at the corresponding time points, and use the 3DCNN module to capture the dynamic features of the video data in the temporal and spatial dimensions.
7. The non-contact aortic pressure measurement method according to claim 1, characterized in that, During the training of the aortic pressure estimation model, the obtained dynamic features are fused for details and global features through the InvertU-Net module. The output of the InvertU-Net module adopts a multi-scale weighted fusion mechanism to adaptively adjust the weights of features at different scales, and completes the training in combination with batch normalization, ReLU activation, and the loss function.
8. A non-contact aortic pressure measurement system, characterized in that, It includes: A video acquisition preprocessing module, configured to: Obtain the video data of the subject's face, and successively perform face detection and recognition, lighting condition judgment, key area detection, and lighting normalization processing to obtain the facial key point image data; An aortic pressure estimation module, configured to: extract hemodynamic features from facial key point image data, use the invasive aortic pressure signal collected at the corresponding time point as the ground truth label, and through training, enable the aortic pressure estimation model to learn the mapping relationship between hemodynamic features and the true aortic pressure; during training, use the 3DCNN module to capture the dynamic features of video data in the temporal and spatial dimensions, and use the InvertU-Net module to fuse the detailed features and global features; A result output module, configured to: use the trained aortic pressure estimation model to obtain the aortic pressure measurement value of the measured subject.
9. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the non-contact aortic pressure measurement method according to any one of claims 1-7 above are implemented.
10. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the non-contact aortic pressure measurement method according to any one of claims 1-7 are implemented.