Pulse condition waveform feature extraction and classification method based on vision and artificial intelligence

By combining multimodal sensors and deep learning technology, pulse features are extracted and classified, solving the problems of high dependence on feature engineering and lack of dynamic changes in existing technologies, and achieving more efficient pulse classification and diagnostic support.

CN121667660APending Publication Date: 2026-03-17QINHUANGDAO SHIDIAN TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511771571.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing machine learning-based pulse classification methods rely heavily on feature engineering, have limited generalization ability, and fail to fully explore the dynamic changes in pulse patterns, making them difficult to meet the needs of clinical applications.

Method used

A high-resolution photoplethysmography (PPG) sensor combined with specific optical imaging techniques is used to acquire multimodal pulse data. Pressure signals are simultaneously acquired by a pressure sensor. Convolutional neural networks (CNNs) are used to extract morphological and texture features. Spatiotemporal attention mechanisms and long short-term memory networks (LSTMs) are combined to capture dynamic changes. Feature fusion and classification are performed using a Transformer architecture classification model.

Benefits of technology

It enables a more comprehensive extraction of pulse characteristics, improves the accuracy and generalization ability of pulse classification, provides support for the objectification and digitization of TCM pulse diagnosis, and provides an accurate diagnostic auxiliary tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121667660A_ABST
    Figure CN121667660A_ABST
Patent Text Reader

Abstract

The invention discloses a pulse condition waveform feature extraction and classification method based on vision and artificial intelligence, and belongs to the technical field of medical diagnosis, and the method comprises the following steps: 1) pulse condition data collection; 2) data preprocessing; 3) feature extraction based on computer vision; 4) feature fusion and classification based on artificial intelligence; 5) evaluating and optimizing the model; by combining computer vision and a multi-modal signal processing technology, multi-dimensional features such as form, texture, time domain, frequency domain and dynamic change can be comprehensively extracted from pulse condition images and signals, essential features of pulse conditions are more accurately reflected, and comprehensive extraction of pulse condition features is realized; by constructing a deep learning model based on a Transform architecture and adopting an effective feature fusion and selection method, the accuracy and generalization ability of pulse condition classification are greatly improved, and powerful support is provided for objectification and digitization of traditional Chinese medicine pulse diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical diagnosis, and particularly relates to a pulse wave feature extraction and classification method based on vision and artificial intelligence. BACKGROUND

[0002] As a traditional diagnostic method, traditional Chinese medicine pulse diagnosis has a long history and rich clinical experience. Pulse can reflect the physiological and pathological state of the human body, and has important guiding significance for the diagnosis and treatment of diseases. Some sensor-based pulse collection devices have appeared, which can convert pulse signals into electrical signals and record them.

[0003] Chinese patent (CN202210042202.7) A pulse identification method, device, equipment and storage medium based on pulse characteristics. The method comprises: acquiring a to-be-measured pulse wave corresponding to a current pulse period; extracting features from the to-be-measured pulse wave corresponding to the current pulse period to obtain at least one to-be-measured pulse feature; using a pre-set pulse representation analysis strategy of opposite pulse, analyzing the at least one to-be-measured pulse feature, obtaining a first target feature score corresponding to a first pulse and a second target feature score corresponding to a second pulse, the first pulse and the second pulse are two pulses in the opposite pulse; obtaining the larger feature score in the first target feature score and the second target feature score, and determining the first pulse or the second pulse corresponding to the larger feature score as the current pulse corresponding to the current pulse period; determining the target pulse according to the current pulse corresponding to all current pulse periods in the target diagnosis time length.

[0004] At present, some existing pulse classification methods based on machine learning have great dependence on feature engineering, and have limited generalization ability when dealing with complex pulse, which is difficult to meet the needs of clinical application. In addition, the current research mostly ignores the dynamic changes of pulse, only analyzes the pulse signal at a certain moment, and fails to fully tap the physiological information contained in the changes of pulse over time. SUMMARY

[0005] The purpose of the present application is to solve the above-mentioned problems, and a pulse waveform feature extraction and classification method based on vision and artificial intelligence is proposed.

[0006] In order to achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows: a pulse waveform feature extraction and classification method based on vision and artificial intelligence, which comprises the following steps:

[0007] 1) Pulse data collection: High-resolution photoelectric plethysmogram (PPG) sensor combined with specific optical imaging technology is used to obtain video stream containing pulse information. By monitoring the reflection and absorption of light of different wavelengths on the surface of human skin, the subtle changes in blood vessel volume caused by heartbeats are captured, thus obtaining more abundant pulse information. Pressure sensors are used to synchronously collect pressure signals of the pulse, forming multi-modal pulse data.

[0008] 2) Data preprocessing: Denoising, grayscale, contrast enhancement and other image preprocessing operations are performed on the collected video stream to improve image quality and facilitate subsequent feature extraction. Digital filtering and wavelet transform techniques are used to remove high-frequency noise and baseline drift and perform signal smoothing on PPG signals and pressure signals to improve signal stability and accuracy.

[0009] 3) Feature extraction based on computer vision: Convolutional Neural Network (CNN) is used to analyze and extract features such as morphology and texture from preprocessed video images. For example, by constructing multiple convolutional layers and pooling layers, the morphological changes of blood vessels and the contour features of pulse waves in pulse images are learned. Temporal and spatial attention mechanisms and Long Short-Term Memory Network (LSTM) are combined to capture dynamic changes in pulse, such as the rising speed, falling speed, and periodic changes of pulse waves at different times.

[0010] 4) Feature fusion and classification based on artificial intelligence: Visual features extracted from video images are fused with time-domain and frequency-domain features of PPG signals and pressure signals. Feature selection algorithms such as ReliefF are used to select key features from the fused feature set, reducing data dimensionality and improving classification efficiency. A pulse classification model based on the Transformer architecture is constructed for pulse classification. Through training on a large-scale pulse dataset, the model's parameters are constantly optimized to improve its generalization ability and classification accuracy.

[0011] 5) Model evaluation and optimization: Cross-validation method and independent test dataset are used to evaluate the pulse classification model, calculating accuracy, recall rate, F1 value and other indicators to comprehensively evaluate the model's performance. Based on the evaluation results, the model structure and parameters are adjusted for optimization, such as increasing or decreasing the number of network layers, adjusting the size of convolution kernel, optimizing the weight distribution of attention mechanism, etc., to further improve the model's classification effect.

[0012] As a further description of the above technical solutions:

[0013] In the step 1), the PPG sensor has a working wavelength range of 500-1000 nm, the pressure sensor has a range of 0-50 kPa and an accuracy of 0.1 kPa, and the image acquisition device has a frame rate of no less than 60 fps and a resolution of no less than 1920x1080.

[0014] As a further description of the above technical solution:

[0015] In the step 2), the video image denoising adopts a Gaussian filtering algorithm, the Gaussian kernel size is (5, 5), the signal digital filtering adopts a low-pass filter with a cutoff frequency of 10 Hz, and the wavelet transform selects a db4 wavelet base with a decomposition layer number of 5.

[0016] As a further description of the above technical solution:

[0017] In the step 3), the CNN model includes multiple convolution layers and pooling layers, the convolution kernel size is 3x3 and 5x5, the pooling kernel size is 2x2, the step length is 2, the spatiotemporal attention mechanism module calculates the attention weight of the image in the spatial and temporal dimensions, and the LSTM network learns the time variation rule of the pulse image.

[0018] As a further description of the above technical solution:

[0019] In the step 4), the feature fusion adopts a splicing method, the feature selection uses a ReliefF algorithm, the iteration number is 100 times, the number of neighbor samples is 10, the Transformer classification model is composed of a multi-head self-attention layer, a feedforward neural network layer and an output layer, and the training adopts a cross-entropy loss function and an Adagrad optimizer.

[0020] As a further description of the above technical solution:

[0021] In the step 5), the evaluation indexes include accuracy, recall rate and F1 value, and the optimization is performed by increasing the number of samples, adjusting the model structure and hyperparameters, etc.

[0022] As described above, due to the adoption of the above technical solution, the present application has the following beneficial effects:

[0023] 1、In the present application, by combining computer vision and multi-modal signal processing technology, the multi-dimensional features such as morphology, texture, time domain, frequency domain and dynamic change can be comprehensively extracted from the pulse image and signal, and the essential features of the pulse can be more accurately reflected, so that the pulse features can be comprehensively extracted.

[0024] 2、In the present application, by constructing a deep learning model based on the Transformer architecture and adopting effective feature fusion and selection methods, the accuracy and generalization ability of pulse classification are greatly improved, which provides strong support for the objectification and digitization of traditional Chinese medicine pulse diagnosis.

[0025] 3、In the application, by using the space-time attention mechanism and the LSTM network, the dynamic change information of pulse condition over time is fully captured, the deficiency of ignoring the dynamic characteristics of pulse condition in the prior art is made up, and pulse condition analysis is more comprehensive and accurate.

[0026] 4、In the application, the device can be integrated into a portable pulse condition diagnosis device, and an objective and accurate pulse condition diagnosis auxiliary tool is provided for clinicians, and the device can also be applied to remote medical treatment, health monitoring and other fields, and has a wide application prospect. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 It is a flowchart of a pulse waveform feature extraction and classification method based on vision and artificial intelligence. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0029] Embodiment:

[0030] S01: Pulse condition data acquisition:

[0031] Device selection and installation:

[0032] A high-resolution PPG sensor is selected, which has a working wavelength range of 500-1000nm and can sensitively detect the light intensity change caused by the change in blood vessel volume on the skin surface. The PPG sensor is fixed at the Cunkou part of the patient's wrist to ensure that the sensor is in close contact with the skin to obtain stable PPG signals.

[0033] At the same time, a high-precision pressure sensor is installed for measuring the pressure change of the pulse. The pressure sensor has a range of 0-50kPa and an accuracy of 0.1kPa. The pressure sensor is placed in the same position as the PPG sensor and is fixed by a special clamp to ensure that the pressure sensor can accurately perceive the pressure fluctuation of the pulse.

[0034] A professional image acquisition device such as a high-speed camera is used, with a frame rate set to 60fps or higher and a resolution of no less than 1920x1080, for shooting videos of the Cunkou part of the wrist. The camera angle should be set to clearly capture the subtle changes on the skin surface when the pulse beats.

[0035] Data acquisition process:

[0036] Let the patient in a quiet, comfortable state, keep the arm natural relaxation, record the basic information of the patient before collecting data, including age, gender, health status, etc.

[0037] Start the data collection system, and start collecting PPG signals, pressure signals and video images, and collect data for 3-5 minutes to obtain enough pulse data for subsequent analysis. In the collection process, the quality of the signal is monitored in real time to ensure the stability and accuracy of the data. If abnormal signals are found, the collection is suspended in time, the equipment connection and patient status are checked, and the collection is restarted;

[0038] S02: Data preprocessing:

[0039] Video image preprocessing:

[0040] The collected video stream is imported into the computer and preprocessed using image processing libraries such as OpenCV. First, denoising is performed using a Gaussian filter algorithm with an appropriate Gaussian kernel size, such as (5, 5), to remove noise interference in the image and make the image smoother;

[0041] Then convert the color image to a grayscale image to simplify the subsequent image processing process, and then perform contrast enhancement by histogram equalization method to expand the gray dynamic range of the image and improve the clarity of the detailed information in the image;

[0042] Signal preprocessing:

[0043] For PPG signals and pressure signals, use the SciPy library in Python for digital filtering, design a low-pass filter with a cutoff frequency of 10 Hz to remove high-frequency noise components in the signal and make the signal smoother. At the same time, use wavelet transform for baseline correction, select an appropriate wavelet basis such as db4 wavelet basis, and set the decomposition layer to 5 layers to effectively remove the baseline drift in the signal and restore the true shape of the signal;

[0044] S03: Feature extraction based on computer vision:

[0045] Build CNN model:

[0046] Use TensorFlow or PyTorch deep learning framework to build a convolutional neural network model. The model structure includes multiple convolutional layers, pooling layers and fully connected layers. The convolutional layers use different sizes of convolutional kernels, such as 3x3 and 5x5, to extract different scale features of the image. The pooling layer uses the maximum pooling method with a pooling kernel size of 2x2 and a step of 2 to reduce the dimension of the feature map and reduce the amount of calculation;

[0047] In the initial stage of the model, the weight parameters are randomly initialized, and then pre-trained on a large-scale pulse image dataset containing images of normal pulse and various common pathological pulse, each sample is labeled by a professional Chinese medicine doctor. During the pre-training process, the cross-entropy loss function and Adam optimizer are used to adjust the parameters of the model, so that the model can initially learn the basic features of the pulse image.

[0048] Spacetime attention mechanism combined with LSTM:

[0049] Based on the CNN model, a spacetime attention mechanism module is introduced. This module calculates the attention weights of the image in the spatial and temporal dimensions, allowing the model to focus on key spatiotemporal regions and moments in the pulse image. For example, higher attention weights are given to the key stages of pulse wave rising and falling.

[0050] The output of the spacetime attention mechanism module is connected to the LSTM network. The LSTM network has a memory function and can learn the temporal variation of the pulse image. Through the hidden layer state transmission of the LSTM network, it can capture the dynamic change characteristics of the pulse, such as the periodic change and amplitude change of the pulse wave.

[0051] S04: Feature fusion and classification based on artificial intelligence:

[0052] Feature fusion:

[0053] Visual features of pulse images are extracted from the CNN-LSTM model, while time-domain features (such as mean, variance, peak time, etc.) and frequency-domain features (such as power spectral density, main frequency, etc.) are extracted from the pre-processed PPG signal and pressure signal.

[0054] These different types of features are fused to form a comprehensive feature vector. The fusion method can use a simple concatenation method to connect the visual feature vector, PPG signal feature vector, and pressure signal feature vector in order.

[0055] Feature selection:

[0056] ReliefF algorithm is used to select the fused features. ReliefF algorithm calculates the correlation between each feature and the class label, and assigns a weight to each feature. The number of iterations is set to 100, and the number of nearest neighbors is set to 10. According to the weight size, the top N features with higher weights are selected as the final feature subset. The value of N can be determined through experimental verification, generally between 50-100, to reduce the data dimension and improve the training efficiency and classification accuracy of the subsequent classification model.

[0057] Construct a Transformer classification model:

[0058] The pulse classification model is constructed based on the Transformer architecture, which mainly consists of a multi-head self-attention layer, a feedforward neural network layer, and an output layer. The multi-head self-attention layer calculates in parallel through multiple attention heads, which can better capture the complex relationships between features. The feedforward neural network layer further transforms and nonlinearly maps the output of the self-attention layer.

[0059] In the model training phase, the labeled pulse data set is used for training. The data set is divided into 70% training set, 15% validation set, and 15% test set. During training, the cross-entropy loss function is used to measure the difference between the model prediction results and the true labels. The Adagrad optimizer is used to adjust the model parameters to gradually reduce the loss function. The number of training rounds is set to 100-200 rounds. According to the performance of the validation set, the optimal model parameters are selected.

[0060] S05: Model evaluation and optimization:

[0061] Model evaluation:

[0062] The trained Transformer classification model is evaluated using the test data set, and the accuracy, recall rate, and F1 value of the model are calculated. The accuracy calculation formula is: accuracy = number of correctly classified samples / total number of samples; the recall rate calculation formula is: recall rate = number of correctly classified positive samples / actual number of positive samples; the F1 value calculation formula is: F1 = 2 × (accuracy × recall rate) / (accuracy + recall rate).

[0063] At the same time, the confusion matrix is drawn to intuitively show the classification of the model in different pulse categories, and to analyze which categories the model's incorrect classification is mainly concentrated in.

[0064] Model optimization:

[0065] According to the evaluation results, the model is optimized. If the accuracy of the model in some categories is low, the number of samples in these categories can be increased for retraining to improve the model's recognition ability for these categories.

[0066] Adjust the structure of the model, such as increasing or decreasing the number of Transformer layers, changing the number of heads of multi-head self-attention, and observing the changes in model performance to select the optimal model structure.

[0067] Optimize the hyperparameters of the model, such as adjusting the learning rate, weight decay coefficient, etc. Methods such as grid search or random search can be used to search for the optimal combination of hyperparameters within a certain range to further improve the classification performance of the model.

[0068] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for extracting and classifying pulse wave features based on vision and artificial intelligence, characterized in that: The method comprises the following steps: 1) Pulse condition data acquisition: using a high-resolution photoelectric plethysmography (PPG) sensor combined with specific optical imaging technology to obtain a video stream containing pulse condition information, and using a pressure sensor to synchronously collect the pressure signal of the pulse condition, forming multi-modal pulse condition data; 2) Data preprocessing: performing image preprocessing operations such as denoising, grayscale, and contrast enhancement on the collected video stream, and using digital filtering and wavelet transform techniques to remove high-frequency noise and baseline drift and perform signal smoothing processing on the PPG signal and the pressure signal; 3) Feature extraction based on computer vision: using a convolutional neural network (CNN) to analyze and extract morphological and textural features from the preprocessed video images, and combining a spatiotemporal attention mechanism and a long short-term memory network (LSTM) to capture the dynamic change characteristics of the pulse condition; 4) Feature fusion and classification based on artificial intelligence: fusing the visual features extracted from the video images with the time-domain and frequency-domain features of the PPG signal and the pressure signal, using a feature selection algorithm to select key features, and constructing a pulse condition classification model based on the Transformer architecture to classify the pulse condition; 5) Model evaluation and optimization: using cross-validation methods and independent test data sets to evaluate the pulse condition classification model, and adjusting the model structure and parameters for optimization according to the evaluation results. 2.The method of claim 1, wherein, In step 1), the PPG sensor has a working wavelength range of 500-1000 nm, the pressure sensor has a range of 0-50 kPa and an accuracy of 0.1 kPa, the image acquisition device has a frame rate of not less than 60 fps and a resolution of not less than 1920×1080. 3.The method of claim 1, wherein, In step 2), the video image denoising uses a Gaussian filter algorithm with a Gaussian kernel size of (5, 5), the signal digital filtering uses a low-pass filter with a cutoff frequency of 10 Hz, and the wavelet transform uses a db4 wavelet basis with 5 layers of decomposition. 4.The method of claim 1, wherein, In step 3), the CNN model includes multiple convolutional and pooling layers, the convolution kernel size is 3×3 and 5×5, the pooling kernel size is 2×2, the step size is 2, the spatiotemporal attention mechanism module calculates the attention weights of the image in the spatial and temporal dimensions, and the LSTM network learns the time variation law of the pulse image. 5.The method of claim 1, wherein, In step 4), the feature fusion uses a concatenation method, the feature selection uses the ReliefF algorithm with 100 iterations and 10 nearest neighbor samples, the Transformer classification model consists of a multi-head self-attention layer, a feedforward neural network layer, and an output layer, and the training uses a cross-entropy loss function and an Adagrad optimizer. 6.The method of claim 1, wherein, In step 5), the evaluation indicators include accuracy, recall, and F1 value, and the optimization is performed by increasing the number of samples, adjusting the model structure and hyperparameters, etc.

Citation Information

Patent Citations

  • Pulse condition recognition method, device and equipment based on pulse characteristics and storage medium

    CN114224297A