A pulse condition recognition method fusing fast and slow time reconstruction and multi-scale convolution
Patent Information
- Application Number
- CN202610917586.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-29
AI Technical Summary
近年来的基于深度学习的智能脉诊方法多将腕脉信号视为一维时间序列,或将信号处理成平均周期的单周期波形表示,无法有效刻画脉象信号在连续搏动过程中呈现的渐变趋势
[0104]本申请通过上述步骤,实现了通过FSMA-PDNet模型智能识别中医脉象,并与其他常用机器学习方法进行对比,对模型的准确性进行评估。
Smart Images

Figure CN122827633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of medicine, signal analysis and processing, and deep learning, specifically to a pulse recognition method based on the fusion of fast and slow temporal reconstruction and multi-scale convolution. Background Technology
[0002] In recent years, with the deep integration of artificial intelligence and sensing technology, the objectification and intelligentization of pulse diagnosis in Traditional Chinese Medicine (TCM) has entered a period of rapid development. Intelligent pulse diagnosis technology, by integrating flexible sensing, signal processing, and deep learning algorithms, aims to transform the traditional TCM experience of pulse diagnosis—characterized by "understanding in the mind but not on the fingertips"—into quantifiable and reproducible digital features. Significant progress has been made in areas such as pulse acquisition equipment development, wrist pulse signal parameter analysis, and disease identification model construction. From a national strategic perspective, the National Key Research and Development Program's "Modernization Research of Traditional Chinese Medicine" key project has listed "Digital Intelligence-Driven Improvement of TCM Four Diagnostic Methods and Differential Diagnosis Ability" as a key research direction. This aims to promote the transformation of TCM diagnosis and treatment from experience-based inheritance to a scientific paradigm by constructing high-quality clinical datasets and breaking through key technologies such as multimodal fusion and small-sample learning. This trend indicates that intelligent pulse diagnosis is not only an inevitable direction of technological evolution but has also become an important component of the national strategy for the inheritance and innovation of TCM.
[0003] From the perspective of signal processing and feature extraction techniques, traditional pulse diagnosis methods mainly focus on the time and frequency domain characteristics of wrist pulse signals, and then combine them with machine learning algorithms such as support vector machines and random forests to build pulse classification models. Recent deep learning-based intelligent pulse diagnosis methods often treat wrist pulse signals as a one-dimensional time series, or process the signal into a single-cycle waveform with an average period, failing to effectively depict the gradual changes in pulse signals during continuous pulsation. Therefore, to more accurately identify pulse types, it is essential to effectively utilize the waveform characteristics (intra-cycle variations) of each cycle of the wrist pulse signal, as well as the rhythmic variations across cycles (inter-cycle variations).
[0004] This invention addresses the aforementioned problems by proposing a pulse recognition method that integrates fast and slow temporal reconstruction with multi-scale convolution. A dual-branch pulse diagnosis network, FSMA-PDNet, is designed, primarily consisting of the MCMSDA network for processing one-dimensional wrist pulse signals, the FaSlTiRe algorithm for reconstructing the one-dimensional wrist pulse signal into a two-dimensional representation, the P2DInception network for processing the two-dimensional representation, and a feature fusion and multi-task classification network. This method fully utilizes the waveform features of each cycle in the one-dimensional representation of the wrist pulse signal and the spatiotemporal features of the two-dimensional representation. Summary of the Invention
[0005] The technical problem this application aims to solve is how to collaboratively utilize one-dimensional temporal features and two-dimensional reconstructed features in wrist pulse signals to simultaneously capture local details and cross-cycle evolution information of the pulse image. This application provides a pulse image recognition method that integrates fast and slow temporal reconstruction with multi-scale convolution to solve the above problem, and evaluates its performance.
[0006] This application achieves its goal through the following technical solution: This application provides a pulse recognition method that integrates fast and slow temporal reconstruction with multi-scale convolution, the method comprising:
[0007] Step 1: Select subjects and collect their clinical pulse data using a three-channel acquisition device based on pressure sensors;
[0008] Step 2: Traditional Chinese medicine experts mark the pulse of the test subjects and the corresponding pulse pattern;
[0009] Step 3: Use wavelet transform to remove high-frequency noise from the pulse signal and use cubic spline interpolation to remove the baseline of the pulse signal;
[0010] Step 4: Construct an intelligent pulse diagnosis network that integrates fast and slow temporal reconstruction with multi-scale convolution, comprising: an MCMSDA network for processing one-dimensional wrist pulse signals, a P2DInception network for representing one-dimensional wrist pulse signals in two dimensions, a feature fusion module, and a multi-task classification network; the preprocessed one-dimensional wrist pulse signals are input into the MCMSDA network and the P2DInception network respectively, and then the output data of the MCMSDA network and the P2DInception network are input into the feature fusion module, and then into the multi-task classification network;
[0011] Step 5: Input the preprocessed one-dimensional wrist pulse signal into the MCMSDA network to extract the multi-scale one-dimensional features of the wrist pulse signal;
[0012] Step 6: Reconstruct the preprocessed one-dimensional wrist pulse signal into a two-dimensional representation, and then extract the two-dimensional features;
[0013] Step 7: Fuse the one-dimensional and two-dimensional features, and use the feature fusion and multi-task classification network to complete the recognition of arterial images;
[0014] Step 8: Intelligently identify TCM pulse patterns using the FSMA-PDNet model.
[0015] Furthermore, the specific method for the MCMSDA network in step 5 is as follows:
[0016] Step 5.1: The MCMSDA network contains three parallel one-dimensional convolutional branches. The kernel size of each branch is designed to have a different temporal receptive field, corresponding to different granularity feature patterns in the pulse signal. Let the preprocessed wrist pulse signal be... Where C0=3 represents the three pulse positions of Cun, Guan, and Chi, and L0 is the signal length; the module uses three parallel branches for feature extraction, denoted as i∈{1,2,3}; each branch uses a different convolution kernel size k. i ∈{3,5,7} and the expansion rate d i The input signal is passed through four one-dimensional convolutional layers ∈{1,2,3} to achieve multi-scale feature modeling under different receptive fields. In each branch, the input signal is passed through four one-dimensional convolutional layers in sequence. After each convolution, batch normalization is performed and ReLU activation function is used to form the basic feature extraction unit. Let j∈{1,2,3,4} represent the layer index inside the branch. To further enhance the feature representation capability, a channel attention (CA) module is introduced after the fourth convolution to adaptively recalibrate the features. The output of the j-th layer in the i-th branch is expressed as:
[0017] ;
[0018] ;
[0019] in This represents a one-dimensional convolution operation, and CA represents the channel attention module. Indicates the batch normalization layer. Indicates the kernel size. This represents the magnitude of the expansion rate; for the i-th branch, its initial input is defined as... After feature extraction from the three branches is complete, the outputs of each branch are concatenated along the channel dimension to obtain the fused feature representation. :
[0020] ;
[0021] To represent the splicing operation, and to reduce channel redundancy after multi-branch fusion and enhance the ability to express features interactively, a method is introduced... Convolution performs channel compression and recoding on the fused features to obtain intermediate feature representations. The dimensionality-reduced features are then input again into the channel attention module for global feature recalibration, resulting in the final output weighted feature map. :
[0022] ;
[0023] Where C1 and L1 represent respectively The number and length of channels, This will be input into the subsequent pooling and classification modules;
[0024] Furthermore, the channel attention module in the MCMSDA network in step 4.1 includes three sub-modules: a compression module, an excitation module, and a recalibration module.
[0025] The specific method of the compression module is as follows: Let the input feature map be... Channel description vectors are obtained through global average pooling. :
[0026] ;
[0027] in, , This indicates the number and length of the input features. Represents the input feature map The Middle The first channel, the first The values at each position are obtained, and the above global average pooling operation is performed on each channel to obtain the channel description matrix. ;
[0028] The specific method of the incentive module is as follows: introducing a weight matrix. and These correspond to channel compression and channel restoration operations, respectively, enabling the modeling of channel dependencies:
[0029] ;
[0030] in and Let represent the ReLU and Sigmoid activation functions respectively, and r be the dimensionality reduction ratio;
[0031] The specific method of the recalibration module is to perform element-wise multiplication along the channel dimension, as follows:
[0032] ;
[0033] in, This represents element-wise multiplication along the channel dimension;
[0034] Step 5.2: Use the time-series pyramid pooling module for pooling;
[0035] Let the number of pyramid levels be S, and the output scale corresponding to the s-th level be n. s Then for input features Adaptive average pooling over the time dimension yields:
[0036] ;
[0037] in, This indicates that the time dimension is compressed to The adaptive average pooling operation, corresponding to the output features Subsequently, the pooling results of each layer are expanded along the time dimension and concatenated along the channel dimension to construct a unified multi-scale temporal feature representation. :
[0038] ;
[0039] Furthermore, the specific method for reconstructing the preprocessed one-dimensional wrist pulse signal into a two-dimensional representation in step 6 is as follows:
[0040] Step A1: Along the channel dimension A reference signal is constructed by averaging. ;
[0041] ;
[0042] in Indicates the number of channels. Indicates the first The signal of each channel is used as a reference signal, and subsequently, all period division operations are based on this reference signal. Complete, and synchronously apply the detected periodic boundaries to each channel to ensure the consistency of multi-channel signals at the periodic level;
[0043] Step A2: Perform spectral analysis on the signal using Fast Fourier Transform and search for the dominant frequency within the heart rate band. :
[0044] ;
[0045] in Indicates the FFT operation. This indicates the heart rate frequency range that conforms to the physiological range, using the dominant frequency. and signal sampling rate This yields an initial estimate of the pulse cycle length. ;
[0046] ;
[0047] Step A3: Invert the original signal to convert the valley values into peak values, and then perform peak detection in the time domain; apply a combined minimum distance interval constraint to the peak detection. :
[0048] ;
[0049] Let the valley values obtained from the final detection be arranged in chronological order as follows: Then the first A complete cycle is defined as:
[0050] ;
[0051] Its corresponding length is To ensure the reliability of subsequent modeling, the signal start segment is an incomplete cycle. Incomplete cycle at the end Need to be discarded;
[0052] Step A4: Introduce two hyperparameters: upper limit of period length With the upper limit of the number of cycles ;in, Used to constrain the maximum length of a single pulse cycle in the time dimension, These parameters are used to limit the number of cycles involved in modeling in each sample. By presetting these two parameters, the original irregular cycle sequence is mapped to a two-dimensional tensor representation of uniform size; through the valley detection and segmentation operations, signal segments of each complete cycle have been obtained.
[0053] Step A5: Use the starting point of each cycle as a unified starting point to maintain alignment on the timeline. One cycle The original length is Uniform mapping to a fixed length is achieved through truncation or zero-padding operations. Then the processed first Each cycle is represented as:
[0054] ;
[0055] Specifically, if Then zero-filling is performed at the end of the cycle; if Then cut off to the front One sampling point;
[0056] Step A6: Introduce a binary mask matrix to explicitly mark whether each time step is a valid signal. The corresponding mask matrix is represented as :
[0057] ;
[0058] Step A7: Arrange the cycles in chronological order, and let the actual number of cycles detected be... Then, the number of cycles can be unified to a certain value through cropping or filling operations. ;like Take the front One cycle; if The tail is filled with all zero cycles; finally, a two-dimensional representation of uniform size is obtained. :
[0059] ;
[0060] Meanwhile, its corresponding mask matrix is represented as :
[0061] ;
[0062] In this representation, the horizontal axis Corresponding to the fast time dimension, it is used to characterize fine-grained morphological changes within a single cycle; the vertical axis... Corresponding to the slow time dimension, it is used to describe the rhythmic evolution process between cycles.
[0063] Furthermore, the method for extracting two-dimensional features in step 6 is as follows: using the P2DInception network to extract two-dimensional features;
[0064] Step B1: The P2DInception network is composed of multiple stacked CBAM-Inception modules, and the CBAM-Inception module includes a CBAM module and an Inception module;
[0065] The Inception module employs a multi-branch parallel structure. Branch 1 downsamples the input features through pooling to enhance the modeling ability of regional statistical information; Branch 2 uses a shallower convolutional structure to extract local detail features; Branch 3 further expands the receptive field by stacking multiple convolutional layers; Branch 4 retains a path based on... The lightweight path of convolution is used to achieve channel mapping and reduce information loss; the outputs of different branches are concatenated along the channel dimension, and feature fusion and dimensionality adjustment are completed through convolution.
[0066] The CBAM module adaptively weights the features. Let the input feature map be represented as... First, the input feature map Perform global max pooling and global average pooling respectively, and then input the two feature vectors obtained into a multilayer perceptron (MLP) with one hidden layer to generate a channel attention map:
[0067] ;
[0068] Compared with the original input features Multiplying yields a weighted feature map. First, Average pooling and max pooling operations are applied along the channel axis, and the results are further concatenated along the channel dimension to generate effective feature descriptors. Then, convolutional layers are used to generate spatial attention maps.
[0069] ;
[0070] in This represents the sigmoid function. This represents a two-dimensional convolution operation; finally, the channel-weighted feature map is... Spatial attention map The final weighted result obtained by multiplication is a two-dimensional feature. ;
[0071] Step B2: Define the Mask-Aware Averaging Operation (MAAP);
[0072] ;
[0073] in This represents element-wise multiplication. For the corresponding mask matrix, Use a small constant to prevent division by zero;
[0074] Step B3: Perform mask-aware averaging of features along the slow time dimension:
[0075] ;
[0076] Then on Perform global average pooling and global max pooling respectively.
[0077] ;
[0078] ;
[0079] in By concatenating the two along the channel dimension, we obtain:
[0080] ;
[0081] Simultaneously, First, calculate the mask-aware average along the fast time dimension to obtain Then perform global average pooling and global max pooling to obtain By combining the two, we can obtain Finally, the feature representations of these two dimensions are concatenated along the channel dimension to form the final fused two-dimensional feature vector. .
[0082] Furthermore, the specific method for step 7 is as follows:
[0083] Step 7.1: The feature representation of the one-dimensional branch output is as follows The feature representation of the two-dimensional branch output is as follows ,in and These represent the channel dimensions of the corresponding features. A lightweight fusion method using direct concatenation is employed to fully preserve and effectively integrate multi-source feature information while maintaining manageable model complexity, resulting in a fused feature representation.
[0084] ;
[0085] in This indicates a splicing operation along the channel dimension;
[0086] Step 7.2: Features after fusion The input is used for discrimination by a multi-task classification network;
[0087] The multi-task classification network contains four multi-classifiers, each classifier head corresponding to a different classification task.
[0088] For each classification task, consider the imbalance in the number of samples from different classes within it; let the first class be... The categories of tasks include The first category, recorded as the... The number of class samples is The initial weights for this category are defined as follows:
[0089] ;
[0090] in, To prevent tiny constants with a denominator of zero;
[0091] The weights are normalized, assigning greater weights to classes with smaller sample sizes and smaller weights to classes with larger sample sizes, and ensuring that... The weight of each category is 1, thus ensuring the consistency of the loss of different tasks on a numerical scale;
[0092] ;
[0093] For input samples Let it be in the first The real labels in each task are The model predicts the class probability as The single-sample weighted cross-entropy loss for this task is:
[0094] ;
[0095] in, This represents the component in the predicted probability that corresponds to the true class. Indicates that the true class in the predicted probability is The corresponding weights; assuming the current batch contains N samples, the loss function for the h-th classification task is:
[0096] ;
[0097] Introducing learnable task weight parameters The losses of different tasks are adaptively weighted; let the original weight parameters of the h-th task be denoted as . via the Softplus function Ensure its non-negativity;
[0098] ;
[0099] For learnable task weight parameters After normalization, the final task weights are:
[0100] ;
[0101] The total loss function of the multi-task classification network is defined as: (This represents the total number of tasks.)
[0102] ;
[0103] The multi-task classification network is trained using this loss, resulting in a well-trained multi-task classification network.
[0104] Through the above steps, this application realizes intelligent identification of traditional Chinese medicine pulse patterns using the FSMA-PDNet model, and compares it with other commonly used machine learning methods to evaluate the accuracy of the model. Attached Figure Description
[0105] Figure 1 This is a diagram of the overall architecture of the FSMA-PDNet model.
[0106] Figure 2 This is a schematic diagram of the MCMSDA module.
[0107] Figure 3 This is a schematic diagram of the timing pyramid pooling module.
[0108] Figure 4 The flowchart is for the FaSlTiRe algorithm.
[0109] Figure 5 This is a schematic diagram of the P2DInception module.
[0110] Figure 6 This is a schematic diagram of mask-aware dual-axis pooling. Detailed Implementation
[0111] Pulse diagnosis is an important step in traditional Chinese medicine diagnosis. Doctors judge a patient's health status by sensing the pulsation of the wrist pulse signal. This application automates this process by using sensors to collect the patient's wrist pulse signal, preprocessing it using signal processing methods, automatically learning the characteristics of different pulse patterns using a deep learning network, and finally using the trained network to achieve intelligent pulse diagnosis in traditional Chinese medicine.
[0112] This application constructs a multiscale attention neural network with fast-slow time reshaping for pulse diagnosis (FSMA-PDNet), which integrates fast and slow time reconstruction with multiscale convolution. Its overall workflow is as follows: Figure 1 As shown. This network mainly consists of the MCMSDA network for processing one-dimensional wrist pulse signals, the FaSlTiRe algorithm for reconstructing one-dimensional wrist pulse signals into two dimensions, the P2DInception network for processing the corresponding two-dimensional representation, and a feature fusion and multi-task classification network. The specific steps are as follows:
[0113] Step 1: The MCMSDA feature extraction network architecture established in this paper is as follows: Figure 2 As shown, this module is used to extract hierarchical feature representations across multiple time scales from the input standardized pulse period matrix. The module contains three parallel one-dimensional convolutional branches, each with a convolutional kernel size designed to correspond to different temporal receptive fields, thus representing feature patterns of different granularities in the pulse signal. Assuming the preprocessed wrist pulse signal is... Where C0=3 represents the three pulse positions of Cun, Guan, and Chi, and L0 is the signal length. The module uses three parallel branches for feature extraction, denoted as i ∈ {1, 2, 3}. Each branch uses a different convolutional kernel size k. i ∈ {3, 5, 7} and the expansion rate d i The input signal is passed through four one-dimensional convolutional layers ∈ {1, 2, 3} to achieve multi-scale feature modeling under different receptive fields. In each branch, the input signal is passed through four one-dimensional convolutional layers sequentially. After each convolution, batch normalization (BN) is performed, and the ReLU activation function is used to form the basic feature extraction unit. Let j ∈ {1, 2, 3, 4} represent the layer index within the branch. To further enhance the feature representation capability, a channel attention (CA) module is introduced after the fourth convolution to adaptively recalibrate the features. Therefore, the output of the j-th layer in the i-th branch can be expressed as:
[0114] ;
[0115] ;
[0116] in Let represent a one-dimensional convolution operation, and CA represent a channel attention module. For the i-th branch, its initial input is defined as... After feature extraction from the three branches is complete, the outputs of each branch are concatenated along the channel dimension to obtain the fused feature, which can be represented as follows:
[0117] ;
[0118] To reduce channel redundancy after multi-branch fusion and enhance feature interaction expression capabilities, a method is introduced. Convolution performs channel compression and recoding on the fused features to obtain intermediate feature representations. The dimensionality-reduced features are then input again into the channel attention module for global feature recalibration, resulting in the final output weighted feature map.
[0119] ;
[0120] Where C1 and L1 represent respectively The number and length of channels, The input will then be used for subsequent pooling and classification modules. Through multi-scale receptive field modeling and an adaptive weighting mechanism based on channel attention, this feature extraction module effectively enhances the feature representation ability of wrist pulse signals at different time scales, thereby improving the overall feature extraction performance.
[0121] The channel attention module mainly consists of three sub-modules: a compression module, an activation module, and a recalibration module. The compression module's function is to compress the spatial information of each channel into channel descriptors, thereby obtaining global semantic information. Let the input feature map be... Channel description vectors are obtained through global average pooling.
[0122] ;
[0123] Then, the above global average pooling operation is performed on each channel to obtain... .
[0124] The activation module's role is to learn the non-linear dependencies between channels through a gating mechanism, thereby generating channel attention weights. Specifically, a weight matrix is introduced. and These correspond to channel compression and channel restoration operations, respectively, thus enabling the modeling of channel dependencies:
[0125] ;
[0126] in and Let represent the ReLU and Sigmoid activation functions, respectively, and r be the dimensionality reduction ratio.
[0127] The recalibration module applies the weights *s* learned by the activation module to the original feature map *F*, achieving adaptive modulation of the feature response through channel-wise weighting. Specifically, this process is accomplished through element-wise multiplication along the channel dimension, and can be represented as follows:
[0128] ;
[0129] in, This represents element-wise multiplication along the channel dimension. Through the above recalibration process, the feature response of each channel will be adaptively scaled according to its corresponding weight.
[0130] The time-series pyramid pooling module established in this paper is as follows: Figure 3 As shown, the temporal location of the wrist pulse signal contains important physiological dynamic information, such as the periodic changes in the wrist pulse waveform and the local peak structure. Therefore, constructing a multi-scale representation in the time dimension is of great significance for fully mining the discriminative features of the wrist pulse signal. Specifically, assuming the number of pyramid layers is S, the output scale corresponding to the s-th layer is n. s Then for input features By performing adaptive average pooling (AAP) over the time dimension, we can obtain...
[0131] ;
[0132] in, This indicates that the time dimension is compressed to The adaptive average pooling operation, corresponding to the output features Subsequently, the pooling results of each layer are expanded along the time dimension and concatenated along the channel dimension to construct a unified multi-scale temporal feature representation.
[0133] ;
[0134] Through the above operations Not only does it retain global statistical information ( ), and also encodes local dynamic change features ( This enhances the discriminative power of feature representation.
[0135] Step 2: This invention further introduces the modeling concepts of "Fast-time" and "Slow-time" in wrist pulse signal analysis. Specifically, Fast-time corresponds to the aforementioned intra-cycle variations, describing the fine-grained waveform evolution along the time axis within a single pulse cycle; Slow-time corresponds to inter-cycle variations, characterizing the rhythmic fluctuations and trend evolution between different pulse cycles. Based on this idea, this paper proposes a Fast–Slow Time Reshaping (FaSlTiRe) algorithm, the process of which is as follows: Figure 4 As shown, this method takes preprocessed multi-channel one-dimensional wrist pulse signals as input and maps the original time series into a two-dimensional representation with clear physical meaning, thereby enabling explicit modeling of intra-cycle and inter-cycle information.
[0136] For multi-channel wrist pulse signals, strict time alignment between channels is crucial. To achieve this, we first align the channels along the channel dimension. To construct a reference signal, we perform an average:
[0137] ;
[0138] in Indicates the number of channels. Indicates the first The signal of each channel. Subsequently, all periodization operations are based on this reference signal. The detected periodic boundaries are then synchronously applied to each channel to ensure consistency of the multi-channel signals at the periodic level. To obtain robust periodic prior information, the signal is first subjected to spectral analysis using Fast Fourier Transform (FFT), and the dominant frequency is searched within a reasonable heart rate frequency band.
[0139] ;
[0140] in Indicates the FFT operation. This indicates the heart rate frequency range that conforms to the physiological range, using the dominant frequency. and signal sampling rate This allows us to obtain an initial estimate of the pulse cycle length:
[0141] ;
[0142] Since the starting point of the wrist pulse signal cycle usually corresponds physiologically to the trough point before the waveform contraction phase, the original signal is inverted to convert the trough into a peak value for easier detection, and then peak detection is performed in the time domain. Simultaneously, to avoid misjudging the internal structures of waveforms such as diphtheria waves and tidal waves as cycle boundaries, this paper applies a combined minimum distance interval constraint to peak detection.
[0143] ;
[0144] in This represents the period length estimated from the dominant frequency. This constraint uses prior information in the frequency domain to limit the peak interval, thereby effectively suppressing the interference of local structural peaks on period segmentation and improving the stability and consistency of period boundary detection.
[0145] Let the valley values obtained from the final detection be arranged in chronological order as follows: Then the first A complete cycle is defined as
[0146] ;
[0147] Its corresponding length is To ensure the reliability of subsequent modeling, the signal start segment is an incomplete cycle. Incomplete cycle at the end It needs to be discarded, and only the complete pulsation cycle should be retained.
[0148] To achieve batch training and structure alignment, this paper introduces two key hyperparameters: upper bound on period length. With the upper limit of the number of cycles .in, Used to constrain the maximum length of a single pulse cycle in the time dimension, These parameters are used to limit the number of cycles involved in modeling for each sample. By presetting these two parameters, the original irregular cycle sequence can be mapped to a uniform two-dimensional tensor representation.
[0149] Based on the valley detection and segmentation operations described above, signal segments for each complete cycle have been obtained. Furthermore, the starting point of each cycle is used as a unified starting point to maintain alignment on the time axis, and the first... One cycle The original length is Uniform mapping to a fixed length is achieved through truncation or zero-padding operations. Then the processed first Each period can be represented as
[0150] ;
[0151] Specifically, if Then zero-filling is performed at the end of the cycle; if Then cut off to the front One sampling point.
[0152] However, zero-padding introduces regions that do not contain real physiological information, which may interfere with model learning if not properly distinguished. Therefore, this paper further introduces a binary mask matrix to explicitly mark whether each time step represents a valid signal. The corresponding mask matrix can be represented as
[0153] ;
[0154] This mask will be used in subsequent pooling operations to shield the filled regions, ensuring that the model focuses only on the true signal portion.
[0155] After completing single-cycle normalization, the cycles are arranged in chronological order. Let the actual number of detected cycles be... Then, the number of cycles can be unified to a certain value through cropping or filling operations. .like Take the front One cycle; if The tail is filled with all zero periods. This results in a uniform two-dimensional representation.
[0156] ;
[0157] Meanwhile, its corresponding mask matrix can be expressed as
[0158] ;
[0159] In this representation, the horizontal axis Corresponding to the fast time dimension, it is used to characterize fine-grained morphological changes within a single cycle; the vertical axis... Corresponding to the slow time dimension, it is used to describe the rhythmic evolution process between cycles.
[0160] In the above description, for ease of explanation, single-channel signals are used as examples. For actual multi-channel wrist pulse signals, this paper performs period truncation and length normalization operations on each channel signal based on a unified period division result, and stacks them in the channel dimension to finally obtain a structured tensor.
[0161] ;
[0162] Since both period division and length normalization operations are based on a unified reference signal, and each channel is strictly aligned in the time dimension, only a shared mask matrix needs to be constructed to mark the effective signal area and mask the filling part.
[0163] Step 3: P2DInception model architecture as follows Figure 5As shown, specifically, the network consists of multiple stacked CBAM-Inception modules. The Inception module employs a multi-branch parallel structure. Branch 1 downsamples the input features through pooling to enhance the modeling ability of regional statistical information; Branch 2 uses a shallower convolutional structure to extract local detail features; Branch 3 further expands the receptive field by stacking multiple convolutional layers; Branch 4 retains a single line based on… Lightweight paths for convolution are used to achieve channel mapping and reduce information loss. The outputs of different branches are concatenated along the channel dimension, and feature fusion and dimensionality adjustment are completed through convolution. To further enhance the model's ability to express key information, this paper introduces CBAM to adaptively weight features, achieving adaptive optimization of input features by calculating attention maps in two independent dimensions: channel and spatial dimension. Let the input feature map be represented as... First, the input feature map Perform global max pooling and global average pooling respectively, and then input the two feature vectors obtained into a multilayer perceptron (MLP) with one hidden layer to generate a channel attention map.
[0164] ;
[0165] Compared with the original input features Multiplying yields a weighted feature map. To compute spatial attention, we first... Average pooling and max pooling operations are applied along the channel axis, and the results are further concatenated along the channel dimension to generate effective feature descriptors. Then, convolutional layers are used to generate spatial attention maps.
[0166] ;
[0167] in This represents the sigmoid function. This represents a two-dimensional convolution operation. Finally, the channel-weighted feature map is... Spatial attention map Multiply them to get the final weighted result.
[0168] In addition, to improve the stability of network training and promote the effective propagation of features, residual connection structures are introduced into the network. This design can alleviate the gradient vanishing problem in deep networks while preserving fine-grained information in low layers, which helps to improve the robustness of overall feature representation.
[0169] After constructing the two-dimensional feature extraction module, the feature tensor output by the network contains rich spatiotemporal information, but its dimensionality is still high and it contains invalid regions introduced by zero-padding. Therefore, this paper further designs a mask-aware biaxial pooling method to achieve robust feature compression and discriminative modeling while preserving effective information, such as... Figure 6 As shown.
[0170] As known from the preceding analysis, the output of the two-dimensional feature extraction module is: ,in This represents the number of feature channels. To avoid interference from the filled region on the statistical features in the FaSlTiRe algorithm, this paper defines Mask-aware Avergae Pooling (MAAP).
[0171] ;
[0172] in This represents element-wise multiplication. For the corresponding mask matrix, A small constant is used to prevent division by zero. Summation is performed along the dimension to be aggregated, and this operation is only performed on the valid signal region, thus ensuring the physical rationality of feature aggregation.
[0173] Based on this, features are decoupled from both fast and slow time dimensions to characterize intra-cycle and cross-cycle information, respectively. First, a mask-perceptual averaging is performed on the features along the slow time dimension.
[0174] ;
[0175] Then on Perform global average pooling and global max pooling respectively.
[0176] ;
[0177] ;
[0178] in By concatenating the two along the channel dimension, we obtain...
[0179] ;
[0180] Simultaneously, First, calculate the mask-aware average along the fast time dimension to obtain Then perform global average pooling and global max pooling to obtain By combining the two, we can obtain Finally, the feature representations from these two dimensions are concatenated along the channel dimension to form the final fused feature vector. .
[0181] Step 4: Wrist pulse signals exhibit significant multi-level feature representation characteristics. One-dimensional temporal features can effectively characterize the dynamic changes of pulse signals, such as waveform fluctuations and rhythm information; while two-dimensional structural features obtained based on temporal reconstruction focus more on describing the morphological distribution and overall structural pattern across cycles. These two types of features are complementary in representation space; therefore, it is necessary to fuse one-dimensional and two-dimensional features to construct a more discriminative unified feature representation.
[0182] As can be seen from the previous analysis, the feature representation of the one-dimensional branch output is as follows: The feature representation of the two-dimensional branch output is as follows ,in and These represent the channel dimensions of the corresponding features. Considering that we have already obtained compact and discriminative feature representations through temporal pyramid pooling and mask-aware biaxial pooling, introducing additional complex structures in the fusion stage may lead to redundant parameters and increase the risk of overfitting. Therefore, this paper adopts a lightweight fusion method of direct concatenation, which, while maintaining controllable model complexity, achieves full preservation and effective integration of multi-source feature information to obtain a fused feature representation.
[0183] ;
[0184] in This indicates a stitching operation along the channel dimension. The fused features. It is then input into a subsequent classifier for discrimination.
[0185] In Traditional Chinese Medicine (TCM) pulse diagnosis theory, pulse characteristics are typically composed of multiple basic elements, such as a floating, slippery, and slow pulse, or a deep, wiry, and thready pulse. A single wrist pulse signal contains various elements, which often exhibit opposing or complementary relationships. Pulses reflecting the same attribute but with opposite characteristics are usually called "opposite pulses." For example, there is a distinction between floating and deep in pulse position, slow and rapid in pulse rate, and full and weak in pulse strength. However, in clinical practice, fullness and weakness are two aspects of pathology and often coexist (mixed fullness and weakness). Therefore, to better reflect the characteristics of clinical data, this paper treats them as two independent and evaluable features. Based on the above analysis, this paper designs a multi-task classification framework containing four classifiers, with each classifier head corresponding to a different classification task.
[0186] For each classification task, consider the imbalance in the number of samples from different classes within it. Let the first class be... The categories of tasks include The first category, recorded as the... The number of class samples is Then the initial weight of this category is defined as:
[0187] ;
[0188] in, To prevent tiny constants with a denominator of zero.
[0189] To avoid introducing additional scale shifts due to differences in the number of categories between different tasks, the weights are normalized, assigning greater weights to categories with smaller sample sizes and smaller weights to categories with larger sample sizes, and ensuring that... The weights of each category are equal to 1, thus ensuring consistency in the numerical scale of the losses across different tasks.
[0190] ;
[0191] Based on this, for the input sample Let it be in the first The real labels in each task are The model predicts the class probability as The single-sample weighted cross-entropy loss for this task is defined as follows:
[0192] ;
[0193] in, This represents the component in the predicted probability corresponding to the true class. Suppose the current batch contains N samples, then the loss function for the h-th classification task can be expressed as:
[0194] ;
[0195] The above method alleviates the class imbalance problem in each classification task by assigning higher weights to minority classes, making the model pay more attention to the difficult-to-learn classes during training.
[0196] In a multi-task learning framework, the importance of different classification tasks to model optimization is not entirely consistent. Directly summing the losses of each task with equal weights might lead to some tasks dominating the training process, thus affecting overall performance. Therefore, we introduce learnable task weight parameters to adaptively weight the losses of different tasks. Let the original weight parameters for the h-th task be... via the Softplus function Guarantee its nonnegativity
[0197] ;
[0198] Further normalization yields the final task weights as follows:
[0199] ;
[0200] Based on this, the total loss function of the entire multi-task learning framework is defined as:
[0201] ;
[0202] This weighting method can dynamically adjust the importance of each task according to the training process, so that the model can take into account the discrimination ability of different pulse elements during the optimization process, thereby improving the overall recognition performance.
[0203] Step 5:
[0204] This step is used to verify the specific performance of the proposed FSMA-PDNet model in actual diagnosis and compare it with several commonly used machine learning methods. To comprehensively evaluate the model's performance in multi-task classification scenarios, this paper selects accuracy (ACC), precision (PR), recall (RE), and F1 score (F1) as evaluation metrics. For the h-th classification task, the evaluation metrics are calculated based on the confusion matrix and defined as follows:
[0205] ;
[0206] ;
[0207] ;
[0208] ;
[0209] in, , , and denoted by , respectively, the number of true positives, true negatives, false positives, and false negatives in the h-th task. Since this paper employs a multi-task learning framework, to comprehensively reflect the overall performance of the model, the average of the evaluation metrics for each task is taken as the final result.
[0210] ;
[0211] Where H represents the total number of classification tasks. The above evaluation metrics (ACC, PR, RE, F1) are calculated for the h-th task. This evaluation method can comprehensively measure the model's discriminative ability across different pulse dimensions.
[0212] The comparative experimental results with different models are shown in Table 1. As can be seen from Table 1, the proposed method outperforms other machine learning algorithms in many metrics. The comparative experimental results of branch ablation are shown in Table 2. As can be seen from Table 2, FSMA-PDNet can effectively complement the limitations of single branches in feature representation and improve the performance of the model.
[0213] Table 1 shows the performance comparison of different models in pulse recognition experiments.
[0214]
[0215] Table 2 shows the performance comparison of the branch ablation experiment.
[0216]
Claims
1. A pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution, the method comprising: Step 1: Select subjects and collect their clinical pulse data using a three-channel acquisition device based on pressure sensors; Step 2: Traditional Chinese medicine experts mark the pulse of the test subjects and the corresponding pulse pattern; Step 3: Use wavelet transform to remove high-frequency noise from the pulse signal and use cubic spline interpolation to remove the baseline of the pulse signal; Step 4: Construct an intelligent pulse diagnosis network that integrates fast and slow temporal reconstruction with multi-scale convolution, comprising: an MCMSDA network for processing one-dimensional wrist pulse signals, a P2DInception network for representing one-dimensional wrist pulse signals in two dimensions, a feature fusion module, and a multi-task classification network; the preprocessed one-dimensional wrist pulse signals are input into the MCMSDA network and the P2DInception network respectively, and then the output data of the MCMSDA network and the P2DInception network are input into the feature fusion module, and then into the multi-task classification network; Step 5: Input the preprocessed one-dimensional wrist pulse signal into the MCMSDA network to extract the multi-scale one-dimensional features of the wrist pulse signal; Step 6: Reconstruct the preprocessed one-dimensional wrist pulse signal into a two-dimensional representation, and then extract the two-dimensional features; Step 7: Fuse the one-dimensional and two-dimensional features, and use the feature fusion and multi-task classification network to complete the recognition of arterial images; Step 8: Intelligently identify TCM pulse patterns using the FSMA-PDNet model.
2. The pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution as described in claim 1, characterized in that, The specific method for the MCMSDA network in step 5 is as follows: Step 5.1: The MCMSDA network contains three one-dimensional convolutional branches set in parallel. The kernel size of each branch is designed to be a different temporal receptive field to correspond to the feature patterns of different granularities in the pulse signal. set up The preprocessed wrist pulse signal is Where C0=3 represents the three pulse positions of Cun, Guan, and Chi, and L0 is the signal length; the module uses three parallel branches for feature extraction, denoted as i∈{1,2,3}; each branch uses a different convolution kernel size k. i ∈{3,5,7} and the expansion rate d i ∈{1,2,3} to achieve multi-scale feature modeling under different receptive fields; in each branch, the input signal passes through four layers of one-dimensional convolution structure in sequence, and batch normalization is performed after each convolution, and ReLU activation function is used to form the basic feature extraction unit; Let j∈{1,2,3,4} represent the layer index within the branch; to further enhance feature representation, a channel attention (CA) module is introduced after the fourth convolutional layer to adaptively recalibrate the features; the output of the j-th layer in the i-th branch is represented as: ; ; in This represents a one-dimensional convolution operation, and CA represents the channel attention module. Indicates the batch normalization layer. Indicates the kernel size. This represents the magnitude of the expansion rate; for the i-th branch, its initial input is defined as... After feature extraction from the three branches is complete, the outputs of each branch are concatenated along the channel dimension to obtain the fused feature representation. : ; To represent the splicing operation, and to reduce channel redundancy after multi-branch fusion and enhance the ability to express features interactively, a method is introduced... Convolution performs channel compression and recoding on the fused features to obtain intermediate feature representations. The dimensionality-reduced features are then input again into the channel attention module for global feature recalibration, resulting in the final output weighted feature map. : ; Where C1 and L1 represent respectively The number and length of channels, This will be input into the subsequent pooling and classification modules; Step 5.2: Use the time-series pyramid pooling module for pooling; Let the number of pyramid levels be S, and the output scale corresponding to the s-th level be n. s Then for input features Adaptive average pooling over the time dimension yields: ; in, This indicates that the time dimension is compressed to The adaptive average pooling operation, corresponding to the output features ; Subsequently, the pooling results of each layer are unfolded along the time dimension and concatenated along the channel dimension to construct a unified multi-scale temporal feature representation. : 。 3. The pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution as described in claim 2, characterized in that, The specific method for reconstructing the preprocessed one-dimensional wrist pulse signal into a two-dimensional representation in step 6 is as follows: Step A1: Along the channel dimension A reference signal is constructed by averaging. ; ; in Indicates the number of channels. Indicates the first The signal of each channel is used as a reference signal, and subsequently, all period division operations are based on this reference signal. Complete, and synchronously apply the detected periodic boundaries to each channel to ensure the consistency of multi-channel signals at the periodic level; Step A2: Perform spectral analysis on the signal using Fast Fourier Transform and search for the dominant frequency within the heart rate band. : ; in Indicates the FFT operation. This indicates the heart rate frequency range that conforms to the physiological range, using the dominant frequency. and signal sampling rate This yields an initial estimate of the pulse cycle length. ; ; Step A3: Invert the original signal to convert the valley values into peak values, and then perform peak detection in the time domain; apply a combined minimum distance interval constraint to the peak detection. : ; Let the valley values obtained from the final detection be arranged in chronological order as follows: Then the first A complete cycle is defined as: ; Its corresponding length is To ensure the reliability of subsequent modeling, the signal start segment is an incomplete cycle. Incomplete cycle at the end Need to be discarded; Step A4: Introduce two hyperparameters: upper limit of period length With the upper limit of the number of cycles ;in, Used to constrain the maximum length of a single pulse cycle in the time dimension, These parameters are used to limit the number of cycles involved in modeling in each sample. By presetting these two parameters, the original irregular cycle sequence is mapped to a two-dimensional tensor representation of uniform size; through the valley detection and segmentation operations, signal segments of each complete cycle have been obtained. Step A5: Use the starting point of each cycle as a unified starting point to maintain alignment on the timeline. One cycle The original length is Uniform mapping to a fixed length is achieved through truncation or zero-padding operations. Then the processed first Each cycle is represented as: ; Specifically, if Then zero-filling is performed at the end of the cycle; if Then cut off to the front One sampling point; Step A6: Introduce a binary mask matrix to explicitly mark whether each time step is a valid signal. The corresponding mask matrix is represented as : ; Step A7: Arrange the cycles in chronological order, and let the actual number of cycles detected be... Then, the number of cycles can be unified to a certain value through cropping or filling operations. ;like Take the front One cycle; if The tail is filled with all zero cycles; finally, a two-dimensional representation of uniform size is obtained. : ; Meanwhile, its corresponding mask matrix is represented as : ; In this representation, the horizontal axis Corresponding to the fast time dimension, it is used to characterize fine-grained morphological changes within a single cycle; the vertical axis... Corresponding to the slow time dimension, it is used to describe the rhythmic evolution process between cycles.
4. The pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution as described in claim 3, characterized in that, The method for extracting two-dimensional features in step 6 is as follows: the P2DInception network is used to extract two-dimensional features; Step B1: The P2DInception network is composed of multiple stacked CBAM-Inception modules, and the CBAM-Inception module includes a CBAM module and an Inception module; The Inception module employs a multi-branch parallel structure. Branch 1 downsamples the input features through pooling operations to enhance the ability to model regional statistical information; Branch 2 uses a shallower convolutional structure to extract local detail features. Branch 3 further expands the receptive field by stacking multiple convolutional layers; Branch 4 retains a single convolutional layer based on... The lightweight path of convolution is used to achieve channel mapping and reduce information loss; the outputs of different branches are concatenated along the channel dimension, and feature fusion and dimensionality adjustment are completed through convolution. The CBAM module adaptively weights the features. Let the input feature map be represented as... First, the input feature map Perform global max pooling and global average pooling respectively, and then input the two feature vectors obtained into a multilayer perceptron (MLP) with one hidden layer to generate a channel attention map: ; Compared with the original input features Multiplying yields a weighted feature map. First, Average pooling and max pooling operations are applied along the channel axis, and the results are further concatenated along the channel dimension to generate effective feature descriptors. Then, convolutional layers are used to generate spatial attention maps. ; in This represents the sigmoid function. This represents a two-dimensional convolution operation; finally, the channel-weighted feature map is... Spatial attention map The final weighted result obtained by multiplication is a two-dimensional feature. ; Step B2: Define the Mask-Aware Averaging Operation (MAAP); ; in This represents element-wise multiplication. For the corresponding mask matrix, Use a small constant to prevent division by zero; Step B3: Perform mask-aware averaging of features along the slow time dimension: ; Then on Perform global average pooling and global max pooling respectively. ; ; in By concatenating the two along the channel dimension, we obtain: ; Simultaneously, First, calculate the mask-aware average along the fast time dimension to obtain Then perform global average pooling and global max pooling to obtain By combining the two, we can obtain Finally, the feature representations of these two dimensions are concatenated along the channel dimension to form the final fused two-dimensional feature vector. .
5. The pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution as described in claim 4, characterized in that, The specific method for step 7 is as follows: Step 7.1: The feature representation of the one-dimensional branch output is as follows The feature representation of the two-dimensional branch output is as follows ,in and These represent the channel dimensions of the corresponding features. A lightweight fusion method using direct concatenation is employed to fully preserve and effectively integrate multi-source feature information while maintaining manageable model complexity, resulting in a fused feature representation. ; in This indicates a splicing operation along the channel dimension; Step 7.2: Features after fusion The input is used for discrimination by a multi-task classification network; The multi-task classification network contains four multi-classifiers, each classifier head corresponding to a different classification task. For each classification task, consider the imbalance in the number of samples from different classes within it; let the first class be... The categories of tasks include The first category, recorded as the... The number of class samples is The initial weights for this category are defined as follows: ; in, To prevent tiny constants with a denominator of zero; The weights are normalized, assigning greater weights to classes with smaller sample sizes and smaller weights to classes with larger sample sizes, and ensuring that... The weight of each category is 1, thus ensuring the consistency of the loss of different tasks on a numerical scale; ; For input samples Let it be in the first The real labels in each task are The model predicts the class probability as The single-sample weighted cross-entropy loss for this task is: ; in, This represents the component in the predicted probability that corresponds to the true class. Indicates that the true class in the predicted probability is The corresponding weights; assuming the current batch contains N samples, the loss function for the h-th classification task is: ; Introducing learnable task weight parameters The losses of different tasks are adaptively weighted; let the original weight parameters of the h-th task be denoted as . via the Softplus function Ensure its non-negativity; ; For learnable task weight parameters After normalization, the final task weights are: ; The total loss function of the multi-task classification network is defined as: (This represents the total number of tasks.) ; The multi-task classification network is trained using this loss, resulting in a well-trained multi-task classification network.
6. The pulse recognition method integrating fast and slow temporal reconstruction and multi-scale convolution as described in claim 2, characterized in that, The channel attention module in the MCMSDA network in step 5.1 includes three sub-modules: compression module, excitation module, and recalibration module. The specific method of the compression module is as follows: Let the input feature map be... Channel description vectors are obtained through global average pooling. : ; in, , This indicates the number and length of the input features. Represents the input feature map The Middle The first channel, the first The values at each position are obtained, and the above global average pooling operation is performed on each channel to obtain the channel description matrix. ; The specific method of the incentive module is as follows: introducing a weight matrix. and These correspond to channel compression and channel restoration operations, respectively, enabling the modeling of channel dependencies: ; in and Let represent the ReLU and Sigmoid activation functions respectively, and r be the dimensionality reduction ratio; The specific method of the recalibration module is to perform element-wise multiplication along the channel dimension, as follows: ; in, This represents element-wise multiplication operations along the channel dimension.