An age estimation network and system based on lead perception state space

CN122642924BActive Publication Date: 2026-09-25HANGZHOU PROTON TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611080405.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-25
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

[0008]其五,在噪声、基线漂移、导联间幅值差异和局部异常波形存在时,模型预测稳定性不足

Benefits of technology

[0064]1、通过导联感知前端得到浅层导联流特征和主干网络输入特征,多级残差U形特征提取模块和多尺度稠密特征融合模块得到深层多尺度时序特征,跨导联注意力模块基于浅层导联流特征和深层多尺度时序特征得到导联引导时序表征,状态空间时序建模模块得到长程时序表征,并经回归输出模块和区间映射模块输出年龄估计值,能够充分利用十二导联心电信号中与年龄相关的形态、节律和导联间差异信息,提高心电年龄估计的准确性和稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122642924B_ABST
    Figure CN122642924B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of electrocardiosignal analysis, and particularly relates to an age estimation network and system based on lead perception state space, comprising a preprocessing module, a lead perception front-end module, a multi-level residual U-shaped feature extraction module, a multi-scale dense feature fusion module, a cross-lead attention module, a state space time sequence modeling module, a regression output module, and an interval mapping module.In the present application, shallow lead flow features and backbone network input features are obtained through the lead perception front-end, deep multi-scale time sequence features are obtained through the multi-level residual U-shaped feature extraction module, lead-guided time sequence representations are obtained through the cross-lead attention module, long-range time sequence representations are obtained through the state space time sequence modeling module, and age estimation values are output through the regression output module and the interval mapping module, which can make full use of age-related morphological, rhythmic and inter-lead difference information in twelve-lead electrocardiosignal, and improve the accuracy and stability of electrocardiosignal age estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electrocardiogram signal analysis technology, specifically to an age estimation network and system based on lead-sensing state space. Background Technology

[0002] Electrocardiography (ECG) is one of the most commonly used non-invasive physiological signals in clinical practice. It can reflect the propagation process of cardiac electrical activity in different spatial directions. With age, the PR interval, QRS morphology, QT-related characteristics, heart rate variability, inter-lead electrical axis changes, and local waveform details in ECG signals may change to varying degrees. The difference between the cardiac age estimated by ECG and the actual age can serve as an indicator of cardiovascular health. Significant differences may reflect vascular aging and increased risk, and can provide auxiliary information for cardiovascular risk screening, abnormal aging indications, and health management.

[0003] Existing methods for estimating age using electrocardiograms typically employ artificial features or conventional convolutional neural networks for whole-segment regression. While these methods can learn certain age-related representations from electrocardiogram segments, they still suffer from the following problems:

[0004] First, the differences and complementarities among the twelve leads are not fully utilized, and the differences in the contributions of different leads to the age estimation task are easily overlooked.

[0005] Secondly, it lacks the ability to model the local waveform morphology, long-term rhythm context, and translead spatial relationships that coexist in electrocardiogram signals.

[0006] Third, ordinary convolutional networks are limited by their receptive field and have limited ability to express long-term time dependencies, while directly using traditional recurrent networks may result in low training efficiency.

[0007] Fourth, there is a lack of a unified and stable output mapping strategy that addresses the differences in age ranges between adults and children;

[0008] Fifth, the model's predictive stability is insufficient when noise, baseline drift, inter-lead amplitude differences, and local abnormal waveforms are present.

[0009] In summary, existing ECG age estimation techniques suffer from problems such as insufficient lead utilization, insufficient multi-scale temporal representation, insufficient long-term dependency modeling, and unstable age interval mapping. Summary of the Invention

[0010] The purpose of this invention is to provide an age estimation network and system based on lead-sensing state space to solve the problems mentioned in the background art.

[0011] To achieve the above objectives, the present invention provides the following technical solution:

[0012] An age estimation network based on lead-sensing state space includes a preprocessing module, a lead-sensing front-end module, a multi-level residual U-shaped feature extraction module, a multi-scale dense feature fusion module, a cross-lead attention module, a state space temporal modeling module, a regression output module, and an interval mapping module.

[0013] The preprocessing module is used to preprocess the twelve-lead ECG signal to obtain a standardized input signal;

[0014] The lead sensing front-end module is used to extract lead-level shallow features. The shallow lead flow features are mapped to lead key features by the lead encoder. At the same time, lead weights are generated based on the shallow lead flow features, and the weighted lead features are mapped to the backbone network input features.

[0015] The multi-level residual U-shaped feature extraction module is used to process the input features of the backbone network to obtain multi-level residual temporal features;

[0016] The multi-scale dense feature fusion module is used to process multi-level residual temporal features to obtain deep multi-scale temporal features;

[0017] The translead attention module can utilize deep multi-scale temporal features and lead key value features to obtain lead guidance temporal representations;

[0018] The state-space timing modeling module is used to perform long-range dependency modeling on the lead guidance timing representation to obtain long-range timing representation.

[0019] The regression output module is used to process long-range time series representations to obtain age prediction values;

[0020] The interval mapping module performs interval mapping on the predicted age values ​​and outputs the estimated age values.

[0021] Furthermore, the preprocessing module performs preprocessing on the twelve-lead ECG signal, including:

[0022] The raw electrocardiogram data is reshaped or organized into a two-dimensional matrix with the shape of twelve leads multiplied by a preset number of sampling points;

[0023] Extract a time series of a preset length from each lead signal;

[0024] Z-score normalization is performed on each lead along the time dimension;

[0025] Bandpass filtering is performed on each lead to suppress baseline drift, high-frequency noise, and power frequency interference;

[0026] Replace abnormal, missing, or infinite values ​​in the preprocessed signal with finite values.

[0027] Furthermore, the lead sensing front-end module includes a deep convolution branch, a lead attention gating branch, and a pointwise convolution branch;

[0028] The depthwise convolution branch is used to perform one-dimensional convolution on each of the twelve leads to extract the local waveform features of each lead;

[0029] The lead attention gating branch is used to globally aggregate the features of each lead along the time dimension, and obtain the lead weights through fully connected mapping and normalization function;

[0030] Pointwise convolutional branches are used to map the weighted lead features to a preset channel dimension to form the input features of the backbone network.

[0031] Furthermore, the multi-level residual U-shaped feature extraction module includes several residual U-shaped blocks, which are connected sequentially, and the encoding depth of the residual U-shaped blocks is different, so as to encode, decode and downsample the input features of the backbone network step by step, and output multi-level residual time-series features containing local waveform information and context rhythm information.

[0032] The residual U-shaped block includes an input convolutional layer, a multi-level coding layer, a bottleneck layer, a multi-level decoding layer, residual connections, and a downsampling layer;

[0033] Input convolutional layers are used to perform local morphological enhancement on input features;

[0034] Multi-level coding layers are used to progressively reduce temporal resolution and expand the receptive field through stride convolutions.

[0035] The bottleneck layer is used to aggregate contextual information in a low-resolution feature space;

[0036] Multi-level decoding layers are used to restore temporal resolution through deconvolution and fuse the features of the corresponding coding layers;

[0037] Residual connections are used to add the decoded features to the block input features;

[0038] The downsampling layer is used to output features with lower temporal resolution to the next layer of residual U-blocks.

[0039] Furthermore, the multi-scale dense feature fusion module includes several multi-scale bottleneck units to process multi-level residual temporal features;

[0040] Multi-scale bottleneck units include one-dimensional pointwise convolution, multiple parallel one-dimensional convolution branches, channel splicing layers, and densely connected layers;

[0041] One-dimensional pointwise convolution is used for input channel compression and nonlinear transformation;

[0042] Multiple parallel one-dimensional convolutional branches with different dilation rates are used to perform multi-scale fusion of multi-level residual temporal features and output deep multi-scale temporal features.

[0043] The channel splicing layer is used to splice the features obtained from convolutional branches with different dilation rates;

[0044] Dense connection layers are used to stitch together the multi-scale features and the input features along the channel dimension.

[0045] Furthermore, the specific steps by which the translead attention module obtains the lead-guided timing representation include:

[0046] Using deep multi-scale temporal features as query input, and using the lead key value features obtained by mapping shallow lead flow features through lead encoders as keys and values, lead guidance temporal representations are obtained through multi-head attention computation.

[0047] Furthermore, the lead encoder includes at least two layers of one-dimensional convolution, normalization layer and nonlinear activation layer. The lead encoder can map the shallow lead flow features that preserve the lead structure to the same channel dimension as the deep multi-scale temporal features to obtain lead key value features.

[0048] Preferably, the regression output module includes an adaptive pooling layer and a fully connected layer. The adaptive pooling layer aggregates long-range time series representations into a global feature vector, and the fully connected layer outputs the age prediction value based on the global feature vector.

[0049] Furthermore, the interval mapping module performs interval mapping on the predicted age values, including:

[0050] Applying the sigmoid function to the age prediction values ​​yields a normalized value between 0 and 1.

[0051] When the object to be estimated belongs to the adult age estimation task, the normalized value is multiplied by a preset first age span and then a preset first age lower limit is added to map the normalized predicted value to the adult age range.

[0052] When the object to be estimated belongs to the child age estimation task, the normalized value is multiplied by a preset second age span to map the normalized predicted value to the child age range.

[0053] An age estimation system based on lead-sensing state space is provided. The system includes a signal acquisition unit, a preprocessing unit, a lead-sensing coding unit, a multi-scale feature extraction unit, a cross-lead attention unit, a state space temporal modeling unit, an age regression unit, an interval mapping unit, and a result output unit.

[0054] The signal acquisition unit is used to acquire twelve-lead electrocardiogram signals;

[0055] The preprocessing unit is used to perform matrix organization, fixed-length truncation, standardization, bandpass filtering, and outlier processing on the twelve-lead ECG signal;

[0056] The lead sensing coding unit is used to extract shallow lead flow features, generate lead adaptive weights, and output the input features of the backbone network.

[0057] The multi-scale feature extraction unit is used to sequentially convert the input features of the backbone network into multi-level residual temporal features and deep multi-scale temporal features;

[0058] The translead attention unit is used to perform translead feature interaction based on shallow lead flow features and deep multi-scale temporal features, and output lead-guided temporal representations;

[0059] The state-space temporal modeling unit is used to perform long-range dependency modeling on the lead guidance timing representation and output long-range temporal representation.

[0060] The age regression unit is used to output predicted age values ​​based on long-range time-series characterization.

[0061] The interval mapping unit is used to perform the sigmoid function and age interval mapping on the age prediction value to obtain the age estimate;

[0062] The results output unit is used to output the age estimate and its corresponding age group mapping results.

[0063] Compared with the prior art, the beneficial effects of the present invention are:

[0064] 1. The system obtains shallow lead flow features and backbone network input features through the lead sensing front end, and obtains deep multi-scale temporal features through the multi-level residual U-shaped feature extraction module and the multi-scale dense feature fusion module. The translead attention module obtains lead-guided temporal representation based on shallow lead flow features and deep multi-scale temporal features. The state-space temporal modeling module obtains long-term temporal representation. The system outputs age estimates through the regression output module and the interval mapping module. This system can make full use of age-related morphological, rhythmic and inter-lead differences in the twelve-lead ECG signal to improve the accuracy and stability of ECG age estimation.

[0065] 2. The lead attention gating mechanism adaptively adjusts the contribution of the twelve leads, improving the ability to utilize the differences and complementary information between leads;

[0066] 3. Through a multi-level residual U-shaped feature extraction module, both local waveform morphology and long-term context features are taken into account;

[0067] 4. Through the multi-scale dense feature fusion module, multi-scale age-related electrocardiogram features are extracted from convolutional branches with different dilation rates;

[0068] 5. By using the cross-lead attention module, shallow lead flow features are reintroduced into the deep temporal representation, enabling the network to utilize both the original lead structure information and deep semantic information simultaneously.

[0069] 6. By capturing long-range time dependencies through the state-space temporal modeling module, the ability to model age-related changes across heartbeats and rhythm cycles is improved;

[0070] 7. By using an age range mapping strategy between adults and children, the network output is restricted to a reasonable age range for the corresponding task, thereby improving prediction stability. Attached Figure Description

[0071] Figure 1 This is a flowchart of the overall network of the present invention;

[0072] Figure 2 This is a flowchart of the electrocardiogram signal preprocessing in this invention;

[0073] Figure 3 This is a flowchart of the lead-level shallow feature extraction process in this invention;

[0074] Figure 4 This is a flowchart illustrating the specific processing of the residual U-shaped block in this invention;

[0075] Figure 5 This is a flowchart of the translead attention module in this invention for obtaining lead-guided timing representation;

[0076] Figure 6 This is a graph showing the age prediction results in this invention. Detailed Implementation

[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0078] Please see Figure 1-6 In this embodiment of the invention, an age estimation network based on lead-sensing state space includes a preprocessing module, a lead-sensing front-end module, a multi-level residual U-shaped feature extraction module, a multi-scale dense feature fusion module, a cross-lead attention module, a state space temporal modeling module, a regression output module, and an interval mapping module.

[0079] The preprocessing module is used to preprocess the twelve-lead ECG signal to obtain a standardized input signal;

[0080] The lead sensing front-end module is used to extract lead-level shallow features. The shallow lead flow features are mapped to lead key features by the lead encoder. At the same time, lead weights are generated based on the shallow lead flow features, and the weighted lead features are mapped to the backbone network input features.

[0081] The multi-level residual U-shaped feature extraction module is used to process the input features of the backbone network to obtain multi-level residual temporal features;

[0082] The multi-scale dense feature fusion module is used to process multi-level residual temporal features to obtain deep multi-scale temporal features;

[0083] The translead attention module can utilize deep multi-scale temporal features and lead key value features to obtain lead guidance temporal representations;

[0084] The state-space timing modeling module is used to perform long-range dependency modeling on lead-guided timing representations to obtain long-range timing representations. Specifically, the state-space timing modeling module uses a selective state-space model or the Mamba module to model timing features with input in the form of batch size, time steps, and number of channels in order to obtain long-range timing representations containing long-range time dependencies.

[0085] The regression output module is used to process long-range time series representations to obtain age prediction values. The regression output module includes an adaptive pooling layer and a fully connected layer. The adaptive pooling layer aggregates long-range time series representations into a global feature vector, and the fully connected layer outputs age prediction values ​​based on the global feature vector. Preferably, adaptive max pooling is used to aggregate the time dimension into global features of length 1, and then the age prediction values ​​are output through the fully connected layer.

[0086] The interval mapping module performs interval mapping on the predicted age values ​​and outputs the estimated age values.

[0087] Specifically, the age estimation network is preferably a LeadMambaNet structure. The age estimation network receives an input tensor with batch size, twelve leads and time length, and outputs a normalized age estimate.

[0088] The system obtains shallow lead flow features and backbone network input features through the lead sensing front end, and then obtains deep multi-scale temporal features through the multi-level residual U-shaped feature extraction module and the multi-scale dense feature fusion module. The translead attention module can obtain lead-guided temporal representation based on shallow lead flow features and deep multi-scale temporal features. The state-space temporal modeling module obtains long-term temporal representation. Finally, the regression output module and interval mapping module output the age estimate, thereby making full use of the age-related morphological, rhythm and inter-lead difference information in the twelve-lead ECG signal to improve the accuracy and stability of ECG age estimation.

[0089] In a preferred embodiment, the state-space temporal modeling module adopts the Mamba module, which models long-sequence ECG features through a selective state-space mechanism, capturing long-term temporal dependencies while maintaining high computational efficiency and outputting long-term temporal representations. Compared with structures that only use local convolutions, the state-space temporal modeling module can better express age-related changes across heartbeats, across rhythm cycles, and across longer time windows.

[0090] like Figure 2 As shown, the preprocessing module performs the following preprocessing steps on the twelve-lead ECG signal:

[0091] S11 reshapes or organizes the raw ECG data into a two-dimensional matrix with the shape of twelve leads multiplied by a preset number of sampling points;

[0092] S12, extract a time series of a preset length from each lead signal;

[0093] S13, Z-score normalization is performed on each lead along the time dimension;

[0094] S14 performs bandpass filtering on each lead to suppress baseline drift, high-frequency noise, and power frequency interference;

[0095] S15 replaces abnormal, missing, or infinite values ​​in the preprocessed signal with finite values.

[0096] Specifically, in S11, the twelve-lead ECG signal of the object to be estimated can come from a resting twelve-lead ECG, a physical examination ECG, a clinical monitoring ECG, or other twelve-lead ECG acquisition devices. After reading the raw ECG data, it is organized into a two-dimensional matrix form of "number of leads × number of sampling points".

[0097] In a preferred embodiment, the raw data is stored in the form of an array file, and after reading, it is reshaped into a 12×5000 two-dimensional matrix. In order to unify the model input length, a preset number of sampling points are extracted from each lead signal as the input segment. In other embodiments, other fixed lengths can be set according to the sampling rate, acquisition duration and computing resources, or a sliding window method can be used to process longer ECG records.

[0098] In S14, the bandpass filter adopts the Butterworth bandpass filter, with a preferred low cutoff frequency of 0.5Hz to 2Hz, a preferred high cutoff frequency of 30Hz to 100Hz, a preferred sampling rate of 250Hz to 1000Hz, and a preferred filter order of 2 to 5. This frequency band can retain information on QRS groups, P waves, T waves, and major age-related morphological changes, while suppressing low-frequency drift and high-frequency interference.

[0099] like Figure 3As shown, the lead sensing front-end module includes a deep convolution branch, a lead attention gating branch, and a pointwise convolution branch;

[0100] The depthwise convolution branch is used to perform one-dimensional convolution on each of the twelve leads to extract the local waveform features of each lead; the lead attention gating branch is used to globally aggregate the features of each lead along the time dimension, and obtain the lead weights through fully connected mapping and normalization functions; the pointwise convolution branch is used to map the weighted lead features to the preset channel dimension to form the input features of the backbone network.

[0101] Specifically, the steps for the lead sensing front-end module to extract lead-level shallow features are as follows:

[0102] S21, use a one-dimensional deep convolution branch grouped by lead to process the twelve leads respectively, so as to extract the local waveform features of each lead, and use the output of the deep convolution as the shallow lead flow features;

[0103] S22, perform global average aggregation of the shallow lead flow characteristics along the time dimension to obtain the statistical response of each lead;

[0104] S23, after the statistical response passes through a fully connected layer and a nonlinear activation layer, it is then processed by a softmax function to generate attention weights for the twelve leads, enabling the age estimation network to adaptively adjust the contribution of each lead to the age estimation task based on the electrocardiogram performance of different individuals.

[0105] S24, multiply the shallow lead flow characteristics with the corresponding lead weights to achieve adaptive weighting at the lead level;

[0106] S25 uses 1×1 pointwise convolution to project the weighted lead features onto the preset channel dimension to form the backbone network input features.

[0107] The backbone network input features are input into the subsequent multi-level residual U-shaped feature extraction module; the shallow lead flow features are input into the subsequent lead encoder, and after mapping, they form lead key value features, which are used as keys and values ​​in the cross-lead attention module. Through this setting, both outputs of the lead sensing front end are used in subsequent modules, respectively serving as the backbone feature extraction entry point and the lead structure guidance entry point.

[0108] like Figure 1 and Figure 4 As shown, in this embodiment, the multi-level residual U-shaped feature extraction module includes several residual U-shaped blocks, which are connected sequentially and have different encoding depths, so as to encode, decode and downsample the input features of the backbone network step by step, and output multi-level residual time-series features containing local waveform information and context rhythm information.

[0109] The residual U-shaped block includes an input convolutional layer, a multi-level coding layer, a bottleneck layer, a multi-level decoding layer, residual connections, and a downsampling layer;

[0110] The input convolutional layer is used to enhance the local morphology of the input features; the multi-level encoding layer is used to gradually reduce the temporal resolution and expand the receptive field through stride convolution; the bottleneck layer is used to aggregate contextual information in the low-resolution feature space; the multi-level decoding layer is used to restore the temporal resolution through deconvolution and fuse the corresponding encoding layer features; the residual connection is used to add the decoded features to the block input features; and the downsampling layer is used to output features with lower temporal resolution to the next layer of residual U-shaped blocks.

[0111] Specifically, the residual U-shaped block processing steps are as follows:

[0112] S31, by performing local convolution processing on the block input features through the input one-dimensional convolutional layer, local waveform enhancement features are obtained;

[0113] S32, input the local waveform enhancement features into the multi-level coding layer. The multi-level coding layer includes several one-dimensional convolutional units that are downsampled step by step along the time dimension. Each coding layer receives the output features of the previous level and reduces the temporal resolution and expands the receptive field through stride convolution to obtain the coding features of the corresponding level.

[0114] S33, the encoded features output from the last level of the encoding layer are input into the bottleneck layer, and the contextual information is further aggregated through one-dimensional convolution, normalization and non-linear activation to obtain the bottleneck features;

[0115] S34. The bottleneck features are input into the multi-level decoding layer. The temporal resolution is restored step by step through deconvolution or upsampling operations. In each decoding process, the features are concatenated and fused with the corresponding layer's encoded features to obtain the decoded fusion features.

[0116] S35, the decoded fusion features are cropped or aligned to the same time length as the local waveform enhancement features, and the residuals of the local waveform enhancement features are added to obtain residual enhancement features, so as to preserve input information and stabilize training;

[0117] S36, perform pooling and 1×1 convolution downsampling on the residual enhancement features to obtain the output features of the residual U-shaped block, and use the output features of the previous residual U-shaped block as the input features of the next residual U-shaped block.

[0118] In a preferred embodiment, the multi-level residual U-shaped feature extraction module includes at least four residual U-shaped blocks with different encoding depths. Through the U-shaped structures of different depths, the network can simultaneously learn short-term waveform morphology, local transmission features, and rhythmic context over a longer time range, and uniformly represent them as multi-level residual temporal features for further processing by the subsequent multi-scale dense feature fusion module.

[0119] like Figure 1 As shown, the multi-scale dense feature fusion module includes several multi-scale bottleneck units to process multi-level residual time-series features and enhance the model's ability to express age-related waveform changes at different time scales.

[0120] Multi-scale bottleneck units include one-dimensional pointwise convolution, multiple parallel one-dimensional convolution branches, channel splicing layers, and densely connected layers;

[0121] Among them, one-dimensional pointwise convolution is used to compress and nonlinearly transform the input channels; multiple parallel one-dimensional convolution branches with different dilation rates are used to perform multi-scale fusion of multi-level residual temporal features and output deep multi-scale temporal features; the channel splicing layer is used to splice the features obtained by convolution branches with different dilation rates; and the dense connection layer is used to splice the spliced ​​multi-scale features with the input features along the channel dimension.

[0122] Specifically, each multi-scale bottleneck unit first performs channel transformation through 1×1 convolution, and then extracts temporal features of different scales through multiple parallel one-dimensional convolution branches. The outputs of branches with different dilation rates are concatenated along the channel dimension and densely connected with the input features of the bottleneck unit. Preferably, the multiple parallel branches use one-dimensional convolutions with dilation rates of 1, 3 and 5.

[0123] In this way, the network can obtain a larger equivalent receptive field without significantly increasing the temporal downsampling depth, and integrate local QRS morphology, P wave and T wave changes, inter-lead amplitude differences and long-term rhythm changes into deep multi-scale temporal features, which are then used as query inputs for subsequent translead attention modules.

[0124] like Figure 1 and Figure 5 As shown, in this embodiment, the specific steps for the translead attention module to obtain the lead-guided timing representation include:

[0125] S41 uses deep multi-scale temporal features as query input.

[0126] S42, the lead key features obtained by mapping the shallow lead flow features through the lead encoder are used as keys and values.

[0127] S43, the lead-guided timing characterization is obtained through multi-head attention calculation.

[0128] The lead encoder includes at least two layers of one-dimensional convolution, normalization, and nonlinear activation. The lead encoder can map the shallow lead flow features that preserve the lead structure to the same channel dimension as the deep multi-scale temporal features to obtain lead key value features.

[0129] Specifically, the lead key value features maintain a correspondence with the shallow lead flow features in the time dimension and are aligned with the deep multi-scale temporal features in the channel dimension, serving as keys and values ​​in the cross-lead attention module;

[0130] The lead encoder comprises a first one-dimensional convolutional layer, a first normalized layer, a first nonlinear activation layer, a second one-dimensional convolutional layer, a second normalized layer, and a second nonlinear activation layer connected in sequence. The specific steps for the lead encoder to obtain lead key value features are as follows:

[0131] S421, the first one-dimensional convolutional layer receives the shallow lead flow features and maps the twelve-lead dimension to the intermediate channel dimension to perform preliminary fusion of the local waveform response of each lead;

[0132] S422, the first normalization layer and the first nonlinear activation layer perform amplitude normalization and nonlinear transformation on the intermediate channel features;

[0133] S423, the second one-dimensional convolutional layer further maps the intermediate channel features to the same channel dimension as the deep multi-scale temporal features;

[0134] S424, the second normalization layer and the second nonlinear activation layer stabilize the mapped features and enhance them nonlinearly, thereby obtaining the lead key value features.

[0135] Through this translead attention structure, the network can reintroduce lead-level original waveform cues into deep multi-scale temporal features to obtain lead-guided temporal representations. These lead-guided temporal representations serve as inputs to the state-space temporal modeling module, enabling age estimation to not only rely on the overall temporal representations but also utilize spatial complementary information between different leads.

[0136] like Figure 1 As shown, in this embodiment, the interval mapping module performs interval mapping on the predicted age value, including:

[0137] Applying the sigmoid function to the age prediction values ​​yields a normalized value between 0 and 1.

[0138] When the object to be estimated belongs to the adult age estimation task, the normalized value is multiplied by a preset first age span and then a preset first age lower limit is added to map the normalized predicted value to the adult age range.

[0139] When the object to be estimated belongs to the child age estimation task, the normalized value is multiplied by a preset second age span to map the normalized predicted value to the child age range.

[0140] Specifically, the network output is a normalized age estimate. During training or inference, the output is subjected to a sigmoid function and mapped to the actual age according to the task's age range.

[0141] Let the first age span be a, the first age lower limit be b, the second age span be c, and the network output be y;

[0142] For the adult age estimation task, the mapping method is as follows:

[0143]

[0144] For the task of estimating children's age, the mapping method is as follows:

[0145]

[0146] This approach can limit the network regression output to an age range that conforms to the task definition, reduce extreme prediction values, and improve training and inference stability.

[0147] After the age estimation network is built, it can be trained using the following steps:

[0148] S51, Construct training and validation datasets, where each sample includes a preprocessed twelve-lead ECG signal and a corresponding real-age label.

[0149] S52, input the electrocardiogram signal into the age estimation network to obtain the age estimate;

[0150] S53 uses age regression loss to optimize model parameters.

[0151] Preferably, the mean absolute error loss is used as the basic loss function:

[0152]

[0153] in, Indicates the first Predicted age for each sample Indicate your real age. Indicates the number of samples in the batch.

[0154] In other implementations, mean squared error loss, Huber loss, age-weighted loss, or a combination thereof can be used. For datasets with uneven age distribution, sample weights can be set according to the frequency of age intervals to improve the prediction performance of samples in a few age groups.

[0155] Preferably, before calculating the loss, the model output and the real age label are organized into the same dimension, for example, both in the form of N×1, to avoid the deep learning framework broadcasting in the batch dimension, which would lead to incorrect cross-sample loss calculation.

[0156] After training, the twelve-lead electrocardiogram signal of the subject to be estimated is acquired, and the preprocessed signal is input into the trained age estimation network to obtain the age estimate.

[0157] In a preferred embodiment, an adult age estimation model and a child age estimation network can be trained separately, and the corresponding network and age mapping strategy can be selected based on the age group of the subject or the source of the task. The final output age estimation result can be used for electrocardiogram age assessment, physiological age analysis, cardiovascular risk screening, or health management systems.

[0158] An age estimation system based on lead-sensing state space is provided. The system includes a signal acquisition unit, a preprocessing unit, a lead-sensing coding unit, a multi-scale feature extraction unit, a cross-lead attention unit, a state space temporal modeling unit, an age regression unit, an interval mapping unit, and a result output unit.

[0159] The signal acquisition unit is used to acquire twelve-lead electrocardiogram signals;

[0160] The preprocessing unit is used to perform matrix organization, fixed-length truncation, standardization, bandpass filtering, and outlier processing on the twelve-lead ECG signal;

[0161] The lead sensing coding unit is used to extract shallow lead flow features, generate lead adaptive weights, and output the input features of the backbone network.

[0162] The multi-scale feature extraction unit is used to sequentially convert the input features of the backbone network into multi-level residual temporal features and deep multi-scale temporal features;

[0163] The translead attention unit is used to perform translead feature interaction based on shallow lead flow features and deep multi-scale temporal features, and output lead-guided temporal representations;

[0164] The state-space temporal modeling unit is used to perform long-range dependency modeling on the lead guidance timing representation and output long-range temporal representation.

[0165] The age regression unit is used to output predicted age values ​​based on long-range time-series characterization.

[0166] The interval mapping unit is used to perform the sigmoid function and age interval mapping on the age prediction value to obtain the age estimate;

[0167] The results output unit is used to output the age estimate and its corresponding age group mapping results.

[0168] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0169] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An age estimation network based on lead-sensing state space, characterized in that, include: The preprocessing module is used to preprocess the twelve-lead ECG signal to obtain a standardized input signal; The lead sensing front-end module is used to extract lead-level shallow features. The shallow lead flow features are mapped to lead key features by the lead encoder. At the same time, lead weights are generated based on the shallow lead flow features, and the weighted lead features are mapped to the backbone network input features. The multi-level residual U-shaped feature extraction module is used to process the input features of the backbone network to obtain multi-level residual temporal features; A multi-scale dense feature fusion module is used to process multi-level residual temporal features to obtain deep multi-scale temporal features; The translead attention module can utilize deep multi-scale temporal features and lead key value features to obtain lead guidance temporal representations; The state-space timing modeling module is used to perform long-range dependency modeling on the lead guidance timing representation to obtain long-range timing representation. The regression output module is used to process long-range time series representations to obtain age prediction values; The interval mapping module performs interval mapping on the predicted age values ​​and outputs the estimated age values.

2. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The preprocessing module performs preprocessing on the twelve-lead ECG signal, including: The raw electrocardiogram data is reshaped or organized into a two-dimensional matrix with the shape of twelve leads multiplied by a preset number of sampling points; Extract a time series of a preset length from each lead signal; Z-score normalization is performed on each lead along the time dimension; Bandpass filtering is performed on each lead to suppress baseline drift, high-frequency noise, and power frequency interference; Replace abnormal, missing, or infinite values ​​in the preprocessed signal with finite values.

3. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The lead sensing front-end module includes: The depthwise convolution branch is used to perform one-dimensional convolution on each of the twelve leads to extract the local waveform features of each lead; The lead attention-gated branch is used to globally aggregate the features of each lead along the time dimension, and obtain the lead weights through fully connected mapping and normalization function; Pointwise convolutional branches are used to map the weighted lead features to a preset channel dimension to form the input features of the backbone network.

4. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The multi-level residual U-shaped feature extraction module includes several residual U-shaped blocks, which are connected sequentially and have different encoding depths. This allows for the step-by-step encoding, decoding, and downsampling of the backbone network input features, outputting multi-level residual time-series features containing local waveform information and contextual rhythm information. The residual U-shaped block includes: Input convolutional layers are used to perform local morphological enhancement on input features; Multi-level coding layers are used to progressively reduce temporal resolution and expand the receptive field through stride convolutions; Bottleneck layer, used to aggregate contextual information in low-resolution feature space; Multi-level decoding layers are used to restore temporal resolution and fuse corresponding coding layer features through deconvolution; Residual connections are used to add the decoded features to the block input features; The downsampling layer is used to output features with lower temporal resolution to the next layer of residual U-blocks.

5. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The multi-scale dense feature fusion module includes several multi-scale bottleneck units to process multi-level residual temporal features; The multi-scale bottleneck unit includes: One-dimensional pointwise convolution is used to compress and perform nonlinear transformations on the input channels; Multiple parallel one-dimensional convolutional branches with different dilation rates are used to perform multi-scale fusion of multi-level residual temporal features and output deep multi-scale temporal features. The channel splicing layer is used to splice the features obtained from convolutional branches with different dilation rates; Dense connection layers are used to stitch together the multi-scale features and the input features along the channel dimension.

6. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The specific steps for the translead attention module to obtain the lead-guided timing representation include: Use deep multi-scale temporal features as query input; The lead key features obtained by mapping the shallow lead flow features through the lead encoder are used as keys and values; Lead guidance timing characterization is obtained through multi-head attention calculation.

7. The age estimation network based on lead-sensing state space according to claim 6, characterized in that, The lead encoder includes at least two layers of one-dimensional convolution, normalization, and nonlinear activation layers. The lead encoder can map the shallow lead flow features that preserve the lead structure to the same channel dimension as the deep multi-scale temporal features to obtain lead key value features.

8. The age estimation network based on lead-sensing state space according to claim 1, characterized in that, The regression output module includes an adaptive pooling layer and a fully connected layer. The adaptive pooling layer aggregates long-range time series representations into a global feature vector, and the fully connected layer outputs an age prediction value based on the global feature vector.

9. The age estimation network based on lead-sensing state space according to claim 8, characterized in that, The interval mapping module performs interval mapping on the age prediction values, including: Applying the sigmoid function to the age prediction values ​​yields a normalized value between 0 and 1. When the object to be estimated belongs to the adult age estimation task, the normalized value is multiplied by a preset first age span and then a preset first age lower limit is added to map the normalized predicted value to the adult age range. When the object to be estimated belongs to the child age estimation task, the normalized value is multiplied by a preset second age span to map the normalized predicted value to the child age range.

10. An age estimation system based on lead-sensing state space, characterized in that, For implementing the age estimation network based on lead-sensing state space as described in any one of claims 1-9, the system comprises: The signal acquisition unit is used to acquire twelve-lead electrocardiogram signals; The preprocessing unit is used to perform matrix organization, fixed-length truncation, standardization, bandpass filtering, and outlier processing on the twelve-lead electrocardiogram signal. Lead sensing coding unit is used to extract shallow lead flow features, generate lead adaptive weights, and output backbone network input features; A multi-scale feature extraction unit is used to sequentially convert the input features of the backbone network into multi-level residual temporal features and deep multi-scale temporal features. The translead attention unit is used to perform translead feature interaction based on shallow lead flow features and deep multi-scale temporal features, and output lead-guided temporal representations. The state-space temporal modeling unit is used to perform long-range dependency modeling on the lead guidance temporal representation and output the long-range temporal representation. An age regression unit is used to output a predicted age value based on the long-range time series characterization. The interval mapping unit is used to perform the sigmoid function and age interval mapping on the age prediction value to obtain the age estimate; The result output unit is used to output the estimated age value and its corresponding age group mapping result.

Citation Information

Patent Citations

  • Biological age prediction method and system based on time-frequency fusion network

    CN122350723A

  • KR20240143480A