Multi-modal physiological signal processing method and wearable device for acute myocardial infarction auxiliary identification

CN122827701APending Publication Date: 2026-09-29GUANGDONG YITONG SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611194159.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明提供面向急性心肌梗死辅助识别的多模态生理信号处理方法及可穿戴设备,用以解决传统高精度模型算力、存储需求大,无法部署在可穿戴微控制器的问题,克服常规量化压缩抹除ST段关键特征、造成模型识别敏感度下降的缺陷,实现在院前、居家场景下,可穿戴设备端侧低功耗、实时、高敏感度地完成急性心肌梗死的辅助识别

Benefits of technology

[0010]本发明实施例提供的面向急性心肌梗死辅助识别的多模态生理信号处理方法,基于简化导联心电信号和光电容积脉搏波信号的预处理,得到规整的当前心电信号序列和当前脉搏波信号序列,依托双向状态空间网络对两类信号序列分别开展特征提取,依托该网络线性计算复杂度的特性,大幅降低整体运算量与内存占用,再结合重建网络由心电信号序列生成胸前导联空间补偿特征与不确定性度量,并依据不确定性度量完成可信度加权约束得到受约束的重建特征,弥补了简化导联缺失的空间信息;接着,通过跨模态注意力机制实现两类特征的跨模态关联与时延对齐得到对齐特征序列,再与重建特征拼接形成多模态融合特征,充分融合多模态有效信息以保障识别精度;最后由分类头输出急性心肌梗死风险概率值并生成风险提示信息,而整套神经网络模型经过ST段感知混合精度量化及ST段时间窗特征对齐蒸馏训练,打破了传统全局统一比特量化的局限,依据网络各层对ST段特征的敏感度差异化分配量化比特,同时通过特征对齐蒸馏守护ST段毫伏级细微电位偏移的关键诊断特征,避免压缩过程中核心特征丢失。既解决了传统高精度模型算力、存储需求大,无法部署在可穿戴微控制器的问题,又克服了常规量化压缩抹除ST段关键特征、造成模型识别敏感度下降的缺陷,最终实现了在院前、居家场景下,可穿戴设备端侧低功耗、实时、高敏感度地完成急性心肌梗死的辅助识别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122827701A_ABST
    Figure CN122827701A_ABST
Patent Text Reader

Abstract

The application provides a multi-modal physiological signal processing method and wearable device for auxiliary identification of acute myocardial infarction, which comprises: based on simplified lead electrocardio signals and photoplethysmography signals, pre-processing is carried out to obtain current electrocardio signal sequences and current pulse wave signal sequences; based on a bidirectional state space network, feature extraction is carried out on the current electrocardio signal sequences and the current pulse wave signal sequences to obtain electrocardio feature sequences and pulse wave feature sequences; based on the current electrocardio signal sequences and a reconstruction network, spatial compensation features of precordial leads and uncertainty measures are generated, and the spatial compensation features are subjected to credibility weighted constraint based on the uncertainty measures to obtain constrained reconstruction features. The application realizes low-power consumption, real-time and high-sensitivity auxiliary identification of acute myocardial infarction by wearable device in pre-hospital and home scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a multimodal physiological signal processing method and wearable device for the auxiliary identification of acute myocardial infarction. Background Technology

[0002] Acute myocardial infarction (AMI), especially ST-segment elevation myocardial infarction (STEMI), is a highly fatal cardiovascular emergency. Clinical treatment is heavily reliant on the reperfusion therapy time window; therefore, early, accurate, and real-time identification of STEMI in pre-hospital or home settings is crucial. However, in these scenarios, limitations in device size and ease of wear typically limit the acquisition of single-lead or simplified-lead ECG signals, often accompanied by photoplethysmography (PPG) signals.

[0003] Existing physiological signal processing methods for acute myocardial infarction typically rely on cloud servers or high-performance computing platforms. This involves complex preprocessing of the acquired raw ECG and pulse wave signals before inputting them into Transformer-based or deep convolutional neural network-based models for feature analysis. Due to the enormous computational complexity and storage requirements of such high-precision models, they cannot be directly deployed on wearable microcontroller hardware with strict power and memory constraints. To meet the needs of edge deployment, existing technologies often employ model compression techniques, such as globally uniform bit-width quantization methods. However, the diagnosis of acute myocardial infarction (especially STEMI) is highly dependent on the subtle millivolt-level potential shifts in the ST segment of the ECG signal. Existing "one-size-fits-all" quantization compression methods do not differentiate the sensitivity of each network layer to ST segment features, often treating weights containing critical ST segment information and redundant weights equally for low-bit quantization. This results in the erasure or distortion of subtle morphological features of the ST segment in the quantization error. As a result, the compression leads to the loss of key diagnostic features, which significantly reduces the model's sensitivity to acute myocardial infarction when running on resource-constrained wearable devices, failing to meet the reliability requirements for clinical auxiliary identification. Summary of the Invention

[0004] This invention provides a multimodal physiological signal processing method and wearable device for the auxiliary identification of acute myocardial infarction, which solves the problem that traditional high-precision models have high computing power and storage requirements and cannot be deployed on wearable microcontrollers. It overcomes the defects of conventional quantization compression that erases key features of the ST segment and causes a decrease in model recognition sensitivity. It enables low-power, real-time and high-sensitivity auxiliary identification of acute myocardial infarction on the wearable device side in pre-hospital and home scenarios.

[0005] In a first aspect, the present invention provides a multimodal physiological signal processing method for the auxiliary identification of acute myocardial infarction, comprising: Based on the simplified lead ECG signal and photoplethysmography pulse wave signal, the current ECG signal sequence and the current pulse wave signal sequence are obtained by preprocessing them respectively. Based on a bidirectional state-space network, features are extracted from the current electrocardiogram (ECG) signal sequence and the current pulse wave signal sequence to obtain ECG feature sequences and pulse wave feature sequences, respectively. Based on the current ECG signal sequence and the reconstruction network, spatial compensation features and uncertainty measures of the precordial leads are generated, and the spatial compensation features are subjected to confidence weighting constraints based on the uncertainty measures to obtain constrained reconstruction features. Based on the cross-modal attention mechanism, the ECG feature sequence and pulse wave feature sequence are cross-modal correlated and time-delay aligned to obtain the aligned feature sequence. The aligned feature sequence is then spliced ​​and fused with the reconstructed features to obtain the multimodal fused features. Based on the classification head, feature mapping is performed on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and risk warning information is generated based on the risk probability value; The neural network model, which consists of the bidirectional state space network, the reconstruction network, the cross-modal attention mechanism, and the classification head, is a lightweight model obtained by training through ST segment perceptual mixed precision quantization and ST segment time window feature alignment distillation.

[0006] In a second aspect, the present invention also provides a multimodal physiological signal processing system for the auxiliary identification of acute myocardial infarction, applied to the multimodal physiological signal processing method for the auxiliary identification of acute myocardial infarction as described in the first aspect; the multimodal physiological signal processing system for the auxiliary identification of acute myocardial infarction includes: The signal preprocessing module is used to preprocess the simplified lead ECG signal and the photoplethysmography pulse wave signal respectively to obtain the current ECG signal sequence and the current pulse wave signal sequence. The dual-branch feature extraction module is used to extract features from the current electrocardiogram signal sequence and the current pulse wave signal sequence based on a bidirectional state space network, so as to obtain an electrocardiogram feature sequence and a pulse wave feature sequence. The constrained reconstruction enhancement module is used to generate spatial compensation features and uncertainty measures of the precordial leads based on the current electrocardiogram signal sequence and the reconstruction network, and to apply a confidence-weighted constraint to the spatial compensation features based on the uncertainty measures to obtain constrained reconstruction features. The multimodal feature fusion module is used to perform cross-modal association and time delay alignment on the electrocardiogram feature sequence and pulse wave feature sequence based on the cross-modal attention mechanism to obtain the aligned feature sequence, and then splice and fuse the aligned feature sequence with the reconstructed features to obtain the multimodal fused features; The risk identification and alert module is used to perform feature mapping on the multimodal fusion features based on the classification head to obtain the risk probability value of acute myocardial infarction, and generate risk alert information based on the risk probability value; The neural network model, which consists of the bidirectional state space network, the reconstruction network, the cross-modal attention mechanism, and the classification head, is a lightweight model obtained by training through ST segment perceptual mixed precision quantization and ST segment time window feature alignment distillation.

[0007] Thirdly, the present invention also provides a wearable device, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby realizing the multimodal physiological signal processing method for assisted identification of acute myocardial infarction as described above.

[0008] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program, which, when executed by a processor, implements the multimodal physiological signal processing method for assisted identification of acute myocardial infarction as described above.

[0009] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal physiological signal processing method for assisted identification of acute myocardial infarction as described above.

[0010] The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction provided in this invention preprocesses simplified lead ECG signals and photoplethysmography (PPG) pulse wave signals to obtain regular current ECG signal sequences and current pulse wave signal sequences. It then uses a bidirectional state-space network to extract features from both signal sequences, leveraging the linear computational complexity of this network to significantly reduce overall computational load and memory consumption. Next, a reconstruction network generates precordial lead spatial compensation features and uncertainty metrics from the ECG signal sequences. Based on the uncertainty metrics, a confidence-weighted constraint is applied to obtain constrained reconstructed features, compensating for the spatial information missing in the simplified leads. Finally, through cross-modal... The state attention mechanism achieves cross-modal association and temporal delay alignment of two types of features to obtain aligned feature sequences, which are then concatenated with reconstructed features to form multimodal fusion features. This fully integrates effective multimodal information to ensure recognition accuracy. Finally, the classification head outputs the probability value of acute myocardial infarction risk and generates risk warning information. The entire neural network model is trained through ST segment perception hybrid precision quantization and ST segment time window feature alignment distillation, breaking the limitations of traditional global unified bit quantization. Quantization bits are allocated differently according to the sensitivity of each layer of the network to ST segment features. At the same time, feature alignment distillation protects the key diagnostic features of subtle millivolt-level potential shifts in the ST segment, avoiding the loss of core features during compression. This solves the problem of high computing power and storage requirements of traditional high-precision models, which cannot be deployed on wearable microcontrollers, and overcomes the defects of conventional quantization compression that erases key ST segment features, causing a decrease in model recognition sensitivity. Ultimately, it enables low-power, real-time, and highly sensitive auxiliary recognition of acute myocardial infarction on wearable devices in pre-hospital and home scenarios. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the multimodal physiological signal processing method for the auxiliary identification of acute myocardial infarction provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the multimodal physiological signal processing system for the auxiliary identification of acute myocardial infarction provided in an embodiment of the present invention; Figure 3 An embodiment diagram of a wearable device provided in this invention; Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0014] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0015] See Figure 1 , Figure 1 This is a flowchart illustrating the multimodal physiological signal processing method for assisting in the identification of acute myocardial infarction provided by this invention. In this embodiment, the execution entity of the multimodal physiological signal processing method for assisting in the identification of acute myocardial infarction is a signal processing system. Therefore, the multimodal physiological signal processing method for assisting in the identification of acute myocardial infarction includes a neural network model composed of a bidirectional state-space network, a reconstruction network, a cross-modal attention mechanism, and a classification head. This neural network model is a lightweight model obtained through ST-segment perception mixed precision quantization and ST-segment time window feature alignment distillation training. The specific steps are as follows: Step 10: Preprocess the simplified lead ECG signal and photoplethysmography pulse wave signal respectively to obtain the current ECG signal sequence and the current pulse wave signal sequence.

[0016] Optionally, the signal processing system receives raw physiological signals from sensors in the wearable device, including simplified lead electrocardiogram (ECG) signals and photoplethysmography (PPG) signals. The simplified lead ECG signals are ECG signals acquired through a single-lead ECG electrode or fewer than twelve standard leads. The PPG signals are light signals reflecting changes in peripheral blood flow volume acquired through a multi-wavelength photoelectric sensor.

[0017] After receiving these two raw physiological signals, the signal processing system preprocesses them, specifically by performing bandpass filtering, baseline drift removal, and lead-by-lead normalization. This eliminates noise interference and standardizes data dimensions, ultimately yielding the current ECG signal sequence and the current pulse wave signal sequence for subsequent network input.

[0018] Specifically, in the bandpass filtering stage, the signal processing system applies a zero-phase bandpass filter with a cutoff frequency of 0.5 Hz to 40 Hz to the current ECG signal sequence. The lower cutoff frequency of 0.5 Hz is used to suppress low-frequency baseline drift caused by respiration or poor electrode contact, while preserving the ST segment slowing component, which characterizes myocardial ischemia. The upper cutoff frequency of 40 Hz is used to suppress high-frequency noise from electromyography interference and power line interference. Zero-phase filtering is used to avoid the phase distortion introduced by traditional causal filtering, which would disrupt the temporal correspondence of the ST segment and pulse wave characteristic points. Simultaneously, a zero-phase bandpass filter with a cutoff frequency of 0.5 Hz to 8 Hz is applied to the current pulse wave signal sequence to preserve the morphological characteristics of the pulse wave, such as the main wave and dicrotic wave, while filtering out high-frequency noise.

[0019] During the baseline drift removal process, the signal processing system uses median filtering or wavelet decomposition to estimate the low-frequency baseline trend term in the current ECG signal sequence and subtracts the trend term from the original filtered signal to ensure that the ST segment potential is comparable with the isoelectric line as a reference.

[0020] During the lead-by-lead normalization phase, the signal processing system performs Z-score normalization on the baseline-removed current ECG signal sequence and the current pulse wave signal sequence, respectively. For any signal sequence x, its normalized sequence... Calculated using the following formula: Where x represents the signal sequence composed of the original sampling points; This represents the mean of the signal sequence within the sliding window; This represents the standard deviation of the signal sequence within the sliding window; To prevent small constants with a denominator of zero (e.g., 10) -8 ).

[0021] Furthermore, in order to accurately locate the ST segment time window required for subsequent steps, the signal processing system also needs to determine the start and end times of the ST segment. This is done by first detecting the peak time of the R wave of the QRS complex and the end time of the J point of the QRS complex. Subsequently, the ST segment window is adaptively determined based on the interval (RR) between adjacent R waves. ST segment start time. and the end time Determined by the following formula: , ;in, The time interval between two adjacent R-wave peaks; The scaling factor, set according to Bazett's heart rate correction principle, ranges from 0.3 to 0.5. The default value for a normal heart rate is 0.4. It is adjusted downwards for tachycardia and upwards for bradycardia, allowing the ST segment window width to adaptively expand and contract with heart rate changes, thus covering a complete and effective ST segment region. This set of ST segment windows is denoted as [missing value]. .

[0022] Finally, the signal processing system slices the preprocessed signal into sliding window segments of fixed duration (e.g., 10 seconds, containing several complete cardiac cycles) to form sample tensors that are fed into the neural network.

[0023] For example, suppose the signal processing system acquires a 10-second single-lead I ECG signal at a sampling rate of 500 Hz, and simultaneously acquires a multi-wavelength red light photoplethysmography (PPG) pulse wave signal at a sampling rate of 250 Hz. During preprocessing, the ECG signal is first subjected to a 0.5-40 Hz zero-phase bandpass filter, and the PPG signal is subjected to a 0.5-8 Hz zero-phase bandpass filter. Then, wavelet decomposition is used to remove low-frequency baseline drift from the ECG signal. Finally, the mean of the 10-second ECG signal is calculated. millivolts, standard deviation millivolts; mean value of pulse wave signal Volts, standard deviation Submerged. Assume. Each sampling point is standardized to obtain the standardized electrocardiogram signal sequence and pulse wave signal sequence.

[0024] Subsequently, the first one was detected. The wave time is 0.8 seconds, and the J point time is... The interval is 0.9 seconds, and the adjacent RR interval is... It is 0.8 seconds. Let the scaling factor be... Then the end time of ST segment Seconds. Therefore, the first heartbeat A period of time window The intervals were determined to be [0.9 seconds, 1.22 seconds]. All subsequent heartbeats will be determined according to this rule. Segment window.

[0025] Step 20: Based on the bidirectional state-space network, feature extraction is performed on the current ECG signal sequence and the current pulse wave signal sequence to obtain the ECG feature sequence and the pulse wave feature sequence.

[0026] Optionally, the signal processing system inputs the current electrocardiogram signal sequence and the current pulse wave signal sequence into two bidirectional state-space network branches with identical structures but independent parameters. Each branch is composed of multiple stacked bidirectional state-space basic blocks, which model long-sequence physiological signals with linear computational complexity and extract deep spatiotemporal features.

[0027] This includes: performing forward and reverse scanning on the current ECG signal sequence and the current pulse wave signal sequence respectively based on a selective state-space mechanism, and adjusting the discretization step size, input matrix, and output matrix based on the current input segment; Optionally, in each bidirectional state-space basic block, the signal processing system first extracts local features from the input signal. The input signal passes through a depth-separable one-dimensional convolutional layer to capture local morphological changes between adjacent sampling points, and then undergoes nonlinear mapping using the SiLU activation function. The expression for the SiLU activation function is as follows: ,in, This is the Sigmoid function. Next, the signal processing system performs forward and reverse scanning of the signal based on a selective state-space mechanism. The core of the selective state-space mechanism lies in making the discretization step size, input matrix, and output matrix functions of the current input signal, thereby achieving adaptive focusing on different cardiac beat morphologies (such as normal sinus beats and ischemic beats). Forward scanning refers to calculating step-by-step from the beginning to the end of the signal, while reverse scanning refers to calculating step-by-step from the end to the beginning of the signal.

[0028] Furthermore, for the k-th time step in the input sequence, the signal processing system dynamically calculates the discretization step size using the following formula. Input projection vector and output projection vector : ; ; ; in, This is the input feature vector at the k-th time step; , , The projection matrix is ​​learnable; It is the bias vector; To smooth the activation function, the step size is kept positive. This mechanism allows the model to selectively adjust the step size to reinforce memory when encountering critical diagnostic segments (such as the ST segment and QRS complex), while compressing the step size to forget irrelevant information in redundant segments.

[0029] Subsequently, the signal processing system updates the hidden state using discretized state-space equations. That is, the hidden state. The updates follow the following recursive relationship: ; ; in, Let N be the hidden state vector at step k, representing the compressed memory of historical information; The discretized state transition matrix is ​​obtained by discretizing the continuous-time state transition matrix A using the zero-order preservation method. ; The input matrix after discretization is usually approximated as follows: ; The output feature at step k; This represents the pass-through jump coefficient. Furthermore, to achieve efficient computation, the state transition matrix A is diagonalized, reducing matrix multiplication to element-wise scalar operations and thus lowering the computational power consumption of the edge hardware.

[0030] The forward and reverse scan outputs are concatenated along the feature dimension, and then processed by residual connection and layer normalization to obtain the ECG feature sequence and pulse wave feature sequence.

[0031] Optionally, the signal processing system performs a forward scan from beginning to end and a reverse scan from end to beginning on the sequence to obtain a forward output sequence. and reverse output sequence Then, the two are concatenated along the feature dimension, and the output of the basic block is obtained through residual connections and layer normalization. This ensures that the model can simultaneously utilize past and future contextual information at a given moment through bidirectional scanning. After stacking multiple bidirectional state-space basic blocks, the signal processing system finally obtains a high-dimensional ECG feature sequence E and a pulse wave feature sequence G. Compared to the Transformer architecture, the computational complexity of the obtained ECG feature sequence E and pulse wave feature sequence G in the bidirectional state-space network increases linearly with the sequence length L. rather than quadratic growth This meets the stringent requirements of wearable microcontroller hardware for low power consumption and real-time performance.

[0032] For example, suppose the current ECG signal sequence has a length of 5000 sampling points (10 seconds @ 500Hz) and a feature dimension of 64. The processing steps are as follows: 1. The signal processing system inputs the sequence into the first layer of bidirectional state space basic blocks. Local features are first extracted using a depthwise separable one-dimensional convolution with a kernel size of 3. 2. For the 1000th sampling point, the input vector x 1000 Enter the selective mechanism module. Using the learned weight matrix... Calculate the step size Seconds. Since this point is located in the elevated region of segment ST, the model automatically adjusts its parameters to enhance its memory of this region.

[0033] 3. The signal processing system uses a diagonalized state transition matrix. and dynamic calculation Update hidden state The forward scan is calculated from point 1 to point 5000, and the reverse scan is calculated from point 5000 to point 1.

[0034] 4. The signal processing system will output a positive signal. and reverse output The concatenation yields a feature vector with dimension 128. After layer normalization and residual connection, the vector is output to the next layer.

[0035] 5. After stacking four such basic blocks, the final output is an ECG feature sequence with a dimension of 128 and a length of 5000. Similarly, the same processing is applied to the pulse wave signal sequence to obtain the pulse wave characteristic sequence. .

[0036] Step 30: Based on the current ECG signal sequence and the reconstruction network, generate spatial compensation features and uncertainty measures for the precordial leads, and apply confidence-weighted constraints to the spatial compensation features based on the uncertainty measures to obtain constrained reconstruction features.

[0037] Optionally, the reconstruction network is a generative bypass network whose task is to estimate the ST-T vector representation of the precordial leads from a limited set of input leads. To compensate for the lack of crucial spatial information carried by the precordial leads in the simplified leads, the signal processing system processes the current ECG signal sequence through the reconstruction network.

[0038] The signal processing system inputs the current ECG signal sequence into the reconstruction network, which simultaneously outputs the reconstructed mean represented by the ST-T vectors of the precordial leads. (i.e., spatial compensation features) and reconstruction variance (i.e., uncertainty measurement). The reconstruction network models heteroscedastic uncertainty, enabling it to assess the reliability of its own predictions while reconstructing waveforms.

[0039] During the training phase, the reconstruction network is optimized using Gaussian negative log-likelihood loss, but during the inference phase, the signal processing system directly utilizes the mean and variance of the output for subsequent processing. Furthermore, the reconstructed mean... and variance It can be expressed by the following formula: ; in, The input is the current ECG signal sequence; To rebuild the network; The parameters for reconstructing the network.

[0040] This includes: determining the uncertainty gating weight vector based on the reconstruction variance of the precordial leads; Optionally, the signal processing system determines the uncertainty gating weight vector based on the reconstruction variance. The larger the variance, the less reliable the reconstruction result (it tends towards the mean of the training set, and there is a risk of "regression to the mean"), therefore, a smaller weight is assigned. Uncertainty-gated weight vector. Calculated using the following formula: ; in, This is an adjustable hyperparameter used to control the intensity of the penalty for uncertainty, with a value range of 5-20; This is a gated weight vector that is negatively correlated with variance, and its value ranges from 0 to 1.

[0041] The constrained reconstruction features are obtained by performing element-wise multiplication of the reconstructed mean of the precordial leads using an uncertainty-gated weight vector.

[0042] Optionally, the signal processing system is based on an uncertainty-gated weight vector. Mean value of reconstruction of precordial leads By performing element-wise multiplication with gated weighting, constrained reconstructed features are obtained. : ; in, This indicates element-wise multiplication. Reliable reconstruction features are fully utilized in this way, while unreliable reconstruction features (high variance components) are automatically suppressed, thus compensating for spatial information while avoiding misleading reconstructions.

[0043] For example, suppose that after the current ECG signal sequence is processed by the reconstruction network, for a certain sample, the output is the reconstructed mean of the ST segment amplitude in lead V2 of the precordial leads. millivolts, corresponding reconstruction variance (millivolt squared). Assume hyperparameters. .

[0044] The signal processing system first calculates the gating weights: Since the variance is small and the weight is close to 1, it indicates that the system has relatively high confidence in the reconstruction result. If the reconstruction variance of another sample is large, for example... ,but At this point, the weight is extremely small, almost completely masking the reconstructed feature. Then, element-wise multiplication is performed: millivolts. Repeat this process for all estimated precordial leads to obtain the complete constrained reconstructed feature vector. .

[0045] Step 40: Based on the cross-modal attention mechanism, perform cross-modal association and time delay alignment on the ECG feature sequence and pulse wave feature sequence to obtain the aligned feature sequence. Then, splice and fuse the aligned feature sequence with the reconstructed features to obtain the multimodal fusion feature.

[0046] Optionally, the signal processing system uses the ECG feature sequence as the query and the pulse wave feature sequence as the key and value to construct a cross-modal attention mechanism. By calculating the correlation between ECG features and pulse wave features in the time dimension, the system automatically learns the pulse conduction time delay between the ECG signal R wave and the pulse wave peak. A dynamic time delay mapping table is constructed, the optimal time delay offset is selected, and the pulse wave feature sequence is time-aligned accordingly. Finally, the aligned feature sequence is output, as detailed in steps 401 to 405.

[0047] Optionally, the signal processing system fuses the alignment feature sequence with the constrained reconstruction feature to integrate diagnostic information from different sources. The alignment feature sequence contains time-delay-corrected complementary information from ECG and pulse waves, reflecting the combined state of cardiac electrical activity and mechanical pumping function; the constrained reconstruction feature supplements the spatial information of the precordial leads lost due to lead absence.

[0048] Therefore, the signal processing system uses a concatenation method to combine these two types of features along the feature dimension. Specifically, for each time step or the feature vector after global aggregation, the signal processing system concatenates the feature vector of the aligned feature sequence with the constrained reconstructed feature vector end to end to form a unified multimodal fusion feature. Through this concatenation and fusion, the multimodal fusion feature retains the temporal dynamic correlation of the original bimodal signal and introduces the spatial dimension information of the reconstruction.

[0049] For example, suppose the obtained aligned feature sequence, after global average pooling, yields a feature vector with dimension 256. Constrained reconstruction features obtained in step 30 It is a 64-dimensional vector (corresponding to several estimated key points in the precordial leads). The signal processing system performs a splicing operation, combining the 256-dimensional vector... With 64-dimensional The connection yields a new multimodal fusion feature vector with dimensions 320 (256+64). The vector This is the final feature representation input into the classification head.

[0050] Step 50: Based on the classification head, perform feature mapping on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and generate risk warning information based on the risk probability value.

[0051] Optionally, the classification head consists of a global average pooling layer (if the preceding layer is not pooled), a linear fully connected layer, and a Sigmoid activation function, used to map high-dimensional features to risk probability values ​​for acute myocardial infarction.

[0052] Signal processing systems fuse multimodal features The input is fed into the classification head. First, global average pooling is performed on the multimodal fusion features to compress the sequence dimension and extract global statistical features. Then, a linear transformation maps the features to a scalar space, and the Sigmoid function compresses them to the [0, 1] interval, yielding the STEMI risk probability p. Its calculation formula is as follows: ; in, This is a global average pooling operation; This is the weight vector of the classification head; For bias terms; It is the Sigmoid activation function. ; This is the output probability value for the risk of acute myocardial infarction.

[0053] The signal processing system generates risk warning information based on the obtained risk probability value p. This risk warning information is for medical personnel's reference and does not constitute a final disease diagnosis. For example, when p exceeds a preset threshold (e.g., 0.5), a "high risk" warning is generated; otherwise, a "low risk" warning is generated. Alternatively, the probability value p and its corresponding confidence interval can be directly output so that clinicians can make a comprehensive judgment in conjunction with other clinical information.

[0054] Specifically, the method for allocating the mixed-precision quantization bit width of the ST segment in a neural network model is as follows: the quantization bit width allocated to each network layer is determined based on the sensitivity of each network layer to the ST segment discrimination loss; where the sensitivity is positively correlated with the quantization bit width.

[0055] Optionally, during quantization, the signal processing system (in the offline training phase) determines the quantization bit width allocated to each network layer based on the sensitivity of each network layer to the ST segment discrimination loss. That is, the sensitivity... For the first Before and after applying quantization perturbation to the layer, only within the ST segment time window The change in the discrimination loss calculated internally. Higher sensitivity requires a larger allocated quantization bit width. The higher the value, the better. Bit allocation follows these principles: ; in, and These are the preset minimum and maximum bit widths (e.g., 4 bits and 8 bits), respectively. For the first ST segment sensitivity of the layer; This represents the maximum sensitivity across all layers. This indicates rounding to the nearest integer.

[0056] Furthermore, during distillation training, the total loss function includes classification cross-entropy loss, temperature-based soft-label distillation loss, and alignment loss of intermediate features between the student and teacher networks within the ST segment time window. Specifically, the ST segment feature alignment loss... Forced student network in ST segment window The intermediate features approximate the teacher network using the following formula: ; in, This is the set of sampling points for segment ST; and These are the intermediate features of the student network and the teacher at step k, respectively. For linear projection of the alignment dimension; The norm is L2 squared. Through this collaborative training, the lightweight model deployed on the edge maintains low storage footprint and low power consumption while maximizing the ability to discriminate subtle ST segment elevation features, thus achieving high-precision auxiliary identification of acute myocardial infarction.

[0057] For example, suppose the signal processing system completes forward inference and obtains multimodal fused features F. The linear layer weights in the classification head... Perform a dot product with F, and add a bias. The logits value is 2.5. The risk probability is calculated using the Sigmoid function. . determination If so, a risk warning message of "high risk of acute myocardial infarction" is generated and sent to medical staff through the display screen or wireless module of the wearable device.

[0058] This invention, based on the preprocessing of simplified lead ECG signals and photoplethysmography (PPG) pulse wave signals, yields regular current ECG and PPG signal sequences. A bidirectional state-space network is used to extract features from both signal sequences. Leveraging the linear computational complexity of this network, the overall computational load and memory usage are significantly reduced. A reconstruction network then generates precordial lead spatial compensation features and uncertainty metrics from the ECG signal sequences. Based on the uncertainty metrics, a confidence-weighted constraint is applied to obtain constrained reconstructed features, compensating for the spatial information missing in the simplified leads. Finally, a cross-modal attention mechanism is used to achieve cross-modal attention between the two types of features. Modal correlation and time delay alignment yield an aligned feature sequence, which is then concatenated with reconstructed features to form a multimodal fusion feature. This fully integrates effective multimodal information to ensure recognition accuracy. Finally, the classification head outputs an acute myocardial infarction risk probability value and generates risk warning information. The entire neural network model undergoes ST-segment perception hybrid precision quantization and ST-segment time window feature alignment distillation training, breaking the limitations of traditional globally unified bit quantization. Quantization bits are allocated differently based on the sensitivity of each network layer to ST-segment features. Simultaneously, feature alignment distillation protects the key diagnostic features of subtle millivolt-level potential shifts in the ST segment, preventing the loss of core features during compression. This solves the problems of high computational and storage requirements of traditional high-precision models, making them unsuitable for deployment on wearable microcontrollers. It also overcomes the shortcomings of conventional quantization compression, which erases key ST-segment features and reduces model recognition sensitivity. Ultimately, it enables low-power, real-time, and highly sensitive auxiliary recognition of acute myocardial infarction on wearable devices in pre-hospital and home settings.

[0059] Furthermore, the multimodal physiological signal processing system for assisting in the identification of acute myocardial infarction provided by the present invention will be described below. The multimodal physiological signal processing system for assisting in the identification of acute myocardial infarction described below can be referred to in correspondence with the multimodal physiological signal processing method for assisting in the identification of acute myocardial infarction described above.

[0060] Optionally, the processes of steps 401 to 405 include: Step 401: Based on the R-wave time index position extracted from the ECG feature sequence, construct a binary time sequence mask vector, and perform element-by-element feature filtering on the binary time sequence mask vector and the ECG feature sequence to obtain a sparse ECG time sequence feature sequence.

[0061] Optionally, the signal processing system locates key trigger points of cardiac electrical activity from the electrocardiogram (ECG) feature sequence, specifically the time index position corresponding to the R-wave peak. Since the pathophysiological changes in acute myocardial infarction mainly manifest as abnormalities in ventricular depolarization and repolarization, and the R-wave marks the beginning of ventricular depolarization, it serves as the benchmark anchor point for subsequent ST-segment morphological analysis and pulse wave conduction time calculation. Therefore, to reduce redundancy in subsequent cross-modal calculations and highlight key timing information, the signal processing system constructs a binary timing mask vector with the same length as the ECG feature sequence.

[0062] The construction rules are as follows: at the position corresponding to the index of the R-wave peak time in the ECG feature sequence, the element value of the mask vector is set to 1; at all other positions not at the R-wave peak time, the element value of the mask vector is set to 0.

[0063] Subsequently, the signal processing system performs element-wise multiplication of the binary time-series mask vector with the original ECG feature sequence. Through this multiplication operation, features at non-R-wave moments are set to zero or significantly suppressed, retaining only the feature information of the R-wave moments and their adjacent key windows, thus obtaining a sparse ECG time-series feature sequence. This sparsification process not only reduces the amount of data involved in subsequent similarity calculations and lowers computational power consumption, but also forces the model to focus on the onset of electrical activity.

[0064] For example, suppose the ECG feature sequence output in step 20 is 1000 time steps long and has a sampling rate of 500 Hz. The signal processing system uses a peak detection algorithm to identify that the sequence contains 5 complete heartbeats, with the R-wave peaks appearing at the indices of the 100th, 300th, 500th, 700th, and 900th time steps, respectively.

[0065] The signal processing system initializes a binary timing mask vector of length 1000, with all initial values ​​being 0. The values ​​at indices 100, 300, 500, 700, and 900 are changed to 1, while the remaining positions remain 0. This mask vector is then element-wise multiplied with an ECG feature sequence of dimension [1000, 64]. The resulting sparse ECG timing feature sequence retains the original feature data only in rows 100, 300, 500, 700, and 900, while the feature data in the remaining rows are all zero vectors.

[0066] Step 402: Based on the preset range of pulse conduction time variation, construct a multi-scale time-delay convolution kernel group containing multiple displacement steps, and perform multi-scale time-delay mapping on the pulse wave feature sequence based on the multi-scale time-delay convolution kernel group to obtain multi-time-delay candidate pulse wave feature tensors.

[0067] Optionally, since there is a physiological pulse conduction time between the electrocardiogram signal (electrical activity) and the photoplethysmography pulse wave signal (mechanical blood flow response), and this time fluctuates within a certain range due to the influence of heart rate, vascular elasticity and myocardial contractility, the signal processing system also needs to perform time delay extension on the pulse wave characteristic sequence to cover all possible physiological time delay situations.

[0068] Specifically, the signal processing system first presets the range of pulse conduction time based on clinical physiological knowledge, for example, setting the minimum conduction time to 100 milliseconds and the maximum conduction time to 300 milliseconds. Combining this with the signal sampling rate, the system converts this time range into a discrete set of displacement steps. For example, if the sampling rate is 250 Hz, then 100 milliseconds corresponds to 25 sampling points, and 300 milliseconds corresponds to 75 sampling points. The signal processing system constructs a multi-scale time-delay convolutional kernel group, which consists of multiple one-dimensional depthwise separable convolutional kernels. Each kernel has a different displacement step size (i.e., a different delay), and the kernel weights are fixed as unit pulses or simple smooth windows, aiming to achieve time-shifting of features rather than complex nonlinear transformations.

[0069] Subsequently, the signal processing system performs parallel convolution on the pulse wave feature sequence using the constructed multi-scale time-delay convolution kernel group. The process is as follows: for each displacement step, the convolution operation shifts the pulse wave feature sequence along the time axis by the corresponding step, generating a time-delay candidate feature sequence. All time-delay candidate feature sequences generated by different displacement steps are stacked on a newly added "time-delay dimension" to form a multi-time-delay candidate pulse wave feature tensor. This tensor's dimensions include the time step, the feature dimension, and the time-delay candidate dimension, where the time-delay candidate dimension represents multiple possible pulse propagation time hypotheses.

[0070] For example, assuming the pulse wave feature sequence has a length of 1000 time steps and a feature dimension of 64, and the preset pulse conduction time range is 100 to 200 milliseconds with a sampling rate of 250 Hz, then the corresponding displacement step size range is 25 to 50 sampling points.

[0071] The signal processing system constructs a multi-scale time-delay convolutional kernel group containing 26 kernels with different shift steps (25, 26, ..., 50). These 26 kernels are used to convolve the pulse wave feature sequences. For example, when using a kernel with a shift step of 25, the features at the first time step of the original sequence are mapped to the 26th time step of the new sequence (the first 25 steps are padded with zeros or truncated); when using a kernel with a shift step of 50, the features at the first time step of the original sequence are mapped to the 51st time step of the new sequence. Stacking these 26 feature sequences mapped with different time delays yields a multi-time-delay candidate pulse wave feature tensor with dimensions [1000, 64, 26]. The third dimension, 26, represents 26 different time-delay hypotheses.

[0072] Step 403: Based on the ECG time-series feature sequence and the multi-time-delay candidate pulse wave feature tensor, calculate the cosine similarity matrix under each time-delay candidate to obtain the time-delay feature correlation cube.

[0073] Optionally, the signal processing system traverses each time-delay slice (i.e., the pulse wave feature sequence corresponding to each specific displacement step) in the multi-time-delay candidate pulse wave feature tensor. For each time-delay slice, its cosine similarity with the sparsed ECG time-series feature sequence in the feature space is calculated. This cosine similarity measures the similarity between the two vectors by calculating the cosine of the angle between them, with a value ranging from -1 to 1. The closer the value is to 1, the more consistent the directions are, and the higher the similarity.

[0074] During the calculation, the signal processing system performs L2 norm normalization on the sparse ECG time-series feature sequence and the pulse wave feature sequence of the current time-delay slice, and then calculates their dot product. Since the ECG feature sequence is already sparse (with values ​​only at the R-wave time), this calculation essentially assesses the correlation strength between the electroactivity characteristics at the R-wave time and the hemodynamic characteristics after a specific time delay. The above calculation is repeated for all time-delay slices, and the results are combined according to the time step and time delay candidate dimensions to form a three-dimensional time-delay feature correlation cube. The three dimensions of this cube are: the time step dimension, the time delay candidate dimension, and the similarity score dimension (usually 1, i.e., scalar similarity). This cube visually demonstrates the electromechanical coupling strength under different assumed time delays at each time point.

[0075] For example, continuing with the embodiment of step 402, the sparse ECG timing feature sequence dimension is [1000, 64], and the multi-delay candidate pulse wave feature tensor dimension is [1000, 64, 26].

[0076] The signal processing system extracts the first slice (corresponding to a displacement of 25 steps) from the candidate time delay dimensions, obtaining a pulse wave feature sequence with dimensions [1000, 64]. The cosine similarity between this sequence and the sparsed ECG time-series feature sequence is calculated at each time step. Since the ECG sequence is non-zero only at the R-wave time, significant similarity values ​​are calculated only at the R-wave time; similarity at other times is close to 0 or at noise levels. This operation is repeated for all 26 time-delay slices, resulting in 26 sets of similarity sequences. These 26 sets of sequences are stacked to form a time-delay feature correlation cube with dimensions [1000, 26] (the feature dimension is omitted here because the similarity has been aggregated into a scalar). In this cube, if the actual pulse conduction time corresponds to a displacement of 30 steps, then in the row corresponding to the R-wave time, the similarity value of the 30th column should be significantly higher than that of the other columns.

[0077] Step 404: Based on the time delay feature correlation cube, perform maximum pooling on the time delay dimension to extract the optimal time delay index corresponding to each heartbeat cycle and construct a dynamic time delay mapping table.

[0078] Optionally, the signal processing system segments the time-delay feature correlation cube in the time dimension, using a single heartbeat cycle as the unit. For each heartbeat cycle (typically defined as the interval between adjacent R waves), a max-pooling operation is performed in the time-delay dimension. Max-pooling involves finding the time-delay index with the highest similarity score in the time-delay dimension within the time range covered by the heartbeat cycle. The displacement step size corresponding to this index is the most likely pulse conduction time within that heartbeat cycle.

[0079] By traversing all heartbeat cycles, the signal processing system obtains a series of optimal time delay indices, which are then arranged in chronological order to construct a dynamic time delay mapping table. This mapping table records the optimal time delay offset corresponding to each heartbeat cycle. Since myocardial ischemia may lead to weakened myocardial contractility or changes in vascular resistance, resulting in beat-by-beat fluctuations in pulse conduction time, this mapping table can reflect subtle changes in physiological state in real time.

[0080] For example, continuing with the embodiment of step 403, the time delay feature correlation cube has dimensions [1000, 26]. Assume the first heartbeat cycle covers time steps 100 to 299.

[0081] The signal processing system extracts a sub-block from time steps 100 to 299 of the cube, with dimensions [200, 26]. Max pooling is performed on the time delay dimension (the second dimension) of this sub-block. It is assumed that near time step 100 (the R-wave moment), the similarity value is highest at a time delay index of 5 (corresponding to a displacement of 29 steps), at 0.95, while the similarity of other time delay indices is below 0.8. The optimal time delay index for this heartbeat cycle is recorded as 5. This process is repeated for all subsequent heartbeat cycles, ultimately resulting in a dynamic time delay mapping table of length 5 (assuming 5 heartbeats), for example: [5, 6, 5, 7, 6], corresponding to the optimal time delay offsets for the 5 heartbeats.

[0082] Step 405: Align and fuse the dynamic time delay mapping table, pulse wave feature sequence and electrocardiogram time sequence to obtain the aligned feature sequence.

[0083] Optionally, the signal processing system performs alignment and fusion based on the dynamic time delay mapping table, pulse wave feature sequence and electrocardiogram time sequence to obtain an aligned feature sequence, as described in steps 4051 to 4054.

[0084] The alignment fusion achieved through a dynamic time-delay mapping table in this invention ensures strict synchronization between ECG and pulse wave signals in the time dimension. This eliminates the intermodal temporal misalignment problem caused by fixed time delays or inaccurate manual feature extraction in traditional methods, enabling subsequent fused features to accurately reflect the instantaneous state of cardiac electromechanical coupling. This significantly improves the robustness and accuracy of the model for assisting in the identification of acute myocardial infarction in variable pre-hospital and home settings, especially when arrhythmias or hemodynamic instability are present.

[0085] Optionally, the processes of steps 4051 to 4054 include: Step 4051: Based on the dynamic time delay mapping table, the pulse wave feature sequence is reconstructed by tensor slices based on index offset to obtain the first pulse wave feature sequence that is synchronized with the ECG time sequence feature sequence in terms of physiological time delay.

[0086] Optionally, the signal processing system traverses each heartbeat period entry in the dynamic delay mapping table. For the i-th heartbeat period, it retrieves its corresponding optimal delay index. Subsequently, in the original pulse wave characteristic sequence, the starting time of that heartbeat cycle is used as a reference to shift forward or backward. Each time step is used to extract pulse wave feature segments corresponding to the electrical activity of the heartbeat cycle.

[0087] The signal processing system reassembles the extracted pulse wave feature segments from all heartbeat cycles according to their original time sequence to form a new feature sequence. During this process, for any missing sequence boundaries caused by time delays, zero-padding or edge duplication strategies are used to fill in the gaps, ensuring the sequence length matches the ECG time-series feature sequence. The resulting sequence is the first pulse wave feature sequence. The feature values ​​at each time step in this sequence have been corrected according to physiological conduction time, ensuring physical synchronization with the electrical activity moments in the ECG time-series feature sequence.

[0088] For example, continuing with the embodiment of step 404, assume the dynamic time delay mapping table is [5, 6, 5, 7, 6], corresponding to 5 heartbeat cycles. The original pulse wave feature sequence length is 1000.

[0089] For the first cardiac cycle (assuming it covers time steps 1-200), the optimal delay index is 5. The signal processing system shifts the features of this cycle in the original sequence forward by 5 time steps, aligning the hemodynamic information that was originally 5 steps behind with the current electrical activity. For the second cardiac cycle (time steps 201-400), the optimal delay index is 6. The features of this cycle are shifted by 6 time steps. This process is repeated for all 5 cardiac cycles, with different shift operations performed. These 5 feature segments, corrected for different delays, are then concatenated end-to-end, and the boundary connections are processed to obtain a pulse wave first feature sequence of length 1000. At this point, the feature at time step 100 (the original R-wave moment) in this sequence actually originates from time step 105 (near the original pulse wave peak) of the original sequence, achieving "electromechanical" synchronization.

[0090] Step 4052: Based on the uncertainty measure in the constrained reconstruction features, generate a confidence weighted vector, and multiply the confidence weighted vector with the aligned pulse wave feature sequence element by element to obtain the second feature sequence of the pulse wave with enhanced confidence.

[0091] Optionally, the signal processing system obtains the uncertainty measure (i.e., reconstruction variance) in the constrained reconstruction features, and generates a confidence-weighted vector with the same length as the feature sequence. The generation logic of this vector is as follows: the smaller the variance, the more the signal morphology conforms to normal physiological patterns or model expectations, and the higher the confidence level; the larger the variance, the more abnormal or unreliable the signal, and the lower the confidence level. Therefore, the confidence-weighted vector... Calculated using the following formula: ; in, As an adjustable hyperparameter, the value range is 1-5. When the synchronization of ECG and pulse wave is good, the value is 1-2. When there is arrhythmia or large pulse conduction fluctuation, the value is 3-5. This represents the reconstruction variance within the corresponding time step or sliding window.

[0092] Subsequently, the signal processing system will generate a confidence-weighted vector. The pulse wave is multiplied element-wise with the first feature sequence. Through this multiplication, the pulse wave features at high confidence times are preserved or enhanced, while the pulse wave features at low confidence times (high uncertainty) are suppressed. The final sequence is the second feature sequence of the pulse wave with enhanced confidence.

[0093] For example, within a certain time period, the reconstruction variance of the constrained reconstruction feature output is high, for example... Assuming The signal processing system calculates the confidence weight at that moment: If there exists another moment when the reconstruction variance is lower, ,but Multiplying these two weight values ​​by the eigenvectors of the corresponding time steps in the first feature sequence of the pulse wave yields the second feature sequence of the pulse wave with enhanced reliability.

[0094] Step 4053: Based on the ECG time-series feature sequence and the pulse wave second feature sequence, the channel dimension is spliced ​​to obtain the preliminary multimodal joint feature tensor.

[0095] Optionally, the signal processing system performs a concatenation operation on the two sequences along the feature channel dimension. The concatenation process is as follows: for each identical time step t, the feature vector of the ECG time-series feature sequence at that time step is extracted. and the eigenvector of the second characteristic sequence of the pulse wave at that moment Concatenate these two vectors end-to-end to form a new joint feature vector. By repeating this operation for all time steps, a preliminary three-dimensional multimodal joint feature tensor is obtained. This tensor contains the time step dimension, the concatenated total feature channel dimension, and the batch dimension (if any). Through concatenation of the channel dimensions, the preliminary multimodal joint feature tensor simultaneously contains the excitation information of cardiac electrical activity (from ECG) and the hemodynamic response information (from PPG) after time delay correction and confidence screening at the same time point.

[0096] For example, suppose the dimension of the sparsed ECG time-series feature sequence is [1000, 64], and the dimension of the confidence-enhanced pulse wave second feature sequence is also [1000, 64].

[0097] In the first time step, the signal processing system concatenates the 64-dimensional ECG feature vector with the 64-dimensional pulse wave feature vector to obtain a 128-dimensional joint feature vector. The same operation is performed for all 1000 time steps. This ultimately yields a preliminary multimodal joint feature tensor with dimensions [1000, 128].

[0098] Step 4054: Perform channel mixing and feature compression transformation based on the preliminary multimodal joint feature tensor to obtain the transformed multimodal joint feature tensor, and perform residual connection based on the transformed multimodal joint feature tensor and the preliminary multimodal joint feature tensor to obtain the aligned feature sequence.

[0099] Optionally, the signal processing system applies a channel mixing layer (e.g., a 1x1 convolution or fully connected layer) to the initial multimodal joint feature tensor. The purpose is to map the concatenated high-dimensional features to a lower-dimensional latent space, while simultaneously capturing the nonlinear interaction between ECG and pulse wave features during the mixing process. After channel mixing and feature compression transformation, a transformed multimodal joint feature tensor is obtained. The feature channel dimension of this tensor is smaller than that of the initial multimodal joint feature tensor, achieving information condensation.

[0100] Subsequently, to alleviate the gradient vanishing problem in deep networks and preserve the original multimodal information, the signal processing system employs a residual connection mechanism. The transformed multimodal joint feature tensor is element-wise added to the original preliminary multimodal joint feature tensor. Furthermore, before addition, if the dimensions of the two are inconsistent, a linear projection layer is used to match the dimensions of the preliminary multimodal joint feature tensor, or the transformed tensor is upsampled / downsampled to match the dimensions. The result of the addition is the final aligned feature sequence, which contains both high-level semantic features refined and compressed nonlinearly, and retains the original, undistorted multimodal details through the residual path.

[0101] For example, continuing with the embodiment of step 4053, the initial multimodal joint feature tensor dimension is [1000, 128].

[0102] The signal processing system uses a 1x1 convolutional layer with 64 output channels (channel mixing and compression) to compress the 128-dimensional features into 64 dimensions, resulting in a transformed multimodal joint feature tensor of dimensions [1000, 64]. For residual connections, the signal processing system first maps the original [1000, 128] tensor to [1000, 64] dimensions using a linear projection layer (or the compression layer output remains 128 dimensions; compression is used as an example here). Assuming a projection layer maps the original tensor to 64 dimensions, the transformed [1000, 64] tensor is added element-wise to the projected original [1000, 64] tensor. Finally, an aligned feature sequence of dimensions [1000, 64] is obtained.

[0103] The alignment feature sequence finally obtained in the embodiments of the present invention is not only strictly synchronized in time and reliable and controllable in quality, but also efficient and compact in expression. It lays a solid data foundation for subsequent fusion with spatial reconstruction features and high-precision identification of acute myocardial infarction, and also significantly improves the model's ability to process complex and dynamic physiological signals on resource-constrained wearable devices.

[0104] Optional, refer to Figure 2 , Figure 2This is a schematic diagram of the multimodal physiological signal processing system for assisting in the identification of acute myocardial infarction provided by the present invention. The multimodal physiological signal processing system for assisting in the identification of acute myocardial infarction includes: The signal preprocessing module 210 is used to preprocess the simplified lead electrocardiogram signal and the photoplethysmography pulse wave signal respectively to obtain the current electrocardiogram signal sequence and the current pulse wave signal sequence. The dual-branch feature extraction module 220 is used to extract features from the current electrocardiogram signal sequence and the current pulse wave signal sequence based on a bidirectional state space network, so as to obtain the electrocardiogram feature sequence and the pulse wave feature sequence. The constrained reconstruction enhancement module 230 is used to generate spatial compensation features and uncertainty measures of the precordial leads based on the current electrocardiogram signal sequence and the reconstruction network, and to apply confidence weighted constraints to the spatial compensation features based on the uncertainty measures to obtain constrained reconstruction features. The multimodal feature fusion module 240 is used to perform cross-modal association and time delay alignment of ECG feature sequences and pulse wave feature sequences based on cross-modal attention mechanism to obtain aligned feature sequences, and then splice and fuse the aligned feature sequences with the reconstructed features to obtain multimodal fused features; The risk identification and alert module 250 is used to perform feature mapping on multimodal fusion features based on the classification head to obtain the risk probability value of acute myocardial infarction, and generate risk alert information based on the risk probability value; The neural network model, consisting of a bidirectional state-space network, a reconstruction network, a cross-modal attention mechanism, and a classification head, is a lightweight model obtained through ST-segment perceptual mixed precision quantization and ST-segment time window feature alignment distillation training.

[0105] This invention, based on the preprocessing of simplified lead ECG signals and photoplethysmography (PPG) pulse wave signals, yields regular current ECG and PPG signal sequences. A bidirectional state-space network is used to extract features from both signal sequences. Leveraging the linear computational complexity of this network, the overall computational load and memory usage are significantly reduced. A reconstruction network then generates precordial lead spatial compensation features and uncertainty metrics from the ECG signal sequences. Based on the uncertainty metrics, a confidence-weighted constraint is applied to obtain constrained reconstructed features, compensating for the spatial information missing in the simplified leads. Finally, a cross-modal attention mechanism is used to achieve cross-modal attention between the two types of features. Modal correlation and time delay alignment yield an aligned feature sequence, which is then concatenated with reconstructed features to form a multimodal fusion feature. This fully integrates effective multimodal information to ensure recognition accuracy. Finally, the classification head outputs an acute myocardial infarction risk probability value and generates risk warning information. The entire neural network model undergoes ST-segment perception hybrid precision quantization and ST-segment time window feature alignment distillation training, breaking the limitations of traditional globally unified bit quantization. Quantization bits are allocated differently based on the sensitivity of each network layer to ST-segment features. Simultaneously, feature alignment distillation protects the key diagnostic features of subtle millivolt-level potential shifts in the ST segment, preventing the loss of core features during compression. This solves the problems of high computational and storage requirements of traditional high-precision models, making them unsuitable for deployment on wearable microcontrollers. It also overcomes the shortcomings of conventional quantization compression, which erases key ST-segment features and reduces model recognition sensitivity. Ultimately, it enables low-power, real-time, and highly sensitive auxiliary recognition of acute myocardial infarction on wearable devices in pre-hospital and home settings.

[0106] Please see Figure 3 , Figure 3 An embodiment diagram of a wearable device provided for an embodiment of the present invention. For example... Figure 3 As shown, an embodiment of the present invention provides a wearable device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it performs the following steps: Based on the simplified lead ECG signal and photoplethysmography pulse wave signal, the current ECG signal sequence and the current pulse wave signal sequence are obtained by preprocessing them respectively. Based on a bidirectional state-space network, features are extracted from the current electrocardiogram (ECG) signal sequence and the current pulse wave signal sequence to obtain ECG feature sequences and pulse wave feature sequences, respectively. Based on the current ECG signal sequence and the reconstruction network, spatial compensation features and uncertainty measures of the precordial leads are generated. The spatial compensation features are then subjected to confidence-weighted constraints based on the uncertainty measures to obtain constrained reconstruction features. Based on the cross-modal attention mechanism, cross-modal correlation and time delay alignment are performed on ECG feature sequences and pulse wave feature sequences to obtain aligned feature sequences. The aligned feature sequences are then spliced ​​and fused with the reconstructed features to obtain multimodal fusion features. Based on the classification head, feature mapping is performed on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and risk warning information is generated based on the risk probability value; The neural network model, consisting of a bidirectional state-space network, a reconstruction network, a cross-modal attention mechanism, and a classification head, is a lightweight model obtained through ST-segment perceptual mixed precision quantization and ST-segment time window feature alignment distillation training.

[0107] Please see Figure 4 , Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with an embodiment of the present invention is shown. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, it performs the following steps: Based on the simplified lead ECG signal and photoplethysmography pulse wave signal, the current ECG signal sequence and the current pulse wave signal sequence are obtained by preprocessing them respectively. Based on a bidirectional state-space network, features are extracted from the current electrocardiogram (ECG) signal sequence and the current pulse wave signal sequence to obtain ECG feature sequences and pulse wave feature sequences, respectively. Based on the current ECG signal sequence and the reconstruction network, spatial compensation features and uncertainty measures of the precordial leads are generated. The spatial compensation features are then subjected to confidence-weighted constraints based on the uncertainty measures to obtain constrained reconstruction features. Based on the cross-modal attention mechanism, cross-modal correlation and time delay alignment are performed on ECG feature sequences and pulse wave feature sequences to obtain aligned feature sequences. The aligned feature sequences are then spliced ​​and fused with the reconstructed features to obtain multimodal fusion features. Based on the classification head, feature mapping is performed on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and risk warning information is generated based on the risk probability value; The neural network model, consisting of a bidirectional state-space network, a reconstruction network, a cross-modal attention mechanism, and a classification head, is a lightweight model obtained through ST-segment perceptual mixed precision quantization and ST-segment time window feature alignment distillation training.

[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the multimodal physiological signal processing method for the auxiliary identification of acute myocardial infarction provided by the above methods, the method including: Based on the simplified lead ECG signal and photoplethysmography pulse wave signal, the current ECG signal sequence and the current pulse wave signal sequence are obtained by preprocessing them respectively. Based on a bidirectional state-space network, features are extracted from the current electrocardiogram (ECG) signal sequence and the current pulse wave signal sequence to obtain ECG feature sequences and pulse wave feature sequences, respectively. Based on the current ECG signal sequence and the reconstruction network, spatial compensation features and uncertainty measures of the precordial leads are generated. The spatial compensation features are then subjected to confidence-weighted constraints based on the uncertainty measures to obtain constrained reconstruction features. Based on the cross-modal attention mechanism, cross-modal correlation and time delay alignment are performed on ECG feature sequences and pulse wave feature sequences to obtain aligned feature sequences. The aligned feature sequences are then spliced ​​and fused with the reconstructed features to obtain multimodal fusion features. Based on the classification head, feature mapping is performed on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and risk warning information is generated based on the risk probability value; The neural network model, consisting of a bidirectional state-space network, a reconstruction network, a cross-modal attention mechanism, and a classification head, is a lightweight model obtained through ST-segment perceptual mixed precision quantization and ST-segment time window feature alignment distillation training.

[0109] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal physiological signal processing method for the auxiliary identification of acute myocardial infarction, characterized in that, include: Based on the simplified lead ECG signal and photoplethysmography pulse wave signal, the current ECG signal sequence and the current pulse wave signal sequence are obtained by preprocessing the simplified lead ECG signal and the current pulse wave signal sequence, respectively. Based on a bidirectional state-space network, features are extracted from the current electrocardiogram (ECG) signal sequence and the current pulse wave signal sequence to obtain ECG feature sequences and pulse wave feature sequences, respectively. Based on the current ECG signal sequence and the reconstruction network, spatial compensation features and uncertainty measures of the precordial leads are generated, and the spatial compensation features are subjected to confidence weighting constraints based on the uncertainty measures to obtain constrained reconstruction features. Based on the cross-modal attention mechanism, the ECG feature sequence and pulse wave feature sequence are cross-modal correlated and time-delay aligned to obtain the aligned feature sequence. The aligned feature sequence is then spliced ​​and fused with the reconstructed features to obtain the multimodal fused features. Based on the classification head, feature mapping is performed on the multimodal fusion features to obtain the risk probability value of acute myocardial infarction, and risk warning information is generated based on the risk probability value; The neural network model, which consists of the bidirectional state space network, the reconstruction network, the cross-modal attention mechanism, and the classification head, is a lightweight model obtained by training through ST segment perceptual mixed precision quantization and ST segment time window feature alignment distillation.

2. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 1, characterized in that, The simplified lead ECG signal includes a single lead ECG signal or an ECG signal with fewer than twelve standard leads; the preprocessing includes bandpass filtering, baseline drift removal, and lead-by-lead normalization.

3. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 1, characterized in that, The cross-modal attention mechanism is used to perform cross-modal correlation and time-delay alignment on the ECG feature sequence and pulse wave feature sequence to obtain an aligned feature sequence, including: Based on the R-wave time index position extracted from the ECG feature sequence, a binary time sequence mask vector is constructed, and the binary time sequence mask vector and the ECG feature sequence are subjected to element-by-element feature filtering to obtain a sparse ECG time sequence feature. Based on a preset range of pulse conduction time variation, a multi-scale time-delay convolution kernel group containing multiple displacement step sizes is constructed, and the pulse wave feature sequence is mapped to multiple scales based on the multi-scale time-delay convolution kernel group to obtain multi-time-delay candidate pulse wave feature tensors. Based on the ECG time-series feature sequence and the multi-time-delay candidate pulse wave feature tensor, the cosine similarity matrix under each time-delay candidate is calculated to obtain the time-delay feature correlation cube. Based on the time delay feature correlation cube, maximum pooling is performed on the time delay dimension to extract the optimal time delay index corresponding to each heartbeat cycle and construct a dynamic time delay mapping table. The aligned feature sequence is obtained by aligning and fusing the dynamic time delay mapping table, the pulse wave feature sequence, and the electrocardiogram timing feature sequence.

4. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 3, characterized in that, The alignment and fusion based on the dynamic time delay mapping table, the pulse wave feature sequence, and the electrocardiogram timing feature sequence to obtain the aligned feature sequence includes: Based on the dynamic time delay mapping table, the pulse wave feature sequence is reconstructed by tensor slices based on index offset to obtain the first pulse wave feature sequence that is synchronized with the electrocardiogram time sequence feature sequence in terms of physiological time delay. Based on the uncertainty measure in the constrained reconstruction features, a confidence weighted vector is generated, and the confidence weighted vector is multiplied element-wise with the aligned pulse wave feature sequence to obtain a second pulse wave feature sequence with enhanced confidence. Based on the electrocardiogram time-series feature sequence and the pulse wave second feature sequence, the channel dimension is spliced ​​to obtain a preliminary multimodal joint feature tensor; Channel mixing and feature compression transformation are performed based on the preliminary multimodal joint feature tensor to obtain the transformed multimodal joint feature tensor. Residual concatenation is then performed based on the transformed multimodal joint feature tensor and the preliminary multimodal joint feature tensor to obtain the aligned feature sequence.

5. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 1, characterized in that, The obtained electrocardiogram feature sequence and pulse wave feature sequence include: Based on the selective state-space mechanism, forward and reverse scanning are performed on the current ECG signal sequence and the current pulse wave signal sequence, respectively, and the discretization step size, input matrix and output matrix are adjusted based on the current input segment; The forward and reverse scan outputs are concatenated along the feature dimension, and the ECG feature sequence and the pulse wave feature sequence are obtained through residual connection and layer normalization.

6. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 1, characterized in that, The spatial compensation feature is the reconstructed mean, the uncertainty measure is the reconstructed variance, and the obtained constrained reconstructed features include: The uncertainty gating weight vector is determined based on the reconstruction variance of the precordial leads; where the larger the variance, the smaller the weight. The constrained reconstruction features are obtained by performing element-wise multiplication of the reconstructed mean of the precordial leads based on the uncertainty-gated weight vector.

7. The multimodal physiological signal processing method for auxiliary identification of acute myocardial infarction according to claim 1, characterized in that, The ST segment time window is a time window whose width is adaptively determined by the QRS complex endpoint of the ECG signal and the interval between adjacent R waves. The ST segment mixed precision quantization bit width allocation method of the neural network model is as follows: the quantization bit width allocated to each network layer is determined based on the sensitivity of each network layer to the ST segment discrimination loss; wherein, the sensitivity is positively correlated with the quantization bit width.

8. A multimodal physiological signal processing system for the auxiliary identification of acute myocardial infarction, characterized in that, The method for multimodal physiological signal processing for the auxiliary identification of acute myocardial infarction as described in any one of claims 1 to 7; The multimodal physiological signal processing system for the auxiliary identification of acute myocardial infarction includes: The signal preprocessing module is used to preprocess the simplified lead ECG signal and the photoplethysmography pulse wave signal respectively to obtain the current ECG signal sequence and the current pulse wave signal sequence. The dual-branch feature extraction module is used to extract features from the current electrocardiogram signal sequence and the current pulse wave signal sequence based on a bidirectional state space network, so as to obtain an electrocardiogram feature sequence and a pulse wave feature sequence. The constrained reconstruction enhancement module is used to generate spatial compensation features and uncertainty measures of the precordial leads based on the current electrocardiogram signal sequence and the reconstruction network, and to apply a confidence-weighted constraint to the spatial compensation features based on the uncertainty measures to obtain constrained reconstruction features. The multimodal feature fusion module is used to perform cross-modal association and time delay alignment on the electrocardiogram feature sequence and pulse wave feature sequence based on the cross-modal attention mechanism to obtain the aligned feature sequence, and then splice and fuse the aligned feature sequence with the reconstructed features to obtain the multimodal fused features; The risk identification and alert module is used to perform feature mapping on the multimodal fusion features based on the classification head to obtain the risk probability value of acute myocardial infarction, and generate risk alert information based on the risk probability value; The neural network model, which consists of the bidirectional state space network, the reconstruction network, the cross-modal attention mechanism, and the classification head, is a lightweight model obtained by training through ST segment perceptual mixed precision quantization and ST segment time window feature alignment distillation.

9. A wearable device, characterized in that, include: Memory, used to store computer software programs; A processor is configured to read and execute the computer software program, wherein when the processor executes the computer software program, it implements the multimodal physiological signal processing method for assisted identification of acute myocardial infarction as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores a computer software program, which, when executed by a processor, implements the multimodal physiological signal processing method for assisted identification of acute myocardial infarction as described in any one of claims 1 to 7.