A non-contact electromagnetic heart monitoring method based on signal semantic decomposition

By using semantic bottleneck modeling and an improved TSMixer model, the problems of low signal-to-noise ratio and cross-modal misalignment in non-contact cardiac monitoring were solved, enabling stable signal extraction and personalized health monitoring under noise and disturbance conditions, and improving the accuracy and robustness of cardiac status identification.

CN121795867BActive Publication Date: 2026-07-03HE FEI ZHONG KE ZHI QI XIN XI KE JI YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511994300.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-07-03
Estimated Expiration
2045-12-26

AI Technical Summary

Technical Problem

Existing non-contact cardiac monitoring methods suffer from low signal-to-noise ratios and frequent phase jumps when faced with posture changes, voluntary movements, multipath interference, and environmental noise. They are difficult to stably extract cardiac structural patterns, and have weak cross-modal alignment and physiological-level interpretability, making it impossible to effectively identify cardiac rhythm drift and individual physiological differences.

Method used

We employ semantic bottleneck modeling, self-supervised perturbation compression, cross-modal semantic alignment, and an improved TSMixer model. By extracting key semantic information through semantic bottleneck modeling, we construct an ECG semantic space and perform cross-modal alignment. We then use the improved TSMixer model to perform multimodal temporal fusion and output an individual cardiac rhythm semantic map.

Benefits of technology

It achieves improved signal stability under noise and disturbance conditions, enhances cross-modal consistency and anti-interference capability, and provides highly accurate and robust cardiac condition identification and individual health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121795867B_ABST
    Figure CN121795867B_ABST
Patent Text Reader

Abstract

The application discloses a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction, comprising the following steps: collecting non-contact reflected electromagnetic signals to generate electromagnetic phase signal data; performing semantic bottleneck modeling to compress and retain key semantic information related to the heart to generate a latent semantic representation tensor; constructing a semantic invariant disturbance signal, which is input into an encoder together with the electromagnetic phase signal for self-supervised reconstruction; constructing an electrocardiogram semantic space and performing cross-modal semantic alignment between the latent semantic representation tensor and the electrocardiogram semantic space; outputting an enhanced latent semantic representation tensor by using an improved TSMixer model; extracting rhythm drift information based on the enhanced latent semantic representation, constructing an individual heart rhythm semantic atlas, and outputting a heart state recognition result. The application realizes semantic-level stable modeling of non-contact heart monitoring, improves the accuracy, robustness and explainability of heart rhythm recognition, and is suitable for intelligent health monitoring and remote medical scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical signal processing and artificial intelligence, and in particular to a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction. Background Technology

[0002] With the rapid development of non-contact vital sign monitoring technology and telemedicine, non-contact cardiac activity monitoring using millimeter-wave radar, electromagnetic arrays, and other methods has gradually become a research hotspot. Existing non-contact cardiac monitoring methods mainly rely on the envelope characteristics of reflected electromagnetic signals, heartbeat harmonic energy, or respiratory-heartbeat separation algorithms to infer cardiac status. However, these methods generally suffer from the following problems in practical applications:

[0003] First, the reflected electromagnetic signal modulated by chest wall micromotions has extremely low amplitude and is significantly affected by posture changes, millimeter-level voluntary movements, multipath interference, and environmental noise, resulting in a low signal-to-noise ratio and frequent phase jumps. Key cardiac cycle features (such as ventricular premature beat precursors, rhythm instability intervals, and myocardial diastolic abnormalities) exhibit strong random fluctuations under noisy conditions, making it difficult for existing models to stably extract physiologically meaningful cardiac structural patterns. Second, traditional methods typically rely on single-channel electromagnetic signals or simple same-frequency demodulation, which lack the ability to express the deep semantic structure in phase signals. Since electromagnetic signals only express surface micromotions and lack the electrophysiological semantics of ECG signals, there is a fundamental misalignment between existing non-contact monitoring methods and standard ECG semantics, resulting in weak cross-modal alignment and physiological-level interpretability. Third, existing signal processing methods mostly use linear or quasi-linear techniques such as spectral peak extraction, wavelet decomposition, and bandpass filtering, which are difficult to address the non-stationarity and individual variability of cardiac signals. In scenarios with strong disturbances (such as coughing, slight movement, and changes in multipath reflexes), phase sequences often exhibit pseudomodalities, information leakage, or energy aliasing, making it difficult to reliably identify cardiac rhythm drift, weak abnormal beats, and individual physiological differences. Finally, most existing multimodal fusion methods rely on simple feature splicing or attention mechanisms, failing to effectively utilize the ECG semantic space or provide a stable rhythm map structure for individual health monitoring. This results in non-contact monitoring results being highly susceptible to environmental influences and lacking robustness, making it difficult to meet the needs of clinical or long-term health management scenarios.

[0004] Therefore, how to provide a non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction. This invention employs semantic bottleneck modeling, self-supervised perturbation compression, cross-modal semantic alignment, and an improved TSMixer model to accurately extract latent semantic features related to cardiac activity. It has the advantages of being non-contact, interpretable, highly robust, and highly versatile, providing a safe, efficient, and structurally innovative technical path for cardiac rhythm state recognition and individual health monitoring.

[0006] A non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction according to an embodiment of the present invention includes the following steps:

[0007] Non-contact reflected electromagnetic signals from the target area are collected, raw co-frequency orthogonal component data are obtained, and phase demodulation and filtering are performed to generate electromagnetic phase signal data.

[0008] Semantic bottleneck modeling is performed on electromagnetic phase signal data, and the electromagnetic phase signal data is compressed while retaining key semantic information related to cardiac physiological semantics, generating a latent semantic representation tensor.

[0009] A semantically invariant perturbation signal is constructed and input into the encoder along with electromagnetic phase signal data for self-supervised reconstruction. The semantic stability of the latent semantic representation tensor is optimized through intramodal semantic invariant compression and consistency constraints.

[0010] An electrocardiogram (ECG) semantic space is constructed based on structured ECG data, ECG semantic feature representations are extracted, and the latent semantic representation tensor is cross-modal semantically aligned with the ECG semantic space to establish a semantic mapping relationship.

[0011] A multimodal temporal tensor containing electromagnetic phase signal data, semantically invariant perturbation signal data, and synchronized electrocardiogram signal data is constructed and input into an improved TSMixer model. The improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head, and outputs an enhanced latent semantic representation tensor.

[0012] Based on the extraction of rhythm drift information using enhanced latent semantic representation tensors, an individual heart rhythm semantic map is constructed, and the heart rhythm state recognition results and individual health monitoring data are output.

[0013] Optionally, the non-contact reflected electromagnetic signal acquisition of the target area, obtaining the original co-frequency orthogonal component data and performing phase demodulation and filtering processing to generate electromagnetic phase signal data specifically includes:

[0014] The non-contact reflected electromagnetic signal is generated by emitting continuous electromagnetic waves in front of the target's chest using a millimeter-wave radar device and receiving the electromagnetic signals reflected back from the chest area.

[0015] The non-contact reflected electromagnetic signal is demodulated at the same frequency to obtain the original orthogonal component data at the same frequency;

[0016] The original orthogonal component data of the same frequency were subjected to low-pass filtering and sampling was synchronized at a set sampling rate to obtain a clean data sequence corresponding to the micro-motion changes on the chest surface.

[0017] The real and imaginary data are constructed into complex numbers according to the sampling points, and the instantaneous phase at each moment is extracted based on the arctangent calculation method to form the original phase sequence. The original phase sequence is normalized, and the phase curve is standardized into time-series data of a set length within a preset time window, and the electromagnetic phase signal data is output.

[0018] Optionally, the step of performing semantic bottleneck modeling on the electromagnetic phase signal data, compressing the electromagnetic phase signal data while retaining key semantic information related to cardiac physiological semantics, and generating a latent semantic representation tensor specifically includes:

[0019] An electromagnetic phase signal data is encoded using a feature encoder consisting of a one-dimensional temporal convolutional network and a Swish activation function, and the output latent semantic representation tensor is generated.

[0020] Based on the information theory modeling method, the mutual information value between the latent semantic representation tensor and the preset target cardiac semantic label is used as the semantic preservation target to perform the maximization operation, and the mutual information value between the latent semantic representation tensor and the original electromagnetic phase signal data is used as the compression constraint target to perform the minimization operation.

[0021] To maximize the semantic preservation objective, adjustable weight coefficients are introduced for joint optimization with the minimization of compression constraint objective.

[0022] Optionally, the construction of the semantically invariant perturbation signal, and its input to the encoder along with the electromagnetic phase signal data for self-supervised reconstruction, optimizes the semantic stability of the latent semantic representation tensor through intramodal semantic invariant compression and consistency constraints, specifically including:

[0023] Randomly select a segment of original orthogonal component data of non-current sample from the training set, apply low-pass filtering to extract background low-frequency components, add them point by point in complex form with the original orthogonal component data of the target sample, and perform phase calculation to form a perturbation phase sequence with background interference characteristics as a semantically invariant perturbation signal.

[0024] Electromagnetic phase signal data and semantically invariant perturbation signal are simultaneously input into a time-series encoder to obtain the original latent semantic representation tensor and the perturbation representation tensor, respectively.

[0025] The original latent semantic representation tensor is input into the reconstruction decoder to recover the reconstructed signal;

[0026] A joint loss function is constructed to simultaneously optimize the semantic invariant compression term and the consistency constraint term. The semantic invariant compression term minimizes the error between the electromagnetic phase signal data and the reconstructed signal, while the consistency constraint term minimizes the structural difference between the latent semantic representation tensors before and after the perturbation.

[0027] The parameters of the time encoder are optimized by minimizing the joint loss function.

[0028] Optionally, the step of constructing an ECG semantic space based on structured ECG data, extracting ECG semantic feature representations, and performing cross-modal semantic alignment between the latent semantic representation tensor and the ECG semantic space to establish a semantic mapping relationship specifically includes:

[0029] Structured electrocardiogram (ECG) data is input into the ECG encoder, which is connected to the ECG decoder and the semantic classifier.

[0030] A β-variable autoencoder structure was used to jointly train the ECG encoder, ECG decoder, and semantic classifier. The training objective function included reconstruction error, KL divergence of latent semantic representation tensor, and disease label-based supervised loss.

[0031] The synchronously acquired structured electrocardiogram data is input into the electrocardiogram encoder to generate electrocardiogram semantic anchor points, and the electromagnetic phase signal data is input into the electromagnetic encoder to extract the electromagnetic latent semantic representation tensor.

[0032] Perform cross-modal semantic alignment operation, which refers to performing consistency matching between the electromagnetic latent semantic representation tensor and the electrocardiogram semantic anchor point, and constructing a cross-modal loss function based on representation alignment constraints and decoding alignment constraints for joint optimization;

[0033] By minimizing the cross-modal loss function, the latent semantic representation tensor can be made to have the ability to express the structure of the ECG semantic space without relying on the ECG signal.

[0034] Optionally, the improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head:

[0035] The modal structure guided encoder receives and encodes electromagnetic phase signal data, semantically invariant perturbation signal data, and structured electrocardiogram data. The modal structure guided encoder includes three parallel sub-paths, which respectively use structural units containing one-dimensional convolutional layers, batch normalization layers, and ReLU activation functions to extract short-term dynamic features, and use a position embedding module for time alignment. Then, a modal structure guided matrix is ​​used to perform channel weighting and output a multimodal embedding feature tensor.

[0036] The dual-channel cross-hybrid module includes a residual fusion channel and an attention interaction channel. The residual fusion channel uses a weighted average and residual connection method to perform fusion on the multimodal embedded feature tensor. The attention interaction channel uses a dot product attention mechanism to calculate the semantic dependency between any two modalities and performs weighted reconstruction. The fused feature representation is normalized using Softmax normalization to obtain a cross-modal joint representation tensor.

[0037] The skip-time structure block uses an exponentially growing skip window structure to perform one-dimensional convolution operations with different interval lengths on the cross-modal joint representation tensor. The size of the convolution kernel is dynamically adjusted according to the time window. The time structure results of all scales are spliced ​​together and the dimensions are compressed through a fully connected linear projection layer to form a time-enhanced representation tensor.

[0038] The joint multi-task semantic output head receives the time-enhanced representation tensor and performs the latent semantic enhancement task and the auxiliary rhythm label prediction task in parallel. The latent semantic enhancement task uses a residual feedforward network and the GELU activation function to extract deep semantic features and outputs the enhanced latent semantic representation tensor. The auxiliary rhythm label prediction task uses a fully connected layer to output the rhythm classification probability.

[0039] Optionally, the step of extracting rhythm drift information based on enhanced latent semantic representation tensors, constructing an individual heart rhythm semantic map, and outputting heart state recognition results specifically includes:

[0040] A sliding time window mechanism is used to segment the enhanced latent semantic representation tensor to generate a sequence of semantic segments arranged in time. For each adjacent semantic segment, the Euclidean distance, cosine distance, element-wise difference, and vector angle change are calculated and fused into rhythmic drift quantity according to preset rules to form a rhythmic drift sequence.

[0041] The rhythm drift sequence is aligned with the window timestamp one by one. The electrophysiological structural features output by the ECG semantic space are extracted within the aligned time period. The parameters of the electrophysiological structural features are numericalized, standardized and encoded. The encoded electrophysiological structural features are combined with the corresponding semantic segments and rhythm drift amount to form a single cardiac state event node and stored in the event list in chronological order.

[0042] The edge weights are calculated for adjacent cardiac state event nodes. The semantic difference, electrophysiological structural difference, mechanical dynamic change rate and rhythm drift amplitude are weighted and summed according to preset weight coefficients to form the edge weights. All nodes and edge weights are organized into a graph structure in chronological order to form a cardiac state semantic graph. The graph aggregation operation is performed within a set sliding aggregation window to obtain a window-level comprehensive semantic representation.

[0043] The window-level comprehensive semantic representation is input into a fully connected classifier. Based on the classifier's output, the cardiac state of the time period to which the window belongs is determined. The output includes cardiac state recognition results, including heart rhythm type, electrophysiological structural state, mechanical activity state, rhythm stability state, abnormal event prompts, and corresponding time location.

[0044] The beneficial effects of this invention are:

[0045] This invention introduces a "signal semantic deconstruction" framework to address the problems of low signal-to-noise ratio, frequent phase jumps, significant cross-modal misalignment, and difficulty in suppressing non-stationary disturbances in existing non-contact cardiac monitoring. It constructs a complete semantic-level modeling link from electromagnetic phase signals to the ECG semantic space, achieving multi-level beneficial effects that cannot be achieved by existing technologies.

[0046] First, this invention forcibly compresses noise components unrelated to the heart in the early stages of feature extraction through semantic bottleneck modeling, ensuring that the representation retains only key semantic information related to cardiac mechano-electrophysiological activities. This fundamentally improves the stability of the signal under conditions of posture interference, micro-motion noise, and multipath variations. Second, it introduces a semantically invariant perturbation learning mechanism, enabling the model to maintain latent semantic consistency even when facing background replacement, low-frequency interference, or non-semantic changes. Self-supervised consistency constraints significantly enhance the representation's anti-interference and generalization capabilities, which is lacking in existing millimeter-wave cardiac monitoring methods. This invention utilizes structured electrocardiogram (ECG) data to construct an ECG semantic space and maps the latent electromagnetic semantics to an ECG semantic dimension with clear physiological meaning through cross-modal semantic alignment. This allows non-contact electromagnetic signals to achieve physiological readability close to the level of ECG interpretation, significantly compensating for the semantic misalignment problem of traditional non-contact methods. This invention employs an improved TSMixer model, achieving multimodal deep temporal fusion through modal structure-guided encoders, dual-channel cross-mixing modules, and skip-time structure blocks. Compared to traditional attention or convolutional structures, it can more efficiently capture cross-scale rhythm variations and fine-grained periodic perturbations, thereby improving the sensitivity of rhythm drift detection and the stability of long-term monitoring. Finally, this invention constructs an individual heart rhythm semantic map by enhancing latent semantic representation, providing interpretable rhythmic event chains and health fluctuation trends, achieving long-term personalized monitoring capabilities far exceeding traditional methods. In summary, this invention achieves substantial breakthroughs in semantic stability, cross-modal consistency, perturbation resistance, temporal fusion capability, and interpretability, significantly improving the accuracy, robustness, and clinical value of non-contact cardiac monitoring. Attached Figure Description

[0047] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0048] Figure 1This is a flowchart of a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction proposed in this invention;

[0049] Figure 2 This is a schematic diagram of a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction proposed in this invention;

[0050] Figure 3 This is a framework diagram of the improved TSMixer model in a non-contact electromagnetic heart monitoring method based on signal semantic deconstruction proposed in this invention. Detailed Implementation

[0051] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0052] refer to Figure 1-3 A non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction includes the following steps:

[0053] Non-contact reflected electromagnetic signals from the target area are collected, raw co-frequency orthogonal component data are obtained, and phase demodulation and filtering are performed to generate electromagnetic phase signal data.

[0054] Semantic bottleneck modeling is performed on electromagnetic phase signal data, and the electromagnetic phase signal data is compressed while retaining key semantic information related to cardiac physiological semantics, generating a latent semantic representation tensor.

[0055] A semantically invariant perturbation signal is constructed and input into the encoder along with electromagnetic phase signal data for self-supervised reconstruction. The semantic stability of the latent semantic representation tensor is optimized through intramodal semantic invariant compression and consistency constraints.

[0056] An electrocardiogram (ECG) semantic space is constructed based on structured ECG data, ECG semantic feature representations are extracted, and the latent semantic representation tensor is cross-modal semantically aligned with the ECG semantic space to establish a semantic mapping relationship.

[0057] A multimodal temporal tensor containing electromagnetic phase signal data, semantically invariant perturbation signal data, and synchronized electrocardiogram signal data is constructed and input into an improved TSMixer model. The improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head, and outputs an enhanced latent semantic representation tensor.

[0058] Based on the extraction of rhythm drift information using enhanced latent semantic representation tensors, an individual heart rhythm semantic map is constructed, and the heart rhythm state recognition results and individual health monitoring data are output.

[0059] In this embodiment, the process of acquiring non-contact reflected electromagnetic signals from the target area, obtaining original co-frequency orthogonal component data, performing phase demodulation and filtering, and generating electromagnetic phase signal data specifically includes:

[0060] The non-contact reflected electromagnetic signal is generated by emitting continuous electromagnetic waves in front of the target's chest using a millimeter-wave radar device and receiving the electromagnetic signal reflected back from the chest area. The electromagnetic signal is modulated by the tiny vibrations of the chest caused by the heartbeat during the reflection process, including weak phase perturbations related to the mechanical movement of the heart.

[0061] The non-contact reflected electromagnetic signal is demodulated at the same frequency to obtain the original orthogonal component data at the same frequency. The demodulation at the same frequency refers to mixing the non-contact reflected electromagnetic signal with the sine wave and cosine wave of the local oscillator signal respectively to obtain two signal components containing real and imaginary part information.

[0062] The original orthogonal component data of the same frequency were subjected to low-pass filtering and sampling was synchronized at a set sampling rate to obtain a clean data sequence corresponding to the micro-motion changes on the chest surface.

[0063] The real and imaginary data are constructed into complex numbers according to the sampling points, and the instantaneous phase at each moment is extracted based on the arctangent calculation method to form the original phase sequence. The original phase sequence is normalized, and the phase curve is standardized into time-series data of a set length within a preset time window, and the electromagnetic phase signal data is output.

[0064] In this embodiment, the step of performing semantic bottleneck modeling on the electromagnetic phase signal data, compressing the electromagnetic phase signal data while retaining key semantic information related to cardiac physiological semantics, and generating a latent semantic representation tensor specifically includes:

[0065] The semantic bottleneck modeling refers to generating a low-dimensional semantic representation tensor through a constrained information compression process, which reduces redundant information in electromagnetic phase signal data while retaining semantic components that are highly relevant to cardiac physiological activities.

[0066] A feature encoder consisting of a one-dimensional temporal convolutional network and a Swish activation function is used to encode electromagnetic phase signal data, which enhances the ability to identify non-stationary rhythmic changes and weak amplitude phase perturbations, and outputs a latent semantic representation tensor.

[0067] Based on the information theory modeling method, the mutual information value between the latent semantic representation tensor and the preset target cardiac semantic label is used as the semantic preservation target and a maximization operation is performed. The mutual information value between the latent semantic representation tensor and the original electromagnetic phase signal data is used as the compression constraint target and a minimization operation is performed. The mutual information value is an index that measures the degree of difference between the joint distribution and independent distribution between two variables, and represents the uncertainty elimination ability of one variable on the other variable.

[0068] To maximize the semantic preservation objective, an adjustable weight coefficient is introduced to jointly optimize the objective of minimizing the compression constraint. The adjustable weight coefficient is used to balance the semantic fidelity and compression capability.

[0069] In this embodiment, the construction of a semantically invariant perturbation signal, which is then input into the encoder along with electromagnetic phase signal data for self-supervised reconstruction, optimizes the semantic stability of the latent semantic representation tensor through intramodal semantic invariant compression and consistency constraints. Specifically, this includes:

[0070] The semantically invariant perturbation signal refers to the perturbation information that is unrelated to the target semantics injected into the input signal without changing the semantics of cardiac physiological activity. It is used to construct a semantically invariant constraint target during the training process. The semantically invariant perturbation signal construction process includes: randomly selecting a segment of original orthogonal component data of non-current samples from the training set, applying a low-pass filter to extract the background low-frequency component, adding it point by point in complex form with the original orthogonal component data of the target sample, and performing phase calculation to form a perturbation phase sequence with background interference characteristics as a semantically invariant perturbation signal.

[0071] Electromagnetic phase signal data and semantically invariant perturbation signal are simultaneously input into a time-series encoder. The time-series encoder uses a one-dimensional convolutional layer to extract local temporal features, combines a dilated convolutional layer to expand the temporal perception range, introduces a GLU gated linear unit for channel control, uses a Swish activation function to introduce nonlinear expression, and constructs a time structure encoding framework through a batch normalization layer and residual connection to obtain the original latent semantic representation tensor and the perturbation representation tensor, respectively.

[0072] The original latent semantic representation tensor is input to the reconstruction decoder, which is a neural network structure consisting of deconvolutional layers, Swish activation functions, batch normalization layers and fully connected linear projection layers, to recover the reconstructed signal, which is used to measure the fidelity of semantic compression.

[0073] A joint loss function is constructed to simultaneously optimize the semantic invariant compression term and the consistency constraint term. The semantic invariant compression term minimizes the error between the electromagnetic phase signal data and the reconstructed signal, while the consistency constraint term minimizes the structural difference between the latent semantic representation tensors before and after the perturbation. The joint loss function is defined as follows:

[0074] ;

[0075] in, For the joint loss function, Electromagnetic phase signal data, It is a semantically invariant perturbation signal. It is a timing encoder. To rebuild the decoder, , which is the weighting coefficient used to balance the optimization ratio of the loss term;

[0076] By minimizing the joint loss function, the parameters of the temporal encoder are optimized, enabling the temporal encoder to generate structurally consistent latent semantic representation tensors even when faced with perturbation inputs. This improves its ability to suppress non-semantic changes and enhances the stability of semantic expression. The final output latent semantic representation tensor possesses high semantic consistency and perturbation invariance, providing a stable feature foundation for ECG semantic space alignment and cross-modal semantic modeling.

[0077] In this embodiment, the step of constructing an ECG semantic space based on structured ECG data, extracting ECG semantic feature representations, and performing cross-modal semantic alignment between the latent semantic representation tensor and the ECG semantic space to establish a semantic mapping relationship specifically includes:

[0078] The ECG semantic space refers to a potential representation space constructed based on structured ECG data, capable of expressing physiological parameters including heart rhythm, PR interval, and QT interval. The ECG semantic space is constructed as follows: structured ECG data is input into an ECG encoder. The structured ECG data refers to multi-lead ECG signals recorded through a standardized protocol. After time-series normalization, noise filtering, and label annotation, a data structure containing waveform data and label metadata is formed. The ECG encoder consists of one-dimensional convolutional layers and fully connected layers to extract low-dimensional semantic features. The ECG encoder connects the ECG decoder and the semantic classifier. The ECG decoder is a waveform reconstruction network composed of multiple deconvolutional layers and normalization layers. The semantic classifier is a fully connected structure that outputs disease label predictions and is used to introduce structured label supervision.

[0079] A β-variable autoencoder structure is used to jointly train the ECG encoder, ECG decoder, and semantic classifier. The training objective function includes reconstruction error, KL divergence of the latent semantic representation tensor, and disease-label-based supervised loss. The objective function is defined as follows:

[0080] ;

[0081] in, Let be the objective function. For structured electrocardiogram data, The latent semantic representation tensor extracted from the ECG encoder. Structured disease labels, including semantic classification indicators for arrhythmia, bradycardia, and QT prolongation. The variational posterior distribution generated by the ECG encoder. The reconstructed likelihood distribution for modeling the ECG decoder. Disease label conditional distribution modeled for semantic classifiers This represents the posterior distribution generated by the ECG encoder. The following pairs of variables Mathematical expectation operation, The KL divergence represents the distance between the variational posterior distribution and the standard normal distribution. It is used to constrain the distribution of the latent semantic representation tensor to approximate the preset prior distribution, thus preventing the encoder from overfitting. and To adjust the weighting coefficients of KL divergence and semantic supervision loss;

[0082] The synchronously acquired structured electrocardiogram data is input into the electrocardiogram encoder to generate electrocardiogram semantic anchor points. The electromagnetic phase signal data is input into the electromagnetic encoder to extract the electromagnetic latent semantic representation tensor. The electromagnetic encoder is a temporal structure-aware encoder model, which consists of a one-dimensional convolutional layer, a dilated convolutional structure, a gating unit and a residual connection module, and is used to extract the temporal structure representation related to heart rhythm.

[0083] A cross-modal semantic alignment operation is performed, which refers to performing consistency matching between the electromagnetic latent semantic representation tensor and the electrocardiogram (ECG) semantic anchor. A cross-modal loss function is constructed based on representation alignment constraints and decoding alignment constraints for joint optimization. The representation alignment constraint achieves semantic fit by minimizing the cosine distance between the electromagnetic latent semantic representation tensor and the ECG semantic anchor. The decoding alignment constraint inputs the electromagnetic semantic representation tensor into the ECG decoder to reconstruct the ECG waveform and aligns it with the original ECG signal. The cross-modal loss function is defined as follows:

[0084] ;

[0085] in, For cross-modal loss function, For structured electrocardiogram data, Electromagnetic phase signal data, It is an electromagnetic encoder. and These are an ECG encoder and an ECG decoder, respectively. , Weighting factors to control semantic similarity and reconstruction consistency;

[0086] By minimizing the cross-modal loss function, the latent semantic representation tensor can be made to have the ability to express the structure of the ECG semantic space without relying on the ECG signal.

[0087] In this embodiment, the improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head.

[0088] The modal structure guided encoder receives and encodes electromagnetic phase signal data, semantically invariant perturbation signal data, and structured electrocardiogram (ECG) data. The modal structure guided encoder includes three parallel sub-paths, which respectively use structural units containing one-dimensional convolutional layers, batch normalization layers, and ReLU activation functions to extract short-term dynamic features, and use a position embedding module for time alignment. Then, channel weighting is performed using the modal structure guided matrix to output a multimodal embedding feature tensor. The modal structure guided matrix is ​​jointly constructed by electromagnetic structure priors and ECG structure priors. The electromagnetic structure priors are derived from the rhythmic fluctuation features of the electromagnetic phase signal, and the ECG structure priors are derived from the labeled ECG event interval information in the structured ECG data.

[0089] The dual-channel cross-hybrid module includes a residual fusion channel and an attention interaction channel. The residual fusion channel uses a weighted average and residual connection method to perform fusion on the multimodal embedded feature tensor. The attention interaction channel uses a dot product attention mechanism to calculate the semantic dependency between any two modalities and performs weighted reconstruction. The fused feature representation is normalized using Softmax normalization to obtain a cross-modal joint representation tensor.

[0090] The skip-time structure block uses an exponentially growing skip window structure to perform one-dimensional convolution operations with different interval lengths on the cross-modal joint representation tensor. The size of the convolution kernel is dynamically adjusted according to the time window. The time structure results of all scales are spliced ​​together and the dimensions are compressed through a fully connected linear projection layer to form a time-enhanced representation tensor.

[0091] The joint multi-task semantic output head receives the time-enhanced representation tensor and executes the latent semantic enhancement task and the auxiliary rhythm label prediction task in parallel. The latent semantic enhancement task uses a residual feedforward network and the GELU activation function to extract deep semantic features and outputs an enhanced latent semantic representation tensor. The auxiliary rhythm label prediction task uses a fully connected layer to output rhythm classification probabilities to assist in supervising the optimization of the main task.

[0092] The improved TSMixer model possesses structure-guided, multimodal cross-fusion, and temporal multi-scale representation capabilities, enabling it to accurately generate enhanced latent semantic representation tensors with ECG semantic consistency and temporal stability under heterogeneous input modal data conditions.

[0093] In this embodiment, the step of extracting rhythm drift information based on enhanced latent semantic representation tensors, constructing an individual heart rhythm semantic map, and outputting heart state recognition results specifically includes:

[0094] A sliding time window mechanism is used to segment the enhanced latent semantic representation tensor to generate a sequence of semantic segments arranged in time. For each adjacent semantic segment, the Euclidean distance, cosine distance, element-wise difference, and vector angle change are calculated and fused into rhythmic drift quantity according to preset rules to form a rhythmic drift sequence.

[0095] The rhythm drift sequence is aligned with the window timestamp one by one. Within the aligned time period, the electrophysiological structural features output by the ECG semantic space are extracted, including the location of the P wave initiation, QRS main wave peak, T wave end, PR interval start and end points and QT interval start and end points according to fixed rules. The parameters of the electrophysiological structural features are numericalized, standardized and encoded. The encoded electrophysiological structural features are combined with the corresponding semantic segments and rhythm drift amount to form a single cardiac state event node and stored in the event list in chronological order.

[0096] Edge weights are calculated for adjacent cardiac state event nodes. Semantic differences are obtained by performing dimension-wise subtraction on the semantic vectors of adjacent event nodes and obtaining the absolute or square value of the difference. Electrophysiological structural differences are obtained by performing subtraction on the parameters of electrophysiological structural features. The rate of change of dynamic components in the enhanced latent semantic representation is calculated to obtain the mechanical dynamic rate of change and the rhythm drift amplitude. The semantic difference, electrophysiological structural difference, mechanical dynamic rate of change and rhythm drift amplitude are weighted and summed according to preset weight coefficients to form edge weights that reflect the strength of node association. All nodes and edge weights are organized into a graph structure in chronological order to form a cardiac state semantic graph. Graph aggregation operation is performed within a set sliding aggregation window to obtain a window-level comprehensive semantic representation.

[0097] The window-level comprehensive semantic representation is input into a fully connected classifier. Based on the classifier's output, the cardiac state of the time period to which the window belongs is determined. The output includes the overall cardiac state recognition result, which includes heart rhythm type, electrophysiological structural state, mechanical activity state, rhythm stability state, abnormal event prompts, and corresponding time location.

[0098] Example 1:

[0099] To verify the feasibility of this invention in practice, it was applied to the scenario of "home-based non-invasive cardiac health monitoring". A group of subjects were selected and their chest non-contact reflected electromagnetic signals were collected by a millimeter-wave radar device without wearable devices. The data were then compared and analyzed with the synchronously recorded structured electrocardiogram data to evaluate the performance of this invention in multiple dimensions such as semantic compression, consistency maintenance, heart rhythm recognition accuracy, and health status monitoring.

[0100] The experimental setup used a millimeter-wave radar module with 24GHz continuous wave transmission and high sampling reception capabilities, positioned 40cm in front of the subject's chest. The sampling frequency was 2kHz, and the time window was set to 60 seconds. During data acquisition, subjects remained in a natural resting state, and ambient light and noise levels were kept stable to ensure signal acquisition quality. Each subject was simultaneously connected to an ECG acquisition device to obtain a standard 12-lead ECG signal, which was used as structured ECG data to construct the ECG semantic space and for subsequent cross-modal semantic alignment.

[0101] In the data preprocessing stage, the non-contact reflected electromagnetic signal is demodulated at the same frequency, low-pass filtered, and constructed using complex numbers. Phase information is then extracted based on the arctangent function to construct a standardized electromagnetic phase signal data sequence. Subsequently, feature compression and semantic preservation are performed through a semantic bottleneck modeling module to generate a latent semantic representation tensor. A semantically invariant perturbation signal is introduced to perform intramodal consistency constraint optimization, thereby improving semantic stability.

[0102] On the electrocardiogram (ECG) channel, a β-VAE-structured ECG semantic space construction module is trained using a disease-labeled dataset to generate semantic anchors for cross-modal semantic alignment. Dual-constraint alignment of representation and decoding is performed on the electromagnetic latent semantic representation tensor to complete the semantic mapping learning from electromagnetic signals to ECG semantics.

[0103] The fused multimodal data is fed into an improved TSMixer model, which includes a modal structure-guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head. The model not only learns the fused representations of electromagnetic and electrocardiographic structures but also enhances cross-modal temporal understanding while maintaining modal independence, ultimately outputting an enhanced latent semantic representation tensor. This tensor is input to a rhythm information extraction module to extract rhythm drift information, which, combined with temporal information, constructs an individual heart rhythm semantic map, enabling the identification of heart rhythm status and the monitoring and analysis of health trends.

[0104] To quantify the performance differences between this invention and existing non-contact cardiac monitoring solutions in key metrics, a control experiment was designed to compare metrics such as rhythm detection accuracy, semantic consistency score, and heart rhythm recognition F1 score. The results are shown in Table 1.

[0105] Table 1 Comparative Experiment Results

[0106]

[0107] As shown in Table 1, the rhythm detection accuracy of this invention is generally higher than that of traditional methods, with an average improvement of nearly 10 percentage points, indicating its significant advantage in capturing subtle changes in heart rhythm. The semantic consistency score of this invention is approximately 0.88, significantly better than the approximately 0.65 of traditional methods, indicating that its latent semantic representation has stronger stability and expressive power for cardiac physiology. The F1 score for heart rhythm state recognition is generally over 0.90, far higher than the 0.70 level of traditional methods, indicating that this invention has superior performance in disease recognition tasks. These improvements are attributed to the semantic bottleneck compression design, intramodal perturbation constraint mechanism, cross-modal semantic mapping learning, and the improved TSMixer model that introduces structure guidance and jump time modeling, effectively enhancing semantic deconstruction and discrimination capabilities.

[0108] This method, without requiring patches, wearable devices, or invasive procedures, accurately extracts deep semantic information related to cardiac activity from reflected electromagnetic waves and performs semantic alignment, state assessment, and health monitoring. It boasts the following significant advantages: strong non-contact capability, suitable for various environments such as homes, medical institutions, and public places; cross-modal alignment enhances interpretability, and semantic stability improves clinical usability; the improved TSMixer possesses structure-guided and time-enhanced capabilities, adapting to multi-source heterogeneous data; and the final cardiac rhythm semantic map supports dynamic monitoring of rhythm change trends, providing effective support for individual health management. Therefore, this method not only improves the accuracy and practicality of non-contact cardiac monitoring but also provides a highly scalable and interpretable new paradigm for intelligent perception of vital signs.

[0109] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A non-contact electromagnetic heart monitoring method based on signal semantic deconstruction, characterized in that, Includes the following steps: Non-contact reflected electromagnetic signals from the target area are collected, raw co-frequency orthogonal component data are obtained, and phase demodulation and filtering are performed to generate electromagnetic phase signal data. Semantic bottleneck modeling is performed on electromagnetic phase signal data, and the electromagnetic phase signal data is compressed while retaining key semantic information related to cardiac physiological semantics, generating a latent semantic representation tensor. A semantically invariant perturbation signal is constructed and input into the encoder along with electromagnetic phase signal data for self-supervised reconstruction. The semantic stability of the latent semantic representation tensor is optimized through intramodal semantic invariant compression and consistency constraints. An electrocardiogram (ECG) semantic space is constructed based on structured ECG data, ECG semantic feature representations are extracted, and the latent semantic representation tensor is cross-modal semantically aligned with the ECG semantic space to establish a semantic mapping relationship. A multimodal temporal tensor containing electromagnetic phase signal data, semantically invariant perturbation signal data, and synchronized electrocardiogram signal data is constructed and input into an improved TSMixer model. The improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head, and outputs an enhanced latent semantic representation tensor. Based on the enhanced latent semantic representation tensor, rhythm drift information is extracted, an individual heart rhythm semantic map is constructed, and the heart state recognition result is output. The improved TSMixer model includes a modal structure guided encoder, a dual-channel cross-mixing module, a skip-time structure block, and a joint multi-task semantic output head. The modal structure guided encoder receives and encodes electromagnetic phase signal data, semantically invariant perturbation signal data, and structured electrocardiogram data. The modal structure guided encoder includes three parallel sub-paths, which respectively use structural units containing one-dimensional convolutional layers, batch normalization layers, and ReLU activation functions to extract short-term dynamic features, and use a position embedding module for time alignment. Then, a modal structure guided matrix is ​​used to perform channel weighting and output a multimodal embedding feature tensor. The dual-channel cross-hybrid module includes a residual fusion channel and an attention interaction channel. The residual fusion channel uses a weighted average and residual connection method to perform fusion on the multimodal embedded feature tensor. The attention interaction channel uses a dot product attention mechanism to calculate the semantic dependency between any two modalities and performs weighted reconstruction. The fused feature representation is normalized using Softmax normalization to obtain a cross-modal joint representation tensor. The skip-time structure block uses an exponentially growing skip window structure to perform one-dimensional convolution operations with different interval lengths on the cross-modal joint representation tensor. The size of the convolution kernel is dynamically adjusted according to the time window. The time structure results of all scales are spliced ​​together and the dimensions are compressed through a fully connected linear projection layer to form a time-enhanced representation tensor. The joint multi-task semantic output head receives the time-enhanced representation tensor and performs the latent semantic enhancement task and the auxiliary rhythm label prediction task in parallel. The latent semantic enhancement task uses a residual feedforward network and the GELU activation function to extract deep semantic features and outputs the enhanced latent semantic representation tensor. The auxiliary rhythm label prediction task uses a fully connected layer to output the rhythm classification probability.

2. The non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction according to claim 1, characterized in that, The non-contact reflected electromagnetic signal acquisition of the target area is used to obtain the original co-frequency orthogonal component data, which is then subjected to phase demodulation and filtering to generate electromagnetic phase signal data. Specifically, this includes: The non-contact reflected electromagnetic signal is generated by emitting continuous electromagnetic waves in front of the target's chest using a millimeter-wave radar device and receiving the electromagnetic signals reflected back from the chest area. The non-contact reflected electromagnetic signal is demodulated at the same frequency to obtain the original orthogonal component data at the same frequency; The original orthogonal component data of the same frequency were subjected to low-pass filtering and sampling was synchronized at a set sampling rate to obtain a clean data sequence corresponding to the micro-motion changes on the chest surface. The real and imaginary data are constructed into complex numbers according to the sampling points, and the instantaneous phase at each moment is extracted based on the arctangent calculation method to form the original phase sequence. The original phase sequence is normalized, and the phase curve is standardized into time-series data of a set length within a preset time window, and the electromagnetic phase signal data is output.

3. The non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction according to claim 1, characterized in that, The step of performing semantic bottleneck modeling on electromagnetic phase signal data, compressing the electromagnetic phase signal data while retaining key semantic information related to cardiac physiological semantics, and generating a latent semantic representation tensor specifically includes: An electromagnetic phase signal data is encoded using a feature encoder consisting of a one-dimensional temporal convolutional network and a Swish activation function, and the output latent semantic representation tensor is generated. Based on the information theory modeling method, the mutual information value between the latent semantic representation tensor and the preset target cardiac semantic label is used as the semantic preservation target to perform the maximization operation, and the mutual information value between the latent semantic representation tensor and the original electromagnetic phase signal data is used as the compression constraint target to perform the minimization operation. To maximize the semantic preservation objective, adjustable weight coefficients are introduced for joint optimization with the minimization of compression constraint objective.

4. The non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction according to claim 1, characterized in that, The construction of semantically invariant perturbation signals, which are then input into the encoder along with electromagnetic phase signal data for self-supervised reconstruction, optimizes the semantic stability of the latent semantic representation tensor through intramodal semantic invariant compression and consistency constraints. Specifically, this includes: Randomly select a segment of original orthogonal component data of non-current sample from the training set, apply low-pass filtering to extract background low-frequency components, add them point by point in complex form with the original orthogonal component data of the target sample, and perform phase calculation to form a perturbation phase sequence with background interference characteristics as a semantically invariant perturbation signal. Electromagnetic phase signal data and semantically invariant perturbation signal are simultaneously input into a time-series encoder to obtain the original latent semantic representation tensor and the perturbation representation tensor, respectively. The original latent semantic representation tensor is input into the reconstruction decoder to recover the reconstructed signal; A joint loss function is constructed to simultaneously optimize the semantic invariant compression term and the consistency constraint term. The semantic invariant compression term minimizes the error between the electromagnetic phase signal data and the reconstructed signal, while the consistency constraint term minimizes the structural difference between the latent semantic representation tensors before and after the perturbation. The parameters of the time encoder are optimized by minimizing the joint loss function.

5. The non-contact electromagnetic cardiac monitoring method based on signal semantic deconstruction according to claim 1, characterized in that, The process of constructing an ECG semantic space based on structured ECG data, extracting ECG semantic feature representations, and aligning the latent semantic representation tensor with the ECG semantic space across modalities to establish a semantic mapping relationship specifically includes: Structured electrocardiogram (ECG) data is input into the ECG encoder, which is connected to the ECG decoder and the semantic classifier. A β-variable autoencoder structure was used to jointly train the ECG encoder, ECG decoder, and semantic classifier. The training objective function included reconstruction error, KL divergence of latent semantic representation tensor, and disease label-based supervised loss. The synchronously acquired structured electrocardiogram data is input into the electrocardiogram encoder to generate electrocardiogram semantic anchor points, and the electromagnetic phase signal data is input into the electromagnetic encoder to extract the electromagnetic latent semantic representation tensor. Perform cross-modal semantic alignment operation, which refers to performing consistency matching between the electromagnetic latent semantic representation tensor and the electrocardiogram semantic anchor point, and constructing a cross-modal loss function based on representation alignment constraints and decoding alignment constraints for joint optimization; By minimizing the cross-modal loss function, the latent semantic representation tensor can be made to have the ability to express the structure of the ECG semantic space without relying on the ECG signal.

6. The non-contact electromagnetic heart monitoring method based on signal semantic deconstruction according to claim 1, characterized in that, The process of extracting rhythm drift information based on enhanced latent semantic representation tensors, constructing an individual heart rhythm semantic map, and outputting heart state recognition results specifically includes: A sliding time window mechanism is used to segment the enhanced latent semantic representation tensor to generate a sequence of semantic segments arranged in time. For each adjacent semantic segment, the Euclidean distance, cosine distance, element-wise difference, and vector angle change are calculated and fused into rhythmic drift quantity according to preset rules to form a rhythmic drift sequence. The rhythm drift sequence is aligned with the window timestamp one by one. The electrophysiological structural features output by the ECG semantic space are extracted within the aligned time period. The parameters of the electrophysiological structural features are numericalized, standardized and encoded. The encoded electrophysiological structural features are combined with the corresponding semantic segments and rhythm drift amount to form a single cardiac state event node and stored in the event list in chronological order. The edge weights are calculated for adjacent cardiac state event nodes. The semantic difference, electrophysiological structural difference, mechanical dynamic change rate and rhythm drift amplitude are weighted and summed according to preset weight coefficients to form the edge weights. All nodes and edge weights are organized into a graph structure in chronological order to form a cardiac state semantic graph. The graph aggregation operation is performed within a set sliding aggregation window to obtain a window-level comprehensive semantic representation. The window-level comprehensive semantic representation is input into a fully connected classifier. Based on the classifier's output, the cardiac state of the time period to which the window belongs is determined. The output includes cardiac state recognition results, including heart rhythm type, electrophysiological structural state, mechanical activity state, rhythm stability state, abnormal event prompts, and corresponding time location.

Citation Information

Patent Citations

  • Electrocardiogram monitoring method based on electromagnetic signals and monitoring model training method and device

    CN118512183A

  • Non-contact electrocardio feature point detection method and system based on millimeter wave radar

    CN119523493A