Report generation method for non-contact radar cardiac oscillogram
Through deep learning technology, a multi-stage deep neural network structure is used to achieve the mapping from radar cardiac waveform to natural language description, solving the problem of automatic analysis and interpretation of non-contact radar cardiac signals, improving the efficiency and accuracy of electrocardiogram report generation, and making it suitable for long-term cardiac monitoring of special patient groups.
Patent Information
- Application Number
- CN202510743084.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, automatic analysis and interpretation of non-contact radar cardiac signals are difficult, resulting in low efficiency and poor accuracy in generating electrocardiogram reports. In particular, long-term continuous monitoring and individualized diagnosis cannot be achieved in special patient groups.
A multi-stage deep neural network structure based on deep learning is adopted, including image feature encoding, time series modeling and attention fusion, and language generation module. The model is trained through the cross-entropy loss function to achieve mapping from radar cardiac waveforms to natural language descriptions.
It realizes automatic analysis of radar cardiac waveforms and generation of accurate text reports, improves diagnostic efficiency and consistency, reduces the workload of medical staff, and is suitable for special patient groups requiring non-contact cardiac monitoring.
Smart Images

Figure CN120673965A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of radar sensing technology, and in particular relates to a method for generating a report of a non-contact radar cardiogram based on deep learning technology. Background Art
[0002] The electrocardiogram (ECG) is a crucial biomedical signal used to describe the activity of the human heart. Continuous ECG monitoring and analysis can improve the diagnosis, management, and prevention of many cardiovascular diseases. ECG monitoring relies on capturing the heart's electrical activity through surface electrodes, detecting the subtle electrical changes caused by myocardial depolarization and repolarization during each cardiac cycle, thereby enabling a comprehensive assessment of the heart's physiological function and pathological status.
[0003] This electrode-based monitoring method enables continuous monitoring of ECG signals, providing strong support for cardiac health monitoring. However, in practical applications, the contact-based skin-attachment nature of this method imposes numerous limitations on monitoring reliability and adaptability. On the one hand, adhesive electrodes and invasive monitoring systems can cause discomfort to patients, negatively impacting their monitoring experience. On the other hand, patients with cardiovascular disease require long-term continuous ECG monitoring to detect incidental cardiac dysfunction events. However, battery limitations and electrode dropouts compromise the continuity and reliability of monitoring. Furthermore, in certain special circumstances, such as burn patients, highly infected patients, and premature infants, traditional electrode adhesion methods are unfeasible, further demonstrating the limitations of this detection method. To address these challenges, wireless sensing technology plays a vital role in contactless cardiac monitoring.
[0004] In recent years, non-contact radar signal monitoring technology has emerged, offering a new approach to addressing these issues. This technology utilizes the interaction between radar waves and the human body to acquire heartbeat signals without direct skin contact. Its advantages include being non-invasive and convenient, making it particularly suitable for long-term continuous monitoring and for specialized patient populations. However, with the development of non-contact radar physiological signal monitoring technology, the challenge remains how to efficiently and accurately analyze and interpret the monitored heartbeat signals, generating clear and accurate text reports.
[0005] In clinical practice, ECG interpretation and report generation typically rely on the experience and expertise of medical professionals. However, faced with a large amount of monitoring data, medical professionals often face a heavy workload, prone to fatigue and negligence, which affects the speed and accuracy of diagnosis. Furthermore, interpretation of ECG signals may vary between different medical professionals, which also introduces a certain degree of uncertainty in clinical diagnosis. Therefore, developing a method that can automatically and accurately generate text reports for non-contact radar cardiac waveforms is of great significance for improving medical efficiency, reducing the workload of medical professionals, and enhancing the accuracy and consistency of diagnosis.
[0006] In response to the above problems, the present invention proposes a report generation method for non-contact radar cardiogram based on deep learning technology. Summary of the Invention
[0007] (1) Technical problems to be solved by the present invention:
[0008] The purpose of the present invention is to propose a report generation method for non-contact radar cardiac waveforms to solve the problems raised in the background technology. The present invention uses advanced algorithms and technical means to realize automatic analysis and interpretation of radar cardiac waveforms, and generate clear, accurate and standardized text reports, providing powerful decision support for clinicians and promoting the widespread application of non-contact radar cardiac signal monitoring technology in the medical field.
[0009] (2) In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0010] A report generation method for a non-contact radar cardiogram waveform graph comprises the following steps:
[0011] S1. Dataset construction: Collect synchronous data of continuous wave radar cardiac signals and electrocardiogram signals to construct an image-text pair dataset for semantic understanding of radar cardiac signals;
[0012] S2. Network Architecture: We propose a multi-stage deep neural network architecture for image semantic generation. The modular design includes an image feature encoding module, a temporal modeling and attention fusion module, and a language generation module. These modules work together to achieve a mapping from image information to natural language expression.
[0013] S3. Loss function design: Cross Entropy Loss is used to quantify the difference between the predicted output and the true label;
[0014] S4. Model training and verification: Combine the operations described in S1 to S3 to build a report generation model for radar cardiac waveforms. During the model training phase, batches of images and their corresponding ECG description statement labels are used as input, and the model gradually learns the image-to-language mapping process. After each round of training, the loss change trend is recorded, and the evaluation function is called within the preset test interval to verify the model performance.
[0015] Preferably, the S1 specifically includes the following contents:
[0016] S1.1 Radar signal processing: Obtaining raw radar signals from the radar system, performing distance reconstruction processing on them, and extracting distance change information related to cardiac activity;
[0017] S1.2 ECG signal acquisition and processing: Collect ECG signals and use a bandpass filter to remove noise and irrelevant frequency components;
[0018] S1.3. Heartbeat detection and signal synchronization: The T-wave end point in the ECG signal is used to detect cardiac events. The radar signal is divided into separate samples according to the heartbeat time interval, ensuring that each sample contains two complete cardiac cycles.
[0019] S1.4. Dataset Annotation: Randomly extract multiple segments of radar cardiac waveform signal images from the synchronized dataset of radar cardiac signals and ECG signals of subjects in a resting state. Manually annotate each segment of data with reference to the ECG signal and provide multiple natural language sentences to describe its ECG characteristics, thereby constructing a data set of image and text pairs.
[0020] S1.5. Dataset partitioning: Divide the dataset into training and testing sets to evaluate the performance of the model.
[0021] Preferably, the image semantic encoding module is used to perform multi-scale semantic extraction on the input image, using a pre-trained ResNet-50 network as a front-end encoder to extract multi-scale semantic information of the input image;
[0022] The temporal modeling and attention fusion module includes a temporal modeling submodule and an attention fusion submodule. The temporal modeling submodule uses a single-layer unidirectional LSTM network to encode the image feature sequence after dimensionality reduction to capture the contextual relationship in the image semantic structure; the attention fusion submodule is used to strengthen the model's attention to salient areas of the image and generate a dynamic semantic context vector in a weighted manner.
[0023] The language generation module uses a recursive decoding network based on the LSTMCell structure to fuse the hidden state of the previous moment, the current attention context vector and the word vector embedding information to complete the generation prediction of the current vocabulary.
[0024] Preferably, the S3 specifically includes the following contents:
[0025] S3.1. In the image captioning task, the output of each time step is a probability distribution, which is used to represent the probability of generating each word. For each word probability distribution generated at time step t, the cross entropy loss between it and the true word label is calculated. The specific calculation formula is as follows:
[0026]
[0027] Where T is the length of the subtitle; V is the size of the vocabulary; y t,i Indicates the true label value at time step t and word i, which is 0 or 1; p t,i Represents the probability of the model predicting the generated word i at time step t;
[0028] S3.2. Subtitle generation is a sequence generation task. The prediction of each time step is based on the output of the previous time step. The loss of the entire sequence is the accumulation of the losses of all time steps. The specific calculation formula is as follows:
[0029]
[0030] Where, L seq Represents the minimization loss function;
[0031] S3.3. During the training process, use the minimization loss function L obtained in S3.2 seq Optimize model parameters using backpropagation and gradient descent.
[0032] Preferably, the performance evaluation indicators of the model validation phase in S4 include BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and METEOR (Metric for Evaluation of Translation with Explicit Ordering).
[0033] (3) The beneficial effects of the present invention include:
[0034] (1) The present invention can automatically extract key physiological features from radar signal images and generate natural language descriptions with medical interpretation significance. This invention effectively expands the application boundaries of artificial intelligence in the field of medical health and provides new ideas for non-contact vital sign analysis in smart medical systems. It not only helps to improve the automation level of remote monitoring and auxiliary diagnosis, but also plays an important role in improving medical efficiency, reducing the workload of medical staff, and enhancing diagnostic accuracy and consistency.
[0035] (2) The present invention enables efficient generation of captions for radar cardiac waveform images. The generated captions are semantically relevant and accurate to the image content, while also exhibiting good fluency and readability. This invention has broad application prospects in radar cardiac image caption generation tasks and can provide strong support for understanding and describing radar cardiac image content. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flow chart of a report generation method for non-contact radar cardiogram waveform graphs proposed by the present invention;
[0037] Figure 2 A flowchart for constructing an image-text pair dataset proposed in Example 2 of the present invention;
[0038] Figure 3 This is a diagram of the network architecture proposed in Example 2 of the present invention;
[0039] Figure 4 This is a flowchart of outputting prediction results proposed in Example 3 of the present invention. DETAILED DESCRIPTION
[0040] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0041] Example 1:
[0042] The present invention proposes a report generation method for a non-contact radar cardiogram waveform, comprising the following contents:
[0043] 1. Dataset Construction
[0044] The dataset used in this paper consists of synchronized data from continuous-wave radar cardiac signals and electrocardiograms (ECGs). The radar cardiac signals were acquired using a 24 GHz six-port continuous-wave radar system operating in the ISM band, while the ECG signals were collected using a Task Force Monitor 3040i with a two-lead electrocardiogram (ECG). The experiment collected radar cardiac signals and synchronized ECG signals from 30 healthy subjects at rest, with a measurement time of 10 minutes.
[0045] 2. Build the network structure.
[0046] This paper proposes a multi-stage deep neural network architecture for image semantic generation. This modular design primarily comprises an image feature encoding module, a temporal modeling and attention fusion module, and a language generation module. These modules work together to map image information to natural language expressions.
[0047] Among them, the image semantic encoding module is used to perform multi-scale semantic extraction on the input image, using the pre-trained ResNet-50 network as the front-end encoder to extract the multi-scale semantic information of the input image; the temporal modeling module uses a single-layer unidirectional LSTM network to encode the image feature sequence after dimensionality reduction to capture the contextual relationship in the image semantic structure; the attention mechanism is used to strengthen the model's attention on the salient areas of the image and generate dynamic semantic context vectors in a weighted manner; the language generation module uses a recursive decoding network based on the LSTMCell structure to fuse the hidden state of the previous moment, the current attention context vector and the word vector embedding information to complete the generation prediction of the current vocabulary.
[0048] To further improve the accuracy and fluency of generated text, the model incorporates a beam search strategy to generate the optimal description sequence during the inference phase. The entire network structure is designed to achieve a deep understanding of image content and accurate generation of natural language descriptions, with high scalability and task generalization capabilities.
[0049] 3. Loss function design.
[0050] In order to improve the supervised training effect of the model, the present invention uses the cross entropy loss function to quantify the difference between the predicted output and the true label.
[0051] 4. Model training and validation
[0052] The present invention designs a complete model training and verification process. During the training phase, the model gradually learns the image-to-language mapping process by taking batches of images and their corresponding electrocardiogram description sentence labels as input. After each round of training, the loss change trend is recorded, and the evaluation function is called within the preset test interval to verify the model performance. The evaluation indicators cover the mainstream image subtitle evaluation standards of BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and METEOR (Metric for Evaluation of Translation with Explicit Ordering), which can measure the quality of generated subtitles from different angles.
[0053] Example 2:
[0054] Based on Example 1, but with the following differences: Figure 1-3 The specific implementation process of the report generation method for non-contact radar cardiogram waveform diagram proposed in the present invention is as follows:
[0055] (1) Dataset construction
[0056] The dataset construction process of the present invention involves acquiring data from the radar system and the Task Force Monitor 3040i acquisition system, and processing the data to construct an image-text pair dataset for semantic understanding of radar cardiac signals.
[0057] like Figure 2 As shown in the figure, the specific steps for constructing the data set are as follows:
[0058] Radar signal processing: The raw data obtained from the radar system includes I (in-phase) and Q (quadrature) signals. These signals first undergo range reconstruction processing to extract distance change information that is closely related to cardiac activity.
[0059] ECG signal acquisition and processing: Lead II ECG signals were acquired using the Task Force Monitor 3040i system (in millivolts). To improve signal quality, a bandpass filter (4-20 Hz) was applied to remove noise and irrelevant frequency components.
[0060] Heartbeat detection and signal synchronization: The T-wave end point in the ECG signal is used to accurately detect cardiac events. The radar signal is then divided into separate samples based on the heartbeat time interval, ensuring that each sample contains two complete cardiac cycles.
[0061] Dataset Annotation: For each healthy subject's resting radar and ECG synchronized signal dataset, four segments of range radar cardiac waveform signals were randomly selected. Each segment was manually annotated by professionals using the ECG signal as a reference, and accompanied by five natural language sentences describing its ECG characteristics. This constituted a dataset of image-text pairs, resulting in a total of 120 radar cardiac waveform images for the 30 subjects.
[0062] Dataset Partitioning: To evaluate the performance of the model, the dataset was divided into a training set and a test set. Specifically, data from 25 subjects were used for model training, and data from the remaining 5 subjects were used for model testing and performance evaluation.
[0063] (2) Build a network structure.
[0064] This paper constructs an end-to-end image semantic description generation network, aiming to achieve accurate understanding of radar physiological signal data and natural language generation. The network structure is built on the "encoder-decoder" framework, integrating convolutional neural networks, attention mechanisms and sequence modeling modules to achieve the conversion from low-level image features to high-level language expressions. The system structure is as follows Figure 3 As shown, it can be divided into the following three main modules:
[0065] 2.1) Image Semantic Encoding Module (Visual Feature Extractor)
[0066] This module extracts deep semantic features from the input image. It uses a deep residual neural network (ResNet-50) pre-trained on ImageNet as the image feature extraction architecture. The convolutional layers (excluding the fully connected layers) are used as the front-end encoder to extract multi-scale semantic information from the input image. This operation preserves the spatial structure and multi-level semantic features of the image, effectively improving the model's ability to model image details.
[0067] To map high-dimensional convolutional features into a unified embedding space, this paper introduces a 1×1 convolutional dimensionality reduction module to make the output feature dimensions compatible with the subsequent LSTM structure. Furthermore, a dropout mechanism is used to regularize the encoder output features to reduce the risk of overfitting and improve the model's generalization ability in small sample scenarios.
[0068] 2.2) Temporal Modeling and Attention Fusion Module (Spatial-Semantic Fusion)
[0069] The image feature sequence after dimensionality reduction is further fed into a single-layer unidirectional LSTM network to realize sequence modeling in the sense of time steps and extract the temporal patterns and spatial dependencies in the image features.
[0070] To further enhance semantic alignment during the generation phase, the present invention introduces a content-based attention mechanism module during the decoding phase. This attention mechanism dynamically focuses on the semantics of key regions by calculating the correlation weights between the current decoding state and the image encoding sequence. This generates a context vector that is both contextually consistent and semantically focused. This vector, combined with original image features through a weighted summation, strengthens the model's information perception during semantic generation.
[0071] 2.3) Language Generation Module (Context-Driven Semantic Generator)
[0072] The language generation module is responsible for converting the above-mentioned attention-enhanced semantic vectors into natural language descriptions. The present invention uses LSTMCell as the basic language decoding unit. Its input consists of the word embedding vector of the current time step and the corresponding image context vector, and outputs the hidden state and predicted distribution of the current time step.
[0073] The word embedding layer is constructed to map discrete tokens in the vocabulary into continuous, dense vectors, capturing semantic similarity and reducing sparsity. During the inference phase, to improve the diversity and quality of generated sentences, the present invention introduces a beam search decoding strategy. This strategy retains multiple candidate paths to search for the optimal sentence, effectively preventing the model from falling into local optima.
[0074] The entire network structure uses an end-to-end training mechanism, with each module collaboratively optimized to achieve efficient mapping from image to language. This structure not only has strong semantic modeling capabilities, but also adapts to the generative tasks required in various scenarios where physiological signals and images are integrated.
[0075] (3) Loss function design.
[0076] In this paper, the design of the loss function is a key component of the model training process, which directly affects the learning effect and final performance of the model. In view of the characteristics of the image caption generation task, this paper adopts the cross-entropy loss function (Cross-Entropy Loss) and combines it with the special requirements of the sequence generation task to design the following loss function:
[0077] 3.1) Cross Entropy Loss Function
[0078] The cross-entropy loss function is one of the most commonly used loss functions in classification tasks. It measures the difference between the probability distribution predicted by the model and the probability distribution of the true label. In the image captioning task, the output at each time step is a probability distribution, indicating the probability of generating each word.
[0079] Specifically, for each time step t, the cross entropy loss between the generated word probability distribution and the true word label is calculated.
[0080] The calculation formula of the cross entropy loss function is as follows:
[0081]
[0082] Where T is the length of the subtitle, V is the size of the vocabulary, and y t,i is the value (0 or 1) of the true label at time step t and word i, p t,i is the probability that the model predicts to generate word i at time step t.
[0083] 3.2) Accumulation of Sequence Loss
[0084] Since caption generation is a sequence generation task, the prediction of each time step is based on the output of the previous time step. Therefore, the loss of the entire sequence is the accumulation of the losses of all time steps:
[0085]
[0086] 3.3) Optimization of loss function
[0087] During the training process, the parameters of the model are adjusted by minimizing the loss function L seq This is usually achieved through backpropagation and gradient descent. In each training step, the loss function is first calculated, then the gradient of the loss function with respect to the model parameters is calculated, and finally the model parameters are updated to reduce the loss function.
[0088] Through this loss function design, the present invention can effectively train the image caption generation model, enabling it to generate captions that are relevant and accurate to the input image content.
[0089] (4) Model training and verification.
[0090] The training and verification process of the present invention includes the following steps:
[0091] 4.1) Training Process
[0092] During the training phase, the model calculates the predicted output through forward propagation and calculates the cross-entropy loss between the model output and the ground truth captions. Backpropagation is then used to calculate the gradient of the loss function with respect to the model parameters. The Adam optimizer is used to update the model parameters based on the calculated gradient. This process is repeated until the preset number of training rounds is reached.
[0093] 4.2) Verification process
[0094] In the verification phase, the model generates subtitles using a beam search strategy and is analyzed using a variety of evaluation metrics that can measure the quality of generated subtitles from different perspectives.
[0095] The BLEU series of indicators (BLEU-1 to BLEU-4) measure the degree of one- to four-tuple (n-gram) matching between the generated subtitles and the reference subtitles, respectively. The BLEU value ranges from 0 to 1, and the higher the value, the closer the machine-generated text is to the reference translation. The ROUGE indicator measures the degree of overlap between the generated subtitles and the reference subtitles. METEOR is a metric for evaluating machine translation and text summarization that takes into account synonyms and sentence structure to more comprehensively evaluate the quality of generated text. The METEOR score is based on precision, recall, and synonym matching between candidate sentences and reference sentences. It uses the harmonic mean to balance precision and recall, and introduces synonym matching based on word alignment to better handle vocabulary diversity.
[0096] 4.3) Experimental Results
[0097] The proposed image captioning algorithm demonstrates excellent performance across multiple evaluation metrics. The experimental results are shown in Table 1. BLEU, ROUGE, and METEOR are commonly used evaluation metrics for natural language generation tasks, measuring the precision, recall, fluency, and overall quality of the generated captions, respectively.
[0098] Table 1 Experimental results
[0099]
[0100] The results show that the proposed method achieved a BLEU-1 score of 0.7737, indicating a high degree of match between the generated subtitles and the reference subtitles at the word level. The BLEU score gradually decreases with increasing n-gram length, a common phenomenon in natural language generation tasks, as longer n-grams are more difficult to match. The proposed method achieved a ROUGE score of 0.6370, indicating good overlap between the generated subtitles and the reference subtitles at the sentence level. The proposed method achieved a METEOR score of 0.7968, indicating that the generated subtitles are generally close to the reference subtitles.
[0101] The proposed image captioning algorithm leverages deep learning techniques, combined with a deep residual neural network (ResNet-50) and a long short-term memory (LSTM) network, to efficiently convert radar cardiogram waveforms into natural language descriptions. The algorithm demonstrates excellent performance across multiple evaluation metrics, demonstrating its effectiveness and superiority in the image captioning task. Through internal baseline testing and multiple rounds of iterative optimization, the proposed algorithm achieved satisfactory results across all evaluation metrics, providing a new approach for semantic understanding and natural language description of radar cardiograms.
[0102] Example 3:
[0103] Based on Example 1-2, but the difference is that, in combination with the report generation method for non-contact radar cardiac waveform diagram proposed in the present invention, during the test, the radar signal diagram is input and the prediction result flow is output as follows: Figure 4 The specific contents include:
[0104] Prediction results: the subject exhibited a rate of 66bpm, a pr interval of 0.13s, and aqrs width of 0.1s.
[0105] Real label 1: recorded heartbeat: 66bpm; atrioventricular interval 0.13s; ventricular depolarization 0.10s.
[0106] Ground truth 2: Clinical evaluation showed a pulse rate of 66 bpm, pr of 0.13 s, and qrs of 0.10 s.
[0107] Real label 3: data indicated a heart rate of 66bpm, a pr delay of 0.13seconds, and a qrs duration of 0.10seconds.
[0108] True label 4: the patient's rhythm strip measured 66 bpm, a pr interval of 0.13 s, and a qrs interval of 0.10 s.
[0109] Ground truth label 5: summary: cardiac rate 66 bpm; pr conduction 0.13 s; qrs activation 0.10 s.
[0110] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solutions and improved concepts of the present invention within the technical scope disclosed by the present invention, and these changes should be covered by the scope of protection of the present invention.
Claims
1. A report generation method for non-contact radar cardiogram waveform, characterized in that: The steps include: S1. Dataset construction: Collect synchronous data of continuous wave radar cardiac signals and electrocardiogram signals to construct an image-text pair dataset for semantic understanding of radar cardiac signals; S2. Network Structure Construction: A multi-stage deep neural network structure for image semantic generation is proposed. The network structure adopts a modular design, including an image feature encoding module, a temporal modeling and attention fusion module, and a language generation module. Each module works together to achieve the mapping from image information to natural language expression; S3. Loss function design: The cross entropy loss function is used to quantify the difference between the predicted output and the true label; S4. Model training and verification: Combine the operations described in S1 to S3 to build a report generation model for radar cardiac waveforms. During the model training phase, batches of images and their corresponding ECG description statement labels are used as input, and the model gradually learns the image-to-language mapping process. After each round of training, the loss change trend is recorded, and the evaluation function is called within the preset test interval to verify the model performance.
2. The method for generating a report for a non-contact radar cardiogram according to claim 1, characterized in that: The S1 specifically includes the following contents: S1.1 Radar signal processing: Obtaining raw radar signals from the radar system, performing distance reconstruction processing on them, and extracting distance change information related to cardiac activity; S1.2 ECG signal acquisition and processing: Collect ECG signals and use a bandpass filter to remove noise and irrelevant frequency components; S1.
3. Heartbeat detection and signal synchronization: The T-wave end point in the ECG signal is used to detect cardiac events. The radar signal is divided into separate samples according to the heartbeat time interval, ensuring that each sample contains two complete cardiac cycles. S1.
4. Dataset Annotation: Randomly extract multiple segments of radar cardiac waveform signal images from the synchronized dataset of radar cardiac signals and ECG signals of the subjects in the resting state. Manually annotate each segment of data with reference to the ECG signal and provide multiple natural language sentences to describe its ECG characteristics. Thus constructing a picture-text pair dataset; S1.
5. Dataset partitioning: Divide the dataset into training and testing sets to evaluate the performance of the model.
3. The method for generating a report for a non-contact radar cardiogram according to claim 2, characterized in that: The image semantic encoding module is used to perform multi-scale semantic extraction on the input image, using a pre-trained ResNet-50 network as a front-end encoder to extract multi-scale semantic information of the input image; The temporal modeling and attention fusion module includes a temporal modeling submodule and an attention fusion submodule. The temporal modeling submodule uses a single-layer unidirectional LSTM network to encode the image feature sequence after dimensionality reduction to capture the contextual relationship in the image semantic structure; the attention fusion submodule is used to strengthen the model's attention to salient areas of the image and generate a dynamic semantic context vector in a weighted manner. The language generation module uses a recursive decoding network based on the LSTMCell structure to fuse the hidden state of the previous moment, the current attention context vector and the word vector embedding information to complete the generation prediction of the current vocabulary.
4. The method for generating a report for a non-contact radar cardiogram according to claim 3, characterized in that: The S3 specifically includes the following contents: S3.
1. In the image captioning task, the output of each time step is a probability distribution, which is used to represent the probability of generating each word. For each word probability distribution generated at time step t, the cross entropy loss between it and the true word label is calculated. The specific calculation formula is as follows: Where T is the length of the subtitle; V is the size of the vocabulary; y t,i Indicates the true label value at time step t and word i, which is 0 or 1; p t,i Represents the probability of the model predicting the generated word i at time step t; S3.
2. Subtitle generation is a sequence generation task. The prediction of each time step is based on the output of the previous time step. The loss of the entire sequence is the accumulation of the losses of all time steps. The specific calculation formula is as follows: Where, L seq Represents the minimization loss function; S3.
3. During the training process, use the minimization loss function L obtained in S3.2 seq Optimize model parameters using backpropagation and gradient descent.
5. The method for generating a report for a non-contact radar cardiogram according to claim 4, characterized in that: The performance evaluation indicators in the model validation phase in S4 include BLEU, ROUGE and METEOR.
Citation Information
Patent Citations
Central arterial pressure waveform reconstruction system based on cross-domain cross-modal migration
CN118844967A
Non-contact electrocardio feature point detection method and system based on millimeter wave radar
CN119523493A
Intelligent creeping type landslide hidden danger identification method combining image processing and semantic understanding
CN119851278A