A method, device and equipment for generating an electrocardiogram report

CN122604388APending Publication Date: 2026-08-21BEIJING XINXINLIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610411773.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

但是,由于心电图类型较多,识别困难,所以,医生手动写心电图报告存在诸多问题,例如,这种相对重复且复杂的工作,消耗医生的时间和人力成本;又例如,人工操作不可避免地会产生一些误差,导致心电图报告不准确;再例如,不同医生之间写报告的习惯不同,使得心电图报告不容易形成统一的规范,不好管理

Benefits of technology

[0037] This application provides a method for generating an electrocardiogram (ECG) report. The method may include, for example,: first, acquiring an acquired ECG; then, based on the ECG and a first sub-model in a generation model, obtaining a first image and a second image corresponding to the ECG, where the first image represents abnormal points in the ECG and the second image represents abnormal segments in the ECG; then, based on the ECG, the first image, the second image, and the second sub-model in the generation model, obtaining an ECG report, which includes the ECG and a first descriptive text related to the abnormal points and abnormal segments. In this way, the generation model can accurately identify abnormal points and abnormal segments in the ECG, and based on this, perform signal analysis and language conversion to obtain prepared descriptive text, thereby forming an ECG report with the ECG. This overcomes the problems of doctors writing ECG reports based on ECGs, which consumes doctors' time and manpower; inaccurate ECG reports due to manual errors; and inconsistent ECG report formats due to different doctors, making it possible to efficiently generate objective and accurate ECG reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122604388A_ABST
    Figure CN122604388A_ABST
Patent Text Reader

Abstract

The application discloses a method, device and equipment for generating an electrocardiogram report. Based on a to-be-processed electrocardiogram and a first sub-model in a generation model, a first image and a second image corresponding to the electrocardiogram are obtained, the first image is used to represent abnormal points in the electrocardiogram, and the second image is used to represent abnormal segments in the electrocardiogram. Then, based on the electrocardiogram, the first image, the second image and a second sub-model in the generation model, an electrocardiogram report is obtained, the electrocardiogram report comprises the electrocardiogram and a first description text, and the first description text is related to the abnormal points and the abnormal segments. In this way, the generation model can accurately identify the abnormal points and the abnormal segments in the electrocardiogram first, and then signal analysis and language conversion are performed on the basis to obtain a prepared description text, so as to form the electrocardiogram report together with the electrocardiogram, and the purpose of efficiently generating an objective and accurate electrocardiogram report is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, specifically to a method, apparatus, and device for generating electrocardiogram (ECG) reports. Background Technology

[0002] An electrocardiogram (ECG) is a technique that uses an electrocardiograph to record the changes in electrical activity of the heart during each cardiac cycle from the body surface. It is a commonly used, non-invasive, easy-to-operate, cost-effective method for screening and diagnosing arrhythmias and cardiovascular diseases in clinical practice. The ECG image (hereinafter referred to as ECG) is a commonly used tool for cardiac diagnosis.

[0003] After an electrocardiogram (ECG) is acquired, doctors typically need to perform signal analysis and identification based on the acquired ECG, and then manually add various identified parameters and conclusions, along with other textual content, to the ECG. Figure 1 Initially, the data was provided to the individual in the form of an electrocardiogram (ECG) report. However, due to the variety of ECG types and the difficulty in identification, manually writing ECG reports by doctors presents several problems. For example, this relatively repetitive and complex task consumes doctors' time and manpower; secondly, manual operation inevitably introduces some errors, leading to inaccurate ECG reports; and thirdly, different doctors have different reporting habits, making it difficult to establish a unified standard for ECG reports and hindering management.

[0004] Therefore, there is an urgent need to provide a solution for generating electrocardiogram (ECG) reports efficiently, accurately, and objectively, to replace doctors manually writing ECG reports. Summary of the Invention

[0005] In view of this, this application provides a method, apparatus, and device for generating electrocardiogram (ECG) reports, which can perform signal analysis on abnormal points and abnormal segments in an ECG using a model, and automatically generate accurate ECG reports by using language analysis on the ECG including the analyzed abnormal points and abnormal segments, thereby improving the efficiency and quality of ECG report generation.

[0006] To address the above problems, the technical solutions provided in this application are as follows:

[0007] In a first aspect, this application provides a method for generating an electrocardiogram (ECG) report, the method comprising:

[0008] Acquire the collected electrocardiogram;

[0009] Based on the electrocardiogram and the first sub-model in the generation model, a first image and a second image corresponding to the electrocardiogram are obtained. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram.

[0010] Based on the electrocardiogram, the first image, the second image, and the second sub-model in the generation model, an electrocardiogram report is obtained. The electrocardiogram report includes the electrocardiogram and a first descriptive text, which is related to the abnormal point and the abnormal segment.

[0011] Optionally, the method further includes:

[0012] The first sub-model is trained based on multiple sets of first training data. Each set of first training data includes a third image, a fourth image, and a fifth image. The third image is the original electrocardiogram, the fourth image is the image in the third image marked with abnormal points, and the fifth image is the image in the third image marked with abnormal segments.

[0013] Optionally, the method further includes:

[0014] Based on the first training data and the spatial transformation network, transformed first training data is obtained, and the multiple sets of first training data include the transformed first training data.

[0015] Optionally, the method further includes:

[0016] The second sub-model is trained based on multiple sets of second training data and the first sub-model that has been trained. Each set of second training data includes a sixth image and a second descriptive text.

[0017] Optionally, the generative model includes a vision module and a language module, wherein the vision module is used to extract and encode features from the input image, and the language module is used to perform semantic analysis on the input features.

[0018] Optionally, the vision module includes any of the following types: deep residual network, densely connected convolutional network, or simple baseline network.

[0019] Optionally, the language module includes any of the following types: recurrent neural network, long short-term memory network, bidirectional long short-term memory network, or Transformer.

[0020] Secondly, this application provides an apparatus for generating an electrocardiogram (ECG) report, the apparatus comprising:

[0021] The acquisition unit is used to acquire the collected electrocardiogram (ECG).

[0022] The first processing unit is used to obtain a first image and a second image corresponding to the electrocardiogram based on the electrocardiogram and a first sub-model in the generation model. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram.

[0023] The second processing unit is configured to obtain an electrocardiogram report based on the electrocardiogram, the first image, the second image, and the second sub-model in the generation model. The electrocardiogram report includes the electrocardiogram and a first descriptive text, wherein the first descriptive text is related to the abnormal point and the abnormal segment.

[0024] Optionally, the device further includes:

[0025] The first training unit is used to train the first sub-model based on multiple sets of first training data. Each set of first training data includes a third image, a fourth image, and a fifth image. The third image is the original electrocardiogram, the fourth image is the image in the third image marked with abnormal points, and the fifth image is the image in the third image marked with abnormal segments.

[0026] Optionally, the device further includes:

[0027] The transformation unit is used to obtain transformed first training data based on the first training data and the spatial transformation network, wherein the plurality of sets of first training data include the transformed first training data.

[0028] Optionally, the device further includes:

[0029] The second training unit is used to train the second sub-model based on multiple sets of second training data and the first sub-model that has been trained. Each set of second training data includes a sixth image and a second descriptive text.

[0030] Optionally, the generative model includes a vision module and a language module, wherein the vision module is used to extract and encode features from the input image, and the language module is used to perform semantic analysis on the input features.

[0031] Optionally, the vision module includes any of the following types: deep residual network, densely connected convolutional network, or simple baseline network.

[0032] Optionally, the language module includes any of the following types: recurrent neural network, long short-term memory network, bidirectional long short-term memory network, or Transformer.

[0033] Thirdly, this application provides an electronic device comprising: a processor and a memory. The memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs stored in the memory, causing the electronic device to perform the methods described in the first aspect above and any one of the first aspects above.

[0034] Fourthly, this application provides a computer-readable medium storing instructions or a computer program that, when executed on a processor, causes the processor to perform the methods described in the first aspect above and any one of the first aspects above.

[0035] Fifthly, this application provides a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the methods described in the first aspect above and any one of the first aspects above.

[0036] Therefore, the embodiments of this application have the following beneficial effects:

[0037] This application provides a method for generating an electrocardiogram (ECG) report. The method may include, for example,: first, acquiring an acquired ECG; then, based on the ECG and a first sub-model in a generation model, obtaining a first image and a second image corresponding to the ECG, where the first image represents abnormal points in the ECG and the second image represents abnormal segments in the ECG; then, based on the ECG, the first image, the second image, and the second sub-model in the generation model, obtaining an ECG report, which includes the ECG and a first descriptive text related to the abnormal points and abnormal segments. In this way, the generation model can accurately identify abnormal points and abnormal segments in the ECG, and based on this, perform signal analysis and language conversion to obtain prepared descriptive text, thereby forming an ECG report with the ECG. This overcomes the problems of doctors writing ECG reports based on ECGs, which consumes doctors' time and manpower; inaccurate ECG reports due to manual errors; and inconsistent ECG report formats due to different doctors, making it possible to efficiently generate objective and accurate ECG reports. Attached Figure Description

[0038] Figure 1 A flowchart illustrating a method for generating an electrocardiogram report provided in an embodiment of this application;

[0039] Figure 2 A schematic diagram of a generative model provided in an embodiment of this application;

[0040] Figure 3 This is a schematic diagram of the structure of a vision module in an embodiment of this application;

[0041] Figure 4 This is a schematic diagram of the structure of a language module in an embodiment of this application;

[0042] Figure 5 This is a schematic diagram of the structure of an electrocardiogram report generation device 500 according to an embodiment of this application. Detailed Implementation

[0043] To make the above-mentioned objectives, features and advantages of the embodiments of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0044] To facilitate understanding and explanation of the technical solutions provided in the embodiments of this application, the background technology of the embodiments of this application will be described below first.

[0045] The electrocardiogram (ECG) in this embodiment can be understood as a 1D image. Each lead ECG signal acquired by the ECG acquisition machine can be formatted as a one-dimensional tensor. Each one-dimensional tensor serves as a numerical representation of the one-dimensional image, resulting in the ECG mentioned in this embodiment.

[0046] High-quality electrocardiograms (ECGs) are crucial for clinical interpretation, diagnosis, and treatment. With the development of artificial intelligence, many deep learning models have been gradually applied to ECG analysis. However, current deep learning models are often designed for specific problems or can only assist in a particular step of the ECG process, resulting in low interpretability and still requiring manual analysis and report writing by doctors. Furthermore, current deep learning models for specific aspects of ECG processing are based on image processing techniques, and the presence of noise and other factors in ECGs makes analysis less accurate.

[0047] It is evident that the current method of generating electrocardiogram (ECG) reports requires doctors to manually write reports, which is an area that needs improvement to meet the needs of saving medical resources, improving medical efficiency, enhancing diagnostic accuracy, and unifying the management of medical records.

[0048] Based on this, embodiments of this application provide a method for generating an electrocardiogram (ECG) report. This method may include, for example,: first, acquiring a collected ECG; then, based on the ECG and a first sub-model in the generation model, obtaining a first image and a second image corresponding to the ECG, wherein the first image is used to characterize abnormal points in the ECG and the second image is used to characterize abnormal segments in the ECG; then, based on the ECG, the first image, the second image, and the second sub-model in the generation model, obtaining an ECG report, wherein the ECG report includes the ECG and a first descriptive text, the first descriptive text being related to the abnormal points and abnormal segments.

[0049] In this way, the generative model can first accurately identify abnormal points and segments in the electrocardiogram (ECG), and then perform signal analysis and language conversion based on this to obtain a prepared descriptive text. This text is then combined with the ECG to form an ECG report, overcoming the problems of doctors writing ECG reports based on ECGs, which consumes doctors' time and manpower, is inaccurate due to manual errors, and has different report formats due to different doctors. This makes it possible to efficiently generate objective and accurate ECG reports.

[0050] It should be noted that the main body implementing this ECG report generation method can be the ECG report generation device provided in the embodiments of this application. This ECG report generation device can be housed in an electronic device or a functional module of an electronic device. For example, this ECG report generation device can be a functional module on a cloud, server, or other network device, or a mobile phone or other terminal device, used to implement the ECG report generation function provided in the embodiments of this application.

[0051] Figure 1 This is a schematic flowchart illustrating a method for generating an electrocardiogram (ECG) report according to an embodiment of this application. This method can be applied to an ECG report generation device, which, for example, can be... Figure 5 The electrocardiogram report generation device 500 shown.

[0052] like Figure 1 As shown, the method may include, for example, the following steps S101 to S103:

[0053] S101, acquire the collected electrocardiogram.

[0054] It is understandable that S101~S103 can be interpreted as the process of automatically generating an ECG report from the collected ECGs using the trained generative model during the inference phase. Prior to the inference phase, the process of training the generative model may also be included.

[0055] Generative models, also known as instance segmentation networks, aim to perform signal analysis on input electrocardiograms (ECGs). Based on the analyzed anomalies and segments, they are encoded and semantically parsed to obtain descriptive text for the ECG, thus outputting an ECG report that includes both the ECG and its corresponding descriptive text. A generative model may include, for example, two parts: a video module and a language module. The video module extracts and encodes features from the input image, while the language module performs semantic analysis on the input features. Both the video module and the language module can be convolutional neural networks (CNNs). Specifically, the vision module can include any of the following types: Deep Residual Network (ResNet), Densely Connected Convolutional Networks (DenseNet), or Simple Baselines for Human Pose Estimation and Tracking (Simple Baseline); the language module can include any of the following types: Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BLSTM), or Transformer.

[0056] For example, the video module is used to analyze the input electrocardiogram (ECG) 0 to obtain an ECG including abnormalities. Figure 1 And the ECG including abnormal segments Figure 2 The video module is also used for ECG monitoring. Figure 1 and electrocardiogram Figure 2 The encoding is performed and fused with the encoding result of ECG 0 to obtain a fused feature vector. The language module receives the fused feature vector output by the video module and extracts text information to output an ECG report including ECG 0 and descriptive text.

[0057] During the training phase of the generative model, a first sub-model in the video module can be trained first, which outputs an electrocardiogram including abnormal points and an electrocardiogram including abnormal segments; then, based on the trained first sub-model, the other parts of the generative model can be trained.

[0058] As an example, training the first sub-model may include, for instance, training the first sub-model based on multiple sets of first training data. Each set of first training data includes a third image, a fourth image, and a fifth image, where the third image is the original electrocardiogram (ECG), the fourth image is an image of the third image with anomalies marked, and the fifth image is an image of the third image with anomaly segments marked. It is understood that for each set of first training data, the third image can be input into the first sub-model, and the first sub-model outputs an ECG including the anomalies. Figure 3 And the ECG including abnormal segments Figure 4 Based on electrocardiogram Figure 3 Differences from the fourth image, and ECG Figure 4 Based on the differences from the fifth image, the parameters of the first sub-model are adjusted. Thus, after training on multiple sets of the first training data, the ECG output of the first sub-model... Figure 3 The difference from the fourth image continued to narrow, ECG Figure 4 The difference from the fifth image continues to decrease, and training of the first sub-model is complete when stopping condition 1 is met. Stopping condition 1 can be, for example, any one of the following conditions: the number of training iterations reaches a predetermined value, the difference no longer decreases, or the difference falls below a predetermined value. It should be noted that the difference in stopping condition 1 can be understood as an electrocardiogram (ECG) reading. Figure 3 Differences from the fourth image 1 and ECG Figure 4 The overall difference determined by difference 2 from the fifth image, for example, can be equal to the sum or average of difference 1 and difference 2.

[0059] Before generalization, a large number of ECGs from users can be collected. Medical professionals then select high-quality images as the third image in the first training dataset. Next, professionals manually label the information of each heartbeat in the signal, including the coordinates and category of abnormal points. These coordinates (also called prompt coordinates) and categories are used as the ground truth of the image. The labeled image is the fourth image corresponding to the third image. The prompt coordinates can be selected from the start of the P wave / QRS wave / TU wave, the peak of the P wave / QRS wave / TU wave, or the end of the P wave / QRS wave / TU wave, etc. Similarly, professionals manually label the bounding boxes and categories of abnormal segments. These boxes can be represented by vertex coordinates or center coordinates + length + width, etc. The labeled image is the fifth image corresponding to the third image. For abnormal segments of the promoter, they can be: PP segment / PR segment / ST segment / QT segment, etc., and can be abnormal heartbeats: atrial type heartbeat (S) / ventricular heartbeat (V) / junctional area heartbeat (J) / atrial flutter (AF) / atrial fibrillation (Af) / atrial pacing (P) / ventricular pacing (A) / atrioventricular sequential pacing (D) / X (artifacts / noise) / etc.

[0060] It should be noted that, in order to increase the richness and quantity of the first training data, after obtaining the limited first training data, the method may further include: obtaining transformed first training data based on the first training data and the spatial transformation network, wherein multiple sets of first training data include the transformed first training data. The spatial transformation network is used to randomly scale the input first training data vertically, randomly flip it horizontally, randomly rotate it vertically, and randomly rotate it horizontally, etc., to expand the training dataset and improve the model's generalization ability.

[0061] Understandably, in order to standardize the format of the first training data collected, signal processing techniques can be used to preprocess the sample signals before generalization, such as removing signal noise and normalization, thereby improving signal quality and model efficiency. After generalization, the signal size can be adjusted to the fixed input signal size required by the model, and the first training data can be cropped to serve as the first training data for training the first sub-model.

[0062] In the above "based on electrocardiogram" Figure 3 Differences from the fourth image, and ECG Figure 4 "Adjusting the parameters of the first sub-model based on the difference from the fifth image," for example, it can also be implemented based on multiple loss functions. Embodiments of this application may use: Binary Cross Entropy (BCE) loss for object categories, Intersection over Union (IOU) loss for bounding boxes, and mask classification BCE loss, where the object category BCE loss indicates which specific category an outlier or outlier segment belongs to; the bounding box IOU loss indicates whether the label size of the outlier or outlier segment is appropriate; and the mask classification BCE loss indicates whether the labeled outlier segment or outlier segment belongs to a true outlier. The loss functions can be as follows: Formula (1) and Formula (2):

[0063] ...Formula (1)

[0064] ...Formula (2)

[0065] In formula (1), L BCE The BCE loss indicates the object category, where p(x) is the model category probability and y is the true cardiac symbol label; L BCE The BCE loss is indicated by the mask classification, where p(x) is the model classification probability and y is the label indicating whether it belongs to the cardiac rhythm. In formula (2), L IOU Indicates bounding box IOU loss, L IOUThe parameter w is the minimum segmentation loss parameter in the segmentation mask generation process. c1 Let be the training scale for channel c1, with a value range of (0, 1); It is a binary mask model; This is a soft prediction mask model.

[0066] It is understandable that after training the first sub-model, the method may further include: training the second sub-model based on multiple sets of second training data and the trained first sub-model, where each set of second training data includes a sixth image and a second descriptive text. The second sub-model can be understood as the part of the generative model other than the first sub-model. This training process may include: inputting the sixth image from the second training data into the generative model; the first sub-model in the generative model outputs a corresponding seventh image including outliers and an eighth image including outlier segments; the second sub-model of the generative model obtains the third descriptive text based on the sixth, seventh, and eighth images; then, based on the difference between the third and second descriptive texts, the second sub-model is adjusted. Thus, after training with multiple sets of second training data, the difference between the third and second descriptive texts output by the second sub-model continuously decreases, and the training of the second sub-model is completed when stopping condition 2 is met. Stopping condition 2 may include, for example, any one of the following conditions: the number of training iterations reaches a predetermined value, the difference no longer decreases, or the difference is lower than a predetermined value. It should be noted that the difference in stopping condition 2 is the difference between the third and second descriptive texts.

[0067] Among them, "adjusting the second sub-model based on the difference between the third and second descriptive texts" can also be implemented based on a loss function. For example, the embodiments of this application can use KL divergence loss (Kullback–Leibler Divergence Loss, KL Div loss). The second sub-model can directly output descriptive text, which belongs to text information, thus reducing the precision loss and improving the accuracy. The loss function can be implemented by the following formula (3):

[0068] ...Formula (3)

[0069] in, Indicating KL divergence loss, This represents the model's predicted value. This represents the actual values ​​of the model (which can also be understood as the model's labeled values), and N represents the number of samples.

[0070] The structure and operation of the generative model can be found below. Figures 2-4 Related descriptions.

[0071] It is evident that the trained generative model, as a tool for automatically generating electrocardiogram (ECG) reports, makes it possible to generate ECG reports efficiently and accurately.

[0072] After obtaining the trained generative model, the collected electrocardiogram (ECG) to be generated can be used as input to the generative model. Based on the following steps S102~S103, the corresponding ECG report will be automatically generated.

[0073] S102, based on the electrocardiogram and the first sub-model in the generative model, obtain the first image and the second image corresponding to the electrocardiogram. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram.

[0074] As an example, S102 may include: inputting an electrocardiogram (ECG) into a first sub-model in the generation model, which performs signal analysis on the ECG to obtain a first image and a second image corresponding to the ECG, which serve as the basis for subsequent video analysis and semantic analysis, making it possible to accurately generate an ECG report.

[0075] S103, based on the electrocardiogram, the first image, the second image, and the second sub-model in the generative model, an electrocardiogram report is obtained. The electrocardiogram report includes the electrocardiogram and a first descriptive text, which is related to abnormal points and abnormal segments.

[0076] As an example, an electrocardiogram, a first image, and a second image can both be input into a second sub-model in the generative model, which outputs an electrocardiogram report.

[0077] In the actual reasoning process, the user is unaware of the first and second images. The first and second images can be understood as intermediate parameters in the process of the generative model automatically generating a report from the electrocardiogram. The user (such as a doctor) only needs to input the electrocardiogram into the generative model, and the generative model can output the corresponding electrocardiogram report.

[0078] The first descriptive text may include, but is not limited to: ECG parameters, a general description of the report, the location and type of abnormal heartbeats, etc.

[0079] As can be seen, this method enables the generative model to accurately identify abnormal points and segments in the electrocardiogram (ECG), and then perform signal analysis and language conversion based on this to obtain a prepared descriptive text. This text, combined with the ECG, forms an ECG report, overcoming the problems of doctors writing ECG reports based on ECGs, which consumes doctors' time and manpower, is inaccurate due to manual errors, and has inconsistent report formats due to different doctors. This makes it possible to efficiently generate objective and accurate ECG reports.

[0080] For the training phase of the generative model, the training data (which can be called the dataset) is divided into a training set, a validation set, and a test set. At the beginning, an initial model can be set up, and random numbers can be used as the initial values ​​of the parameters of the initial model. The parameters of the initial model are trained using sample images from the training set, and the error of the initial model is detected by the validation set. The parameters of the initial model can be adjusted by gradient descent to optimize the parameters of the initial model and improve the accuracy of the initial model. After the initial model is trained, the performance of the initial model can be evaluated based on the test set. This phase can be called the evaluation phase. For example, the evaluation phase can use the Metric for Evaluation of Translation with Explicit Ordering (METEOR) or the Bilingual Evaluation Understudy (BLEU) to evaluate the effect of the initial model after training. Taking METEOR as an example, it can be evaluated based on the following formula (4):

[0081] ...Formula (4)

[0082] in, Recall rate is the ratio of the number of matched terms to the number of reference description terms. This represents a penalty factor used to account for operations involving the insertion, deletion, and replacement of words. This indicates the fragment matching degree, used to account for trimming and expansion operations; This is a parameter, which can be set to 0.5 to balance the impact of recall and penalty factor.

[0083] For example, generative models can be found in [reference needed]. Figure 2 As shown. Generative model 1 may include a video module 11 and a language module 12. The video module 11 may include a first sub-model 111 and a video sub-model 112. The language module 12 and the video sub-model 112 are collectively referred to as the second sub-model 13. Therefore, in other words, generative model 1 can also be described as including a first sub-model 111 and a second sub-model 13. The video sub-model 112 may include, for example, an image patching unit 1121, a prompt encoder unit 1122, and a feature encoding unit 1123.

[0084] In the first stage of the training phase, the first sub-model 111 can be trained separately. After the first sub-model 111 is trained, the second sub-model 13 is trained. Taking the original electrocardiogram 0 and descriptive text 0 in the second training data as an example, the process of training the second sub-model 13 may include: inputting the original electrocardiogram 0 into the first sub-model 111, and the first sub-model 111 outputs image 1 including abnormal points and image 2 including abnormal segments; on the one hand, the original electrocardiogram 0 can be input into the patch image unit 1121, and output signal 1; on the other hand, image 1 and image 2 can be input into the prompt encoder 1122, and output signal 2; then, signal 1 and signal 2 are input into the feature encoding unit 1123, and the feature encoding unit 1123 outputs signal 3; then, signal 3 is input into the language module 12, and the language module 12 outputs descriptive text 1. The second sub-model 13 is adjusted based on the difference between descriptive text 0 and descriptive text 1.

[0085] Similarly, in the inference phase, taking the original ECG 0 to be analyzed as an example, the inference process can include: inputting the original ECG 0 into the first sub-model 111, which outputs image 1 and image 2; on one hand, inputting the original ECG 0 into the patchimage unit 1121, which outputs signal 4; on the other hand, inputting image 1 and image 2 into the prompt encoder 1122, which outputs signal 5; then, inputting signals 4 and 5 into the feature encoding unit 1123, which outputs signal 6; and finally, inputting signal 6 into the language module 12, which outputs descriptive text 2. Because the training data has good quality and generalization ability, the trained generative model 1 has good quality. Therefore, the descriptive text 2 output in the inference phase is the same as or very close to the ideal descriptive text 0 corresponding to the original ECG 0. That is, based on the trained generative model 1, high-quality ECG descriptive text can be obtained, thus making it possible to automatically generate high-quality ECG reports.

[0086] For the vision module 11, one possible structure is shown in Figure 3. This vision module 11 can be considered to consist of multiple convolutional layers, multiple downsampling layers, and fully connected layers. For example, the size of the original signal input to the vision module 11 can be 4096 (length) × 1. This original signal is first subjected to four blocks of sequential convolution and downsampling operations: after the convolution operation of Block 1, a 2048 × 1 signal is output to Block 2; after the convolution operation of Block 2, a 1024 × 1 signal is output to Block 3; after the convolution operation of Block 3, a 512 × 1 signal is output to Block 4; after the convolution operation of Block 4, the signal is processed to obtain the final signal; the final signal passes through two fully connected layers (Linear) to obtain the output signal of the vision module 11.

[0087] like Figure 3 As shown, each block can consist of multiple 3x3 convolutional layers (conv), and blocks are connected by conv and pooling layers (Pooling). Alternatively, the conv and pooling layers can be deployed inside the block. Figure 3 The diagram shows a visual module 11 consisting of four blocks. In practical applications, the number of blocks and the internal structure of each block can be flexibly set according to actual needs.

[0088] For the language module 12, a possible structure can be seen in Figure 4. This language module 12 adopts an encoder-decoder structure, progressively extracting and reconstructing features through downsampling (i.e., convolution) and upsampling (i.e., deconvolution) modules. Skip connections are introduced during the upsampling stage to enhance the recovery of detailed information. It can be considered to consist of multiple convolutional layers, multiple deconvolutional layers, and fully connected layers. For example, the raw signal input to the language module 12 can first undergo a convolutional downsampling operation once sequentially using six encoders; then, it can undergo a deconvolutional upsampling operation once sequentially using six decoders. The input to each decoder includes the output of the previous decoder and the output of the last encoder. See Figure 4 for details. Figure 4 As shown. Thus, this continues until the 6th decoder ( Figure 4 After the signal is processed by the decoder 6 (output), the resulting image is the descriptive text corresponding to the original signal.

[0089] Figure 4Each encoder in the language module can contain multiple 3x3 convolutions, each followed by a group normalization (GN) layer and a leaky rectified linear unit (LER) activation function. Encoders are connected by max pooling. Each decoder can contain multiple 3x3 convolutions, each followed by a GN and a leaky ReLU activation function. Decoders are connected by transposed convolutions. In practical applications, the number of encoders and decoders can be flexibly set according to actual needs. It should be noted that language module 12 uses the leaky ReLU activation function, and the use of GN can prevent overfitting.

[0090] In this embodiment, signal analysis in electrocardiograms (ECGs) and video and language analysis based on abnormal signals are integrated into a generative model. This model can automatically generate ECG reports, reducing repetitive manual reporting by doctors and other professionals, and also reducing patient waiting time for handwritten reports. It should be noted that the method provided in this embodiment can also be applied to the automatic generation of other medical image reports, such as automatically generating corresponding reports for chest CT images. The difference lies in the fact that abnormal points and segments in chest CT images differ from those in ECGs, and the principles underlying the analysis and processing of the first training data and the first sub-model are different.

[0091] Thus, through this method, the generative model can first accurately identify abnormal points and segments in the electrocardiogram (ECG), and then perform signal analysis and language conversion based on this to obtain prepared descriptive text. This text is then combined with the ECG to form an ECG report, overcoming the problems of doctors writing ECG reports based on ECGs, which consumes doctors' time and manpower, is inaccurate due to manual errors, and has inconsistent report formats due to different doctors. This makes it possible to efficiently generate objective and accurate ECG reports.

[0092] See Figure 5 As shown, this figure is a structural schematic diagram of an electrocardiogram report generation device 500 provided in an embodiment of this application. Figure 5 As shown, the device 500 may include at least:

[0093] Acquisition unit 501 is used to acquire the collected electrocardiogram;

[0094] The first processing unit 502 is used to obtain a first image and a second image corresponding to the electrocardiogram based on the electrocardiogram and a first sub-model in the generation model. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram.

[0095] The second processing unit 503 is configured to obtain an electrocardiogram report based on the electrocardiogram, the first image, the second image, and the second sub-model in the generation model. The electrocardiogram report includes the electrocardiogram and a first descriptive text, wherein the first descriptive text is related to the abnormal point and the abnormal segment.

[0096] Optionally, the device 500 further includes:

[0097] The first training unit is used to train the first sub-model based on multiple sets of first training data. Each set of first training data includes a third image, a fourth image, and a fifth image. The third image is the original electrocardiogram, the fourth image is the image in the third image marked with abnormal points, and the fifth image is the image in the third image marked with abnormal segments.

[0098] Optionally, the device 500 further includes:

[0099] The transformation unit is used to obtain transformed first training data based on the first training data and the spatial transformation network, wherein the plurality of sets of first training data include the transformed first training data.

[0100] Optionally, the device 500 further includes:

[0101] The second training unit is used to train the second sub-model based on multiple sets of second training data and the first sub-model that has been trained. Each set of second training data includes a sixth image and a second descriptive text.

[0102] Optionally, the generative model includes a vision module and a language module, wherein the vision module is used to extract and encode features from the input image, and the language module is used to perform semantic analysis on the input features.

[0103] Optionally, the vision module includes any of the following types: deep residual network, densely connected convolutional network, or simple baseline network.

[0104] Optionally, the language module includes any of the following types: recurrent neural network, long short-term memory network, bidirectional long short-term memory network, or Transformer.

[0105] In addition, embodiments of this application also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method as described in any of the above embodiments.

[0106] This application also provides a computer-readable storage medium storing instructions that, when executed on a processor, cause the processor to perform the method described in any of the preceding embodiments.

[0107] This application also provides a computer program product, including computer program instructions that, when executed on a computer, cause the computer to perform the method described in any of the above embodiments.

[0108] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

[0109] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0110] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0111] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0112] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating an electrocardiogram (ECG) report, characterized in that, The method includes: Acquire the collected electrocardiogram; Based on the electrocardiogram and the first sub-model in the generation model, a first image and a second image corresponding to the electrocardiogram are obtained. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram. Based on the electrocardiogram, the first image, the second image, and the second sub-model in the generation model, an electrocardiogram report is obtained. The electrocardiogram report includes the electrocardiogram and a first descriptive text, which is related to the abnormal point and the abnormal segment.

2. The method according to claim 1, characterized in that, The method further includes: The first sub-model is trained based on multiple sets of first training data. Each set of first training data includes a third image, a fourth image, and a fifth image. The third image is the original electrocardiogram, the fourth image is the image in the third image marked with abnormal points, and the fifth image is the image in the third image marked with abnormal segments.

3. The method according to claim 2, characterized in that, The method further includes: Based on the first training data and the spatial transformation network, transformed first training data is obtained, and the multiple sets of first training data include the transformed first training data.

4. The method according to claim 2, characterized in that, The method further includes: The second sub-model is trained based on multiple sets of second training data and the first sub-model that has been trained. Each set of second training data includes a sixth image and a second descriptive text.

5. The method according to any one of claims 1-4, characterized in that, The generative model includes a vision module and a language module. The vision module is used to extract and encode features from the input image, and the language module is used to perform semantic analysis on the input features.

6. The method according to claim 5, characterized in that, The vision module includes any of the following types: deep residual network, densely connected convolutional network, or simple baseline network.

7. The method according to claim 5, characterized in that, The language module includes any of the following types: recurrent neural network, long short-term memory network, bidirectional long short-term memory network, or Transformer.

8. An apparatus for generating an electrocardiogram (ECG) report, characterized in that, The device includes: The acquisition unit is used to acquire the collected electrocardiogram (ECG). The first processing unit is used to obtain a first image and a second image corresponding to the electrocardiogram based on the electrocardiogram and a first sub-model in the generation model. The first image is used to characterize abnormal points in the electrocardiogram, and the second image is used to characterize abnormal segments in the electrocardiogram. The second processing unit is configured to obtain an electrocardiogram report based on the electrocardiogram, the first image, the second image, and the second sub-model in the generation model. The electrocardiogram report includes the electrocardiogram and a first descriptive text, wherein the first descriptive text is related to the abnormal point and the abnormal segment.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a processor, cause the processor to perform the method as described in any one of claims 1-7.