Method, device, and computer program for obtaining neural network model for predicting disease on basis of electrocardiogram data
By employing generative learning and contrastive learning for a neural network model, the method effectively learns latent features of ECG signals with limited data, addressing the challenges of predicting heart diseases and improving diagnostic efficiency and accuracy.
Patent Information
- Application Number
- PCT/KR2024/020323
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-13
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Existing methods for analyzing electrocardiogram (ECG) data to predict heart diseases face challenges due to the variability of signal patterns based on patient-specific physiological characteristics, health status, and disease history, leading to difficulties in securing diverse and high-quality data, which in turn affects the performance and generalization of artificial intelligence models.
A method that combines generative learning and contrastive learning for a neural network model, utilizing self-supervised learning to effectively learn latent features of ECG signals even with limited data, thereby improving the accuracy and reliability of heart disease prediction.
This approach enables the neural network model to overcome limitations in learning data, enhancing the efficiency and usability of heart disease prediction in medical diagnosis, while ensuring accurate and reliable results.
Smart Images

Figure KR2024020323_19062025_PF_FP_ABST
Abstract
Description
Method, device and computer program for obtaining a neural network model for predicting diseases based on electrocardiogram data
[0001] The present disclosure relates to artificial intelligence technology in the medical field, and more particularly, to a method, device, and computer program for performing training on a neural network model to predict a disease.
[0002] Recent advancements in information and communication technology and artificial intelligence are enabling the analysis and utilization of diverse medical data. In particular, the electrocardiogram (ECG), a vital biosignal recording the heart's electrical activity, is essential for the diagnosis and monitoring of heart disease. While ECG data analysis plays a crucial role in determining the type of heart disease, existing analytical methods have relied on a limited workforce, such as specialist physicians and researchers, limiting efficiency and speed.
[0003] With the advent of electrocardiogram analysis methods utilizing AI, attempts are being made to automate the prediction of heart disease. However, unresolved issues still exist. Because electrocardiogram data signal patterns vary significantly depending on a patient's physiological characteristics, health status, and disease history, data reflecting diverse conditions and situations is essential. However, securing sufficient such data is challenging. High-quality electrocardiogram data can only be collected through specialized medical environments and equipment, making data collection time-consuming and costly. Furthermore, a lack of labeled data can degrade the performance of AI models, and without sufficient data diversity, model generalization performance is also difficult to ensure.
[0004] Although various approaches have been attempted to address these issues, a suitable method for effectively learning the latent features of electrocardiogram signals and accurately predicting diseases using limited data has not yet been sufficiently presented.
[0005] The present disclosure has been made in response to the aforementioned background technology, and aims to provide a method, device, and computer program for obtaining a neural network model for predicting a disease based on electrocardiogram data.
[0006] However, the problems to be solved in this disclosure are not limited to the problems mentioned above, and other problems not mentioned can be clearly understood based on the description below.
[0007] A method for obtaining a neural network model for predicting a disease based on time series data, which is performed by a computing device including at least one processor for realizing a task as described above, includes the steps of: dividing first time series data to obtain a plurality of first segmented data; obtaining first learning data corresponding to the time series data by masking at least one of the obtained plurality of first segmented data; and performing self-supervised learning for the first neural network model based on the first learning data.
[0008] Alternatively, the step of performing learning on the first neural network model may include the step of performing generative learning and contrastive learning on the first neural network model based on the first learning data.
[0009] Alternatively, the step of performing generative learning and contrastive learning for the first neural network model may include the step of inputting the first learning data into the first neural network model to obtain output data in which masked segmentation data of the first learning data is restored, and performing generative learning for the first neural network model based on the output data and the time series data.
[0010] Alternatively, the method includes the step of obtaining a plurality of second segmented data by dividing second time series data corresponding to positive pairs or negative pairs for the first time series data, and the step of obtaining second learning data corresponding to the second time series data by masking at least one of the obtained plurality of second segmented data, and the step of performing generative learning and contrastive learning for the first neural network model may include the step of performing contrastive learning for the first neural network model by using the first learning data and the second learning data as input pairs.
[0011] Alternatively, the method may include, after generative learning and contrastive learning for the first neural network model are completed, a step of extracting an encoder included in the first neural network model, a step of combining the extracted encoder and a classifier to obtain a second neural network model, and a step of performing supervised learning-based fine-tuning for the second neural network model to predict a patient's disease based on third time series data to which a preset class is assigned.
[0012] Alternatively, the step of obtaining the plurality of second segmented data may include a step of setting the second time series data as corresponding to a positive pair if the first time series data and the second time series data are obtained from the same object, and a step of setting the second time series data as corresponding to a negative pair if the first time series data and the second time series data are obtained from different objects.
[0013] Alternatively, the step of obtaining the plurality of second segmented data may include the step of adjusting the first time series data to obtain augmented data corresponding to the first time series data, and setting the obtained augmented data to correspond to a positive pair for the first time series data.
[0014] Alternatively, the step of performing generative learning and contrastive learning for the first neural network model may apply a first loss function corresponding to the generative learning and a second loss function corresponding to the contrastive learning in combination to the first neural network model.
[0015] Alternatively, the first neural network model may include an encoder that extracts latent features of the first learning data, a decoder that is connected to the encoder and restores masked segmentation data of the first learning data based on the extracted latent features, and a representer that is connected to the encoder and obtains a representation for the contrastive learning based on the extracted latent features.
[0016] A computing device for obtaining a neural network model for predicting a disease based on time series data for realizing the task described above includes a processor including at least one core and a memory including program codes executable by the processor, wherein the processor divides first time series data to obtain a plurality of first division data, performs masking processing on at least one of the obtained plurality of first division data to obtain first learning data corresponding to the time series data, and performs self-supervised learning for the first neural network model based on the first learning data.
[0017] A computer program stored in a computer-readable storage medium for realizing the task described above, wherein the computer program, when executed on one or more processors, performs operations for obtaining a neural network model for predicting a disease based on time series data, the operations including an operation for obtaining a plurality of first segmented data by dividing first time series data, an operation for obtaining first learning data corresponding to the time series data by masking at least one of the obtained plurality of first segmented data, and an operation for performing self-supervised learning on the first neural network model based on the first learning data.
[0018] A method for acquiring a neural network model for predicting disease based on electrocardiogram (ECG) data, according to one embodiment of the present disclosure, combines generative learning and contrastive learning using ECG data. This allows the neural network model to effectively learn latent features of ECG signals even in situations where labels are insufficient or data is limited. This overcomes training data limitations while enhancing the accuracy and reliability of heart disease prediction, significantly improving efficiency and usability in the field of medical diagnosis.
[0019] FIG. 1 is a block diagram of a computing device according to one embodiment of the present disclosure.
[0020] FIG. 2 is a flowchart of a method for obtaining a neural network model for predicting a disease based on electrocardiogram data according to an embodiment of the present disclosure.
[0021] FIG. 3 is an exemplary diagram of a method for obtaining a neural network model for predicting a disease based on electrocardiogram data according to one embodiment of the present disclosure.
[0022] FIG. 4 is a flowchart illustrating a method for performing generative learning and contrastive learning based on self-supervised learning for a first neural network model according to an embodiment of the present disclosure.
[0023] FIG. 5 is a flowchart illustrating a method for performing generative learning for a second neural network model based on labeled electrocardiogram data according to an embodiment of the present disclosure.
[0024] FIG. 6 is an exemplary diagram illustrating a method for performing supervised learning for a second neural network model based on labeled electrocardiogram data according to an embodiment of the present disclosure.
[0025] FIG. 7 is an exemplary diagram of a method for generating a learning data set for contrastive learning using electrocardiogram data obtained from different subjects according to one embodiment of the present disclosure.
[0026] FIG. 8 is an exemplary diagram of a method for generating a learning data set for contrastive learning using multiple electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0027] FIG. 9 is an exemplary diagram of a method for generating learning data for contrast learning using multiple target electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0028] FIG. 10 is an exemplary diagram of a method for generating a learning data set for contrastive learning by using augmented data obtained by adjusting multiple electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0029] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. The embodiments presented in this disclosure are provided to enable those skilled in the art to utilize or implement the contents of the present disclosure. Accordingly, various modifications to the embodiments of the present disclosure will be apparent to those skilled in the art. That is, the present disclosure may be implemented in various different forms and is not limited to the embodiments described below.
[0030] Throughout the specification of this disclosure, identical or similar drawing numbers refer to identical or similar components. Furthermore, for the purpose of clearly describing the disclosure, drawing numbers for parts in the drawings that are not relevant to the description of the disclosure may be omitted.
[0031] The term "or" as used herein is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified herein or clear from context, "X employs A or B" should be understood to mean either of its natural inclusive permutations. For example, unless otherwise specified herein or clear from context, "X employs A or B" can be interpreted to mean either X employs A, X employs B, or X employs both A and B.
[0032] The term "and / or" as used herein should be understood to refer to and include all possible combinations of one or more of the related concepts listed.
[0033] The terms "comprises" and / or "comprising" as used herein should be understood to mean the presence of certain features and / or components. However, it should be understood that the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other features, other components, and / or combinations thereof.
[0034] Unless otherwise specified in this disclosure or unless the context makes it clear that the singular form is being referred to, the singular should generally be construed to include “one or more.”
[0035] The term "Nth (N is a natural number)" used in this disclosure can be understood as an expression used to distinguish components of this disclosure from each other based on a predetermined standard such as a functional perspective, a structural perspective, or convenience of explanation. For example, components performing different functional roles in this disclosure can be distinguished as a first component or a second component. However, components that are substantially the same within the technical spirit of this disclosure but must be distinguished for convenience of explanation may also be distinguished as a first component or a second component.
[0036] The term "acquisition" as used in this disclosure may be understood to mean not only receiving data through a wired or wireless communication network with an external device or system, but also generating data in an on-device form.
[0037] Meanwhile, the term "module" or "unit" used in the present disclosure can be understood as a term referring to an independent functional unit that processes computing resources, such as a computer-related entity, firmware, software or a part thereof, hardware or a part thereof, or a combination of software and hardware. At this time, the "module" or "unit" may be a unit composed of a single element, or a unit expressed as a combination or set of multiple elements. For example, as a narrow concept, a "module" or "unit" may refer to a hardware element of a computing device or a set thereof, an application program that performs a specific function of software, a processing process implemented through software execution, or a set of instructions for program execution, etc. In addition, as a broad concept, a "module" or "unit" may refer to the computing device itself that constitutes the system, or an application running on the computing device, etc. However, since the above-described concept is only an example, the concept of “module” or “part” may be defined in various ways within a range understandable to those skilled in the art based on the contents of the present disclosure.
[0038] The term "model" as used herein may be understood as a system implemented using mathematical concepts and language to solve a specific problem, a set of software units to solve a specific problem, or an abstract model of a processing process to solve a specific problem. For example, a neural network "model" may refer to the entire system implemented as a neural network that has problem-solving capabilities through learning. In this case, the neural network can have problem-solving capabilities by optimizing the parameters connecting nodes or neurons through learning. A neural network "model" may include a single neural network or a set of neural networks that are a combination of multiple neural networks.
[0039] The term "data" used in this disclosure may include "images," signals, and the like. The term "image" used in this disclosure may refer to multidimensional data composed of discrete image elements. In other words, "image" may be understood as a term referring to a digital representation of an object visible to the human eye. For example, "image" may refer to multidimensional data composed of elements corresponding to pixels in a two-dimensional image. "Image" may refer to multidimensional data composed of elements corresponding to voxels in a three-dimensional image.
[0040] The explanation of the above terms is intended to aid understanding of the present disclosure. Therefore, unless explicitly stated as limiting the contents of the present disclosure, it should be noted that the above terms are not intended to limit the technical ideas of the contents of the present disclosure.
[0041] FIG. 1 is a block diagram of a computing device (100) according to one embodiment of the present disclosure.
[0042] Referring to FIG. 1, a computing device (100) includes a processor (110) (hereinafter, processor (110)) including at least one core and a memory (120).
[0043] However, since FIG. 1 is only an example, the computing device (100) may further include other configurations for implementing a computing environment. Furthermore, only some of the disclosed configurations may be included in the computing device (100).
[0044] A processor (110) according to an embodiment of the present disclosure may be understood as a configuration unit including hardware and / or software for performing computing operations. For example, the processor (110) may read a computer program to perform data processing for machine learning. The processor (110) may process computational processes such as processing input data for machine learning, feature extraction for machine learning, and error calculation based on backpropagation. The processor (110) for performing such data processing may include a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The above-described type of processor (110) is only one example, and thus, the type of processor (110) may be configured in various ways within a range understandable to those skilled in the art based on the contents of the present disclosure.
[0045] A computing device (100) can acquire training data necessary for training a neural network model from a plurality of objects (e.g., patients, etc.). Specifically, the computing device (100) can configure training data using time series data acquired from each object. The time series data may be data sequentially recorded according to the time measured and acquired from the object. The time series data may include electrocardiography (ECG) data, photoplethysmography (PPG) data, electroencephalogram (EEG) data, electromyogram (EMG) data, etc. However, for the convenience of explanation of the present disclosure, the time series data will be described below assuming electrocardiogram data.
[0046] The processor (110) is electrically connected to other components of the computing device (100) (e.g., memory (120), etc.) and controls the overall operation of the computing device (100).
[0047] The memory (120) according to one embodiment of the present disclosure may be understood as a configuration unit including hardware and / or software for storing and managing data processed in the computing device (100). That is, the memory (120) may store any type of data generated or determined by the processor (110) and any type of data received by the network unit of the computing device (100). For example, the memory (120) may include at least one type of storage medium among a flash memory (120) type, a hard disk type, a multimedia card micro type, a card type memory (120), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory (120), a magnetic disk, and an optical disk. Additionally, the memory (120) may include a database system that controls and manages data in a predetermined system. The types of memory (120) described above are merely examples, and thus, the types of memory (120) may be configured in various ways within a range understandable to those skilled in the art based on the contents of the present disclosure.
[0048] The memory (120) can structure and organize and manage data, combinations of data, and program codes executable by the processor (110) required for the processor (110) to perform operations. For example, the memory (120) can store a neural network model and a learning data set used to train the neural network model. In addition, the memory (120) can store a program code for performing self-supervised learning, supervised learning, generative learning, or contrastive learning on the neural network model based on the learning data set, or for operating a neural network model for which learning has been completed, and processed data generated as the program code is executed.
[0049] FIG. 2 is a flowchart illustrating a method for obtaining a neural network model for predicting a disease based on electrocardiogram data according to an embodiment of the present disclosure. FIG. 3 is an exemplary diagram illustrating a method for obtaining a neural network model for predicting a disease based on electrocardiogram data according to an embodiment of the present disclosure.
[0050] According to one embodiment of the present disclosure, the processor (110) segments the first electrocardiogram data (610) to obtain a plurality of first segmented data (S310). The processor (110) may acquire the first electrocardiogram data (610) from another external electronic device via a communication interface or may acquire the first electrocardiogram data (610) stored in the memory (120). Here, the first electrocardiogram data (610) includes a plurality of electrocardiogram data acquired from a plurality of different patients. Therefore, the first electrocardiogram data (610) may also be referred to as a first learning data set. In particular, the first electrocardiogram data (610) may be electrocardiogram data that is not assigned a label, such as specific class information.
[0051] When the first electrocardiogram data (610) is acquired, the processor (110) may divide the first electrocardiogram data (610) into units of a preset size to acquire a plurality of first segmented data. Specifically, the processor (110) may divide the plurality of electrocardiogram data included in the first electrocardiogram data (610) into units of a preset size. In this process, the processor (110) may divide the first electrocardiogram data (610) into consecutive patch units along the time axis and process them. For example, the processor (110) may divide the first electrocardiogram data (610) acquired at a sample rate of 250 Hz, which is 10 seconds long, into units of 1 second (every 250 points) to acquire 10 patches. Meanwhile, the processor (110) may set the patch size so that each patch includes the PQRST cycle of the electrocardiogram signal corresponding to the electrocardiogram data, or may divide the data so that a certain portion overlaps between adjacent patches. In this way, the processor (110) can obtain a plurality of first electrocardiogram data (610) corresponding to each electrocardiogram data included in the first electrocardiogram data (610).
[0052] In addition, the processor (110) can obtain first learning data corresponding to electrocardiogram data by masking at least one of the acquired plurality of first segmented data (S320).
[0053] Specifically, the processor (110) can randomly select some of the plurality of first segmented data (i.e., the plurality of patches) or select and mask the first segmented data at a certain ratio. At this time, the processor (110) can perform masking on the first segmented data by covering or invalidating all or a certain portion of the selected first segmented data. The masking can be set to a continuous section or a non-consecutive section, and the processor (110) can mask the first segmented data by applying a zero value or adding noise to the set section. Through this masking process, the processor (110) can provide the first ECG data (610) as input data in an incomplete form so that the neural network model can learn the potential features of the first ECG data. In addition, the processor (110) can additionally assign a masked token indicating a masked position and maintain the order of the original data. The processor (110) can acquire a plurality of first segmented data including the masked first segmented data as first learning data. At this time, the first learning data includes a plurality of masked first segmented data corresponding to a plurality of electrocardiogram data included in the first electrocardiogram data (610).
[0054] When the first learning data is acquired, the processor (110) performs self-supervised learning on the first neural network model (10) based on the first learning data (S330). The processor (110) can input a plurality of masked first segmented data included in the first learning data into the first neural network model (10) to train the first neural network model (10). The processor (110) can perform self-supervised learning on the first neural network model (10) based on the first learning data acquired from the first unlabeled electrocardiogram data (610). The processor (110) can perform self-supervised learning on the first neural network model using the first unlabeled electrocardiogram data (610). For example, the processor (110) may train the first neural network model (10) to extract latent features of the first electrocardiogram data (610) and to predict or restore masked first segmentation data among a plurality of first segmentation data corresponding to the first electrocardiogram data (610).
[0055] To this end, according to one embodiment of the present disclosure, the processor (110) may train the first neural network model (10) to restore the masked first segmentation data based on a generative learning method based on self-supervised learning. Specifically, the processor (110) may input the first learning data into the first neural network model (10) to obtain first output data in which the masked segmentation data of the first learning data is restored. In addition, the processor (110) may perform generative learning for the first neural network model (10) based on the first output data and the first electrocardiogram data (610).
[0056] At this time, referring to FIG. 3, the first neural network model (10) according to an embodiment of the present disclosure may include an encoder (710) and a decoder (720). The first neural network model (10) may extract latent feature information of the first learning data input to the encoder (710) to generate latent representation tokens. For example, the encoder (710) may also generate latent representation tokens for the masked first segmentation data based on features of other first segmentation data adjacent to the masked first segmentation data included in the first learning data.
[0057] The encoder (710) may include a 1D convolution block (711) and a transformer block (712). The input terminal of the encoder (710) may include a 1D convolution block (711). The 1D convolution block (711) processes a plurality of first segmented data included in the first learning data along the time axis and may learn local features of an electrocardiogram signal corresponding to the first electrocardiogram data (610). Through this, each first segmented data may be converted into a compressed feature map through a filtering process. The feature map may be transmitted to the transformer block (712) included in the output terminal of the encoder (710). At this time, a classification token may be transmitted together to the transformer block (712). The classification token input to the transformer block (712) may be set to an initial value that can be learned. The transformer block (712) can learn temporal correlations and global patterns between the masked first segment data and adjacent other segment data based on a self-attention mechanism. Through this, the transformer block (712) can ultimately output latent representation tokens and classification tokens that summarize the global characteristics of the first learning data. At this time, the latent representation tokens can restore the masked data or include feature representations that can be utilized for contrastive learning.
[0058] The potential expression tokens extracted through the encoder (710) are transmitted to the decoder (720), and the decoder (720) can restore the masked first segment data and the remaining first segment data to output data (i.e., first output data) corresponding to the first electrocardiogram data (610). At this time, a mask token that can identify the location of the masked first segment data can be input to the decoder (720). The plurality of potential expression tokens (and mask tokens) received from the encoder (710) are processed through the transformer block (721) of the decoder (720), and the transformer block (721) can learn the temporal relationship between the input potential expression tokens and mask tokens by utilizing a self-attention mechanism.
[0059] Meanwhile, the first output data obtained from the decoder (720) may be in the same form as the first electrocardiogram data (610), that is, in the form of entire electrocardiogram data in which multiple segmented data restored through multiple potential expression tokens are merged. The processor (110) may calculate the difference between the first output data obtained in this way and the first electrocardiogram data (610) corresponding to the original data as a reconstruction loss, and may improve the restoration performance of the masked electrocardiogram data by repeatedly training the first neural network model (10) to minimize the reconstruction loss. At this time, the loss function of the generative learning (hereinafter, the first loss function) may be as follows and represented by Equation 1.
[0060] [Formula 1]
[0061]
[0062] Here, x i is the i-th value of the original first electrocardiogram data (610), is the i-th value of the restored first output data, and N may be the total number of points of the first electrocardiogram data (610).
[0063] According to one embodiment of the present disclosure, the processor (110) may perform self-supervised learning-based contrastive learning for the first neural network model (10) based on the first learning data at step S330. In particular, the processor (110) may selectively perform generative learning and contrastive learning for the first neural network model (10), or may perform them simultaneously and in parallel.
[0064] Referring to FIG. 3, a first neural network model (10) according to an embodiment of the present disclosure may include an expressor (730) that obtains a representation for contrastive learning based on latent features extracted by being connected to an encoder (710). Specifically, the expressor (730) may receive potential representation tokens (and classification tokens) transmitted from the encoder (710) and generate an embedding vector for performing contrastive learning. At this time, the expressor (730) may be configured to generate an embedding vector capable of learning the similarity between a plurality of electrocardiogram data by utilizing both features of the masked first segmentation data and the unmasked first segmentation data. Meanwhile, the expressor (730) may include a non-linear block (Non-Linear Block) (731). The non-linear block (731) may non-linearly transform the potential representation tokens and classification tokens generated by the encoder (710) to suppress unnecessary features and emphasize only key features.
[0065] FIG. 4 is a flowchart illustrating a method for performing generative learning and contrastive learning based on self-supervised learning for a first neural network model (10) according to an embodiment of the present disclosure. Steps S410 and S420 illustrated in FIG. 4 correspond to steps S210 and S220 illustrated in FIG. 2, and step S430 is identically applied to the description of generative learning described above, so a detailed description thereof will be omitted.
[0066] The processor (110) can obtain a plurality of second segmented data by dividing second electrocardiogram data (620) corresponding to positive or negative pairs of first electrocardiogram data (610) (S440).
[0067] Contrastive learning is a self-supervised learning method that learns the relative representation between different data, and optimizes the neural network model to position similar data close together in the embedding space and position different data far apart in the embedding space.
[0068] In the case of contrastive learning, positive pairs and negative pairs are formed between multiple learning data, and for multiple learning data forming a positive pair, similarity is maximized, and for multiple learning data forming a negative pair, difference is maximized. Accordingly, the processor (110) can obtain learning data forming a positive pair and second learning data forming a negative pair for each electrocardiogram data included in the first electrocardiogram data (610). At this time, the processor (110) can match specific electrocardiogram data included in the second learning data forming a positive pair with specific electrocardiogram data included in the first electrocardiogram data (610), and can match specific electrocardiogram data with at least one electrocardiogram data included in the second learning data forming a negative pair. In addition, the processor (110) can train (i.e., contrastive training) the first neural network model (10) to extract embedding vectors corresponding to the plurality of learning data so that they are positioned close to each other in the embedding space for the plurality of learning data that constitute a positive pair, and to extract embedding vectors (or latent representation tokens) corresponding to the plurality of learning data so that they are positioned far from each other in the embedding space for the plurality of learning data that constitute a negative pair.
[0069] According to one embodiment of the present disclosure, if the first electrocardiogram data (610) and the second electrocardiogram data (620) are obtained from the same subject (e.g., a patient), the processor (110) may set the second electrocardiogram data (620) as corresponding to a positive pair, and if the first electrocardiogram data (610) and the second electrocardiogram data (620) are obtained from different subjects, the processor (110) may set the second electrocardiogram data (620) as corresponding to a negative pair.
[0070] Alternatively, the processor (110) may adjust the first electrocardiogram data (610) to obtain augmented data corresponding to the first electrocardiogram data (610) and set the obtained augmented data to correspond to a positive pair for the first electrocardiogram data (610).
[0071] The method of configuring learning data for positive pairs and learning data for negative pairs is described in detail in FIGS. 7 to 9.
[0072] Meanwhile, when the processor (110) obtains second electrocardiogram data (620) corresponding to a positive pair or a negative pair of the first electrocardiogram data (610), the processor (110) may divide the second electrocardiogram data (620) to obtain a plurality of second divided data. The second electrocardiogram data (620) may be data including a plurality of electrocardiogram data constituting a positive pair or a negative pair of the plurality of electrocardiogram data included in the first electrocardiogram data (610). Therefore, the second electrocardiogram data (620) may also be referred to as a second learning data set. The second electrocardiogram data (620) may be divided into the same size and number as the first electrocardiogram data (610). Since the description of the method of dividing the second electrocardiogram data (620) is equally applicable to the method of dividing the first electrocardiogram data (610), a detailed description thereof will be omitted.
[0073] In addition, the processor (110) may obtain second learning data corresponding to the second electrocardiogram data (620) by masking at least one of the acquired plurality of second segmented data (S450). At this time, the masked portion of the plurality of second segmented data may be the same as the masked portion of the plurality of first segmented data. For example, some of the plurality of second segmented data may be masked at the same temporal position as the masked first segmented data with the same masking ratio. Since the description of the method of masking the plurality of second segmented data is equally applicable to the description of the method of masking the plurality of first segmented data, a detailed description thereof will be omitted.
[0074] The processor (110) may obtain a plurality of second segmented data including masked second segmented data as second learning data. At this time, the second learning data includes a plurality of masked second segmented data corresponding to a plurality of electrocardiogram data included in the second electrocardiogram data (620). The processor (110) may perform contrastive learning on the first neural network model (10) by inputting the first learning data and the second learning data as input pairs (S460). At this time, the processor (110) may input the first learning data and the second learning data matched as negative pairs or the first learning data and the second learning data matched as positive pairs (more specifically, the first learning data and the second learning data obtained from the electrocardiogram data included in the first electrocardiogram data (610) and the second electrocardiogram data (620) included in the second electrocardiogram data (620) that are matched as negative pairs or positive pairs, respectively) into the first neural network model (10) to perform contrastive learning.
[0075] Specifically, the processor (110) may input the first learning data and the second learning data as input pairs into the encoder (710) of the first neural network model (10) to obtain potential expression tokens corresponding to the first learning data and the second learning data, respectively. Then, the processor (110) may input the potential expression tokens corresponding to the first learning data and the second learning data, respectively, output from the encoder (710), into the expressor (730), to generate an expression vector (embedding vector). The expressor (730) may refine the information of the potential expression tokens and output an expression vector optimized for contrastive learning based on the learned features. For example, the expressor (730) may remove unnecessary features and extract only key features in the process of generating an expression vector optimized for contrastive learning based on the potential expression tokens. The expression vectors obtained from the expressor (730) may serve as a standard for comparing similarities or differences between data. The processor (110) can train the first neural network model (10) to increase the similarity between expression vectors for positive pairs and to decrease the similarity between expression vectors for negative pairs during the contrastive learning process. At this time, the similarity can be calculated mainly based on cosine similarity. Meanwhile, the processor (110) can perform optimization through contrastive loss. At this time, the loss function of contrastive learning (hereinafter, referred to as the second loss function) can be as shown in the second equation below.
[0076] [Formula 2]
[0077]
[0078] Here, sim(z i ,z j ) is the cosine similarity between two representation vectors of a positive pair, and sim(z i ,z k) is the cosine similarity between two representation vectors of a negative pair, τ is the temperature scaling parameter, and N is the total number of input data.
[0079] At this time, the processor (110) can apply a first loss function corresponding to generative learning and a second loss function corresponding to contrastive learning in combination to the first neural network model (10). Specifically, the processor (110) can set an integrated loss function that combines the first loss function and the second loss function as shown in the third equation below.
[0080] [Formula 3]
[0081]
[0082] λ is the weight of the contrastive loss function and can be a parameter that adjusts the balance between generative learning and contrastive learning.
[0083] In conclusion, the processor (110) applies the first loss function and the second loss function in combination to the first neural network model (10) so that generative learning and contrastive learning can be performed simultaneously, thereby improving the restoration performance of the first neural network model (10) for the masked first electrocardiogram data (610) and training the first neural network model (10) so that it can precisely distinguish between the potential feature expressions of positive and negative pairs.
[0084] FIG. 5 is a flowchart illustrating a method for performing supervised learning for a second neural network model (20) based on labeled electrocardiogram data according to an embodiment of the present disclosure. FIG. 6 is an exemplary diagram illustrating a method for performing supervised learning for a second neural network model (20) based on labeled electrocardiogram data according to an embodiment of the present disclosure.
[0085] Steps S510 to S560 illustrated in FIG. 5 may correspond to steps S410 to S460 illustrated in FIG. 4, respectively. Therefore, a detailed description thereof will be omitted.
[0086] Referring to FIG. 5, when generative learning or contrastive learning for the first neural network model (10) is completed, the processor (110) can extract an encoder (710) included in the first neural network model (10) (S570). The encoder (710) can be trained to perform the role of converting input electrocardiogram data into potential expression tokens and classification tokens, having learned both local features and global patterns of the electrocardiogram signal based on the first electrocardiogram data (610) (and the second electrocardiogram data (620)).
[0087] And, the processor (110) can obtain a second neural network model (20) by combining the extracted encoder (710) and classifier (820) (S580). Referring to FIG. 6, the processor (110) can extract an encoder (710) from a first neural network model (10) for which generative learning and contrastive learning have been completed, and obtain a second neural network model (20) by combining the extracted encoder (710) and classifier (820). Here, the classifier (820) can be formed with a multi-layer perceptron (MLP) structure, and can be designed to output a class probability value of electrocardiogram data input to the encoder (710) based on a classification token output from the encoder (710).
[0088] The processor (110) may perform supervised learning-based fine-tuning on the second neural network model (20) to predict a patient's disease based on third electrocardiogram data (630) assigned with a preset class (S590). Here, the third electrocardiogram data (630) includes a plurality of electrocardiogram data acquired from a plurality of different patients. Therefore, the third electrocardiogram data (630) may also be referred to as a third learning data set. In particular, the third electrocardiogram data (630) may be electrocardiogram data assigned with a label such as specific class information. The specific class may include information indicating the patient's health status or disease, for example, disease types such as Normal, Myocardial Infarction, Arrhythmia, and Hypertrophy.
[0089] The processor (110) can divide the third electrocardiogram data (630) into patch units and input them into the encoder (710) of the second neural network model (20) to generate classification tokens. The third electrocardiogram data (630) can be divided into patch units of the same size and number as the first electrocardiogram data (610) used to train the encoder (710). The classification tokens extracted through the encoder (710) are transmitted to the classifier (820) (MLP), and the classifier (820) outputs a probability value for each class based on the input classification tokens. The processor (110) can calculate the classification loss between the predicted class probability value and the actual label (640) assigned to the third electrocardiogram data (630), and optimize it by iteratively adjusting the parameters of the second neural network model (20) to minimize the classification loss. Through this, the processor (110) can perform fine-tuning on the second neural network model (20) to predict a disease of a patient corresponding to a specific class assigned to the third electrocardiogram data (630).
[0090] According to one embodiment of the present disclosure, when the processor (110) acquires a plurality of electrocardiogram data corresponding to a specific object, the processor (110) may selectively combine two of the plurality of electrocardiogram data acquired from the specific object to acquire one or more positive pairs. In addition, the processor (110) may selectively combine one of the plurality of electrocardiogram data acquired from the specific object and one of the plurality of electrocardiogram data acquired from another object to acquire one or more negative pairs.
[0091] That is, the processor (110) can select two pieces of electrocardiogram data from a plurality of electrocardiogram data acquired from a subject, and acquire the selected two pieces of electrocardiogram data as a positive pair of first learning data and second learning data. In particular, two pieces of electrocardiogram data acquired as positive phases from the same subject can be selected regardless of the time or order in which the electrocardiogram data were acquired. In addition, the processor (110) can select one piece of electrocardiogram data from a plurality of electrocardiogram data acquired from a specific subject, and select one piece of electrocardiogram data from a plurality of electrocardiogram data acquired from another subject, and acquire the first piece of electrocardiogram data as a negative pair of the second learning data. At this time, the acquired negative pair can also be acquired as a negative pair of learning data for another subject. That is, the negative pair acquired for a specific subject can be shared with another subject corresponding to the electrocardiogram data included in the negative pair.
[0092] FIG. 7 is an exemplary diagram of a method for generating a learning data set for contrastive learning using electrocardiogram data obtained from different subjects according to one embodiment of the present disclosure.
[0093] Referring to FIG. 7, the processor (110) can obtain a plurality of electrocardiogram data (11, 12-B, 12-C, and 12-D) from patient A (200-A), patient B (200-B), patient C (200-C), and patient D (200-D), respectively. At this time, the processor (110) can select two of the plurality of electrocardiogram data (11) obtained from patient A (200-A) to obtain a positive pair of the learning data set in order to obtain a learning data set for contrastive learning for patient A (200-A). In addition, the processor (110) can obtain a negative pair included in the learning data set for patient A (200-A) by combining one of the plurality of electrocardiogram data (11) acquired from patient A (200-A) with one of the plurality of electrocardiogram data (12-B and 12-C) acquired from patient B (200-B) or patient C (200-C), respectively. This also applies equally to patient B (200-B) and patient C (200-C).
[0094] At this time, according to one embodiment of the present disclosure, another object corresponding to the electrocardiogram data of a specific subject and the other electrocardiogram data forming the negative pair may have similar biological information. Here, the biological information may include at least one of the age, gender, height, and weight of the subject. To this end, the processor (110) may identify the biological information of the subject and another object, and obtain electrocardiogram data for obtaining a negative pair from another object having biological information similar to the biological information of the subject. The processor (110) may determine that a plurality of objects having at least one of the age, gender, height, and weight that matches have similar biological information.
[0095] Referring back to FIG. 7, the processor (110) can identify that patient A (200-A), patient B (200-B), and patient C (200-C) have matching ages and genders, while patient D (200-D) has different ages and genders. Therefore, the processor (110) may not utilize the electrocardiogram data (12-D) of patient D (200-D) when acquiring the negative pairs of the learning data set of patient A (200-A).
[0096] To this end, the processor (110) can calculate the similarity between the biological information of the subject and the biological information of other subjects. For example, the processor (110) can extract a plurality of vectors corresponding to the biological information of each subject and calculate the similarity of the biological information based on the distance between the extracted vectors. In addition, the processor (110) can obtain a negative pair of the learning data set by combining the electrocardiogram data of a plurality of subjects whose similarity is greater than a preset value.
[0097] Meanwhile, according to an embodiment of the present disclosure, the processor (110) may acquire positive and negative pairs of a learning data set from a plurality of electrocardiogram data acquired from the same subject. In particular, when the number of positive pairs acquired based on a plurality of electrocardiogram data acquired from a specific subject is less than a preset number, or the number of negative pairs acquired based on electrocardiogram data acquired from another subject is less than a preset number, the processor (110) may acquire positive and negative pairs of a learning data set from a plurality of electrocardiogram data acquired from the same subject. Hereinafter, embodiments of the present disclosure related thereto will be described.
[0098] According to one embodiment of the present disclosure, the processor (110) may acquire a plurality of electrocardiogram data corresponding to a subject multiple times and then set reference data among the plurality of electrocardiogram data. Here, the reference data may be electrocardiogram data that serves as a basis for selecting electrocardiogram data included in negative pairs (and positive pairs) among the plurality of electrocardiogram data acquired from the same subject.
[0099] According to one embodiment of the present disclosure, among a plurality of electrocardiogram data acquired for a subject, electrocardiogram data first acquired from the subject can be set as reference data.
[0100] At this time, the processor (110) can selectively combine two of the plurality of electrocardiogram data acquired within a preset time from the time point at which the reference data was acquired with the reference data to obtain one or more positive pairs, and can selectively combine the reference data with one of the plurality of electrocardiogram data acquired after a preset time from the time point at which the reference data was acquired to obtain one or more negative pairs.
[0101] FIG. 8 is an exemplary diagram of a method for generating a learning data set for contrastive learning using multiple electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0102] Specifically, the processor (110) can identify the time point at which electrocardiogram data set as reference data for the subject is acquired. At this time, the processor (110) can select a plurality of electrocardiogram data acquired within a preset time from the time point at which the reference data is acquired. In addition, the processor (110) can selectively combine two of the reference data and the selected plurality of electrocardiogram data to acquire a positive pair. For example, referring to FIG. 8, the processor (110) can set the first acquired electrocardiogram data (i.e., the electrocardiogram data acquired on October 13, 2024) (11-1) among the plurality of electrocardiogram data (11-1 to 11-8, hereinafter 11) as the reference data, and select a plurality of electrocardiogram data (11-1 to 11-3) acquired on the same date as the date at which the reference data is acquired (October 13, 2024) to acquire a positive pair. Multiple electrocardiogram data (11-1 to 11-3) of the same date can be acquired by patient A (200-A) visiting the hospital and measuring the electrocardiogram signal multiple times. Furthermore, the processor (110) can select two of the electrocardiogram data (11-1 to 11-3) acquired on the same date as the date on which the reference data was acquired by the patient visiting the hospital, thereby acquiring a positive pair.
[0103] In addition, the processor (110) can select a plurality of electrocardiogram data acquired after a preset time from the time point at which the reference data was acquired. In addition, the processor (110) can obtain a negative pair by combining the reference data and one electrocardiogram data selected from the plurality of selected electrocardiogram data. For example, referring again to FIG. 8, the processor (110) can select a plurality of electrocardiogram data (11-4 to 11-8) acquired on a different date (October 27, 2024 and November 27, 2024) from the date at which the reference data was acquired (October 13, 2024) in order to obtain a negative pair. That is, the processor (110) can select the electrocardiogram data (11-4 to 11-8) acquired when the patient visits the hospital again after the reference data is acquired when the patient visits the hospital, thereby obtaining a negative pair.
[0104] Meanwhile, the processor (110) may acquire multiple sets of electrocardiogram data by segmenting electrocardiogram data acquired from the subject. Here, the electrocardiogram data to be segmented may correspond to an electrocardiogram signal that is longer than the electrocardiogram signal corresponding to the electrocardiogram data. Hereinafter, for the convenience of describing the present disclosure, the electrocardiogram data corresponding to the electrocardiogram signal measured over a long period of time to be segmented will be referred to as target electrocardiogram data.
[0105] The processor (110) can acquire multiple pieces of electrocardiogram data by dividing the target electrocardiogram data at preset time intervals. For example, if the electrocardiogram data is acquired by sampling an electrocardiogram signal having a length of 10 seconds, the target electrocardiogram data can be acquired by sampling an electrocardiogram signal having a length exceeding 10 seconds. That is, by dividing the electrocardiogram signal having a length exceeding 10 seconds into 10-second intervals, electrocardiogram data corresponding to the 10-second electrocardiogram signal can be acquired. Alternatively, the processor (110) can acquire multiple pieces of electrocardiogram data corresponding to the 10-second electrocardiogram signal by dividing the target electrocardiogram data corresponding to the electrocardiogram signal having a length exceeding 10 seconds.
[0106] Meanwhile, the processor (110) can divide a plurality of target electrocardiogram data into each target electrocardiogram data and obtain a plurality of electrocardiogram data corresponding to each target electrocardiogram data.
[0107] At this time, according to one embodiment of the present disclosure, the processor (110) may set reference data among a plurality of electrocardiogram data. The processor (110) may set the first acquired electrocardiogram data among the plurality of electrocardiogram data as the reference data as described above. Alternatively, the processor (110) may identify the target electrocardiogram data having the largest number of segmented electrocardiogram data among the plurality of target electrocardiogram data, and may set reference data from the plurality of electrocardiogram data corresponding to the identified target electrocardiogram data. At this time, among the plurality of electrocardiogram data corresponding to the identified target electrocardiogram data, the electrocardiogram data located in the center of the target electrocardiogram data or the first acquired electrocardiogram data may be set as the reference data.
[0108] In addition, the processor (110) can selectively combine a plurality of electrocardiogram data obtained by segmenting from target electrocardiogram data identical to the reference data and two from the reference data to obtain one or more positive pairs, and can selectively combine one of the plurality of electrocardiogram data obtained by segmenting from target electrocardiogram data different from the reference data and the reference data to obtain one or more negative pairs.
[0109] FIG. 9 is an exemplary diagram of a method for generating learning data for contrast learning using multiple target electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0110] Referring to FIG. 9, the processor (110) may divide the target electrocardiogram data (13-1) acquired on October 13, 2024 and the target electrocardiogram data (13-2) acquired on October 27, 2024, respectively, to obtain a plurality of electrocardiogram data (11-1 to 11-5). At this time, the processor (110) may set the first acquired electrocardiogram data (11-1) among the plurality of electrocardiogram data (11-1 to 11-3) acquired by dividing the target electrocardiogram data (13-1) acquired on October 13, 2024, as reference data, and may obtain a positive pair of a learning data set by combining the reference data with the remaining electrocardiogram data (11-2 and 11-3) acquired from the target electrocardiogram data (13-1) that is the same as the reference data. And, the processor (110) can obtain a negative pair of the learning data set by combining the reference data (11-1) with one selected from the remaining electrocardiogram data (11-4 and 11-5) obtained from the target electrocardiogram data (target electrocardiogram data obtained on October 27, 2024) (13-2) that is different from the reference data.
[0111] FIG. 10 is an exemplary diagram of a method for generating a learning data set for contrastive learning by using augmented data obtained by adjusting multiple electrocardiogram data obtained from the same subject according to one embodiment of the present disclosure.
[0112] Alternatively, according to an embodiment of the present disclosure, the processor (110) may adjust the set reference data to obtain a plurality of augmented data (hereinafter, first augmented data) corresponding to the reference data, and may adjust at least one remaining electrocardiogram data other than the reference data to obtain a plurality of augmented data (hereinafter, second augmented data) corresponding to the remaining electrocardiogram data. The processor (110) may adjust each electrocardiogram data to obtain augmented data corresponding to each electrocardiogram data. For example, the processor (110) may adjust the electrocardiogram data by adding baseline noise or adding muscle artifact to the electrocardiogram signal corresponding to the electrocardiogram data to change the baseline of the waveform of the electrocardiogram signal. This may be performed by applying a baseline noise addition pattern or a muscle artifact addition pattern to the electrocardiogram data. Alternatively, the processor (110) may adjust the electrocardiogram data by adding white noise to the electrocardiogram data or applying a partial zero padding method. In this way, the processor (110) can adjust each of the plurality of electrocardiogram data to obtain augmented data corresponding to each electrocardiogram data.
[0113] At this time, the processor (110) can selectively combine two of the reference data and the plurality of first augmented data corresponding to the reference data to obtain one or more positive pairs, and can selectively combine the reference data with one of the plurality of second augmented data corresponding to the plurality of electrocardiogram data other than the reference data to obtain one or more negative pairs.
[0114] Referring to FIG. 10, the processor (110) can acquire a plurality of pieces of electrocardiogram data (11-1 to 11-3) by dividing the target electrocardiogram data (13-1) acquired on October 13, 2024. At this time, the processor (110) can set the electrocardiogram data (11-1) acquired first among the plurality of pieces of electrocardiogram data (11-1 to 11-3) acquired by dividing the target electrocardiogram data (13-1) acquired on October 13, 2024 as reference data, and can acquire a plurality of first augmented data (31-1 and 31-2) by adjusting the reference data. In addition, the processor (110) can acquire a plurality of second augmented data (32-1 to 32-4) by adjusting the electrocardiogram data (11-2 and 11-3) other than the reference data. And, the processor (110) can obtain a positive pair of the learning data set by combining the reference data with one selected from among the plurality of first augmented data (31-1 and 31-2), and can obtain a negative pair of the learning data set by combining the reference data with one selected from among the plurality of second augmented data (32-1 to 32-4).
[0115] At this time, the first augmentation method for acquiring multiple first augmentation data and the second augmentation method for acquiring multiple second augmentation data may be different. For example, if the first augmentation method is a white noise addition method, the second augmentation method may be a partial zero padding method.
[0116] Meanwhile, the processor (110) may perform contrastive learning on the first neural network model (10) based on a learning data set used for contrastive learning, which is composed of the generated first learning data and the second learning data set. That is, the processor (110) may perform contrastive learning on the first neural network model (10) based on a learning data set including the acquired positive pairs and negative pairs. In particular, the processor (110) may perform contrastive learning on the first neural network model (10) so that a plurality of embedding vectors corresponding to a plurality of electrocardiogram data included in a positive pair become closer in the embedding space, and a plurality of embedding vectors corresponding to a plurality of electrocardiogram data included in a negative pair become farther apart in the embedding space. To this end, the processor (110) may perform learning in a manner of minimizing the distance between the embedding vectors of the positive pairs and maximizing the distance between the embedding vectors of the negative pairs by using a contrastive loss function or an InfoNCE loss function.
[0117] The second neural network model (20), including an encoder (710) that has completed contrastive learning, can be effectively utilized to analyze electrocardiogram (ECG) data and predict a patient's health status. The second neural network model (20), including an encoder (710) that has learned the similarities and differences between various ECG patterns through contrastive learning, can more precisely interpret a patient's ECG data by forming an embedding space that distinguishes between normal and abnormal patterns. This enables personalized disease prediction and monitoring, and enables more rapid prediction of heart disease risk.
[0118] The second neural network model (20) including the encoder (710) for which contrastive learning has been completed can perform analysis of basic electrocardiogram patterns without new labeling through zero-shot learning, and can increase learning and prediction accuracy for specific diseases even with a small amount of data through few-shot learning. The second neural network model (20) including the encoder (710) for which contrastive learning has been completed can be expanded into a model specialized in diagnosing specific diseases by performing fine-tuning with data related to specific diseases.
[0119] FIG. 11 is a block diagram of a computing device (1100) according to another embodiment of the present disclosure.
[0120] Referring to FIG. 11, a computing device (1100) according to an embodiment of the present disclosure includes a processor (1110), a memory (1120), a communication interface (1130), a sensing unit (1140), a display (1150), a user interface (1160), a camera (1170), and a speaker (1180). Among the configurations illustrated in FIG. 10, the processor (1110) and the memory (1120) correspond to the configurations of the processor (110) and the memory (120) of the computing device (100) illustrated in FIG. 1, and thus a detailed description thereof will be omitted.
[0121] A communication interface (1130) according to an embodiment of the present disclosure may be understood as a component that transmits and receives data through any known wired or wireless communication system. For example, the communication interface (1130) may perform data transmission and reception using a wired or wireless communication system such as a local area network (LAN), wideband code division multiple access (WCDMA), long term evolution (LTE), wireless broadband internet (WiBro), fifth generation mobile communication (5G), ultrawide-band, ZigBee, radio frequency (RF) communication, wireless LAN, wireless fidelity, near field communication (NFC), or Bluetooth. Since the above-described communication systems are only examples, the wired and wireless communication system for data transmission and reception of the communication interface (1130) may be applied in various ways other than the above-described examples.
[0122] The communication interface (1130) can receive data necessary for the processor (1110) to perform calculations through wired or wireless communication with any system or any client, etc. In addition, the communication interface (1130) can transmit data generated through calculations of the processor (1110) through wired or wireless communication with any system or any client, etc. For example, the communication interface (1130) can receive medical data through communication with a cloud server that performs tasks such as standardization of databases and medical data in a hospital environment, or a computing device (1100), etc. The communication interface (1130) can transmit output data of the second neural network model, intermediate data, processed data, etc. derived from the calculation process of the processor (1110), etc. through communication with the aforementioned database, server, or computing device (1100). For example, the processor (1110) can obtain a plurality of electrocardiogram data for each subject's biometric data of the subject (1) from an external computing device (e.g., an external server device or an external biometric signal measuring device) through a communication interface (1130).
[0123] The sensing unit (1140) can obtain biometric data of a subject. For example, the sensing unit (1140) can include a plurality of electrodes (e.g., 12 leads). At this time, the processor (1110) can obtain the user's electrocardiogram signal as biometric data through at least one electrode. In addition, the sensing unit (1140) can include an image sensor or an optical sensor. At this time, the processor (1110) can obtain the user's optical blood flow signal as biometric data through the image sensor (or optical sensor).
[0124] The display (1150) can display various images. Here, the images include both still images and moving images. The display (1150) can output guide information regarding activities generated based on the user status. The display (1150) can be implemented as various types of displays, such as an LCD (Liquid Crystal Display Panel), an OLED (Organic Light Emitting Diodes), an LCoS (Liquid Crystal on Silicon), a DLP (Digital Light Processing), etc. In addition, the display (1150) can also include a driving circuit, a backlight unit, etc., which can be implemented in a form such as an a-si TFT, an LTPS (low temperature poly silicon) TFT, an OTFT (organic TFT), etc.
[0125] Meanwhile, the display (1150) may be implemented as a touch screen by being combined with a touch panel. In this case, the display (1150) may not only function as an output interface that outputs images through the touch screen, but also as an input interface that receives a user's touch input. The display (1150) may display the results of a judgment on the possibility of heart disease predicted through a pre-trained second neural network model, a treatment plan, and management information for the subject (1).
[0126] The user interface (1160) is a component used by the computing device (1100) to perform interaction with the user, and may include at least one of a touch sensor, a motion sensor, a button, a jog dial, and a switch, but is not limited thereto. The processor (1110) may receive the user's biological information (occupation, age, gender, etc.) through the user interface (1160).
[0127] The camera (1170) captures images of objects around the user. Specifically, the camera (1170) can capture images of food consumed by the user. At this time, the processor (1110) can determine the nutritional status of the user based on the user's condition and the image of the food consumed by the user, and provide recommended dietary information related to heart disease as guide information. To this end, the camera (1170) can be implemented with an imaging device such as an imaging device having a CMOS structure (CIS, CMOS Image Sensor) or an imaging device having a CCD structure (Charge Coupled Device). However, the present invention is not limited thereto, and the camera (1170) can be implemented with a camera module having various resolutions capable of capturing an object. Meanwhile, the camera (1170) can be implemented with a depth camera (e.g., an IR depth camera), a stereo camera, an RGB camera, etc.
[0128] The speaker (1180) is a component that outputs various audio data on which various processing operations, such as decoding, amplification, and noise filtering, have been performed by an audio processing unit (not shown). The speaker (1180) can output various notification sounds or voice messages. According to one embodiment of the present disclosure, the processor (1110) can convert an electrical signal received from an external device into a user voice and output it through the speaker (1180). For example, the speaker (1180) can output a voice message warning of or suggesting diagnosis of a heart disease based on the judgment result on the possibility of a heart disease identified through a second neural network model.
[0129] The various embodiments of the present disclosure described above can be combined with additional embodiments and modified within the scope understood by those skilled in the art in light of the detailed description above. It should be understood that the embodiments of the present disclosure are illustrative in all respects and not restrictive. For example, each component described as a single component may be implemented in a distributed manner, and likewise, components described as distributed may be implemented in a combined manner. Accordingly, all changes or modifications derived from the meaning, scope, and equivalent concepts of the claims of the present disclosure should be construed as being included within the scope of the present disclosure.
Claims
1. A method for obtaining a neural network model for predicting a disease based on time series data, performed by a computing device including at least one processor, A step of dividing first time series data to obtain multiple first division data; A step of obtaining first learning data corresponding to the first time series data by masking at least one of the acquired plurality of first segmented data; and A step of performing self-supervised learning for a first neural network model based on the first learning data; comprising; method 2. In paragraph 1, The step of performing learning for the above first neural network model is: A step of performing generative learning or contrastive learning for the first neural network model based on the first learning data; comprising; method.
3. In paragraph 2, The steps of performing generative learning and contrastive learning for the above first neural network model are as follows: A step of inputting the first learning data into the first neural network model to obtain output data in which the masked segmentation data of the first learning data is restored, and performing generative learning for the first neural network model based on the output data and the first time series data; comprising; method.
4. In paragraph 2, A step of obtaining a plurality of second segmented data by dividing second time series data corresponding to positive or negative pairs for the first time series data; and A step of obtaining second learning data corresponding to the second time series data by masking at least one of the acquired plurality of second segmented data; comprising; The steps of performing generative learning and contrastive learning for the above first neural network model are as follows: A step of performing contrastive learning on the first neural network model by using the first learning data and the second learning data as input pairs; comprising; method.
5. In paragraph 2, When generative learning and contrastive learning for the first neural network model are completed, a step of extracting an encoder included in the first neural network model; A step of obtaining a second neural network model by combining the extracted encoder and classifier; and A step of performing supervised learning-based fine-tuning on the second neural network model to predict the patient's disease based on the third time series data to which a preset class is assigned; comprising; method.
6. In paragraph 4, The step of obtaining the above plurality of second segmented data is: If the first time series data and the second time series data are obtained from the same object, the second time series data is set to correspond to a positive pair, If the first time series data and the second time series data are obtained from different objects, a step of setting the second time series data as corresponding to a negative pair is included; method.
7. In paragraph 4, The step of obtaining the above plurality of second segmented data is: A step of adjusting the first time series data to obtain augmented data corresponding to the first time series data, and setting the obtained augmented data to correspond to a positive pair for the first time series data; comprising; method.
8. In paragraph 2, The steps of performing generative learning and contrastive learning for the above first neural network model are as follows: Applying a first loss function corresponding to the generative learning and a second loss function corresponding to the contrastive learning in combination to the first neural network model. method.
9. In paragraph 2, The above first neural network model is, An encoder for extracting latent features of the first learning data; A decoder connected to the encoder and restoring the masked segmentation data of the first learning data based on the extracted latent features; and A representation device, which is connected to the encoder and obtains a representation for the contrastive learning based on the extracted latent features; method.
10. In a computing device for obtaining a neural network model for predicting a disease based on time series data, a processor comprising at least one core; and A memory including program codes executable by the processor; The above processor, By dividing the first time series data, a plurality of first division data are obtained, and at least one of the obtained plurality of first division data is masked to obtain first learning data corresponding to the time series data, and based on the first learning data, self-supervised learning for the first neural network model is performed. Computing device.
11. A computer program stored in a computer-readable storage medium, wherein the computer program, when executed on one or more processors, performs operations for obtaining a neural network model for predicting a disease based on time series data. The above actions are, An operation of dividing first time series data to obtain multiple first division data; An operation of obtaining first learning data corresponding to the time series data by masking at least one of the acquired plurality of first segmented data; and An operation of performing self-supervised learning for a first neural network model based on the first learning data; Computer program.
Citation Information
Patent Citations
Al-air secondary, and method of fabricating of the same
KR1020220132476A
Computer program and method for artificial neural network model learning based on time series bio-signals
KR102197112B1
Electrocardiogram created apparatus base on generative adversarial network and method thereof
KR102412974B1
ECG search and interpretation based on a dual ECG and text embedding model
US20230238133A1