Method, program and device for quantifying the quality of a biological signal

The method and apparatus address the limitation of existing signal quality indices by using labeled electrocardiogram datasets to generate a machine learning model that evaluates clinical readability, ensuring biological signals are usable for medical diagnostics.

JP2025540871AActive Publication Date: 2025-12-16MEDICAL AI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025534926
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-11-14
Publication Date
2025-12-16
Estimated Expiration
2043-11-14

AI Technical Summary

Technical Problem

Existing signal quality indices in the medical field do not adequately reflect the clinical readability of biological signals, as they focus solely on noise and artifacts, disregarding the signal's usefulness for clinical interpretation.

Method used

A method and apparatus that quantify signal quality by generating a machine learning model using electrocardiogram datasets labeled based on morphological features and domain expert analysis, enabling the evaluation of clinical readability.

Benefits of technology

Provides a reliable method to assess the clinical readability of biological signals, ensuring they contain sufficient information for clinical use, rather than just noise levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025540871000001_ABST
    Figure 2025540871000001_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a method, a program, and an apparatus for quantifying quality of a biological signal are disclosed, which may include obtaining at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled based on a domain expert's analysis of an interpretation of the electrocardiogram signal, and generating a machine learning model for quantifying quality of the electrocardiogram signal based on the at least one of the first electrocardiogram data set or the second electrocardiogram data set.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to deep learning technology in the medical field, and more particularly to a method and apparatus that can quantify the readability of biological signals by reflecting it in quality analysis. [Background technology]

[0002] In order to utilize signal-based data in applications, it is important to evaluate the quality of the collected signals. Therefore, metrics for evaluating signal quality have been studied in various application fields such as acoustics, communications, optics, and medicine. These metrics for evaluating signal quality that have been studied in various application fields are commonly called signal quality index (SQI).

[0003] Most of the signal quality indices that have been studied to date reflect the signal quality based on the amount of noise or artifacts that occur in the signal itself. However, in the medical field, there are many cases where a signal containing a lot of noise or artifacts is still usable for clinical judgment, such as diagnosing or predicting a disease. In other words, even if a signal is evaluated as having poor quality due to noise or artifacts by a signal quality index that has been studied to date, it should still be evaluated as usable data in the medical field as long as it contains sufficient information necessary for clinical interpretation. Therefore, in the medical field, when evaluating signal quality, it is necessary to reflect whether the signal is clinically readable.

[0004] For example, when the quality of an electrocardiogram signal is evaluated in order to use the electrocardiogram to diagnose cardiac disease A, even if an electrocardiogram signal is judged to be of poor quality based on an existing signal quality index, it may still contain all the information necessary to diagnose cardiac disease A. Therefore, even if the signal quality is poor based on a reference signal quality index, it is not conforming to evaluate the signal as difficult to use because it lacks the information necessary to unconditionally read cardiac disease A. In other words, since signal quality must be evaluated according to the intended use of the signal, it is necessary to have a signal quality index that is optimized for the intended use in the medical field. Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure aims to provide a method and apparatus for quantifying signal quality that can reflect the clinical readability of a biological signal, and a method and apparatus that can generate a reliable machine learning model for the above-mentioned quality quantification.

[0006] However, the problems to be solved by the present disclosure are not limited to those mentioned above, and other problems not mentioned will be clearly understood from the following description. [Means for solving the problem]

[0007] To achieve the above object, an embodiment of the present disclosure discloses a method for quantifying quality of a biological signal, performed by a computing device, including the steps of obtaining at least one of a first electrocardiogram dataset labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram dataset labeled based on a domain expert's analysis of an interpretation of the electrocardiogram signal, and generating a machine learning model for quantifying quality of the electrocardiogram signal based on at least one of the first electrocardiogram dataset or the second electrocardiogram dataset.

[0008] Alternatively, the first electrocardiogram data set may be labeled into a first class indicating that the electrocardiogram signal is a readable signal or into a second class indicating that the electrocardiogram signal is an unreadable signal, based on noise identified based on morphological characteristics of the electrocardiogram signal.

[0009] Alternatively, the second class may correspond to at least one of the cases where noise that makes it impossible to identify at least one of the start and end points of the waveform of the electrocardiogram signal is present at a predetermined ratio or more, or where noise that makes it impossible to identify the R peak of the electrocardiogram signal is present at a predetermined ratio or more.

[0010] Alternatively, the second electrocardiogram data set may be labeled based on the mode of the domain expert's analysis of whether noise present in the electrocardiogram signal affects disease readings.

[0011] Alternatively, one of the first electrocardiogram data set and the second electrocardiogram data set may be divided into a training data set, a validation data set, and a test data set and used to generate the machine learning model, and the other of the first electrocardiogram data set and the second electrocardiogram data set may be used as a test data set to generate the machine learning model.

[0012] Alternatively, the machine learning model may include a first model based on a neural network that estimates the readability of the electrocardiogram signal based on the electrocardiogram dataset, and a second model based on regression analysis that estimates the readability of the electrocardiogram signal based on the electrocardiogram dataset.

[0013] Alternatively, generating a machine learning model for quantifying the quality of an electrocardiogram signal based on at least one of the first electrocardiogram dataset or the second electrocardiogram dataset may include training a candidate model of the first model based on a first training dataset included in the first electrocardiogram dataset; verifying performance of the trained candidate model based on a first validation dataset included in the first electrocardiogram dataset; evaluating performance of at least one candidate model selected by the validation based on a first test dataset included in the first electrocardiogram dataset and a second test dataset that is the second electrocardiogram dataset; and generating the first model based on a specification of the candidate model identified by the evaluation.

[0014] Alternatively, generating a machine learning model for quantifying the quality of an electrocardiogram signal based on at least one of the first electrocardiogram dataset or the second electrocardiogram dataset may include extracting information about a signal quality index (SQI) used to evaluate noise in an electrocardiogram signal from the first electrocardiogram dataset to generate an index dataset; training candidate models for the second model based on a third training dataset included in the generated index dataset; verifying performance of the trained candidate models based on a third validation dataset included in the generated index dataset; evaluating performance of at least one candidate model selected by the validation based on a third test dataset included in the generated index dataset; and generating the second model based on specifications of the candidate models identified by the evaluation.

[0015] To achieve the above-described object, an embodiment of the present disclosure discloses a method for quantifying the quality of a biological signal, the method being performed by a computer device. The method includes the steps of acquiring read target data and inputting the acquired read target data into a machine learning model to calculate a score indicating the readability of the read target data. The machine learning model may be generated based on at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled based on a domain expert's analysis of the reading of the electrocardiogram signal.

[0016] Alternatively, the method may further include a step of estimating whether or not a signal contained in the acquired read target data is missing, and a step of determining whether or not the acquired read target data is readable data by combining the calculated score and the estimated whether or not a signal is missing.

[0017] Alternatively, the presence or absence of signal leakage can be estimated based on whether the signal values ​​included in the acquired read target data are blank for a predetermined ratio or more, or whether the waveform of the signal included in the acquired read target data is flat.

[0018] Alternatively, the step of determining whether the acquired data to be read is readable data by combining the calculated score and the estimated signal leakage may include the steps of: comparing the calculated score with a threshold value to determine whether noise is present in the signal included in the acquired data to be read; and determining whether the acquired data to be read is readable data based on at least one of the presence or absence of signal noise determined based on the calculated score or the estimated signal leakage.

[0019] Alternatively, the step of comparing the calculated score with a threshold value to determine whether noise exists in the signal contained in the acquired read target data may include a step of determining that noise exists in the signal of the specific lead if the score calculated based on the specific lead is equal to or greater than the threshold value.

[0020] Alternatively, the step of determining whether the acquired read target data is readable data based on at least one of the presence or absence of noise in the signal determined based on the calculated score or the presence or absence of leakage of the estimated signal may include a step of determining that noise is present based on the calculated score or that the signal of a specific lead estimated to have leakage of a signal is unreadable data.

[0021] To achieve the above object, one embodiment of the present disclosure provides a computer program stored on a computer-readable storage medium. The computer program, when executed by one or more processors, performs operations for quantifying quality of a biological signal, including acquiring at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled according to the empirical judgment of a domain expert related to reading the electrocardiogram signal, and generating a machine learning model for quantifying quality of the electrocardiogram signal based on at least one of the acquired first electrocardiogram data set or the acquired second electrocardiogram data set.

[0022] To achieve the above object, one embodiment of the present disclosure provides a computing device for quantifying the quality of a biological signal. The device may include a processor including at least one core, a memory including program code executable by the processor, and a network unit for acquiring at least one of a first electrocardiogram data set labeled based on morphological features of an electrocardiogram signal or a second electrocardiogram data set labeled by empirical judgment of a domain expert related to reading electrocardiogram signals. Here, the processor may generate a machine learning model for quantifying the quality of an electrocardiogram signal based on at least one of the acquired first electrocardiogram data set or the acquired second electrocardiogram data set. [Effects of the Invention]

[0023] The present disclosure aims to provide a quality quantification method and apparatus that can reflect the readability of a biological signal, and also aims to provide a method and apparatus that can generate a reliable machine learning model for the above-mentioned quality quantification. [Brief explanation of the drawings]

[0024] [Figure 1] FIG. 1 is a block diagram of a computing device according to one embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating a process for generating a machine learning model according to an embodiment of the present disclosure. [Figure 3] FIG. 10 is a block diagram illustrating a process for generating a machine learning model according to an alternative embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a process for quantifying the quality of a biosignal in a computing device according to one embodiment of the present disclosure. [Figure 5] 1 is a flowchart illustrating a method for generating a machine learning model for quantifying the quality of a biological signal according to one embodiment of the present disclosure. [Figure 6] 1 is a flowchart illustrating a method for quantifying the quality of a biological signal according to one embodiment of the present disclosure. [Figure 7] 1 is a flowchart summarizing a process for quantifying the quality of a biosignal according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement the present disclosure. The embodiments presented in this disclosure are provided to enable those skilled in the art to use or practice the contents of the present disclosure. Therefore, various modifications to the embodiments of the present disclosure will be apparent to those skilled in the art. That is, the present disclosure may be embodied in various different forms and is not limited to the following embodiments.

[0026] Throughout the specification of the present disclosure, the same or similar reference numerals refer to the same or similar components. In addition, in order to clearly explain the present disclosure, reference numerals of parts that are not relevant to the explanation of the present disclosure may be omitted from the drawings.

[0027] The term "or" as used in this disclosure is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or otherwise clear from the context in this disclosure, "X uses A or B" should be understood to mean one of the natural inclusive permutations. For example, unless otherwise specified or otherwise clear from the context in this disclosure, "X uses A or B" can be interpreted as either X uses A, X uses B, or X uses both A and B.

[0028] The term "and / or" as used in this disclosure must be understood to indicate and include all possible combinations of one or more of the associated listed concepts.

[0029] The terms "comprises" and / or "comprising" as used in this disclosure should be understood to mean that the specified features and / or components are present. However, the terms "comprises" and / or "comprising" should not be understood to exclude the presence or addition of one or more other features, other components and / or combinations thereof.

[0030] In this disclosure, unless otherwise specified or clear from the context as referring to the singular form, the singular should generally be construed as including "one or more."

[0031] The term "nth (n is a natural number)" used in this disclosure can be understood as an expression used to distinguish components of the present disclosure from one another based on a predetermined criterion, such as functional, structural, or convenience of description. For example, in this disclosure, components that perform different functional roles can be classified as a first component or a second component. However, components that are substantially identical within the technical concept of the present disclosure but must be distinguished for convenience of description can also be classified as a first component or a second component.

[0032] The term "acquire" as used in this disclosure may be understood to refer to generating or receiving data in an on-device form, as well as receiving data from an external device or system via a wireless communication network.

[0033] Meanwhile, the terms "module" or "unit" used in this disclosure may be understood to refer to an independent functional unit that processes computing resources, such as a computer-related entity, firmware, software or a portion thereof, hardware or a portion thereof, or a combination of software and hardware. Here, a "module" or "unit" may refer to a unit composed of a single element or a unit expressed as a combination or collection of multiple elements. For example, as a concept of connotation, a "module" or "unit" may refer to a hardware element or a collection of hardware elements of a computing device, an application program that achieves a specific software function, a processing procedure implemented by the execution of software, or a collection of instructions for executing a program. Furthermore, as a broad concept, a "module" or "unit" may refer to a computing device itself that constitutes a system, or an application executed on a computing device. However, the above concepts are merely examples, and the concepts of "module" and "unit" may be defined in various ways within the scope of understanding of those skilled in the art based on the contents of this disclosure.

[0034] The term "model" used in this disclosure may be understood as a system implemented using mathematical concepts and language to solve a specific problem, a set of software units for solving a specific problem, or an abstract model of a processing process for solving a specific problem. For example, a machine learning "model" may refer to a system that performs calculations based on a machine learning algorithm. Here, machine learning algorithms may include classification algorithms such as naive Bayes and decision trees, regression analysis algorithms such as linear regression and logistic regression, and deep learning algorithms such as convolutional neural networks. The types of machine learning algorithms disclosed herein are not limited to the above examples and may be configured in various ways within the scope that one skilled in the art can understand based on the above examples.

[0035] As used in this disclosure, the term "image" refers to multidimensional data composed of discrete image elements. In other words, "image" can be understood as a term referring to a digital representation of an object that can be viewed by the human eye. For example, "image" can refer to multidimensional data composed of elements that correspond to pixels in a two-dimensional image. "Image" can refer to multidimensional data composed of elements that correspond to voxels in a three-dimensional image.

[0036] The explanations of the above terms are intended to facilitate understanding of the present disclosure. Therefore, unless the above terms are explicitly stated as matters limiting the contents of the present disclosure, care should be taken not to use them in a way that limits the technical ideas of the contents of the present disclosure.

[0037] FIG. 1 is a block diagram of a computing device according to one embodiment of the present disclosure.

[0038] The computing device 100 according to an embodiment of the present disclosure may be a hardware device or part of a hardware device that performs comprehensive data processing and calculations, or may be a software-based computing environment connected via a communication network. For example, the computing device 100 may be a server that performs intensive data processing functions and shares resources, or a client that shares resources by interacting with the server. The computing device 100 may also be a cloud system that enables multiple servers and clients to interact with each other to comprehensively process data. The above description is merely an example of a type of computing device 100, and various types of computing device 100 may be configured within the scope that can be understood by those skilled in the art based on the contents of the present disclosure.

[0039] 1, a computing device 100 according to an embodiment of the present disclosure may include a processor 110, a memory 120, and a network unit 130. However, since FIG. 1 is merely an example, the computing device 100 may include other components for implementing a computer environment. Also, the computing device 100 may include only some of the disclosed components.

[0040] The processor 110 according to an embodiment of the present disclosure may be understood as a component including hardware and / or software for performing computing operations. For example, the processor 110 may read a computer program and perform data processing for machine learning. The processor 110 may perform computational processes such as preprocessing of input data for machine learning, error calculation based on backpropagation, etc. The processor 110 for performing such data processing may include a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc. The types of processor 110 described above are merely examples, and various configurations of the processor 110 are possible within the scope of what one skilled in the art would understand based on the present disclosure.

[0041] The processor 110 may generate a machine learning model for quantifying the quality of an electrocardiogram signal based on electrocardiogram data. The processor 110 may generate a machine learning model for quantifying the quality of an electrocardiogram signal using an electrocardiogram dataset labeled with factors that affect the results of an electrocardiogram signal reading that serves as the basis for clinical judgment. Here, the electrocardiogram dataset used to generate the machine learning model may include at least one of an electrocardiogram set labeled based on morphological features of the waveform that affect the reading of the electrocardiogram signal or an electrocardiogram dataset labeled through analysis by a domain expert of the reading of the electrocardiogram signal. A domain expert may be understood as a group or a member of the group that can interpret electrocardiogram signals and make clinical judgments, such as diagnosing a specific disease. That is, the processor 110 may generate a machine learning model that can provide a quantitative indicator of whether an electrocardiogram signal is of a quality that can be used for clinical reading by utilizing features identifiable in the waveform of the electrocardiogram signal and the empirical basis and judgment used to read the electrocardiogram signal.

[0042] The processor 110 may estimate the quality of the ECG data to be read using the machine learning model generated as described above. Here, the quality of the ECG data may indicate whether the ECG data is clinically readable. The processor 110 may then determine whether the ECG data to be read is clinically readable based on the quality of the ECG data to be read. Specifically, the processor 110 may input the ECG data to be read into the machine learning model and generate a quantitative indicator for the lead-specific signal quality of the ECG data to be read. The processor 110 may also analyze the waveform of the ECG data to determine whether signal leakage exists for each lead of the ECG data to be read. The processor 110 may then combine the quantitative indicator generated by the machine learning model and the determination result on the lead-specific signal leakage generated by the waveform analysis to determine whether the ECG data to be read is usable for clinical reading for each lead.

[0043] The memory 120 according to an embodiment of the present disclosure may be understood as a component including hardware and / or software for storing and managing data processed by the computing device 100. That is, the memory 120 may store any type of data generated or determined by the processor 110 and any type of data received by the network unit 130. For example, the memory 120 may include at least one type of storage medium selected from the group consisting of flash memory, hard disk, multimedia card micro, card-type memory, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. The memory 120 may also include a database system that manages data in a predetermined manner. The types of memory 130 described above are merely examples, and various configurations of the memory 120 are possible within the scope of what would be understood by one skilled in the art based on the present disclosure.

[0044] The memory 120 may structure and organize and manage data, a combination of data, and program code executable by the processor 110 required for the processor 110 to perform calculations. For example, the memory 120 may store electrocardiogram data acquired via the network unit 130 (described below). The memory 120 may store program code that causes the processor 110 to generate a machine learning model, program code that causes the processor 110 to estimate the quality of the electrocardiogram data using the generated machine learning model, and various data calculated by executing the program code.

[0045] The network unit 130 according to an embodiment of the present disclosure may be understood as a component that transmits and receives data via any type of known wired or wireless communication system. For example, the network unit 130 may transmit and receive data using a wired or wireless communication system such as a local area network (LAN), wideband code division multiple access (WCDMA), long term evolution (LTE), wireless broadband internet (WiBro), 5th generation mobile communication (5G), ultra wideband, ZigBee, radio frequency (RF) communication, wireless LAN, wireless fidelity (WiFi), near field communication (NFC), or Bluetooth. The above-described communication systems are merely examples, and various other wired or wireless communication systems for transmitting and receiving data by the network unit 130 may be used.

[0046] The network unit 130 may receive data necessary for the processor 110 to perform calculations via wired or wireless communication with any system, server, client, etc. The network unit 130 may also transmit data generated by calculations by the processor 110 via wired or wireless communication with any system, server, client, etc. For example, the network unit 130 may receive an electrocardiogram data set via wired or wireless communication with an electrocardiogram sensing device, a database in a medical environment, etc. The network unit 130 may transmit various data generated by calculations by the processor 110 based on the electrocardiogram data via wired or wireless communication with an electrocardiogram sensing device, a database in a medical environment, etc.

[0047] FIG. 2 is a block diagram illustrating a process for generating a machine learning model according to an embodiment of the present disclosure.

[0048] 2 , a computer device 100 according to an embodiment of the present disclosure can generate a machine learning model 200 for quantifying the quality of an electrocardiogram signal based on a first electrocardiogram data set 10 labeled based on morphological features of the electrocardiogram signal and a second electrocardiogram data set 20 labeled through domain expert analysis of electrocardiogram signal interpretation. The computer device 100 can generate the machine learning model 200 that estimates the clinical readability of an electrocardiogram signal and provides a quantitative indicator by training a candidate model based on machine learning based on the first electrocardiogram data set 10 and the second electrocardiogram data set 20 and tuning and evaluating hyperparameters of the candidate model. Here, the first electrocardiogram data set 10 can be a data set labeled according to whether the electrocardiogram signal is readable, analyzed by noise recognized based on the morphological features of the electrocardiogram signal. The second electrocardiogram data set 20 may then be a data set labeled according to the results of an analysis by a clinical domain expert as to whether an electrocardiogram signal can be used to read a particular disease.

[0049] For example, the first electrocardiogram data set 10 may include labels based on noise distinguishable by morphological features of the waveform of the electrocardiogram signal. The labels included in the first electrocardiogram data set 10 may include a first class indicating that the electrocardiogram signal is a readable signal due to noise determined based on the morphological features of the electrocardiogram signal, and a second class indicating that the electrocardiogram signal is an unreadable signal due to noise determined based on the morphological features of the electrocardiogram signal. Here, the first class or the second class may be distinguished based on whether noise is present at a level at which feature points used for clinical reading of the electrocardiogram signal can be identified. Specifically, if noise that makes it impossible to identify at least one of the start point or end point of the waveform of the electrocardiogram signal is present at a predetermined ratio or above, or if noise that makes it impossible to identify the R peak of the electrocardiogram signal is present at a predetermined ratio or above, data including the electrocardiogram signal may be labeled as the second class. Here, the predetermined ratio may be a ratio set by the manufacturer or user of the computer device 100 to achieve the purpose of quantifying quality. Furthermore, if the waveform of the electrocardiogram signal is flat, the data including the electrocardiogram signal may be labeled as Class 2. If none of the above three conditions is met, the data including the electrocardiogram signal may be labeled as Class 1. In this way, the electrocardiogram data included in the first electrocardiogram data set 10 may be labeled according to the effect that noise detected based on the morphological characteristics of the electrocardiogram signal has on the clinical readability of the electrocardiogram signal.

[0050] The second electrocardiogram dataset 20 may include labels based on expert analysis of the impact of noise contained in the electrocardiogram signal on clinical readability. The labels included in the second electrocardiogram dataset 20 may include a third class indicating that the signal is usable for reading according to the clinical domain expert analysis, and a fourth class indicating that the signal is unusable for reading according to the clinical domain expert analysis. Since the third class or the fourth class is determined based on the empirical and intuitive analysis of the domain expert, the second electrocardiogram dataset 20 may be labeled based on the most frequent value of the domain expert's analysis of whether noise present in the electrocardiogram signal affects disease reading. In other words, to increase the reliability of the labels, the electrocardiogram data included in the second electrocardiogram dataset 20 may be labeled into the third or fourth class based on the most frequent result in big data corresponding to a collection of analysis results of the domain experts.

[0051] The computer device 100 can generate a high-quality model by combining datasets labeled in different manners for training, validation, and testing to generate the machine learning model 200. Referring to Fig. 2, the computer device 100 can use either the first electrocardiogram dataset 10 or the second electrocardiogram dataset 20 for training, validation, and testing to generate the machine learning model 200. The computer device 100 can then use the other of the first electrocardiogram dataset 10 or the second electrocardiogram dataset 20 for testing the machine learning model 200. In other words, by combining the first electrocardiogram dataset 10 and the second electrocardiogram dataset 20 to generate the machine learning model 200, the computer device 100 can generate a model whose performance is verified according to various criteria and which is optimized for quantifying the quality of an electrocardiogram signal.

[0052] Specifically, the computer device 100 can divide the first electrocardiogram dataset 10 into a first training dataset 11, a first validation dataset 15, and a first test dataset 19 for use. The computer device 100 can then use the second electrocardiogram dataset 20 as a second test dataset 25 to generate the machine learning model 200. That is, the computer device 100 can use a portion of the first electrocardiogram dataset 10 for training and validating candidate models for generating the machine learning model 200, and can use the second electrocardiogram dataset 20 together with the remaining portion of the first electrocardiogram dataset 10 for evaluating at least one candidate model selected through the training and validation. The computer device 100 can then generate the machine learning model 200 based on the evaluation results for at least one candidate model performed using the remaining portion of the first electrocardiogram dataset 10 and the second electrocardiogram dataset 20.

[0053] For example, the computer device 100 can divide the first electrocardiogram dataset 10 so that the ratio of the first learning dataset 11, the first validation dataset 15, and the first test dataset 19 is 8:1:1 and use the divided data to generate the machine learning model 200. The computer device 100 can train candidate models designed with various parameters for generating the machine learning model based on the first learning dataset 11 included in the first electrocardiogram dataset 10. Here, the candidate models can be models that receive electrocardiogram data for each lead and estimate the readability of lead-specific electrocardiogram signals included in the electrocardiogram data. The computer device 100 can perform performance validation on the trained candidate models based on the first validation dataset 15 included in the first electrocardiogram dataset 10 and select at least one candidate model through the performance validation. The computer device 100 can evaluate the performance of at least one candidate model selected through the performance validation using the second electrocardiogram dataset 20 as a second test dataset 25 together with the first test dataset 19 included in the first electrocardiogram dataset 10. The computer device 100 may determine candidate models whose evaluation values ​​are equal to or greater than a predetermined standard as final candidate models for generating the machine learning model 200. The computer device 100 may generate the machine learning model 200 based on specifications such as the structure and parameters of the final candidate model. Here, the machine learning model 200 may be a model that receives electrocardiogram data for each lead and outputs an electrocardiogram score 30 indicating the readability of lead-specific signals included in the electrocardiogram data. The electrocardiogram score 30 is an index that quantitatively indicates a noise evaluation result that reflects a probability value for the readability of an electrocardiogram signal, and may be expressed as a number, symbol, or the like. However, the above-described ratio of data sets is merely an example, and the present disclosure is not limited thereto.

[0054] In this way, the computer device 100 can use the entire second electrocardiogram dataset 20 together with the portion of the first electrocardiogram dataset 10 to evaluate the model generated based on the portion of the first electrocardiogram dataset 10, so that the effect of signal noise on clinical readings, such as disease diagnosis, is reflected in quantifying the signal quality. In other words, by using the entire second electrocardiogram dataset 20 to evaluate the model generated based on the first electrocardiogram dataset 10, the computer device 100 can ensure that the signal quality estimated by the machine learning model 200 reflects whether the electrocardiogram signal is clinically readable. By generating such a machine learning model 200, the computer device 100 can quantify the signal quality so that the signal quality indicates the signal's clinical readability, rather than simply evaluating the signal quality based on signal noise or artifacts.

[0055] FIG. 3 is a block diagram illustrating a process for generating a machine learning model according to an alternative embodiment of the present disclosure.

[0056] 3, a machine learning model according to an alternative embodiment of the present disclosure may include a first neural network-based model 210 for estimating readability of an electrocardiogram signal based on an electrocardiogram dataset, and a second regression analysis-based model 220 for estimating readability of an electrocardiogram signal based on an electrocardiogram dataset. Specifically, according to the alternative embodiment of the present disclosure, the computer device 100 may ensemble the first neural network-based model 210 and the second regression analysis-based model 220 to generate the machine learning model.

[0057] For example, the first model 210 may include a convolutional neural network based on ResNet. The first model 210 may be generated by the following process. First, a candidate model of the first model 210 may be trained using a first training dataset 42. Here, the training may be performed by supervised learning using labels based on morphological features of ECG signals included in the first training dataset 42. Once the training of the candidate model of the first model 210 is complete, the candidate model of the first model 210 may be subjected to performance verification using a first validation dataset 43. The candidate model of the first model 210 selected through the performance verification may be subjected to performance evaluation using a first test dataset 44 and a second test dataset 55. Then, the first model 210 may be generated based on the specifications of a candidate model that has good performance in both the first test dataset 44 and the second test dataset 55 through the performance evaluation. Here, the first model 210 may be a model that receives electrocardiogram data and outputs a first electrocardiogram score 61 that indicates the readability of lead-specific signals included in the electrocardiogram data. The first electrocardiogram score 61 may be a concept corresponding to the electrocardiogram score 30 in FIG. 2 described above.

[0058] The second model 220 may include a model based on logistic regression analysis. Such a second model 220 may be generated by the following process: First, a candidate model of the second model 220 may use, as input, an index dataset 41 extracted from the first electrocardiogram dataset 40. The index dataset 41 may be a dataset generated by extracting information about signal quality indices (such as zerocrossSQI, minSQI, maxSQI, powerSQI, q1SQI, q3SQI, sSQI, kSQI, highfreqSQI, baseSQI, and pSQI) that can be used to evaluate noise in an electrocardiogram signal. That is, the candidate model of the second model 220 may be trained by receiving a third training dataset 45 of the index dataset 41 that includes signal features that can be mathematically clearly expressed by signal quality indices for noise evaluation. Once training of the candidate model of the second model 220 is complete, the candidate model of the second model 220 may be subjected to performance verification by receiving a third validation dataset 46. Candidate models for the second model 220 selected through performance verification may undergo performance evaluation using the third test data set 47. Then, the second model 220 may be generated based on specifications of candidate models that have good performance on the third test data set 47 through performance evaluation. Here, the second model 220 may be a model that receives index data extracted from electrocardiogram data and outputs a second electrocardiogram score 65 indicating the readability of lead-specific signals included in the electrocardiogram signal. The second electrocardiogram score 65 may be a concept corresponding to the electrocardiogram score 30 of FIG. 2 described above.

[0059] In this way, the computer device 100 can expect generalized performance even for external data other than the data set on which learning and evaluation were performed by using the second model 220 based on regression analysis together to minimize the overfitting problem, which is a limitation of the neural network-based first model 210. Furthermore, since the second model 220 receives features extracted by a clear mathematical formula and can confirm the decision process and basis for what criteria a decision was made based on, the computer device 100 can overcome the limitation of the first model 210, which is that it is difficult to confirm the basis for a decision, by utilizing the second model 220 together.

[0060] Furthermore, the first model 210 based on a neural network is considered unreliable because the basis for its judgment is unknown. The computer device 100 can provide minimal safety measures for reliability issues by generating a machine learning model by ensembling the first model 210 and the second model 220. Generally, the first model 210 based on a neural network is considered to have high accuracy but low evidence, while the second model 220 based on regression analysis is considered to have low accuracy but clear evidence. Therefore, the computer device 100 can customize and generate a final result value according to the user's desired purpose by ensembling the first model 210 and the second model 220.

[0061] FIG. 4 is a block diagram illustrating a process of quantifying the quality of a biosignal in a computer device according to one embodiment of the present disclosure.

[0062] 4, a computer device 100 according to an embodiment of the present disclosure may input read target data 70 to a first model 210 included in a machine learning model to calculate a first electrocardiogram score 81 indicating readability for each lead of the read target data 70. Here, the first model 210 may be generated using at least one of an electrocardiogram dataset labeled based on morphological features of an electrocardiogram signal or an electrocardiogram dataset labeled through analysis by a domain expert on the reading of an electrocardiogram signal. Although not shown in FIG. 4, the computer device 100 may extract index data including information on a signal quality index relative to noise in the electrocardiogram signal from the read target data 70. The computer device 100 may input the index data extracted from the read target data 70 to a second model 220 included in the machine learning model to calculate a second electrocardiogram score 85 indicating readability for each lead of the read target data 70. The second model 220 can be generated based on an index data set extracted from an electrocardiogram data set that has been labeled based on morphological features of the electrocardiogram signal.

[0063] The computer device 100 can estimate whether or not there is a signal missing in the data to be read 70. The computer device 100 can estimate whether or not there is a signal missing for all leads included in the data to be read 70 and generate a missingness estimation result 89 for each lead. Specifically, whether or not there is a signal missing can be estimated based on whether or not a signal value for each lead of the data to be read 70 is blank for a predetermined ratio or more, or whether or not the waveform of the signal for each lead of the data to be read 70 is flat. Here, the predetermined ratio may be a ratio set by a manufacturer or user of the computer device 100 to achieve a purpose of quantifying quality.

[0064] The computer device 100 may determine whether the target data 70 is readable by combining the first electrocardiogram score 81, the second electrocardiogram score 85, and the leak presence / absence estimation result 89. The computer device 100 may generate a determination result 90 regarding whether the target data 70 is readable for each lead by comprehensively analyzing the output value of the machine learning model and the estimation result regarding the presence / absence of leaks in the lead-specific signals of the target data 70. For example, the computer device 100 may compare the first electrocardiogram score 81 with a first threshold value and determine that noise exists in the signal of a lead that is equal to or greater than the first threshold value. The computer device 100 may also compare the second electrocardiogram score 85 with a second threshold value and determine that noise exists in the signal of a lead that is equal to or greater than the second threshold value. Here, the first threshold value and the second threshold value are values ​​predetermined by a manufacturer or user of the computer device 100 according to their purpose, and may be the same value or different values. The computer device 100 can determine that noise is present in the signal of a lead that is determined to have leaked based on the lead-specific leak presence / absence estimation result 89, and can determine that noise is not present in the signal of a lead that is determined not to have leaked based on the lead-specific leak presence / absence estimation result 89. If noise is determined to be present in the signal of a lead based on at least one of the determination results based on the first electrocardiogram score 81, the determination result based on the second electrocardiogram score 85, or the determination result based on the lead-specific leak presence / absence estimation result 89, the computer device 100 can determine that the signal of the lead is unreadable data. In this way, the computer device 100 can provide quantitative information about the quality of the signal with high reliability by mutually complementarily using analysis results that can be utilized to determine readability.

[0065] FIG. 5 is a flowchart illustrating a method for generating a machine learning model for quantifying the quality of a biological signal according to one embodiment of the present disclosure.

[0066] 5, a computer device 100 according to an embodiment of the present disclosure may acquire at least one of a first electrocardiogram data set labeled based on morphological features of an electrocardiogram signal or a second electrocardiogram data set labeled based on a domain expert's analysis of an electrocardiogram signal reading (S110). Here, the first electrocardiogram data set may be labeled into a first class indicating that the electrocardiogram signal is a readable signal or a second class indicating that the electrocardiogram signal is an unreadable signal, depending on noise detected by the morphological features of the electrocardiogram signal. The second class may correspond to at least one of a case where noise that makes it impossible to identify at least one of the start and end points of the electrocardiogram signal waveform is present at a predetermined ratio or more, or a case where noise that makes it impossible to identify the R peak of the electrocardiogram signal is present at a predetermined ratio or more. The second electrocardiogram data set may be labeled based on the mode of the domain expert's analysis of whether noise present in the electrocardiogram signal affects disease reading. For example, the computer device 100 may acquire at least one of the first electrocardiogram data set and the second electrocardiogram data set via wired or wireless communication with a client for labeling electrocardiogram data. The computer device 100 may also include an input / output unit to directly label the electrocardiogram data and generate at least one of the first electrocardiogram data set and the second electrocardiogram data set.

[0067] The computer device 100 may generate a machine learning model for quantifying the quality of an electrocardiogram signal based on at least one of the first electrocardiogram data set or the second electrocardiogram data set acquired in step S110 (S120). To generate the machine learning model, the computer device 100 may divide the first electrocardiogram data set into a training data set, a validation data set, and a test data set. To generate the machine learning model, the computer device 100 may use the second electrocardiogram data set as a test data set. Here, the machine learning model is a model that outputs a quantified index of the readability of an electrocardiogram signal based on the lead-specific electrocardiogram data set, and may include a first model based on a neural network and a second model based on regression analysis.

[0068] For example, the computer device 100 may train a first candidate model based on a first training dataset included in a first electrocardiogram dataset. The computer device 100 may verify the performance of the trained candidate model based on a first validation dataset included in the first electrocardiogram dataset. The computer device 100 may evaluate the performance of at least one candidate model selected through validation based on a first test dataset included in the first electrocardiogram dataset and a second test dataset, which is a second electrocardiogram dataset. The computer device 100 may identify a model that performs well on data labeled with different features and criteria by using both the first test dataset and the second test dataset for evaluation. Here, the model with good performance identified by the computer device 100 may be a model with the highest performance evaluation index or a model that meets or exceeds a specific reference value. The computer device 100 may generate a first model based on specifications of the candidate model identified through evaluation. Here, the model specifications are information about parameters for configuring a neural network, and may include kernel size, depth, width, learning rate, etc.

[0069] The computer device 100 may generate an index dataset by extracting information about a signal quality index used to evaluate noise in an electrocardiogram signal from the first electrocardiogram dataset. That is, the computer device 100 may generate the index dataset based on features indicative of a signal quality index due to noise in an electrocardiogram signal in the first electrocardiogram dataset. The computer device 100 may train a second candidate model based on a third training dataset included in the generated index dataset. The computer device 100 may verify the performance of the trained candidate model based on a third validation dataset included in the generated index dataset. The computer device 100 may evaluate the performance of at least one candidate model selected through validation based on a third test dataset included in the generated index dataset. The computer device 100 may identify a model with good performance based on the third test dataset. Here, the model with good performance identified by the computer device 100 may be a model with the highest performance evaluation index or a model that meets or exceeds a specific reference value. The computer device 100 may generate a second model based on specifications of the candidate model identified through evaluation. Here, the model specifications may be information about model parameters for performing logistic regression analysis.

[0070] FIG. 6 is a flowchart illustrating a method for quantifying the quality of a biosignal according to one embodiment of the present disclosure.

[0071] 6, a computer device 100 according to an embodiment of the present disclosure may acquire data to be read (S210). The data to be read may be understood as electrocardiogram data generated for use in clinical readings such as diagnosing or predicting a specific disease. For example, the computer device 100 may receive the data to be read generated by the electrocardiogram sensing device via wired or wireless communication with the electrocardiogram sensing device.

[0072] The computer device 100 may input the read target data acquired in step S210 into a machine learning model and calculate a score indicating the readability of the read target data (S220). Here, the machine learning model may be a model generated based on the electrocardiogram data set labeled based on the morphological features of the electrocardiogram signal in FIG. 5 and the electrocardiogram data set labeled based on a domain expert's analysis of the reading of the electrocardiogram signal. For example, the computer device 100 may input the read target data into a first model based on a neural network included in the machine learning model and calculate a first score indicating noise reflecting readability. The computer device 100 may also input the read target data into a second model based on regression analysis included in the machine learning model and calculate a second score indicating noise reflecting readability.

[0073] Meanwhile, the computer device 100 may estimate whether or not a signal included in the read target data acquired in step S210 is leaking. The estimation of whether or not a signal is leaking may be performed in parallel with step S220, in which a score is calculated through machine learning. For example, if a signal value is blank for 50% or more based on a specific lead of the read target data, the computer device 100 may estimate that a signal of the corresponding lead is leaking. Also, if a signal waveform is flat based on a specific lead of the read target data, the computer device 100 may estimate that a signal of the corresponding lead is leaking. The value of 50 mentioned above is merely an example, and the ratio value for determining whether or not a blank exists may be changed depending on the intended use of the computer device 100.

[0074] The computer device 100 may determine whether the data to be read is readable data by combining the score calculated in step S220 and the presence or absence of signal leakage estimated in the above process. The computer device 100 may determine whether noise exists in the signal included in the data to be read by comparing the score calculated in step S220 with a threshold value. The computer device 100 may then determine whether the data to be read is readable data based on at least one of the presence or absence of signal noise or signal leakage determined based on the score. For example, the computer device 100 may determine whether noise exists in the read-specific signals of the data to be read by comparing a first score generated through a first model based on a neural network with a threshold value. The computer device 100 may determine whether noise exists in the read-specific signals of the data to be read by comparing a second score generated through a second model based on regression analysis with a threshold value. The computer device 100 may then determine whether noise exists in the read-specific signals based on the presence or absence of lead-specific signal leakage. The computer device 100 may determine whether a clinical read is possible for each read of the data to be read by combining the results of the determinations on the presence or absence of noise.

[0075] 7 is a flowchart summarizing a process of quantifying the quality of a biological signal according to an embodiment of the present disclosure. Steps S310 and S320 in FIG. 7 correspond to the steps in FIG. 6 described above, and therefore will not be described below.

[0076] Referring to FIG. 7, a computer device 100 according to an embodiment of the present disclosure may determine whether a score calculated using a machine learning model for data to be read is equal to or greater than a threshold value (S330). If the score for a particular lead is less than the threshold value, the computer device 100 may determine that noise is not present in the signal for that lead (S341). Conversely, if the score for a particular lead is equal to or greater than the threshold value, the computer device 100 may determine that noise is present in the signal for that lead (S345). If the machine learning model includes a first model based on a neural network and a second model based on regression analysis, the computer device 100 may individually compare the score for each model with a threshold value to determine whether noise is present. Here, the threshold value compared with the score calculated by the first model and the threshold value compared with the score calculated by the second model may be the same or different.

[0077] The computer device 100 can determine whether signal leakage exists based on whether the read-specific signals of the data to be read are blank for a predetermined ratio or more or whether the waveform of the read-specific signals is flat (S350). If the signal of a specific read of the data to be read is blank for a predetermined ratio or more or the waveform of the signal of the specific read is flat, the computer device 100 can determine that the read has signal leakage (S361). The computer device 100 can then determine that noise exists in the signal of the read that has signal leakage. Conversely, if the signal of a specific read of the data to be read is blank for less than a predetermined ratio or the waveform of the signal of the specific read is not flat, the computer device 100 can determine that the read has no signal leakage (S365). The computer device 100 can then determine that noise does not exist in the signal of the read that has no signal leakage.

[0078] If it is determined that noise is present in a particular lead in the above process, the computer device 100 may determine that the signal of the particular lead is clinically unreadable (S370). That is, if it is determined that noise is present because the score is estimated to be equal to or greater than the critical value for the particular lead (S345) or if it is determined that a signal is missing (S361), the computer device 100 may determine that the signal of the particular lead is unreadable (S370). Conversely, if it is determined that noise is not present in a particular lead in the above process, the computer device 100 may determine that the signal of the particular lead is clinically readable (S380). That is, if it is determined that noise is not present because the score is estimated to be less than the critical value for the particular lead (S341) or if it is determined that a signal is not missing (S365), the computer device 100 may determine that the signal of the particular lead is readable (S380).

[0079] The various embodiments of the present disclosure described above can be combined with additional embodiments and can be modified within the scope that can be understood by those skilled in the art based on the above detailed description. It should be understood that the embodiments of the present disclosure are illustrative in all respects and are not limiting. For example, each component described as a single type can also be implemented in a distributed form, and similarly, components described as distributed can be implemented in a combined form. Therefore, all modifications and variations derived from the meaning, scope, and equivalent concepts of the claims of the present disclosure should be construed as being within the scope of the present disclosure.

Claims

1. 1. A method for quantifying a quality of a biological signal, the method being performed by a computing device including at least one processor, the method comprising: obtaining at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled based on a domain expert's analysis of the reading of the electrocardiogram signal; generating a machine learning model for quantifying electrocardiogram signal quality based on at least one of the first electrocardiogram data set or the second electrocardiogram data set; A method comprising:

2. 2. The method of claim 1, wherein the first electrocardiogram data set is labeled into a first class indicating that the electrocardiogram signal is a readable signal or a second class indicating that the electrocardiogram signal is an unreadable signal, based on noise recognized based on morphological characteristics of the electrocardiogram signal.

3. The second class is When noise that makes it impossible to identify at least one of the start point or end point of the waveform of the electrocardiogram signal exists at a predetermined ratio or more, or When noise that prevents the R peak of the electrocardiogram signal from being identified is present at a predetermined ratio or more, The method of claim 2 , corresponding to at least one of:

4. The method of claim 1 , wherein the second electrocardiogram data set is labeled based on a mode of the domain expert's analysis of whether noise present in an electrocardiogram signal affects disease readings.

5. one of the first electrocardiogram data set and the second electrocardiogram data set is divided into a training data set, a validation data set, and a test data set and used to generate the machine learning model; The method of claim 1 , wherein the other of the first electrocardiogram data set and the second electrocardiogram data set is used as a test data set to generate the machine learning model.

6. The machine learning model a first neural network-based model that estimates the readability of an electrocardiogram signal based on the electrocardiogram data set; a second regression-based model that estimates the readability of the electrocardiogram signal based on the electrocardiogram data set; The method of claim 1 , comprising:

7. generating a machine learning model for quantifying electrocardiogram signal quality based on at least one of the first electrocardiogram data set or the second electrocardiogram data set, training a candidate model of the first model based on a first training data set included in the first electrocardiogram data set; validating performance of the trained candidate model based on a first validation dataset included in the first electrocardiogram dataset; evaluating performance of at least one candidate model selected by the validation based on a first test dataset included in the first electrocardiogram dataset and a second test dataset that is the second electrocardiogram dataset; generating the first model based on a specification of candidate models identified by the evaluation; The method of claim 6, comprising:

8. generating a machine learning model for quantifying electrocardiogram signal quality based on at least one of the first electrocardiogram data set or the second electrocardiogram data set, extracting information about a signal quality index (SQI) used to evaluate noise in an electrocardiogram signal from the first electrocardiogram data set to generate an index data set; training a candidate model for the second model based on a third training dataset included in the generated index dataset; verifying performance of the trained candidate model based on a third validation dataset included in the generated index dataset; evaluating the performance of at least one candidate model selected by the validation based on a third test dataset included in the generated index dataset; generating the second model based on specifications of candidate models identified by the evaluation; The method of claim 6, comprising:

9. 1. A method for quantifying a quality of a biological signal, performed by a computing device including at least one processor, comprising: acquiring data to be read; inputting the acquired read target data into a machine learning model to calculate a score indicating the readability of the read target data; Including, the machine learning model is generated based on at least one of a first electrocardiogram data set labeled based on morphological features of electrocardiogram signals or a second electrocardiogram data set labeled based on a domain expert's analysis of readings of electrocardiogram signals; method.

10. estimating whether or not there is a missing signal included in the acquired read target data; determining whether the acquired read target data is readable data by combining the calculated score and the estimated signal leakage; 10. The method of claim 9, further comprising:

11. The method of claim 10, wherein the presence or absence of the signal leakage is estimated based on whether the signal values ​​included in the acquired read target data are blank for a predetermined ratio or more, or whether the waveform of the signal included in the acquired read target data is flat.

12. determining whether the acquired read target data is readable data by combining the calculated score and the estimated signal leakage, comparing the calculated score with a threshold value to determine whether noise exists in the signal included in the acquired read target data; determining whether the acquired read target data is readable data based on at least one of whether noise exists in the signal determined based on the calculated score or whether leakage of the estimated signal exists; The method of claim 10, comprising:

13. The step of comparing the calculated score with a threshold value to determine whether noise exists in a signal included in the acquired read target data includes: determining that noise exists in the signal of a specific lead when the score calculated based on the specific lead is equal to or greater than a threshold value; 13. The method of claim 12, comprising:

14. determining whether the acquired read target data is readable data based on at least one of whether noise exists in the signal determined based on the calculated score or whether leakage of the estimated signal exists, determining that noise exists or that the signal of a specific read that is estimated to have leaked is unreadable data based on the calculated score; 13. The method of claim 12, comprising:

15. a computer program stored on a computer-readable storage medium, the computer program, when executed by one or more processors, performing operations for quantifying a quality of a biological signal; The operation is obtaining at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled by empirical judgment of a domain expert related to reading the electrocardiogram signal; generating a machine learning model for quantifying electrocardiogram signal quality based on at least one of the acquired first electrocardiogram data set or the acquired second electrocardiogram data set; a computer program comprising:

16. 1. A computing device for quantifying a quality of a biosignal, comprising: a processor including at least one core; a memory containing program code executable by the processor; a network unit for obtaining at least one of a first electrocardiogram data set labeled based on morphological features of the electrocardiogram signal or a second electrocardiogram data set labeled according to the empirical judgment of a domain expert related to the interpretation of the electrocardiogram signal; Including, the processor generates a machine learning model for quantifying electrocardiogram signal quality based on at least one of the acquired first electrocardiogram data set or the acquired second electrocardiogram data set. Device.

Citation Information

Patent Citations

  • Method for classifying the quality of biological sensor data - Patents.com

    JP2024534073A

  • How to diagnose a disease

    JP2025501130A

  • Semicondutor device for generating a reference current or votlage in a temperature change

    KR1020230112326A

  • Information processing device, authentication system, information processing method, non-transitory computer-readable medium, trained model, and method for generating trained model

    WO2023228731A1