Method and apparatus for analyzing medical data

The method generates integrated synthetic medical datasets using advanced algorithms to address data sharing challenges, ensuring secure and efficient analysis across institutions, overcoming data insufficiency and error issues.

WO2025147097A1PCT designated stage expired Publication Date: 2025-07-10EVIDNET CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/000040
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2025-01-02
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The challenge of analyzing medical data across different institutions is hindered by personal information leakage risks and strict access controls, making it difficult to share and utilize medical data for research, while existing synthetic data solutions do not adequately address data insufficiency and error issues.

Method used

A method and device for generating an integrated synthetic dataset using Generative Adversarial Networks (GAN) and Synthetic Minority Oversampling Technique (SMOTE) algorithms, combining medical datasets from multiple institutions to derive analysis results, including independent and dependent variables, and applying Bayesian inference for probability distributions.

Benefits of technology

Enables secure and efficient analysis of medical data without direct storage or export, reducing time and economic costs, and providing reliable analysis results by generating synthetic datasets that reflect actual medical data distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025000040_10072025_PF_FP_ABST
    Figure KR2025000040_10072025_PF_FP_ABST
Patent Text Reader

Abstract

According to one embodiment of the present disclosure, a method for analyzing medical data of a medical institution is provided. The method comprises the steps of: acquiring a first medical data set of a first medical institution and a second medical data set of a second medical institution from the first medical institution and the second medical institution, respectively; generating a total synthetic data set from the first medical data set and the second medical data set; and acquiring an analysis result for a target medical data set from the total synthetic data set on the basis of the target medical data set of a target medical institution, wherein the first medical data set, the second medical data set and the total synthetic data set can include a preset independent variable and a preset dependent variable.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for analyzing medical data

[0001] The present disclosure relates to a method and apparatus for analyzing medical data, and more particularly, to a method and apparatus for analyzing a medical dataset of a target medical institution from a synthetic dataset.

[0002] Patient medical data is considered personal and sensitive information under the Personal Information Protection Act. Utilizing this data carries the risk of personal information leakage or personal identification. Consequently, access to medical data outside of hospitals is limited, and strict standards govern its use and application. Recently, the medical field has seen a growing trend of applying synthetic medical data generated from medical data to research. Synthetic medical data allows research to proceed without the need for direct storage or export of medical data, enabling research results to be obtained without the security issues associated with personal and sensitive information leakage. Furthermore, synthetic data generation technology has recently been proposed as a solution to address data entry errors and data insufficiency, significantly reducing the time and financial costs associated with medical data analysis.

[0003] In relation to this, Korean patent registration number 10-2403461 has been issued.

[0004] This disclosure addresses the aforementioned background technology and addresses the problem of analyzing medical data. For example, the present disclosure addresses the problem of obtaining analysis results for a target medical dataset from an integrated synthetic dataset based on a target medical dataset from a target medical institution.

[0005] Meanwhile, the technical task to be achieved by the present disclosure is not limited to the technical task mentioned above, and may include various technical tasks within a scope obvious to a person skilled in the art from the contents described below.

[0006] The present disclosure has been devised in response to the aforementioned background technology, and aims to provide a method and device for analyzing medical data of a medical institution.

[0007] According to one aspect of the present disclosure, a method for analyzing medical data of a medical institution, performed by a computing device, may be provided. The method includes the steps of: obtaining a first medical dataset of a first medical institution and a second medical dataset of a second medical institution from a first medical institution and a second medical institution; generating an integrated (total) synthetic dataset from the first medical dataset and the second medical dataset; and obtaining an analysis result for a target medical dataset from the integrated synthetic dataset based on a target medical dataset of a target medical institution, wherein the first medical dataset, the second medical dataset, and the integrated synthetic dataset may include predetermined independent variables and predetermined dependent variables.

[0008] In one embodiment, the step of generating the integrated synthetic dataset may include the step of obtaining clinical prior information including information on a correlation between the predetermined independent variable and the predetermined dependent variable; and the step of generating the integrated synthetic dataset from the clinical prior information, the first medical dataset, and the second medical dataset.

[0009] In one embodiment, the step of generating the integrated synthetic dataset may include the step of generating a first synthetic dataset corresponding to the first medical dataset and a second synthetic dataset corresponding to the second medical dataset from the first medical dataset and the second medical dataset; and the step of generating the integrated synthetic dataset by integrating the first synthetic dataset and the second synthetic dataset.

[0010] In one embodiment, the step of generating the first synthetic dataset and the second synthetic dataset may include the step of generating the first synthetic dataset and the second synthetic dataset from the first medical dataset and the second medical dataset using a pre-trained Generative Adversarial Networks (GAN) model.

[0011] In one embodiment, the step of generating the first synthetic dataset and the second synthetic dataset may include the step of generating the first synthetic dataset and the second synthetic dataset from the first medical dataset and the second medical dataset using a Synthetic Minority Oversampling Technique (SMOTE) algorithm.

[0012] In one embodiment, the step of generating the first synthetic dataset and the second synthetic dataset may include the step of obtaining clinical prior information including information on a correlation between the predetermined independent variable and the predetermined dependent variable; and the step of generating the first synthetic dataset and the second synthetic dataset by reflecting the first medical dataset and the second medical dataset in the clinical prior information using a Bayesian algorithm based on Bayesian inference.

[0013] In one embodiment, the method further includes a step of determining, from the first medical data set and the second medical data set, a first independent variable that is correlated with the dependent variable and is not included in the predetermined independent variables, and the analysis result may include information on the correlation between the first independent variable and the predetermined dependent variable.

[0014] In one embodiment, the integrated synthetic dataset includes a first weight corresponding to the preset independent variable, and the step of obtaining an analysis result for the target medical dataset may include: obtaining a second weight corresponding to the preset independent variable from the target medical dataset; obtaining a third weight corresponding to the preset independent variable based on the first weight and the second weight; and obtaining the analysis result from the integrated synthetic dataset based on the preset independent variable, the preset dependent variable, and the third weight.

[0015] In one embodiment, the step of obtaining a third weight corresponding to the preset independent variable may include: determining a weighting degree of each of the first weight and the second weight based on the number of first medical data of the first medical data set, the number of second medical data of the second medical data set, and the number of target medical data of the target medical data set; and obtaining the third weight based on the weighting degrees of each of the first weight and the second weight.

[0016] In one embodiment, the step of determining the weighting degree of each of the first weight and the second weight may include the step of obtaining clinical prior information including clinical statistical data in which clinical values ​​are recorded; and the step of determining the weighting degree of each of the first weight and the second weight based on the number of clinical medical data of the clinical statistical data, the number of first medical data of the first medical data set, the number of second medical data of the second medical data set, and the number of target medical data of the target medical data set.

[0017] In one embodiment, the step of obtaining the analysis result for the target medical data set may include the step of generating a target synthetic data set corresponding to the target medical data set from the integrated synthetic data set based on the preset independent variable, the preset dependent variable, and the third weight; and the step of obtaining the analysis result from the target synthetic data set.

[0018] In one embodiment, the step of obtaining an analysis result for the target medical dataset may include: determining, from the first medical dataset and the second medical dataset, a first independent variable that is correlated with the dependent variable and is not included in the preset independent variables; generating, from the integrated synthetic dataset, a target synthetic dataset corresponding to the target medical dataset based on the preset independent variable, the first independent variable, the preset dependent variable, and the third weight; and obtaining the analysis result from the target synthetic dataset.

[0019] In one embodiment, the analysis results for the target medical data set may include information about the predetermined independent variable, the predetermined dependent variable, the third weight, and the correlation between the predetermined independent variable and the predetermined dependent variable.

[0020] In one embodiment, the first medical dataset and the second medical dataset may be medical datasets converted from the first raw medical dataset and the second raw medical dataset into a data structure of a Common Data Model (CDM), wherein the first medical dataset corresponds to the first raw medical dataset, and the second medical dataset corresponds to the second raw medical dataset.

[0021] In one embodiment, the step of obtaining an analysis result for the target medical dataset includes: generating a first integrated synthetic dataset from the first medical dataset and the second medical dataset; generating a second integrated synthetic dataset from the first medical dataset and the second medical dataset; and obtaining an analysis result for the target medical dataset from the first integrated synthetic dataset and the second integrated synthetic dataset based on the target medical dataset, wherein the first integrated synthetic dataset includes a first-first weight corresponding to the preset independent variable, and the second integrated synthetic dataset includes a first-second weight corresponding to the preset independent variable, and the first-first weight and the first-second weight may be different from each other.

[0022] In one embodiment, the predetermined dependent variables included in each of the first integrated synthetic dataset and the second integrated synthetic dataset may have probability distributions based on different means and standard deviations.

[0023] In one embodiment, the step of obtaining an analysis result for the target medical dataset may include: obtaining a first analysis result for the target medical dataset from the first integrated synthetic dataset; obtaining a second analysis result for the target medical dataset from the second integrated synthetic dataset; and obtaining an analysis result for the target medical dataset based on the first analysis result and the second analysis result.

[0024] In one embodiment, the method further comprises the step of determining the required number of data of the medical data set for generating the integrated synthetic data set based on at least one of a preset effect size, a preset power, and a preset significance level, and each of the first medical data set and the second medical data set may include medical data greater than the required number of data.

[0025] According to one aspect of the present disclosure, a computing device for analyzing medical data of a medical institution may be provided. The computing device includes at least one processor; and a memory, wherein the at least one processor obtains a first medical dataset of a first medical institution and a second medical dataset of the second medical institution from a first medical institution and a second medical institution, generates a total synthetic dataset from the first medical dataset and the second medical dataset, and obtains an analysis result for the target medical dataset from the total synthetic dataset based on a target medical dataset of a target medical institution, and the first medical dataset, the second medical dataset, and the total synthetic dataset may include a preset independent variable and a preset dependent variable.

[0026] According to one aspect of the present disclosure, a computer program stored in a computer-readable storage medium may be provided. When the computer program is executed by one or more processors, the computer program causes the one or more processors to perform operations for analyzing medical data of a medical institution, the operations including: obtaining a first medical dataset of the first medical institution and a second medical dataset of the second medical institution from a first medical institution and a second medical institution; generating an integrated (total) synthetic dataset from the first medical dataset and the second medical dataset; and obtaining an analysis result for the target medical dataset from the integrated synthetic dataset based on a target medical dataset of a target medical institution, and the first medical dataset, the second medical dataset, and the integrated synthetic dataset may include predetermined independent variables and predetermined dependent variables.

[0027] According to some embodiments of the present disclosure, analysis results for a target medical dataset can be obtained from an integrated synthetic dataset based on a target medical dataset of a target medical institution.

[0028] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0029] Various aspects are now described with reference to the drawings, wherein like reference numerals are used to refer to similar elements generally. In the following examples, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of one or more aspects. However, it will be apparent that such aspects may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to facilitate the description of one or more aspects.

[0030] FIG. 1 is a block diagram of a computing device that obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0031] FIG. 2 is a schematic diagram illustrating a network function according to one embodiment of the present disclosure.

[0032] FIG. 3 is a flowchart illustrating a method by which a computing device obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0033] FIG. 4 is a schematic diagram illustrating a process in which a computing device obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0034] FIG. 5 is a schematic diagram illustrating a process for calculating weights of target medical institutions according to one embodiment of the present disclosure.

[0035] FIG. 6 is a simplified, general schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented.

[0036] Various embodiments are now described with reference to the drawings. In this specification, various descriptions are provided to facilitate understanding of the present disclosure. However, it will be apparent that these embodiments may be practiced without these specific details.

[0037] As used herein, the terms "component," "module," "system," and the like refer to computer-related entities, hardware, firmware, software, a combination of software and hardware, or an execution of software. For example, a component may be, but is not limited to, a procedure running on a processor, a processor, an object, a thread of execution, a program, and / or a computer. For example, both an application running on a computing device and the computing device may be a component. One or more components may reside within a processor and / or a thread of execution. A component may be localized within a single computer. A component may be distributed between two or more computers. Furthermore, these components may execute from various computer-readable media having various data structures stored therein. Components may communicate via local and / or remote processes, for example, by signals comprising one or more data packets (e.g., data from one component interacting with another component in a local system, a distributed system, and / or data transmitted to another system via a network such as the Internet via signals).

[0038] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from context, "X employs A or B" is intended to mean either of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, "X employs A or B" can apply to any of these cases. Furthermore, the terms "or" and "and / or" as used herein should be understood to refer to and include all possible combinations of one or more of the associated items listed.

[0039] Additionally, the terms "comprises" and / or "comprising" should be understood to imply the presence of the features and / or components in question. However, it should be understood that the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other features, components, and / or groups thereof. Furthermore, unless otherwise specified or clear from the context to refer to the singular form, the singular in the specification and claims should generally be construed to mean "one or more."

[0040] And, the term "at least one of A or B" should be interpreted to mean "if it includes only A", "if it includes only B", or "if it is combined in the composition of A and B".

[0041] Those skilled in the art should further appreciate that the various illustrative logical blocks, configurations, modules, circuits, means, logics, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate the interchangeability of hardware and software, various illustrative components, blocks, configurations, means, logics, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. However, such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0042] The description of the disclosed embodiments is provided to enable a person skilled in the art to make or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art. The general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present invention is not limited to the embodiments disclosed herein. The present invention is to be construed in the widest scope consistent with the principles and novel features disclosed herein.

[0043] In the present disclosure, terms expressed as N, such as first, second, or third, are used to distinguish at least one entity. For example, entities expressed as first and second may be the same or different from each other.

[0044] FIG. 1 is a diagram illustrating a computing device that obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0045] The configuration of the computing device (100) illustrated in FIG. 1 is merely a simplified example. In one embodiment of the present disclosure, the computing device (100) may include other configurations for performing the computing environment of the computing device (100), and only some of the disclosed configurations may constitute the computing device (100).

[0046] A computing device (100) according to some embodiments of the present disclosure may be a device for performing analysis on a target medical dataset. For example, the computing device (100) may be a device for generating a total synthetic dataset and performing analysis on the target medical dataset from the total synthetic dataset. The computing device (100) may include any type of server or user terminal.

[0047] A computing device (100) may include a processor (110), memory (130), and network unit (150).

[0048] The processor (110) may be composed of one or more cores and may include a processor for performing operations related to data processing, such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU) of the computing device (100).

[0049] According to one embodiment of the present disclosure, the processor (110) may perform operations for learning a neural network. For example, the processor (110) may perform calculations for learning a neural network, such as processing input data for learning in deep learning (DL), extracting features from the input data, calculating errors, and updating weights of the neural network using backpropagation. At least one of the CPU, GPGPU, and TPU of the processor (110) may process learning of a network function. For example, the CPU and GPGPU may together process learning of a network function and classification of data using the network function. Furthermore, in one embodiment of the present disclosure, processors of a plurality of computing devices may be used together to process learning of a network function and classification of data using the network function.

[0050] The processor (110) can typically control the overall operation of the computing device (100). The processor (110) can process signals, data, information, etc. input or output through components included in the computing device (100) or run application programs stored in the memory (130), thereby providing or processing appropriate information or functions to the user.

[0051] In one embodiment of the present disclosure, the memory (130) may store any form of information generated or determined by the processor (110) and / or any form of information received by the network unit (150). In one embodiment, the memory (130) may store a database. The database may be a collection of data stored in a form that can be processed by the computing device (100). For example, the memory (130) may include a medical data database. The medical data database may store a plurality of medical data.

[0052] In one embodiment of the present disclosure, the memory (130) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and / or an optical disk. The computing device (100) may also operate in relation to web storage that performs the storage function of the memory (130) on the Internet. The description of the memory described above is merely an example, and the present disclosure is not limited thereto. The memory (130) may be operated by the processor (110).

[0053] According to one embodiment of the present disclosure, the network unit (150) may include any wired or wireless communication network capable of transmitting and receiving any type of data and signals, etc., as represented in the network of the present disclosure. The technologies described herein may be used not only in the networks mentioned above, but also in other networks.

[0054] FIG. 2 is a diagram illustrating a network function according to one embodiment of the present disclosure.

[0055] Throughout this specification, the terms artificial intelligence-based model (e.g., summary information generation model, information classification model, etc.), computational model, neural network, network function, and neural network may be used with the same meaning (interchangeably).

[0056] A neural network can be composed of a set of interconnected computational units, generally referred to as nodes. These nodes can also be referred to as neurons. A neural network consists of at least one node. The nodes (or neurons) that make up a neural network can be interconnected by one or more links.

[0057] Within a neural network, one or more nodes connected via links can form a relationship between input nodes and output nodes. The concept of input nodes and output nodes is relative, meaning that any node that is in an output node relationship with one node can also be in an input node relationship with another node, and vice versa. As described above, the relationship between input nodes and output nodes can be created based on links. One input node can be connected to one or more output nodes via links, and vice versa.

[0058] In a relationship between input nodes and output nodes connected through a single link, the data of the output node can have its value determined based on the data input to the input node. Here, the link interconnecting the input nodes and output nodes can have a weight. The weight can be variable and can be varied by the user or an algorithm so that the neural network can perform a desired function. For example, when one or more input nodes are interconnected to one output node through each link, the output node can determine the output node value based on the values ​​input to the input nodes connected to the output node and the weight set on the link corresponding to each input node.

[0059] As described above, a neural network is a network in which one or more nodes are interconnected through one or more links, forming input and output node relationships within the network. The characteristics of a neural network can be determined based on the number of nodes and links within the network, the relationships between the nodes and links, and the weights assigned to each link. For example, if two neural networks have the same number of nodes and links but different weight values ​​for the links, the two neural networks can be perceived as different from each other.

[0060] A neural network can be composed of a set of one or more nodes. A subset of the nodes comprising the neural network can form a layer. Some of the nodes comprising the neural network can form a layer based on their distances from the initial input node. For example, a set of nodes that are n distances from the initial input node can form n layers. The distance from the initial input node can be defined by the minimum number of links required to reach the node from the initial input node. However, this definition of a layer is arbitrary for illustrative purposes, and the degree of a layer within a neural network can be defined in a different way than described above. For example, a layer of nodes can be defined by its distance from the final output node.

[0061] In one embodiment of the present disclosure, a set of neurons or nodes may be defined as a layer.

[0062] An initial input node may refer to one or more nodes within a neural network into which data is directly input without going through links with other nodes. Alternatively, within a neural network, it may refer to nodes that do not have other input nodes connected by links in the relationship between nodes based on links. Similarly, a final output node may refer to one or more nodes within a neural network that do not have output nodes in their relationship with other nodes. Furthermore, a hidden node may refer to nodes that constitute a neural network other than the initial input node and the final output node.

[0063] A neural network according to one embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be the same as the number of nodes in an output layer, and the number of nodes decreases and then increases as it progresses from the input layer to a hidden layer. In addition, a neural network according to another embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be less than the number of nodes in an output layer, and the number of nodes decreases as it progresses from the input layer to the hidden layer. In addition, a neural network according to another embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be greater than the number of nodes in an output layer, and the number of nodes increases as it progresses from the input layer to the hidden layer. A neural network according to another embodiment of the present disclosure may be a neural network in a combined form of the neural networks described above.

[0064] A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to input and output layers. Using a deep neural network, one can identify latent structures in data. A deep neural network can include a convolutional neural network (CNN), a recurrent neural network (RNN), an autoencoder, a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a Q network, a U network, a Siamese network, and the like. The description of the above-described deep neural network is merely an example, and the present disclosure is not limited thereto.

[0065] For example, the artificial intelligence-based model of the present disclosure may include a transformer, a generative pre-trained transformer (GPT), a bidirectional encoder representation from transformers (BERT), etc.

[0066] The artificial intelligence-based model of the present disclosure can be represented by a network structure of any structure described above, including an input layer, a hidden layer, and an output layer.

[0067] The neural network that can be used in the artificial intelligence-based model of the present disclosure may be trained using at least one of supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, federated learning for distributed deep learning, and incremental learning. Training of the neural network may be a process of applying knowledge to the neural network to perform a specific operation.

[0068] Neural networks can be trained to minimize output errors. This process involves repeatedly inputting training data into the neural network, calculating the neural network output and target error for the training data, and backpropagating the neural network error from the output layer to the input layer to update the weights of each node in the neural network to reduce the error. In supervised learning, training data with the correct answer for each training data is used (i.e., labeled training data). In unsupervised learning, the correct answer may not be labeled for each training data. For example, in supervised learning for data classification, the training data may be data with each category labeled. Labeled training data is input into the neural network, and the error can be calculated by comparing the output (category) of the neural network with the labels of the training data.

[0069] As another example, in unsupervised learning for data classification, errors can be calculated by comparing input training data with the output of a neural network. The calculated errors are backpropagated through the neural network (i.e., from the output layer to the input layer), and the connection weights of each node in each layer of the neural network can be updated based on the learning rate. The amount of change in the updated connection weights of each node can be determined by the learning rate. The neural network's calculation of the input data and backpropagation of the error can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of iterations of the neural network's learning cycle. For example, a high learning rate can be used early in the neural network's training process to quickly achieve a certain level of performance, thereby improving efficiency. A lower learning rate can be used later in the training process to improve accuracy.

[0070] In neural network training, training data can typically be a subset of real-world data (i.e., the data to be processed using the trained neural network). Therefore, there can be a learning cycle where errors on the training data decrease but errors on the real-world data increase. Overfitting is a phenomenon where excessive training on the training data leads to increased errors on the real-world data. For example, a neural network trained on yellow cats may fail to recognize cats when shown non-yellow colors, a type of overfitting. Overfitting can increase errors in machine learning algorithms. Various optimization methods can be used to prevent overfitting. These methods include increasing the training data, regularization, dropout, which disables some nodes in the network during the learning process, and utilizing batch normalization layers.

[0071] In one embodiment, the computing device (100) may utilize a pre-trained Generative Adversarial Networks (GAN) model. The pre-trained GAN model may include a generator neural network and a discriminator neural network, and the generator neural network may correspond to a pre-trained artificial intelligence model that generates synthetic data from real data and the discriminator neural network may discriminate between real data and synthetic data. The generator neural network may be pre-trained through the above-described learning process to make it difficult for the discriminator neural network to discriminate between real data and generated synthetic data. The computing device (100) may use the pre-trained GAN model to generate a synthetic data set having the probability distribution and characteristics of the real medical data set from the real medical data set.

[0072] Hereinafter, a method is disclosed in which a computing device (100) generates an integrated synthetic dataset from medical datasets of each of a plurality of medical institutions, and obtains analysis results for a target medical dataset of a target medical institution from the integrated synthetic dataset, according to one embodiment of the present disclosure.

[0073] FIG. 3 is a flowchart illustrating a method by which a computing device obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0074] In step S301, the computing device (100) can acquire a first medical dataset of the first medical institution and a second medical dataset of the second medical institution from the first medical institution and the second medical institution. The computing device (100) can acquire a medical dataset from each of a plurality of medical institutions. For convenience of explanation, the following description will be given as an example of acquiring medical datasets from the first medical institution and the second medical institution. However, the present invention is not limited thereto, and a medical dataset can be acquired from each of a plurality of medical institutions, and an integrated synthetic dataset can be generated from the plurality of medical datasets.

[0075] In one embodiment, a medical dataset may refer to a collection of medical data. Medical data may include information related to a specific individual's medical care. For example, medical data may include variables related to at least one disease in the medical field. For example, medical data may include data that includes a specific individual's physical information, health information, and treatment information, and may include a specific individual's CT image data, MRI test result data, electrocardiogram measurement result data, and Electronic Medical Record (EMR) data. Health information may include information about an individual's disease, such as the presence or absence of a specific individual's disease, the name of the disease, and the stage of the disease. Medical information may include information about an individual's medical record, such as the presence or absence of a specific individual's prescribed medication, the name of the prescribed medication, the duration of the prescribed medication, the dosage of the prescribed medication, whether or not the individual has undergone surgery, the name of the surgery, and the details of the surgery.

[0076] In one embodiment, medical data may include variables related to a specific person's physical information, health information, and treatment information. Furthermore, medical data may include disease-related variables such as lifestyle habits (drinking, smoking, etc.), family history, age, gender, cholesterol levels, and genetic variables. The computing device (100) may set research subjects and objectives. Furthermore, the computing device (100) may set independent and dependent variables among the variables included in the medical data according to the research subjects and objectives. In one embodiment, the independent variable may refer to a variable that is a cause, and the dependent variable may refer to a variable whose value is determined by the independent variable. Each of the independent and dependent variables may include at least one variable. The first medical dataset of the first medical institution and the second medical dataset of the second medical institution may include preset independent variables and preset dependent variables.

[0077] For example, if the research objective is whether side effects of pioglitazone occur in patients with non-alcoholic fatty liver disease, the computing device (100) may set the research subjects as patients with non-alcoholic fatty liver disease who took pioglitazone, and may obtain a medical dataset including medical data of patients with non-alcoholic fatty liver disease who took pioglitazone from a first medical institution and a second medical institution to create a synthetic dataset. In this case, the computing device (100) may set the presence or absence of non-alcoholic fatty liver disease, the stage of non-alcoholic fatty liver disease, the period of occurrence of non-alcoholic fatty liver disease, whether or not pioglitazone was taken, the dosage of pioglitazone, the period of taking pioglitazone, etc. as independent variables. The computing device (100) may set variables related to side effects such as weight gain and edema (e.g., weight, presence or absence of edema, etc.) as dependent variables.

[0078] In one embodiment, the first medical dataset and the second medical dataset may be medical datasets converted from the first raw medical dataset and the second raw medical dataset into a data structure of a common data model (CDM). At this time, the first medical dataset may correspond to the first raw medical dataset, and the second medical dataset may correspond to the second raw medical dataset. In one embodiment, the raw medical dataset may be an unstructured medical dataset with different data structures that are each held by a plurality of medical institutions. In one embodiment, the common data model may mean a data model that sets the same data structure and specifications for the medical datasets with different data structures that are each held by a plurality of medical institutions. For example, the common data model may include OMOP-CDM, Sentinel-CDM, PCORnet CDM, etc. Variables contained in raw medical data can be converted or categorized into standardized data formats, labels, and classes, transforming them into medical data with a common data model data structure. This allows for batch processing of medical datasets from multiple medical institutions.

[0079] In one embodiment, the computing device (100) can convert the first raw medical dataset and the second raw medical dataset into the first medical dataset and the second medical dataset having the data structure of the common data model. In another embodiment, the computing device (100) can obtain the first medical dataset and the second medical dataset having the data structure of the common data model from the first medical institution and the second medical institution.

[0080] In one embodiment, the computing device (100) may determine the required number of data of the medical dataset for generating the integrated synthetic dataset based on at least one of a preset effect size, a preset power, and a preset significance level. Each of the first medical dataset and the second medical dataset may include medical data greater than the required number of data. For example, the required number of data may refer to the required number of data of a sample group to obtain meaningful results for the population. In one embodiment, the effect size may be a value obtained by dividing the difference in the mean values ​​of the first medical dataset and the second medical dataset under study by the standard deviation of the first medical dataset or the standard deviation of the second medical dataset. In one embodiment, the power may be a value representing the degree of probability sufficient to conclude that the result is statistically significant. In one embodiment, the significance level may be a value obtained by subtracting the confidence level from 1. For example, if the confidence level is 95%, the significance level may be 0.05, which is 1-0.95. The computing device (100) can set an effect size, a test power, and a significance level, and determine the required number of data for each of the first medical data set and the second medical data set based on the set effect size, test power, and significance level.

[0081] In one embodiment, the computing device (100) may not utilize a corresponding medical dataset for generating an integrated synthetic dataset for a medical institution among a plurality of medical institutions that includes medical data including predetermined independent variables and predetermined dependent variables according to the research subject and purpose in a number less than the required number of data.

[0082] In one embodiment, the computing device (100) may delete, from the first medical data set, first-1 medical data that does not include at least one of a predetermined dependent variable and a predetermined independent variable among the first medical data included in the first medical data set. Similarly, the computing device (100) may delete, from the second medical data set, second-1 medical data that does not include at least one of a predetermined dependent variable and a predetermined independent variable among the second medical data included in the second medical data set.

[0083] In one embodiment, the computing device (100) may delete, from the first medical data set, first-second medical data that includes a value that deviates from a preset threshold from a clinical value included in clinical prior information among the first medical data included in the first medical data set. Similarly, the computing device (100) may delete, from the second medical data set, second-second medical data that includes a value that deviates from a preset threshold from a clinical value included in clinical prior information among the second medical data included in the second medical data set. For example, the medical data that includes a value that deviates from a preset threshold may be medical data that includes an input error value.

[0084] In step S302, the computing device (100) may generate an integrated synthetic dataset from the first medical dataset and the second medical dataset. The integrated synthetic dataset may include preset independent variables and preset dependent variables. In one embodiment, the integrated synthetic dataset may be a synthetic dataset generated to follow a probability distribution of a preset dependent variable according to preset independent variables derived from the first medical dataset and the second medical dataset. For example, when the preset dependent variable includes multiple dependent variables, the integrated synthetic dataset may be generated to follow the probability distribution of each of the multiple dependent variables according to the preset independent variables. In one embodiment, the computing device (100) may transmit the integrated synthetic dataset to an external device or an external server. Through this, a technical effect may be achieved in that a research institution or researcher who has requested an analysis may be provided with a synthetic dataset based on an actual medical dataset.

[0085] In one embodiment, the computing device (100) may obtain clinical prior information including information on the correlation between predetermined independent variables and predetermined dependent variables. In one embodiment, the clinical prior information may include prior clinical studies on predetermined research subjects and objectives, clinical statistics, etc. For example, the clinical prior information may include information on the research period, disease (or drug) code information of the target disease that is the research objective, information on the prevalence and incidence of the target disease by year / medical institution, and statistical data (e.g., statistical amounts, statistical methods, statistical data, statistical ranges, statistical units, etc.) on predetermined independent variables and predetermined dependent variables. In one embodiment, the clinical prior information may include clinical statistics in which clinical values ​​are recorded. For example, the clinical statistics may include information on the number of clinical medical data.

[0086] In one embodiment, information regarding the correlation between a predetermined independent variable and a predetermined dependent variable may be information regarding the numerical change of the predetermined dependent variable according to the predetermined independent variable. For example, the predetermined independent variable may include multiple independent variables, and each of the multiple independent variables may be assigned a corresponding weight. The weight corresponding to each of the multiple independent variables may reflect the correlation between the predetermined independent variable and the predetermined dependent variable. For example, if the predetermined independent variable and the multiple dependent variable are linearly proportional, the weight corresponding to each of the multiple independent variables may be a proportional constant. Based on the weights for the predetermined independent variables, a probability distribution of the predetermined dependent variable according to the predetermined independent variable may be obtained.

[0087] In one embodiment, the computing device (100) can generate an integrated synthetic dataset from the acquired clinical prior information, the first medical dataset, and the second medical dataset. For example, the computing device (100) can generate an integrated synthetic dataset that follows the probability distribution of a predetermined dependent variable according to a predetermined independent variable in clinical statistics included in the clinical prior information and the probability distribution of a predetermined dependent variable according to a predetermined independent variable in the first medical dataset and the second medical dataset. Through this, the technical effect of being able to generate a synthetic dataset that follows a probability distribution based on clinical values ​​using the clinical prior information can be achieved. In addition, the technical effect of being able to generate a synthetic dataset that follows a probability distribution based on actual data using medical data from an actual medical institution, rather than simply based on clinical values, can be achieved.

[0088] In one embodiment, when the predetermined independent variable and the predetermined dependent variable are numerical variables having numerical values, the computing device (100) can calculate a probability distribution of the numerical values ​​of the predetermined dependent variable for each numerical interval of the predetermined independent variable from the first medical dataset and the second medical dataset. In addition, the computing device (100) can calculate a probability distribution of the numerical values ​​of the predetermined dependent variable for each numerical interval of the predetermined independent variable from the clinical prior information, the first medical dataset, and the second medical dataset. The computing device (100) can generate an integrated synthetic dataset that follows the probability distribution of the numerical values ​​of the predetermined dependent variable for each numerical interval of the calculated predetermined independent variable. In this case, the integrated synthetic dataset may have a weight corresponding to each numerical interval of the predetermined independent variable according to the probability distribution of the numerical values ​​of the predetermined dependent variable. For example, the weight may correspond to the average value of the predetermined dependent variable according to the numerical interval of the predetermined independent variable.

[0089] For example, if the research objective is to study the correlation between obesity and diabetes, the independent variable may be set as the Body Mass Index (BMI) value, and the dependent variable may be set as the blood sugar level. The computing device (100) may calculate the probability distribution of blood sugar levels for each BMI index value range from the first medical dataset and the second medical dataset, each of which includes the BMI index and blood sugar level. The computing device (100) may generate an integrated synthetic dataset having the calculated probability distribution.

[0090] In one embodiment, if the preset independent variable is a numerical variable and the preset dependent variable is a logical variable having a value of 1 (True) or 0 (False) depending on whether a disease (or side effect) occurs, the computing device (100) can calculate a probability distribution for the occurrence probability of the preset dependent variable according to the values ​​of the preset independent variable from the first medical data set and the second medical data set. In addition, the computing device (100) can calculate a probability distribution for the occurrence probability of the preset dependent variable according to the values ​​of the preset independent variable from the clinical prior information, the first medical data set, and the second medical data set. For example, if the preset dependent variable is 1 (True), it can mean that the disease (or side effect) corresponding to the preset dependent variable has occurred, and if the preset dependent variable is 0 (False), it can mean that the disease (or side effect) corresponding to the preset dependent variable has not occurred, etc. The computing device (100) can generate an integrated synthetic dataset that follows a probability distribution for the occurrence probability of a predetermined dependent variable based on the numerical values ​​of the predetermined independent variables. In this case, the integrated synthetic dataset can have weights corresponding to the predetermined independent variables based on the probability distribution for the occurrence probability of the predetermined dependent variables. For example, the weights can correspond to the occurrence probability of the predetermined dependent variable based on the numerical values ​​of the predetermined independent variables.

[0091] For example, if the research objective is to study the correlation between obesity and diabetes, the independent variable may be set as the Body Mass Index (BMI) value, and the dependent variable may be set as the presence or absence of diabetes. The computing device (100) may calculate a probability distribution for the probability of diabetes occurrence according to the BMI index from the first medical dataset and the second medical dataset, each of which includes the BMI index and the presence or absence of diabetes. The computing device (100) may generate an integrated synthetic dataset having the calculated probability distribution.

[0092] In one embodiment, when a predetermined independent variable is a logical variable having a value of 1 (True) or 0 (False) depending on whether a disease occurs or whether a prescription drug is present, and a predetermined dependent variable is a numerical variable, the computing device (100) can calculate a probability distribution of a predetermined dependent variable value according to a predetermined independent variable having a value of 1 (True) from the first medical data set and the second medical data set. In addition, the computing device (100) can calculate a probability distribution of a predetermined dependent variable value according to a predetermined independent variable having a value of 1 (True) from clinical prior information, the first medical data set, and the second medical data set. The computing device (100) can generate an integrated synthetic data set that follows the probability distribution of the predetermined dependent variable value according to the calculated predetermined independent variable. In this case, the integrated synthetic data set can have a weight corresponding to the predetermined independent variable according to the probability distribution of the predetermined dependent variable value. For example, the weights may correspond to the mean of a predetermined dependent variable based on predetermined independent variables.

[0093] For example, when studying the side effects of pioglitazone on patients with nonalcoholic fatty liver disease, the independent variables may be set as the presence or absence of nonalcoholic fatty liver disease and the use of pioglitazone, and the dependent variable may be set as weight gain. The computing device (100) may calculate a probability distribution of weight gain according to pioglitazone use in patients with nonalcoholic fatty liver disease from a first medical dataset and a second medical dataset including medical data on patients with nonalcoholic fatty liver disease who took pioglitazone. The computing device (100) may generate an integrated synthetic dataset having the calculated probability distribution.

[0094] In one embodiment, when the predetermined independent variable and the predetermined dependent variable are logical variables, the computing device (100) can calculate the occurrence probability of the predetermined dependent variable according to the predetermined independent variable from the first medical data set and the second medical data set. In addition, the computing device (100) can calculate the occurrence probability of the predetermined dependent variable according to the predetermined independent variable from the clinical prior information, the first medical data set, and the second medical data set. The computing device (100) can generate an integrated synthetic data set having the occurrence probability of the predetermined dependent variable according to the calculated predetermined independent variable. In this case, the integrated synthetic data set can have a weight corresponding to the predetermined independent variable according to the occurrence probability of the predetermined dependent variable. For example, the weight can correspond to the occurrence probability of the predetermined dependent variable according to the predetermined independent variable.

[0095] For example, if the research objective is to study the correlation between obesity and diabetes, the independent variable may be set as obesity and the dependent variable may be set as diabetes. The computing device (100) may calculate a probability distribution for the probability of diabetes occurrence according to obesity from the first medical dataset and the second medical dataset, each of which includes information on obesity and diabetes. The computing device (100) may generate an integrated synthetic dataset having the calculated probability distribution.

[0096] In one embodiment, each of the preset independent variables and the preset dependent variables may include at least one of a numeric variable and a logical variable. For example, the computing device (100) may calculate a weight corresponding to each of the plurality of independent variables according to a method for calculating a probability distribution (or occurrence probability) of the preset dependent variables according to the preset independent variables. In one embodiment, the average value of each of the plurality of dependent variables may be calculated as a weighted average obtained by multiplying each of the plurality of independent variables by a plurality of weights corresponding to each of the plurality of independent variables. For example, when W1,…,Wn are weights corresponding to each independent variable, and X1,…Xn are a plurality of independent variables, the average value of each dependent variable may be calculated as (W1*X1 + W2*X2 +…+ Wn*Xn) / (W1 + W2 +…+ Wn).

[0097] In one embodiment, the computing device (100) can generate a first synthetic dataset corresponding to the first medical dataset and a second synthetic dataset corresponding to the second medical dataset from the first medical dataset and the second medical dataset. A detailed description of a method in which the computing device (100) generates a synthetic dataset from each of a plurality of medical institutions and generates an integrated synthetic dataset from the synthetic datasets corresponding to each of the plurality of medical institutions will be described later with reference to FIG. 4.

[0098] In step S303, the computing device (100) may obtain analysis results for the target medical dataset from the integrated synthetic dataset based on the target medical dataset of the target medical institution. In one embodiment, the target medical institution may be a medical institution for which analysis results are to be obtained. For example, the computing device (100) may obtain information about the target medical institution from a research institute or researcher who wishes to obtain analysis results for the target medical institution. The computing device (100) may obtain information about the target medical institution from an external device or an external server. In one embodiment, the target medical institution may be one of the first medical institution and the second medical institution. In another embodiment, the target medical institution may be a medical institution that is not included in the first medical institution and the second medical institution. The analysis results for the target medical dataset may include information about the correlation between a predetermined independent variable and a predetermined dependent variable in the target medical dataset. In one embodiment, the computing device (100) can transmit analysis results for a target medical dataset to an external device or server. This enables the research institution or researcher requesting the analysis to receive analysis results reflecting the target medical dataset in an integrated synthetic dataset, resulting in the technical benefit of:

[0099] For example, the computing device (100) can obtain an analysis result for the target medical dataset by reflecting the probability distribution of the preset dependent variable in the target medical dataset to the probability distribution of the preset dependent variable in the integrated synthetic dataset. For example, the computing device (100) can obtain an analysis result for the target medical dataset by reflecting the second weight corresponding to the preset independent variable in the target medical dataset to the first weight corresponding to the preset independent variable in the integrated synthetic dataset generated based on clinical prior information, the first medical dataset, and the second medical dataset. Through this, it is possible to obtain an analysis result that reflects the probability distribution of the target medical dataset while not being overfitted to the target medical dataset.

[0100] In one embodiment, the computing device (100) can determine, from the first medical dataset and the second medical dataset, a first independent variable that is correlated with a predetermined dependent variable and is not included in the predetermined independent variables. The computing device (100) can calculate a relationship for each of a plurality of variables included in the first medical dataset and the second medical dataset, and determine a variable that is correlated with the predetermined dependent variable as the first independent variable. At this time, the analysis result for the target medical dataset can include information on the correlation between the first independent variable and the predetermined dependent variable. Through this, a technical effect of being able to identify an independent variable that a research institution or researcher has not considered among a plurality of variables that are correlated with the predetermined dependent variable can be achieved.

[0101] In one embodiment, the computing device (100) can obtain weights to be reflected in the analysis results for the target medical dataset from clinical prior information, the integrated synthetic dataset, and the target medical dataset. A detailed description of the method for obtaining weights to be reflected in the analysis results for the target medical dataset will be described below with reference to FIG. 5.

[0102] In one embodiment, the computing device (100) can generate multiple integrated synthetic datasets from a first medical dataset and a second medical dataset. For convenience of explanation, a method for the computing device (100) to generate a first integrated synthetic dataset and a second integrated synthetic dataset from the first medical dataset and the second medical dataset will be described as an example.

[0103] In one embodiment, the computing device (100) can generate a first integrated synthetic dataset from a first medical dataset and a second medical dataset. In addition, the computing device (100) can generate a second integrated synthetic dataset from the first medical dataset and the second medical dataset. The computing device (100) can obtain analysis results for the target medical dataset from the first integrated synthetic dataset and the second integrated synthetic dataset based on the target medical dataset. The first integrated synthetic dataset may include a first-first weight corresponding to a preset independent variable, and the second integrated synthetic dataset may include a first-second weight corresponding to a preset independent variable. In this case, the first-first weight and the first-second weight may be different from each other. For example, when generating an integrated synthetic dataset from a first medical dataset and a second medical dataset, the values ​​of the predetermined independent variables and the predetermined dependent variables included in each generated synthetic dataset may differ from each other within a predetermined probability distribution. Accordingly, each time an integrated synthetic dataset is generated from the first medical dataset and the second medical dataset, the values ​​of the synthetic data included in the integrated synthetic dataset and the first weight of the integrated synthetic dataset may differ from each other.

[0104] In one embodiment, the predetermined dependent variables included in each of the first integrated synthetic dataset and the second integrated synthetic dataset may have probability distributions based on different means and standard deviations. For example, if the first-first weight and the first-second weight corresponding to each of the first integrated synthetic dataset and the second integrated synthetic dataset are different from each other, the probability distributions of the first integrated synthetic dataset and the second integrated synthetic dataset may be different from each other.

[0105] In one embodiment, the computing device (100) may obtain a first analysis result for the target medical dataset from the first integrated synthetic dataset. In addition, the computing device (100) may obtain a second analysis result for the target medical dataset from the second integrated synthetic dataset. The first analysis result and the second analysis result for the target medical dataset may have different third weights and probability distributions for a predetermined dependent variable. That is, the first analysis result and the second analysis result for the target medical dataset may be different analysis results. The computing device (100) may obtain the analysis result for the target medical dataset based on the first analysis result and the second analysis result. For example, the analysis result for the target medical dataset may be an analysis result including the first analysis result and the second analysis result. In one embodiment, the computing device (100) may transmit the analysis result for the target medical dataset including the first analysis result and the second analysis result to an external device or an external server. Through this, it is possible to achieve the technical effect of generating multiple synthetic datasets from the first medical dataset and the second medical dataset, performing multiple simulations on the target medical dataset based on the multiple synthetic datasets, and providing highly reliable analysis results to the research institute or researcher who requested the analysis.

[0106] FIG. 4 is a schematic diagram illustrating a process in which a computing device obtains analysis results for a target medical data set according to one embodiment of the present disclosure.

[0107] Referring to FIG. 4, a computing device (100) according to one embodiment can obtain a first medical dataset (401a), a second medical dataset (401b), and a third medical dataset (401c) from a first medical institution, a second medical institution, and a third medical institution. The computing device (100) can generate a first synthetic dataset (402a) from the first medical dataset (401a). The computing device (100) can generate a second synthetic dataset (402b) from the second medical dataset (401b). The computing device (100) can generate a third synthetic dataset (402c) from the third medical dataset (401c). The computing device (100) can generate an integrated synthetic dataset (403) by integrating the first synthetic dataset (402a), the second synthetic dataset (402b), and the third synthetic dataset (402c).

[0108] In one embodiment, the computing device (100) may obtain an analysis result (405) for the target medical data set (404) from the integrated synthetic data set (403) based on the target medical data set (404) of the target medical institution. For example, the computing device (100) may obtain an analysis result (405) for the target medical data set (404) by reflecting the probability distribution of the preset dependent variable in the target medical data set (404) to the probability distribution of the preset dependent variable in the integrated synthetic data set (403). Hereinafter, for the convenience of explanation, a method in which the computing device (100) generates the integrated synthetic data set (403) from the first medical data set (401a) and the second medical data set (401b) will be described as an example. However, without being limited thereto, the computing device (100) can generate an integrated synthetic data set (403) from the first medical data set (401a), the second medical data set (401b), and the third medical data set (401c) in the same manner.

[0109] In one embodiment, the computing device (100) can generate a first synthetic dataset (402a) corresponding to the first medical dataset (401a) and a second synthetic dataset (402b) corresponding to the second medical dataset (401b) from the first medical dataset (401a) and the second medical dataset (401b). The computing device (100) can generate an integrated synthetic dataset (403) by integrating the first synthetic dataset (402a) and the second synthetic dataset (402b).

[0110] In one embodiment, the computing device (100) can generate a first synthetic dataset (402a) and a second synthetic dataset (402b) from a first medical dataset (401a) and a second medical dataset (401b) using a pre-trained GAN model. For example, the computing device (100) can generate a first synthetic dataset (402a) having a probability distribution of a predetermined dependent variable in the first medical dataset (401a) using a pre-trained GAN model. In addition, the computing device (100) can generate a second synthetic dataset (402b) having a probability distribution of a predetermined dependent variable in the second medical dataset (401b) using a pre-trained GAN model.

[0111] In one embodiment, the computing device (100) can generate a clinical medical dataset that follows a probability distribution of a predetermined dependent variable in clinical prior information. The computing device (100) can generate a first synthetic dataset (402a) from the first medical dataset (401a) and the clinical medical dataset using a pre-trained GAN model. The computing device (100) can determine the number of clinical medical datasets to be generated based on a predetermined first reflection ratio between the clinical prior information and the first medical dataset (401a). For example, the number of clinical medical datasets to be generated can be determined based on the number of the first medical dataset (401a) and the predetermined first reflection ratio. The computing device (100) can generate a second synthetic data set (402b) and a third synthetic data set (402c) from the second medical data set (401b) and the third medical data set (401c) using a pre-trained GAN model in the same manner.

[0112] In one embodiment, the computing device (100) may generate a first synthetic dataset (402a) and a second synthetic dataset (402b) from a first medical dataset (401a) and a second medical dataset (401b) using a SMOTE (Synthetic Minority Oversampling Technique) algorithm. In one embodiment, the SMOTE algorithm may mean an algorithm that calculates K nearest neighbor data belonging to the same class for data sampled from a dataset using a K-NN (K-Nearest Neighbor) algorithm, and generates data corresponding to an arbitrary point located on a line segment connected to each of the sampled data and the K nearest neighbor data as synthetic data.

[0113] For example, the computing device (100) can select an arbitrary point located on a line segment connecting the first medical data included in the first medical data set (401a) and each of the nearest neighboring medical data on a plane or space with a plurality of independent variables as axes. The computing device (100) can generate first synthetic data having the numerical values ​​of each of the plurality of independent variables corresponding to the selected point. At this time, the dependent variable of the first synthetic data can have the same class as the dependent variable class of the first medical data and the nearest neighboring medical data. For example, the dependent variable class can be a class indicating the occurrence or absence of a disease (or side effect). The computing device (100) can generate the first synthetic data set (402a) from the plurality of first synthetic data. The computing device (100) can generate a second synthetic data set (402b) and a third synthetic data set (402c) from the second medical data set (401b) and the third medical data set (401c) using the SMOTE algorithm in the same manner.

[0114] In one embodiment, the computing device (100) may generate a first synthetic dataset (402a) and a second medical dataset (402b) by reflecting the first medical dataset (401a) and the second medical dataset (401b) on clinical prior information using a Bayesian algorithm based on Bayesian inference. For example, the computing device (100) may generate a first synthetic dataset (402a) by reflecting the first medical dataset (401a) on clinical prior information using a Bayesian algorithm based on Bayesian inference. Similarly, the computing device (100) may generate a second medical dataset (402b) by reflecting the second medical dataset (401b) on clinical prior information using a Bayesian algorithm based on Bayesian inference. In one embodiment, a Bayesian algorithm may mean an algorithm that infers a posterior probability distribution by applying medical data to a prior probability distribution.

[0115] For example, the computing device (100) can obtain a prior probability distribution of a predetermined dependent variable according to a predetermined independent variable based on clinical prior information. The computing device (100) can set a probability distribution such as a beta distribution, a gamma distribution, a normal distribution, etc. for the first medical data included in the first medical data set (401a), obtain a likelihood, and apply the obtained likelihood to the prior probability distribution. The computing device (100) can obtain a posterior probability distribution of a predetermined dependent variable according to a predetermined independent variable by applying the first medical data set (401a) to the prior probability distribution. The computing device (100) can generate a first synthetic data set (402a) based on the mean and standard deviation of the obtained posterior probability distribution. The computing device (100) can generate the first synthetic data set (402a) that follows a probability distribution based on a mean and standard deviation that are identical or similar to the mean and standard deviation of the obtained posterior probability distribution. The computing device (100) can generate a second synthetic data set (402b) and a third synthetic data set (402c) from the second medical data set (401b) and the third medical data set (401c) using a Bayesian algorithm based on Bayesian inference in the same manner.

[0116] For example, if the research objective is a study on the blood pressure control effect of furosemide on hypertensive patients, the computing device (100) can set independent variables such as presence or absence of hypertension, whether or not furosemide is taken, the period of taking furosemide, and the dosage, and set systolic blood pressure (SBP) as the dependent variable. The computing device (100) can obtain a prior probability distribution of systolic blood pressure according to furosemide use based on previous clinical studies or clinical statistical data included in the clinical prior information. The computing device (100) can obtain a posterior probability distribution of systolic blood pressure according to furosemide use by applying the first medical dataset and the second medical dataset to the prior probability distribution. The computing device (100) can generate a first synthetic dataset (402a) that follows a probability distribution based on the mean and standard deviation of the obtained posterior probability distribution.

[0117] FIG. 5 is a schematic diagram illustrating a process for calculating weights of target medical institutions according to one embodiment of the present disclosure.

[0118] In one embodiment, the computing device (100) may obtain a third weight (504) corresponding to a preset independent variable to be applied to the analysis result for the target medical data set (503) based on the clinical prior information (501), the integrated synthetic data set (502), and the target medical data set (503). In one embodiment, when the clinical prior information (501) is reflected in the generation of the integrated synthetic data set (502), the computing device (100) may obtain a third weight (504) corresponding to a preset independent variable to be applied to the analysis result for the target medical data set (503) based on the integrated synthetic data set (502) and the target medical data set (503) of the target medical institution.

[0119] In one embodiment, the integrated synthetic dataset (502) may include a first weight corresponding to a preset independent variable. The computing device (100) may obtain a second weight corresponding to the preset independent variable from the target medical dataset (503). Based on the first weight and the second weight, the computing device (100) may obtain a third weight (504) corresponding to the preset independent variable. Based on the preset independent variable, the preset dependent variable, and the third weight (504), the computing device (100) may obtain an analysis result for the target medical dataset (503) from the integrated synthetic dataset (502).

[0120] In one embodiment, the computing device (100) may obtain the third weight (504) based on the statistics, range and units of the preset independent variables (and the preset dependent variables) in the clinical prior information (501), the number of data in the target medical data set (503), the standard deviation of the probability distribution, the numerical values ​​of the preset independent variables and the preset dependent variables in the target medical data set (503), the deviation and the likelihood ratio compared to the clinical prior information (501), etc. For example, the computing device (100) may determine the third weight (504) to be large as the likelihood ratio of the target medical data set (503) compared to the clinical prior information (501) is large, the number of data is large, and the variance is small. That is, the clinical prior information may act as a global parameter for obtaining the third weight (504), and the target medical data set may act as a local parameter for obtaining the third weight (504).

[0121] In one embodiment, the computing device (100) may determine the weighting degree of each of the first weight and the second weight based on the number of first medical data of the first medical data set, the number of second medical data of the second medical data set, and the number of target medical data of the target medical data set (503). For example, the weighting degree of each of the first weight and the second weight may mean a reflection ratio to be reflected in the third weight (504). The computing device (100) may determine the weighting degree of each of the first weight and the second weight based on the number of synthetic data of the integrated synthetic data set (502) and the number of target medical data of the target medical data set (503) so that the probability distribution of the data set with a larger number of data is more reflected in the third weight (504). In this case, the number of synthetic data of the integrated synthetic data set (502) may be equal to the sum of the number of first medical data and the number of second medical data. The computing device (100) can obtain the third weight (504) based on the weighting degree of each of the first weight and the second weight.

[0122] In one embodiment, the computing device (100) may obtain clinical prior information (501) including a clinical medical dataset in which clinical values ​​are recorded. The computing device (100) may determine the weighting degree of each of the first weight and the second weight based on the number of clinical medical data of the clinical medical dataset, the number of first medical data of the first medical dataset, the number of second medical data of the second medical dataset, and the number of target data of the target medical dataset. The computing device (100) may compare the number of synthetic data of the integrated synthetic dataset (502) and the number of target medical data of the target medical dataset (503), and determine the weighting degree of each of the first weight and the second weight so that the probability distribution of the dataset with a larger number of data is more reflected. In this case, the number of synthetic data of the integrated synthetic dataset (502) may be equal to the sum of the number of clinical medical data, the number of first medical data, and the number of second medical data. The computing device (100) can obtain the third weight (504) based on the weighting degree of each of the first weight and the second weight.

[0123] In one embodiment, the computing device (100) can generate a target synthetic dataset corresponding to a target medical dataset (503) from the integrated synthetic dataset (502) based on a preset independent variable, a preset dependent variable, and a third weight (504). The probability distribution of the preset dependent variable in the target synthetic dataset can be determined based on the third weight (504). The computing device (100) can obtain analysis results for the target medical dataset (503) from the target synthetic dataset.

[0124] In one embodiment, the computing device (100) can determine, from the first medical dataset and the second medical dataset, a first independent variable that is correlated with the dependent variable and is not included in the preset independent variables. The computing device (100) can generate a target synthetic dataset corresponding to the target medical dataset from the integrated synthetic dataset (502) based on the preset independent variable, the first independent variable, the preset dependent variable, and the third weight (504). The computing device (100) can obtain an analysis result for the target medical dataset (503) from the target synthetic dataset. Through this, a technical effect of being able to generate a synthetic dataset including an independent variable that has not been considered by a research institute or researcher among a plurality of variables that are correlated with the preset dependent variable can be achieved.

[0125] In one embodiment, the analysis results for the target medical data set (503) may include information regarding a predetermined independent variable, a predetermined dependent variable, a third weight (504), and a correlation between the predetermined independent variable and the predetermined dependent variable. For example, the correlation between the independent variable and the dependent variable included in the analysis results may be obtained based on the third weight (504).

[0126] According to one embodiment of the present disclosure, a computer-readable medium storing a data structure is disclosed. A data structure may refer to the organization, management, and storage of data that enables efficient access and modification of the data. A data structure may refer to the organization of data to solve a specific problem (e.g., data retrieval, data storage, or data modification in the shortest possible time). A data structure may also be defined as a physical or logical relationship between data elements designed to support a specific data processing function. A logical relationship between data elements may include a connection relationship between user-defined data elements. A physical relationship between data elements may include an actual relationship between data elements physically stored in a computer-readable storage medium (e.g., a persistent storage device). Specifically, a data structure may include a set of data, relationships between data, and functions or commands applicable to the data. An effectively designed data structure enables a computing device (100) to perform operations while minimizing the use of its resources. Specifically, an effectively designed data structure enables the computing device (100) to increase the efficiency of operations, reading, insertion, deletion, comparison, exchange, and searching.

[0127] Data structures can be categorized as linear or nonlinear, depending on their form. A linear data structure can be a structure in which only one data item is linked to the next. Linear data structures can include lists, stacks, queues, and deques. A list can refer to a series of data sets with an internal order. Lists can also include linked lists. A linked list is a data structure in which data is linked in a single line, each item having a pointer. In a linked list, a pointer can contain information about the next or previous item. Linked lists can be expressed as singly linked lists, doubly linked lists, or circular linked lists, depending on their form. A stack can be a data listing structure with limited data access. A stack can be a linear data structure in which data operations (e.g., insertion or deletion) can only be performed at one end of the data structure. Data stored in a stack can be a Last-in-First-out (LIFO) data structure. A queue is a data structure with limited access to data. Unlike a stack, it can be a first-in, first-out (FIFO) data structure, with later data being retrieved later. A deck can be a data structure that can process data at both ends.

[0128] A nonlinear data structure can be a structure in which multiple pieces of data are connected behind a single piece of data. Nonlinear data structures can include graph data structures. A graph data structure can be defined by vertices and edges, and an edge can include a line connecting two different vertices. Graph data structures can include tree data structures. A tree data structure can be a data structure in which there is only one path connecting two different vertices among multiple vertices included in the tree. In other words, it can be a data structure that does not form a loop in a graph data structure.

[0129] Throughout this specification, the terms computational model, neural network, network function, and neural network may be used interchangeably. Hereinafter, they are collectively referred to as neural networks. The data structure may include a neural network. And the data structure including the neural network may be stored on a computer-readable medium. The data structure including the neural network may also include preprocessed data for processing by the neural network, data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with each node or layer of the neural network, loss functions for learning the neural network, etc. The data structure including the neural network may include any of the components disclosed above. That is, the data structure including the neural network may be configured to include all or any combination of preprocessed data for processing by the neural network, data input to the neural network, weights of the neural network, hyperparameters of the neural network, data obtained from the neural network, activation functions associated with each node or layer of the neural network, loss functions for learning the neural network, etc. In addition to the aforementioned configurations, a data structure including a neural network may include any other information that determines the characteristics of the neural network. Furthermore, the data structure may include any form of data used or generated in the computational process of the neural network, and is not limited to the aforementioned. The computer-readable medium may include a computer-readable recording medium and / or a computer-readable transmission medium. A neural network may be composed of a set of interconnected computational units, which may generally be referred to as nodes. These nodes may also be referred to as neurons. A neural network is composed of at least one node.

[0130] The data structure may include data input to a neural network. The data structure including the data input to the neural network may be stored on a computer-readable medium. The data input to the neural network may include training data input during the neural network training process and / or input data input to the neural network after training has been completed. The data input to the neural network may include data that has undergone preprocessing and / or data that is the target of preprocessing. Preprocessing may include a data processing process for inputting data to the neural network. Accordingly, the data structure may include data that is the target of preprocessing and data generated by the preprocessing. The above-described data structure is merely an example, and the present disclosure is not limited thereto.

[0131] The data structure may include weights of the neural network. (In this specification, the terms "weight" and "parameter" may be used interchangeably.) The data structure including the weights of the neural network may be stored in a computer-readable medium. The neural network may include a plurality of weights. The weights may be variable and may be varied by a user or an algorithm so that the neural network can perform a desired function. For example, when one or more input nodes are interconnected to one output node by respective links, the output node may determine a data value output from the output node based on values ​​input to the input nodes connected to the output node and weights set for links corresponding to each input node. The above-described data structure is merely an example, and the present disclosure is not limited thereto.

[0132] By way of example and not limitation, the weights may include weights that vary during the neural network training process and / or weights that have completed neural network training. The weights that vary during the neural network training process may include weights at the start of the training cycle and / or weights that vary during the training cycle. The weights that have completed neural network training may include weights that have completed the training cycle. Accordingly, a data structure including the weights of a neural network may include a data structure including weights that vary during the neural network training process and / or weights that have completed neural network training. Therefore, the above-described weights and / or combinations of each weight are included in the data structure including the weights of a neural network. The above-described data structures are merely examples and the present disclosure is not limited thereto.

[0133] A data structure including the weights of a neural network can be stored in a computer-readable storage medium (e.g., memory, hard disk) after going through a serialization process. Serialization may be a process of converting a data structure into a form that can be stored in the same or another computing device (100) and reconstructed and used later. The computing device (100) can serialize the data structure to transmit and receive data over a network. The data structure including the weights of a serialized neural network can be reconstructed in the same computing device (100) or another computing device (100) through deserialization. The data structure including the weights of a neural network is not limited to serialization. Furthermore, the data structure including the weights of a neural network may include a data structure (e.g., a B-Tree, a Trie, an m-way search tree, an AVL tree, a Red-Black Tree in a nonlinear data structure) for increasing computational efficiency while minimizing the use of resources of the computing device (100). The foregoing is merely an example, and the present disclosure is not limited thereto.

[0134] The data structure may include hyperparameters of a neural network. Furthermore, the data structure including the hyperparameters of the neural network may be stored on a computer-readable medium. The hyperparameters may be variables that can be varied by the user. The hyperparameters may include, for example, a learning rate, a cost function, the number of learning cycle repetitions, weight initialization (e.g., setting a range of weight values ​​to be subject to weight initialization), and the number of hidden units (e.g., the number of hidden layers, the number of nodes in the hidden layer). The above-described data structure is merely an example, and the present disclosure is not limited thereto.

[0135] FIG. 6 is a simplified, general schematic diagram of an exemplary computing environment in which embodiments of the present disclosure may be implemented.

[0136] Although the present disclosure has been described above as being generally implemented by a computing device (100), those skilled in the art will appreciate that the present disclosure may be implemented in combination with computer-executable instructions and / or other program modules that may be executed on one or more computers and / or as a combination of hardware and software.

[0137] Generally, program modules include routines, programs, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Furthermore, those skilled in the art will appreciate that the methods of the present disclosure can be implemented with other computer system configurations, including single-processor or multiprocessor computer systems, minicomputers, mainframe computers, as well as personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, and the like, each of which may be operatively connected to one or more associated devices.

[0138] The described embodiments of the present disclosure can also be practiced in distributed computing environments, where certain tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.

[0139] Computers typically include a variety of computer-readable media. Computer-readable media can be any media that can be accessed by a computer, and includes both volatile and nonvolatile media, transitory and non-transitory media, removable and non-removable media. By way of example, and not limitation, computer-readable media can include computer-readable storage media and computer-readable transmission media. Computer-readable storage media includes both volatile and nonvolatile media, transitory and non-transitory media, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital video disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be accessed by a computer and used to store the desired information.

[0140] Computer-readable transmission media typically includes any information delivery media that embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism. The term modulated data signal means a signal that has one or more of its characteristics set or changed so as to encode information in the signal. By way of example, and not limitation, computer-readable transmission media includes wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, RF, infrared, or other wireless media. Combinations of any of the above are also intended to be included within the scope of computer-readable transmission media.

[0141] An exemplary environment (1100) implementing various aspects of the present disclosure is illustrated, including a computer (1102) comprising a processing unit (1104), system memory (1106), and a system bus (1108). The system bus (1108) connects system components, including but not limited to the system memory (1106), to the processing unit (1104). The processing unit (1104) may be any of a variety of commercially available processors. Dual processors and other multiprocessor architectures may also be utilized as the processing unit (1104).

[0142] The system bus (1108) may be any of several types of bus structures that may be additionally interconnected to a memory bus, a peripheral bus, and a local bus using any of a variety of commercial bus architectures. The system memory (1106) includes read-only memory (ROM) (1110) and random access memory (RAM) (1112). A basic input / output system (BIOS) is stored in non-volatile memory (1110), such as ROM, EPROM, or EEPROM, and includes basic routines that help transfer information between components within the computer (1102), such as during start-up. The RAM (1112) may also include high-speed RAM, such as static RAM, for caching data.

[0143] The computer (1102) also includes an internal hard disk drive (HDD) (1114) (e.g., EIDE, SATA) - which may also be configured for external use within a suitable chassis (not shown), a magnetic floppy disk drive (FDD) (1116) (e.g., for reading from or writing to a removable diskette (1118)), and an optical disk drive (1120) (e.g., for reading from or writing to a CD-ROM disk (1122) or other high-capacity optical media such as a DVD). The hard disk drive (1114), the magnetic disk drive (1116), and the optical disk drive (1120) may be connected to the system bus (1108) by a hard disk drive interface (1124), a magnetic disk drive interface (1126), and an optical drive interface (1128), respectively. The interface (1124) for implementing an external drive includes at least one or both of Universal Serial Bus (USB) and IEEE 1394 interface technologies.

[0144] These drives and their associated computer-readable media provide non-volatile storage of data, data structures, computer-executable instructions, and the like. In the case of the computer (1102), the drives and media correspond to storing any data in a suitable digital format. While the description of computer-readable media above refers to HDDs, removable magnetic disks, and removable optical media such as CDs or DVDs, those of ordinary skill in the art will appreciate that other types of computer-readable media, such as zip drives, magnetic cassettes, flash memory cards, cartridges, and the like, may also be used in the exemplary operating environment, and that any such media may contain computer-executable instructions for performing the methods of the present disclosure.

[0145] A number of program modules, including an operating system (1130), one or more application programs (1132), other program modules (1134), and program data (1136), may be stored in the drive and RAM (1112). All or portions of the operating system, applications, modules, and / or data may also be cached in RAM (1112). It will be appreciated that the present disclosure may be implemented in various commercially available operating systems or combinations of operating systems.

[0146] A user may enter commands and information into the computer (1102) via one or more wired / wireless input devices, such as a keyboard (1138) and a pointing device such as a mouse (1140). Other input devices (not shown) may include a microphone, an IR remote control, a joystick, a game pad, a stylus pen, a touch screen, and the like. These and other input devices are often connected to the processing unit (1104) via an input device interface (1142) that is connected to the system bus (1108), but may be connected by other interfaces such as a parallel port, an IEEE 1394 serial port, a game port, a USB port, an IR interface, and the like.

[0147] A monitor (1144) or other type of display device is also connected to the system bus (1108) via an interface, such as a video adapter (1146). In addition to the monitor (1144), the computer typically includes other peripheral output devices (not shown), such as speakers, a printer, and so on.

[0148] The computer (1102) may operate in a networked environment using logical connections to one or more remote computers, such as remote computer(s) (1148), via wired and / or wireless communications. The remote computer(s) (1148) may be a workstation, a computing device computer, a router, a personal computer, a portable computer, a microprocessor-based entertainment device, a peer device, or other conventional network node, and generally include many or all of the components described for the computer (1102), although for simplicity, only the memory storage device (1150) is shown. The logical connections shown include wired / wireless connections to a local area network (LAN) (1152) and / or a larger network, such as a wide area network (WAN) (1154). Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks, such as intranets, all of which may be connected to a worldwide computer network, such as the Internet.

[0149] When used in a LAN networking environment, the computer (1102) is connected to a local network (1152) via a wired and / or wireless communication network interface or adapter (1156). The adapter (1156) may facilitate wired or wireless communications to the LAN (1152), which may also include a wireless access point installed therein for communicating with the wireless adapter (1156). When used in a WAN networking environment, the computer (1102) may include a modem (1158), be connected to a communications computing device on the WAN (1154), or have other means of establishing communications over the WAN (1154), such as via the Internet. The modem (1158), which may be internal or external and wired or wireless, is connected to the system bus (1108) via a serial port interface (1142). In a networked environment, program modules or portions thereof described for the computer (1102) may be stored in a remote memory / storage device (1150). It will be appreciated that the network connections depicted are exemplary and other means of establishing a communications link between the computers may be used.

[0150] The computer (1102) operates to communicate with any wireless device or object that is arranged and operates via wireless communication, such as a printer, a scanner, a desktop and / or portable computer, a portable data assistant (PDA), a communication satellite, any equipment or location associated with a radio-detectable tag, and a telephone. This includes at least Wi-Fi and Bluetooth wireless technologies. Accordingly, the communication may be a predefined structure as in a conventional network, or may simply be an ad hoc communication between at least two devices.

[0151] Wi-Fi (Wireless Fidelity) enables connections to the Internet and other devices without wires. Wi-Fi is a wireless technology that allows devices, such as computers, to send and receive data anywhere within the coverage area of ​​a base station, both indoors and outdoors, similar to cell phones. Wi-Fi networks use wireless technologies called IEEE 802.11 (a, b, g, etc.) to provide secure, reliable, and high-speed wireless connections. Wi-Fi can be used to connect computers to each other, to the Internet, and to wired networks (using IEEE 802.3 or Ethernet). Wi-Fi networks can operate in the unlicensed 2.4 and 5 GHz radio bands, at data rates of, for example, 11 Mbps (802.11a) or 54 Mbps (802.11b), or in products that include both bands (dual-band).

[0152] Those skilled in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips referenced in the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0153] Those skilled in the art will appreciate that the various illustrative logical blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, various forms of programs or design code (referred to herein, for convenience, as software), or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0154] The various embodiments presented herein can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term article of manufacture includes a computer program, carrier, or media accessible from any computer-readable storage device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Furthermore, various storage media presented herein include one or more devices and / or other machine-readable media for storing information.

[0155] It should be understood that the specific order or hierarchy of steps in the presented processes is merely an example of exemplary approaches. It should be understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of the present disclosure based on design priorities. The appended method claims provide elements of various steps in a sample order, but are not intended to be limited to the specific order or hierarchy presented.

[0156] The description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments disclosed herein, but is to be construed in the broadest scope consistent with the principles and novel features disclosed herein.

[0157] As described above, the relevant contents have been described in the best form for carrying out the invention.

[0158] It can be used in devices and systems that analyze medical data.

Claims

1. A method for analyzing medical data of a medical institution, performed by a computing device, A step of obtaining a first medical dataset of a first medical institution and a second medical dataset of a second medical institution from a first medical institution and a second medical institution; A step of generating a total synthetic data set from the first medical data set and the second medical data set; and A step of obtaining analysis results for the target medical dataset from the integrated synthetic dataset based on the target medical dataset of the target medical institution; Including, The above first medical dataset, the above second medical dataset and the above integrated synthetic dataset include preset independent variables and preset dependent variables, The above integrated synthetic data set includes a first weight corresponding to the above preset independent variable, and The steps for obtaining analysis results for the above target medical data set are: A step of obtaining a second weight corresponding to the preset independent variable from the target medical data set; A step of obtaining a third weight corresponding to the preset independent variable based on the first weight and the second weight; and A step of obtaining the analysis result from the integrated synthetic data set based on the above-described independent variable, the above-described dependent variable, and the third weight; Including, method.

2. In paragraph 1, The steps for generating the above integrated synthetic dataset are: A step of obtaining clinical prior information including information on the correlation between the above-described independent variable and the above-described dependent variable; and A step of generating the integrated synthetic dataset from the above clinical prior information, the first medical dataset, and the second medical dataset; Including, method.

3. In paragraph 1, The steps for generating the above integrated synthetic dataset are: A step of generating a first synthetic data set corresponding to the first medical data set and a second synthetic data set corresponding to the second medical data set from the first medical data set and the second medical data set; and A step of generating the integrated synthetic dataset by integrating the first synthetic dataset and the second synthetic dataset; Including, method.

4. In paragraph 3, The step of generating the first synthetic data set and the second synthetic data set comprises: A step of generating the first synthetic dataset and the second synthetic dataset from the first medical dataset and the second medical dataset using a pre-trained Generative Adversarial Networks (GAN) model; Including, method.

5. In paragraph 3, The step of generating the first synthetic data set and the second synthetic data set comprises: A step of generating the first synthetic dataset and the second synthetic dataset from the first medical dataset and the second medical dataset using the SMOTE (Synthetic Minority Oversampling Technique) algorithm; Including, method.

6. In paragraph 3, The step of generating the first synthetic data set and the second synthetic data set comprises: A step of obtaining clinical prior information including information on the correlation between the above-described independent variable and the above-described dependent variable; and A step of generating the first synthetic dataset and the second synthetic dataset by reflecting the first medical dataset and the second medical dataset into the clinical prior information using a Bayesian algorithm based on Bayesian inference; Including, method.

7. In paragraph 1, A step of determining a first independent variable from the first medical data set and the second medical data set that is correlated with the dependent variable and is not included in the preset independent variables; Including more, and The above analysis results include information on the correlation between the first independent variable and the predetermined dependent variable. method.

8. In paragraph 1, The step of obtaining the third weight corresponding to the above-described independent variable is: A step of determining the weighting degree of each of the first weight and the second weight based on the number of first medical data of the first medical data set, the number of second medical data of the second medical data set, and the number of target medical data of the target medical data set; and A step of obtaining the third weight based on the weighting degree of each of the first weight and the second weight; Including, method.

9. In paragraph 8, The step of determining the weighting degree of each of the first weight and the second weight is: A step of obtaining clinical pre-information including clinical statistics in which clinical values ​​are recorded; and A step of determining the weighting degree of each of the first weight and the second weight based on the number of clinical medical data of the clinical statistics, the number of first medical data of the first medical data set, the number of second medical data of the second medical data set, and the number of target medical data of the target medical data set; Including, method.

10. In paragraph 1, The steps for obtaining analysis results for the above target medical data set are: A step of generating a target synthetic data set corresponding to the target medical data set from the integrated synthetic data set based on the above-described independent variable, the above-described dependent variable, and the third weight; and A step of obtaining the analysis results from the target synthetic data set; Including, method.

11. In paragraph 1, The steps for obtaining analysis results for the above target medical data set are: A step of determining a first independent variable from the first medical data set and the second medical data set that is correlated with the dependent variable and is not included in the preset independent variables; A step of generating a target synthetic data set corresponding to the target medical data set from the integrated synthetic data set based on the above-described independent variable, the first independent variable, the above-described dependent variable, and the third weight; and A step of obtaining the analysis results from the target synthetic data set; Including, method.

12. In paragraph 1, The analysis results for the above target medical data set are: Including information about the above-described independent variable, the above-described dependent variable, the third weight, and the correlation between the above-described independent variable and the above-described dependent variable. method.

13. In paragraph 1, The first medical dataset and the second medical dataset are medical datasets converted from the first raw medical dataset and the second raw medical dataset into a data structure of a common data model (CDM), wherein the first medical dataset corresponds to the first raw medical dataset, and the second medical dataset corresponds to the second raw medical dataset. method.

14. A method for analyzing medical data of a medical institution, performed by a computing device, A step of obtaining a first medical dataset of a first medical institution and a second medical dataset of a second medical institution from a first medical institution and a second medical institution; A step of generating a total synthetic data set from the first medical data set and the second medical data set; and A step of obtaining analysis results for the target medical dataset from the integrated synthetic dataset based on the target medical dataset of the target medical institution; Including, The above first medical dataset, the above second medical dataset and the above integrated synthetic dataset include preset independent variables and preset dependent variables, The steps for obtaining analysis results for the above target medical data set are: A step of generating a first integrated synthetic dataset from the first medical dataset and the second medical dataset; A step of generating a second integrated synthetic data set from the first medical data set and the second medical data set; and A step of obtaining analysis results for the target medical dataset from the first integrated synthetic dataset and the second integrated synthetic dataset based on the target medical dataset; Including, and The above first integrated synthetic data set includes the first-first weight corresponding to the above preset independent variable, The above second integrated synthetic data set includes first and second weights corresponding to the above preset independent variables, The above 1-1 weight and the above 1-2 weight are different from each other. method.

15. In paragraph 14, The predetermined dependent variables included in each of the first integrated synthetic data set and the second integrated synthetic data set have probability distributions based on different means and standard deviations. method.

16. In paragraph 14, The step of obtaining the analysis results for the above target medical data set is: A step of obtaining a first analysis result for the target medical dataset from the first integrated synthetic dataset; A step of obtaining a second analysis result for the target medical data set from the second integrated synthetic data set; and A step of obtaining an analysis result for the target medical data set based on the first analysis result and the second analysis result; Including, method.

17. In paragraph 1, A step of determining the required number of data of a medical data set for generating the integrated synthetic data set based on at least one of a preset effect size, a preset power, and a preset significance level; Including more, and Each of the first medical data set and the second medical data set includes medical data greater than the required number of data. method.

18. A computing device for analyzing medical data of a medical institution, at least one processor; and memory; Including, At least one processor of the above, Obtaining a first medical dataset of the first medical institution and a second medical dataset of the second medical institution from the first medical institution and the second medical institution, Generating a total synthetic data set from the first medical data set and the second medical data set, and Based on the target medical dataset of the target medical institution, analysis results for the target medical dataset are obtained from the integrated synthetic dataset. The above first medical dataset, the above second medical dataset and the above integrated synthetic dataset include preset independent variables and preset dependent variables, The above integrated synthetic data set includes a first weight corresponding to the above preset independent variable, and Obtaining analysis results for the above target medical data set is as follows: Obtain the second weight corresponding to the preset independent variable from the above target medical data set, Based on the first weight and the second weight, a third weight corresponding to the preset independent variable is obtained, and Obtaining the analysis result from the integrated synthetic data set based on the above-described independent variable, the above-described dependent variable and the third weight. Computing device.

19. A computer program stored in a computer-readable storage medium, wherein when the computer program is executed by one or more processors, the computer program causes the one or more processors to perform operations for analyzing medical data of a medical institution, the operations comprising: An operation of acquiring a first medical dataset of a first medical institution and a second medical dataset of a second medical institution from a first medical institution and a second medical institution; An operation of generating a total synthetic data set from the first medical data set and the second medical data set; and An operation of obtaining analysis results for a target medical dataset from the integrated synthetic dataset based on a target medical dataset of a target medical institution; Including, The above first medical dataset, the above second medical dataset and the above integrated synthetic dataset include preset independent variables and preset dependent variables, The above integrated synthetic data set includes a first weight corresponding to the above preset independent variable, and The operation to obtain the analysis results for the above target medical data set is: An operation of obtaining a second weight corresponding to the preset independent variable from the target medical data set; An operation of obtaining a third weight corresponding to the preset independent variable based on the first weight and the second weight; and An operation of obtaining the analysis result from the integrated synthetic data set based on the above-described independent variable, the above-described dependent variable, and the third weight; Including, A computer program stored on a computer-readable storage medium.

Citation Information

Patent Citations

  • Selecting a training dataset to train the model

    JP2023533587A

  • Method, apparatus and computer program for training artificial intelligence model that judges danger in work site

    KR1020220092361A

  • Electronic device for providing exercise route information and method for controlling thereof

    KR1020240047871A

  • Method to analyze medical data

    KR102230660B1

  • Method and apparatus for analyzing medical data

    KR102661357B1