Phenotype prediction system and method

WO2026182321A1PCT designated stage Publication Date: 2026-09-03LG MANAGEMENT DEV INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/013746
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-06-12
Filing Date
2025-09-05
Publication Date
2026-09-03

Smart Images

  • Figure KR2025013746_03092026_PF_FP_ABST
    Figure KR2025013746_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a phenotype prediction system comprising: a memory storing one or more instructions; and at least one processor that executes the one or more instructions stored in the memory, wherein an operation performed by the one or more instructions comprises a step of learning phenotype prediction from omics data, and the step of learning the phenotype prediction comprises the steps of: acquiring at least one data including the omics data; generating an omics feature on the basis of the omics data; generating, by a reconstruction model, a latent vector on the basis of the omics feature; and predicting, by a phenotype prediction model, a phenotype on the basis of the latent vector.
Need to check novelty before this filing date? Find Prior Art

Description

Phenotype prediction system and method

[0001] The present disclosure relates to a phenotype prediction system and method. Specifically, the present disclosure relates to an improved system and method capable of efficiently predicting phenotypes from omics data.

[0002] Omics data generated by analyzing the various biomolecules possessed by living organisms contains diverse information, so analyzing this data can identify the disease state or risk level of an organism, as well as the biomolecules that contribute to the disease.

[0003] Due to recent advancements in life science technology, the time and cost required to produce omics data have decreased exponentially, leading to a massive increase in the volume of available omics data. Consequently, there is a growing number of attempts to understand phenomena occurring in living organisms by training artificial intelligence models with omics data.

[0004] However, while omics data contains a vast number of variables, most variables have little influence on specific phenotypes, making it difficult to efficiently predict phenotypes from omics data. As such, omics data suffers from high dimensionality and sparsity, and the complexity of the data leads to various problems such as overfitting, reduced model interpretability, and computational difficulties.

[0005] The present disclosure aims to provide a system and method capable of predicting phenotypes from omics data, and to provide a phenotype prediction system and method that performs predictions efficiently from vast amounts of omics data while offering excellent performance.

[0006] The present disclosure provides a phenotype prediction system as a means for solving the above problem, wherein the phenotype prediction model includes the step of predicting a phenotype from the latent vector.

[0007] The present disclosure comprises a memory for storing one or more instructions; and

[0008] At least one that executes the one or more instructions stored in the memory.

[0009] Includes processors,

[0010] The operation performed by the above one or more instructions is

[0011] The method includes a step of learning phenotype predictions from omics data, wherein the step of learning phenotype predictions is

[0012] A step of acquiring at least one piece of data including omics data,

[0013] A step of generating omics features from the above omics data,

[0014] A step in which a reconstruction model generates a latent vector from the above omics features, and

[0015] The present invention provides a phenotype prediction system in which a phenotype prediction model includes the step of predicting a phenotype from the above-mentioned latent vector.

[0016] The present disclosure describes a reconstruction model comprising a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector.

[0017] The step of learning the above phenotype prediction is,

[0018] The step of the first reconstruction sub-neural network generating the latent vector from the omics features,

[0019] The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector,

[0020] A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and

[0021] A phenotype prediction system is provided that includes a step in which the reconstruction model learns based on the above evaluation.

[0022] The present disclosure describes the step of learning the phenotype prediction above.

[0023] (1) The above reconstruction model and the above phenotype prediction model each learn, or

[0024] (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or

[0025] (3) After the above reconstruction model and the above phenotype prediction model are trained alternately, the above phenotype prediction model is further trained. A phenotype prediction system is provided.

[0026] The present disclosure provides a phenotype prediction system in which the phenotype prediction model additionally receives one or more individual features as input to predict a phenotype.

[0027] The present disclosure provides a phenotype prediction system in which the phenotype is associated with a disease.

[0028] The present disclosure provides a phenotype prediction system in which the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

[0029] The present disclosure provides a phenotype prediction system that further comprises the step of obtaining disease-associated biomarker information or the step of obtaining disease-associated phenotype information of an individual.

[0030] The present disclosure is a phenotype prediction method performed by at least one processor, wherein

[0031] The above method includes a step of learning phenotype predictions from omics data, and the step of learning phenotype predictions is

[0032] A step of acquiring at least one piece of data including omics data,

[0033] A step of generating omics features from the above omics data,

[0034] A step in which a reconstruction model generates a latent vector from the above omics features, and

[0035] The present invention provides a phenotype prediction method in which a phenotype prediction model includes the step of predicting a phenotype from the above-mentioned latent vector.

[0036] The present disclosure describes a reconstruction model comprising a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector.

[0037] The step of learning the above phenotype prediction is,

[0038] The step of the first reconstruction sub-neural network generating the latent vector from the omics features,

[0039] The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector,

[0040] A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and

[0041] A phenotype prediction method is provided that includes a step in which the reconstruction model learns based on the above evaluation.

[0042] The present disclosure describes a step of learning the phenotype prediction.

[0043] (1) The above reconstruction model and the above phenotype prediction model each learn, or

[0044] (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or

[0045] (3) A method for predicting phenotypes is provided in which the reconstruction model and the phenotype prediction model learn alternately, and then the phenotype prediction model learns additionally.

[0046] The present disclosure provides a phenotype prediction method in which the phenotype prediction model additionally receives one or more individual features as input to predict a phenotype.

[0047] The present disclosure provides a method for predicting a phenotype in which the phenotype is associated with a disease.

[0048] The present disclosure provides a phenotype prediction method in which the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

[0049] The present disclosure provides a phenotype prediction method that further comprises the step of obtaining disease-associated biomarker information or obtaining disease-associated phenotype information of an individual.

[0050] The present disclosure provides a program stored on a computer-readable recording medium to execute the above method on a computer.

[0051] The present disclosure is a computerized method for predicting a phenotype,

[0052] The above method includes a step of learning phenotype predictions from omics data, and the step of learning phenotype predictions is

[0053] A step of acquiring at least one piece of data including omics data,

[0054] A step of generating omics features from the above omics data,

[0055] A step in which a reconstruction model generates a latent vector from the above omics features, and

[0056] The present invention provides a phenotype prediction method in which a phenotype prediction model includes the step of predicting a phenotype from the above-mentioned latent vector.

[0057] The present disclosure describes a reconstruction model comprising a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector.

[0058] The step of learning the above phenotype prediction is,

[0059] The step of the first reconstruction sub-neural network generating the latent vector from the omics features,

[0060] The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector,

[0061] A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and

[0062] A phenotype prediction method is provided that includes a step in which the reconstruction model learns based on the above evaluation.

[0063] The present disclosure describes a step of learning the phenotype prediction.

[0064] (1) The above reconstruction model and the above phenotype prediction model each learn, or

[0065] (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or

[0066] (3) A method for predicting phenotypes is provided in which the reconstruction model and the phenotype prediction model learn alternately, and then the phenotype prediction model learns additionally.

[0067] The present disclosure provides a phenotype prediction method in which the phenotype prediction model additionally receives one or more individual features as input to predict a phenotype.

[0068] The present disclosure provides a method for predicting a phenotype in which the phenotype is associated with a disease.

[0069] The present disclosure provides a phenotype prediction method in which the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

[0070] The present disclosure provides a phenotype prediction method that further comprises the step of obtaining disease-associated biomarker information or obtaining disease-associated phenotype information of an individual.

[0071] The phenotype prediction system and method of the present disclosure have superior phenotype prediction performance compared to conventional methods and can perform learning more efficiently compared to conventional methods.

[0072] By using the phenotype prediction system and method of the present disclosure, more reliable prediction results can be obtained efficiently. The phenotype prediction system and method of the present disclosure can be utilized to obtain information for diagnosing the disease state of an individual, to obtain information for evaluating the disease risk of an individual, and to discover targets for new drug development.

[0073] FIG. 1 is a block diagram of an apparatus capable of implementing a phenotype prediction system and method according to one embodiment of the present disclosure.

[0074] FIG. 2 is a flowchart of a phenotype prediction system and method according to one embodiment of the present disclosure.

[0075] FIG. 3 is an exemplary diagram illustrating a method for predicting a phenotype according to a phenotype prediction system according to one embodiment of the present disclosure.

[0076] FIG. 4 is an exemplary diagram illustrating a method in which a reconstruction model according to one embodiment of the present disclosure performs learning.

[0077] FIG. 5 is an exemplary diagram illustrating how a phenotype prediction model according to one embodiment of the present disclosure performs learning.

[0078] FIG. 6 is an exemplary diagram illustrating a method for obtaining disease-associated phenotype information of an individual using a phenotype prediction model according to one embodiment of the present disclosure.

[0079] FIG. 7 is an exemplary diagram illustrating a method for obtaining disease-associated biomarker information using a phenotypic prediction model according to one embodiment of the present disclosure.

[0080] Figure 8 is a diagram showing the relationship between omics data, object data, and phenotype.

[0081] To clarify the technical concept of the present disclosure, embodiments of the present disclosure will be described in detail with reference to the attached drawings. In describing the present disclosure, detailed descriptions of related known functions or components will be omitted if it is determined that such detailed descriptions would unnecessarily obscure the essence of the present disclosure. Components having substantially the same function or configuration among the drawings have been assigned the same reference numerals and symbols as much as possible, even if they are shown in different drawings. For convenience of explanation, devices and methods are described together where necessary. Each operation of the present disclosure does not necessarily have to be performed in the order described and may be performed in parallel, selectively, or individually.

[0082] The terms used in the embodiments of this disclosure have been selected to be as widely used and general as possible, taking into account the functions of this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, terms used in this specification should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0083] Throughout this disclosure, singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms such as "comprising" or "having" are intended to specify the existence of features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, throughout this disclosure, when a part is described as "comprising" a certain component, it means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0084] Expressions such as "at least one" modify the entire list of components and do not modify the components of the list individually. For example, "at least one of A, B, and C" and "at least one of A, B, or C" refer to only A, only B, only C, both A and B, both B and C, both A and C, all of A, B, and C, or any combination thereof.

[0085] Throughout the entire disclosure, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other elements interposed between them. Furthermore, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0086] Throughout this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware. Instead, in some situations, the expression “system configured to” may mean that the system is “capable of” in conjunction with other devices or components. For example, the phrase “processor configured to perform A, B, and C” may mean a dedicated processor for performing the said operations (e.g., an embedded processor) or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in memory.

[0087] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0088] Artificial intelligence is a field of computer science and information technology that studies methods to enable computers to perform thinking, learning, and self-development tasks typically accomplished by human intelligence; it refers to the ability of computers to mimic intelligent human behavior. Furthermore, artificial intelligence does not exist in isolation but is closely related, directly or indirectly, to many other areas of computer science. Particularly in the modern era, there is active research being conducted across various fields of information technology to introduce AI elements and utilize them for problem-solving.

[0089] Machine learning is a field of artificial intelligence that enables computers to learn without explicit programming. Specifically, machine learning can be defined as a technology that studies and builds systems and algorithms capable of learning, making predictions, and improving their own performance based on empirical data. Rather than executing strictly defined static program commands, machine learning algorithms adopt an approach of constructing specific models to derive predictions or decisions based on input data. The term 'machine learning' may be used interchangeably with 'machine learning'.

[0090] Many machine learning algorithms have been developed to address how to classify data in machine learning. Representative examples of these algorithms include Decision Trees, Bayesian Networks, Support Vector Machines (SVM), and Artificial Neural Networks (ANN).

[0091] A decision tree is an analytical method that performs classification and prediction by plotting decision rules in a tree structure. A Bayesian network is a model that represents the probabilistic relationships (conditional independence) between multiple variables in a graph structure. Bayesian networks are suitable for data mining through unsupervised learning.

[0092] Support Vector Machines are supervised learning models for pattern recognition and data analysis, primarily used for classification and regression analysis. Artificial neural networks model the operating principles of biological neurons and the relationships between them; they are information processing systems in which multiple neurons, referred to as nodes or processing elements, are connected in a layered structure.

[0093] Artificial neural networks are models used in machine learning, serving as statistical learning algorithms in machine learning and cognitive science that draw inspiration from biological neural networks (particularly the brain within the animal central nervous system). Specifically, an artificial neural network can refer to a model in which artificial neurons (nodes), forming a network through synaptic connections, change the strength of these connections through learning to possess problem-solving capabilities. The term artificial neural network may be used interchangeably with neural network.

[0094] An artificial neural network may include multiple layers, and each layer may include multiple neurons. Additionally, an artificial neural network may include synapses connecting neurons. An artificial neural network can generally be defined by the following three factors: the connection patterns between neurons in different layers, a learning process that updates the weights of the connections, and an activation function that generates an output value from a weighted sum of inputs received from the previous layer.

[0095] Artificial neural networks may include, but are not limited to, network models such as Deep Neural Networks (DNN), Recurrent Neural Networks (RNN), Bidirectional Recurrent Deep Neural Networks (BRDNN), Multilayer Perceptrons (MLP), and Convolutional Neural Networks (CNN). In this specification, the term 'layer' may be used interchangeably with the term 'layer'.

[0096] Artificial neural networks are classified into single-layer neural networks and multi-layer neural networks depending on the number of layers. A typical single-layer neural network consists of an input layer and an output layer. Additionally, a typical multi-layer neural network consists of an input layer, one or more hidden layers, and an output layer.

[0097] The input layer is a layer that receives external data, and the number of neurons in the input layer is equal to the number of input variables. The hidden layer is located between the input layer and the output layer, receives signals from the input layer, extracts features, and transmits them to the output layer. The output layer receives signals from the hidden layer and outputs an output value based on the received signals. Input signals between neurons are multiplied by their respective connection strengths (weights) and then summed; if this sum is greater than the neuron's threshold, the neuron is activated and outputs the value obtained through the activation function.

[0098] Meanwhile, a deep neural network containing multiple hidden layers between the input layer and the output layer can be a representative artificial neural network that implements deep learning, a type of machine learning technique. Meanwhile, the term 'deep learning' may be used interchangeably with the term 'deep learning'.

[0099] The machine learning workflow consists of a series of processes involving collecting data for learning and validation, modeling, and training the model, and may include the processes of collecting training data, checking and exploring data, data preprocessing and cleaning, modeling, and training.

[0100] 1. Collect Training Data

[0101] The phenotype prediction system of the present disclosure acquires at least one omics data (311) and uses it for learning to predict the phenotype. Omics data refers to data used to understand life phenomena by analyzing biomolecules such as genomics, transcriptomics, proteomics, metabolomics, and liposomes.

[0102] The omics data used in the system of the present disclosure is not particularly limited, but as an example, omics data selected from genomic data, transcriptomic data, and combinations thereof may be used. Genomic data that may be used in the system of the present disclosure includes DNA base sequence data, and mutation data of DNA base sequences may also be used. Mutations of DNA base sequences appear in the form of substitutions, insertions, and deletions, and the mutation data of DNA base sequences is not particularly limited as long as it is data containing mutation information of DNA base sequences, but includes mutation information resulting from substitution, insertion, or deletion of DNA base sequences, and may include, for example, one or more selected from single nucleotide polymorphisms (SNPs), indels, recombination, and combinations thereof.

[0103] The phenotype prediction system of the present disclosure may acquire one or more individual data (331) and use them for learning to predict phenotypes. In the present disclosure, individual data refers to data of an individual identical to the individual constituting the omics data acquired for learning, and the type thereof is not particularly limited as long as it is information that can be used for phenotype prediction. For example, it may include information such as the individual's diet data or gender data.

[0104] Phenotype data may be used for training the phenotype prediction system of the present disclosure and may be used to compare the similarity between the phenotype predicted by the phenotype prediction system and the actual phenotype. Here, a phenotype refers to an actual trait that appears according to genetic information as a result of the expression of a genotype, and is a concept that includes not only physical traits but also characteristics such as behavior. The phenotypes that can be predicted in the present disclosure are not particularly limited, but may be disease-related phenotypes. A disease-related phenotype refers to a manifestation observable in an individual carrying a disease and may include various aspects such as symptoms, changes in appearance, behavior, and biochemical characteristics. For example, a decline in memory observed in an individual with dementia is a disease-related phenotype.

[0105] In the present disclosure, different types of training data sets may be used for training, and each training data may further include one or more experimental results used as feature labels. At least a portion of the training data set may be used to train a learning model, and another portion may be used to evaluate the performance of the learned learning model.

[0106] 2. Inspection and Exploration of Training Data

[0107] Once training data for training a learning model is collected, the collected training data can be examined and explored regarding its structure, noise data, and data cleaning methods for machine learning applications.

[0108] This stage of data inspection and exploration is called Exploratory Data Analysis (EDA), which can be described as the process of observing and understanding collected data from various angles. Before training the data, independent variables, dependent variables, variable types, and data types are examined using visualizations such as graphs and statistical tests, allowing the characteristics of the data and inherent structural relationships to be identified in advance. Through this EDA, examining the distribution and values ​​of the data enables a better understanding of the phenomena represented by the data and the discovery of potential problems. Furthermore, by examining the data from various angles, diverse patterns that might not have been identified during the problem definition stage can be discovered, allowing for the modification of existing hypotheses or the formulation of new ones. Exploratory data analysis can broadly encompass the process of searching for data outliers and analyzing the relationships between data attributes.

[0109] The process of detecting outliers involves verifying whether they exist in the data and can include sampling methods, statistical methods, and visualization methods. Sampling methods involve drawing random samples from the data to identify overall trends and anomalies in the data values. Statistical methods may utilize summary statistics, such as the mean, median, and mode to identify the center of the data, or range and variance to check the dispersion. Visualization methods utilize probability density functions, histograms, dot plots, word clouds, time series charts, and maps to determine which statistical indicators are appropriate for the individual attributes of the collected data. However, when using statistical indicators, caution should be exercised as the mean reflects all data values ​​within a set and is therefore affected by outliers, whereas the median uses only the single value in the middle, allowing for representative results even in the presence of outliers.

[0110] The process of analyzing relationships between data attributes involves identifying combinations of attributes within the data that possess meaningful correlations. Relationship analysis can be conducted differently depending on the combination of attributes between qualitative attributes (Categorical Variables; Qualitative), which cannot be expressed numerically but can be arbitrarily quantified, and quantitative attributes (Numeric Variables; Quantitative), which can be quantified. Categorical-categorical relationships can display the number of values ​​corresponding to each pair of attribute values ​​using cross-tabulation tables or mosaic plots; Numeric-categorical relationships can be visually represented through box plots or by observing statistical values ​​by category (mean, median, etc.); and Numeric-numeric relationships can analyze the association between two attributes using correlation coefficients. It can be confirmed that a correlation coefficient of -1 indicates a negative correlation where the two attributes change in opposite directions, 0 indicates no correlation, and 1 indicates a positive correlation where the two attributes always change in the same direction. The relationship between two attributes with a correlation coefficient can also exhibit various aspects, which can be visually represented using a scatter plot.

[0111] 3. Preprocessing of Training Data

[0112] Data that has completed inspection and exploration undergoes data preprocessing to transform it into a format suitable for machine learning training models. Data preprocessing involves cleaning the data and converting it into a form that the model can understand; it generally includes handling missing data, outlier removal, data scaling, categorical data encoding, feature selection and extraction, and data transformation. The detailed processes of data preprocessing may be performed in whole or in part selectively, and a separate machine learning model may be used for this purpose.

[0113] Handling Missing Data is the process of handling missing values ​​when they exist in the data; these values ​​can be displayed as NaN (Not a Number) or empty, or deleted. Filling in or deleting missing values ​​improves data completeness, and values ​​such as the mean, median, or mode may be used when filling in missing values.

[0114] Outlier removal is the process of eliminating outliers, which are values ​​that deviate from typical data patterns. Since outliers can degrade model performance, they must be removed or replaced; this involves identifying outliers and deleting the corresponding rows or columns or replacing them with other values.

[0115] Data scaling is the process of adjusting the size of data; through data scaling, the range of the data is adjusted, which can improve model performance or accelerate convergence. Data scaling allows data characteristics to be aligned within a similar range, and generally, standardization and normalization can be applied.

[0116] Categorical Data Encoding is the process of converting categorical variables, which are represented as string or integer values ​​and cannot be directly input into a model, into a numeric type that can be input. Generally, one-hot encoding or label encoding can be used to convert categorical variables into numeric types.

[0117] Feature selection and extraction is intended to improve the performance of a model by selecting the most useful features for model training or extracting new features. Through this process, the complexity of the model can be reduced and overfitting can be prevented.

[0118] Data transformation involves converting data to extract new information or enable a model to understand it better, and may include the tokenization of text data or the preprocessing of image data. Through data transformation, model performance can be improved by extracting useful features from original data or converting data into an appropriate format.

[0119] Through data preprocessing as described above, it is possible to achieve the effects of improving the performance and ensuring the stability of machine learning models.

[0120] Meanwhile, if the collected protein amino acid sequence-based data has not been preprocessed according to requirements, tokenization, cleaning, and normalization can be performed to suit the intended use of the data.

[0121] In order for computers to understand and process text, it must be appropriately converted into numbers. Since the performance of natural language processing varies significantly depending on how words are represented, many techniques have been proposed to quantify words. Currently, word embedding, which vectorizes each word through artificial neural network learning, is the most widely used method, and this word embedding approach can also be applied to omics data.

[0122] 4. Learning of Artificial Neural Networks

[0123] Artificial neural networks can be trained using training data sets. Here, training refers to the process of determining the parameters of an artificial neural network using training data to achieve objectives such as classifying input data, performing regression analysis, or clustering. Typical examples of artificial neural network parameters include weights assigned to synapses or biases applied to neurons.

[0124] An artificial neural network trained on training data can classify or cluster input data according to the patterns of the input data.

[0125] The following explains the learning methods of artificial neural networks. The learning methods of artificial neural networks can be broadly classified into supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0126] Supervised learning is a method of machine learning designed to infer a function from training data. Among the functions inferred in this way, outputting a continuous value is called regression, and predicting and outputting the class of an input vector is called classification.

[0127] In supervised learning, an artificial neural network is trained with labels for the training data. Here, a label refers to the correct answer (or result value) that the artificial neural network must infer when training data is input into the artificial neural network. In this specification, the correct answer (or result value) that the artificial neural network must infer when training data is input is referred to as a label or labeling data. Furthermore, in this specification, setting labels on the training data for the training of the artificial neural network is referred to as labeling the training data. In this case, the training data and the corresponding labels constitute a single training set, and can be input to the artificial neural network in the form of a training set.

[0128] Meanwhile, training data represents multiple features, and labeling the training data implies that labels are attached to the features represented by the training data. In this case, the training data can represent the features of the input object in the form of a vector. An artificial neural network can infer a function regarding the association between the training data and the labeled data by utilizing the training data and the labeled data. Furthermore, the parameters of the artificial neural network can be determined (optimized) through the evaluation of the function inferred by the network.

[0129] Unsupervised learning is a type of machine learning in which no labels are provided for the training data. Specifically, unsupervised learning may be a learning method in which an artificial neural network is trained to find and classify patterns within the training data itself, rather than the relationship between the training data and the corresponding labels. Examples of unsupervised learning include clustering or Independent Component Analysis. In this specification, the term 'clustering' may be used interchangeably with the term 'clustering'.

[0130] Examples of artificial neural networks that utilize unsupervised learning include Generative Adversarial Networks (GANs) and Autoencoders (AEs).

[0131] Generative Adversarial Networks (GANs) are machine learning methods in which two distinct artificial intelligence models—a generator and a discriminator—compete to improve performance. In this context, the generator is a model that creates new data, capable of generating new data based on original data. The discriminator, on the other hand, is a model that recognizes data patterns, performing the role of distinguishing whether input data is original data or new data generated by the generator. Furthermore, the generator learns by receiving input data that failed to deceive the discriminator, while the discriminator learns by receiving input data that was deceived by the generator. Consequently, the generator can evolve to deceive the discriminator as effectively as possible, and the discriminator can evolve to better distinguish between original data and data generated by the generator.

[0132] An autoencoder is a neural network that aims to reproduce the input itself as the output. An autoencoder includes an input layer, at least one hidden layer, and an output layer. In this case, since the number of nodes in the hidden layer is less than the number of nodes in the input layer, the dimensionality of the data is reduced, and accordingly, compression or encoding is performed. Additionally, the data output from the hidden layer enters the output layer. In this case, since the number of nodes in the output layer is greater than the number of nodes in the hidden layer, the dimensionality of the data is increased, and accordingly, decompression or decoding is performed.

[0133] Meanwhile, an autoencoder represents input data as hidden layer data by adjusting the connection strengths of neurons through learning. In the hidden layer, information is represented with fewer neurons than in the input layer, and the fact that input data can be reproduced as output implies that the hidden layer has discovered and represented hidden patterns from the input data.

[0134] Semi-supervised learning is a type of machine learning that refers to a learning method utilizing both labeled and unlabeled training data. One technique within semi-supervised learning involves inferring labels from unlabeled training data and then performing learning using those inferred labels; this method can be particularly useful when the cost of labeling is high.

[0135] Reinforcement learning is a theory that states that if an agent is provided with an environment where it can determine the best action to take at every moment, it can find the optimal path through experience alone, without relying on data. Reinforcement learning is primarily executed via a Markov Decision Process (MDP). To explain the MDP, first, an environment is provided containing the information necessary for the agent to take its next action; second, the agent's behavior within that environment is defined; third, rewards are determined for success and penalties for failure; and fourth, the optimal policy is derived through repeated experience until future rewards reach their peak.

[0136] The structure of an artificial neural network is determined by the configuration of the model, activation function, loss function or cost function, learning algorithm, optimization algorithm, etc., and hyperparameters are set in advance before learning, and model parameters are set through learning thereafter, so the content can be determined.

[0137] For example, factors determining the structure of an artificial neural network may include the number of hidden layers, the number of hidden nodes included in each hidden layer, the input feature vector, the target feature vector, etc.

[0138] Hyperparameters include various parameters that must be initially set for training, such as the initial values ​​of model parameters. Model parameters, on the other hand, include various parameters intended to be determined through training. For example, hyperparameters may include initial values ​​for inter-node weights, initial values ​​for inter-node bias, mini-batch size, number of training iterations, and learning rate. Additionally, model parameters may include inter-node weights and inter-node bias.

[0139] A loss function can be used as an indicator (criterion) to determine optimal model parameters during the learning process of an artificial neural network. In an artificial neural network, learning refers to the process of manipulating model parameters to reduce the loss function, and the objective of learning can be viewed as determining model parameters that minimize the loss function. The loss function can primarily be the Mean Squared Error (MSE) or the Cross Entropy Error (CEE), but the present invention is not limited thereto. The Cross Entropy Error can be used when the correct label is one-hot encoded. One-hot encoding is an encoding method in which the correct label value is set to 1 only for neurons corresponding to the correct answer, and the correct label value is set to 0 for neurons that are not the correct answer.

[0140] In machine learning or deep learning, learning optimization algorithms can be used to minimize the loss function, and learning optimization algorithms include Gradient Descent (GD), Stochastic Gradient Descent (SGD), Momentum, Nesterov Accelerate Gradient (NAG), Adagrad, AdaDelta, RMSProp, Adam, Nadam, etc.

[0141] Gradient Descent is a technique that adjusts model parameters in a direction that reduces the loss function value by considering the gradient of the loss function from the current state. The direction in which model parameters are adjusted is called the step direction, and the magnitude of the adjustment is called the step size. In this context, the step size can refer to the learning rate. Gradient Descent obtains the gradient by taking the partial derivative of the loss function with respect to each model parameter, and updates the model parameters by changing them in the direction of the obtained gradient by the learning rate.

[0142] Stochastic Gradient Descent is a technique that divides training data into mini-batches and performs gradient descent on each mini-batch to increase the frequency of gradient descent.

[0143] Adagrad, AdaDelta, and RMSProp are techniques that improve optimization accuracy in SGD by adjusting the step size. In SGD, Momentum and NAG are techniques that improve optimization accuracy by adjusting the step direction. Adam is a technique that improves optimization accuracy by combining Momentum and RMSProp to adjust both the step size and the step direction. Nadam is a technique that improves optimization accuracy by combining NAG and RMSProp to adjust both the step size and the step direction.

[0144] The learning speed and accuracy of artificial neural networks are characterized by being heavily dependent on hyperparameters, as well as the network structure and the type of learning optimization algorithm. Therefore, to obtain a good learning model, it is important to set appropriate hyperparameters in addition to determining a suitable network structure and learning algorithm.

[0145] Typically, hyperparameters are experimentally set to various values ​​while training the artificial neural network, and then set to the optimal value that provides stable training speed and accuracy based on the training results.

[0146] The phenotype prediction system of the present disclosure learns to predict phenotypes from the omics data of an individual using a training dataset containing omics data.

[0147] FIG. 3 is an exemplary diagram showing the learning steps of the phenotype prediction system of the present disclosure. Referring to FIG. 3, the phenotype prediction system (300) of the present disclosure includes a reconstruction model (310) and a phenotype prediction model (330), wherein the reconstruction model, which receives omics features (313) as input, generates a latent vector (315), and the phenotype prediction model, which receives the latent vector as input, predicts a disease-associated phenotype (335).

[0148] The reconstruction model of the phenotype prediction system of the present disclosure performs the function of identifying significant features from omics features. While the biological information contained in omics data is vast, it also includes a large amount of information irrelevant to specific phenotypes; therefore, training efficiency is significantly reduced when training using omics data containing such vast biological information as is. In the phenotype prediction system of the present disclosure, the reconstruction model is trained to extract only significant features from omics features, and by utilizing the trained reconstruction model for phenotype prediction, the efficiency and prediction reliability of the phenotype prediction system can be improved.

[0149] FIG. 4 is an example diagram showing the learning process of a reconstruction model included in the phenotype prediction system of the present disclosure.

[0150] Referring to FIG. 4, the reconstruction model (310) includes a first reconstruction sub-neural network (314) and a second reconstruction sub-neural network (316). The first reconstruction sub-neural network compresses the dimensions of input omics features (313) to generate a latent vector (315). The second reconstruction sub-neural network expands the compressed latent vector to generate a reconstructed omics feature (317).

[0151] The reconstruction model may additionally include a performance evaluation unit (318) of the reconstruction model, which receives omics features (313) input to the reconstruction model and reconstructed omics features (317) generated using the reconstruction model, evaluates whether the reconstruction model has reconstructed omics features similarly to the input omics features, and outputs the evaluated result. The reconstruction model may be trained to reconstruct omics features similarly to the input omics features by receiving the value output from the performance evaluation unit of the reconstruction model.

[0152] FIG. 5 is an example diagram showing the learning process of a phenotype prediction model included in the phenotype prediction system of the present disclosure.

[0153] Referring to FIG. 5, the phenotype prediction model includes a phenotype prediction neural network (334). The phenotype prediction neural network receives a latent vector (315) generated by the reconstruction model as input and predicts the phenotype (335). Additionally, the phenotype prediction neural network can predict the phenotype (335) by additionally receiving an object feature (333) generated from object data (331) as input. By additionally using object data for training, the efficiency and performance of the phenotype prediction system can be further improved.

[0154] The phenotype prediction model may additionally include a performance evaluation unit (337) of the phenotype prediction model, and may additionally receive phenotype features for evaluating the performance of the phenotype prediction model. The performance evaluation unit of the phenotype prediction model evaluates whether the phenotype prediction neural network predicted the phenotype similarly to the phenotype data (336) and outputs the evaluated result. The phenotype prediction system of the present disclosure may be trained to predict the phenotype (335) similarly to the phenotype data (336) by receiving the value output from the performance evaluation unit of the phenotype prediction model.

[0155] The phenotype prediction system of the present disclosure may be trained at least once. The training order of the reconstruction model and the phenotype prediction model is not particularly limited, but the reconstruction model and the phenotype prediction model may be trained separately, the reconstruction model and the phenotype prediction model may be trained alternately, or the phenotype prediction model may be trained additionally after the reconstruction model and the phenotype prediction model have been trained alternately.

[0156] In one embodiment of the present disclosure, after the reconstruction model completes learning, the latent vector generated by the reconstructed model that has completed learning is stored in a storage device (320), and the phenotype prediction model that receives the stored latent vector proceeds with learning, thereby allowing the reconstruction model and the phenotype prediction model to learn respectively.

[0157] In another embodiment, the reconstruction model receives a latent vector generated during the learning process and the phenotype prediction model proceeds with learning, and then receives a latent vector generated by the learned reconstruction model and proceeds with learning again, so that the reconstruction model and the phenotype prediction model can learn alternately.

[0158] In another embodiment, a reconstructing model receives a latent vector generated during the learning process and a phenotype prediction model proceeds with learning. Then, the phenotype prediction model receives a latent vector generated by the learned reconstructing model and proceeds with learning again. After the reconstructing model and the phenotype prediction model learn alternately, the latent vector generated by the reconstructing model that has completed learning is stored in a storage device, and the phenotype prediction model that receives the stored latent vector may proceed with additional learning.

[0159] MoE (Mixture of Experts) Architecture

[0160] In one embodiment of the present disclosure, a system for predicting phenotypes may be performed by utilizing a model architecture such as MoE. Here, MoE may refer to an architecture of a machine learning model that solves complex problems by combining multiple expert models.

[0161] Such MoE may include expert models, which are multiple small networks designed to learn different parts and / or different features of a given data and perform data processing operations accordingly, and a gating network that evaluates the performance of each expert model and determines which expert model is most suitable for assigning a specific task based on the given data based on this evaluation.

[0162] Thus, according to the MoE architecture, a gating network that acquires predetermined input data determines probabilistic or deterministic task assignments for each expert model, and the selected expert models perform their respective tasks and return the results, thereby enabling data processing for a specific task.

[0163] According to one embodiment of the present disclosure, a MoE model used may refer to a specific MoE model implemented according to a common method known in the art. For example, the MoE model may include a Switch Transformer, Conditional Computation in Neural Networks, Sparse Mixture of Experts, and / or a Megatron-LM.

[0164] Additionally, in one embodiment of the present disclosure, a MoE model based on the combination of a plurality of Specialized Models (SM) and Routers (Gating Network, RT) may be included, and a MoE model based on domain-specific Specialized Models may also be included.

[0165] By utilizing such MoE, the overall efficiency and performance of the phenotype prediction system of the present disclosure can be enhanced by activating only specific parts and concentrating computational resources in cases such as handling complex tasks or large datasets.

[0166] 5. Computing System Device

[0167] Embodiments of the present invention may be implemented as application-specific integrated circuits (ASICs) designed to suit specific application fields and special functions of devices.

[0168] Custom integrated circuits are also referred to as custom semiconductors. Unlike standard semiconductors, which have fixed specifications and can be applied to any electronic product or application as long as certain requirements are met, custom semiconductors are used for specific products or functions and are integrated circuits manufactured by semiconductor companies to meet specific orders. In other words, custom semiconductors are designed and manufactured to perform only the functions necessary for a specific device or feature. Custom semiconductors are broadly classified according to their design method into Full Custom ICs, which design and manufacture circuits from scratch to meet user requirements, and Semi-Custom ICs, which design and manufacture circuits using parts of a standardized design.

[0169] Application-specific semiconductors are primarily used in communication systems, high-performance computing systems, consumer electronics, automobiles, industrial automation, medical devices, the military, and the aerospace industry; recently, they are being applied to AI semiconductors that execute the large-scale computations required for AI implementation with high performance and power efficiency.

[0170] Application-specific semiconductors (ASICs) are used as core components in communication systems, such as network routers, switches, and modems, performing data packet processing, protocol conversion, and signal processing to provide high throughput and low latency. In high-performance computing systems, ASICs serve as key components for high-speed and parallel processing, while in consumer electronics—including digital cameras, smartphones, tablets, and game consoles—ASICs provide high-performance and low-power solutions required to perform specific functions. In the automotive industry, ASICs are used to control various electronic systems within vehicles, and in industrial automation systems, they provide solutions for high-precision control and high-performance processing.

[0171] The application-specific integrated circuit to which the embodiment of the present invention is applied includes a memory in which an individual memory interface (I / F) is implemented, and may include a plurality of function blocks that request memory access. Each function block may be a Direct Memory Access (DMA) function block, a processor, a video processor, a cache controller, a decompression block, or a data path block. The basic configuration of the application-specific integrated circuit may include a transistor that amplifies or switches an electrical signal, a logic gate which is a circuit that performs a logical function by combining transistors, a memory cell that stores data, an analog circuit which is a circuit that processes continuous voltage or current by combining transistors, and an Intellectual Property Core (IP Core) such as a microprocessor, DSP, or graphics core that is pre-designed to perform a specific function.

[0172] The ASIC may include an individual memory I / F that interfaces with individual memory and an embedded memory I / F that interfaces with embedded memory. The individual memory I / F is connected to each function block to receive memory access signals (e.g., control signals, address signals, and data signals) and, based on these input signals, can generate signals to control the individual memory. The embedded memory I / F is connected to each function block to receive memory access signals (e.g., control signals, address signals, and data signals) and, based on these input signals, can generate modified memory access signals to control the embedded memory. The individual memory I / F and the embedded memory I / F are designed within the memory control block of the ASIC to provide a memory control structure that can be flexibly applied to both the individual memory and the embedded memory.

[0173] Additionally, an application-specific integrated circuit (ASIC) for an artificial neural network is composed of multiple neurons arranged in an array and multiple synapse circuits, each neuron being composed of a register, a microprocessor, and at least one input, and each synapse circuit being configured to include memory for storing synapse weights. Here, each neuron of the ASIC may be connected to at least one other neuron through one of the multiple synapse circuits.

[0174] Although the present disclosure has been described as generally being implementable by a computing device, a person skilled in the art will be well aware that the present disclosure may be implemented in combination with computer-executable instructions and / or other program modules that can be executed on one or more computers and / or as a combination of hardware and software.

[0175] Those skilled in the art of the present disclosure will understand that information and signals may be represented using any various different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced in the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0176] Those skilled in the art will understand that the various exemplary logic blocks, modules, processors, means, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented by electronic hardware, various forms of programs or design code (referred to herein as software for convenience), or a combination of all such. To clearly illustrate this interoperability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in relation to their functions. Whether such functions are implemented as hardware or software depends on the design constraints imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementation decisions should not be interpreted as being outside the scope of this disclosure.

[0177] The various embodiments presented herein may be implemented as methods, devices, or manufactured articles using standard programming and / or engineering techniques. The term manufactured article includes a computer program, carrier, or medium accessible from any computer-readable storage device. For example, computer-readable storage media include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical discs (e.g., CDs, DVDs, etc.), smart cards, and flash memory devices (e.g., EEPROMs, cards, sticks, key drives, etc.). Additionally, the various storage media presented herein include one or more devices and / or other machine-readable media for storing information.

[0178] It should be understood that the specific order or hierarchy of steps in the presented processes is an example of exemplary approaches. It should be understood that the specific order or hierarchy of steps in the processes may be rearranged within the scope of this disclosure based on design priorities. The appended method claims provide elements of various steps in a sample order, but do not imply being limited to the specific order or hierarchy presented.

[0179] FIG. 1 illustrates an example of a block diagram of a computing system device (100) implementing a phenotype prediction service according to one embodiment of the present disclosure.

[0180] Various operations of the phenotype prediction system and method of the present disclosure may be performed by any suitable means capable of performing corresponding functions. Such means may include various hardware, software components, modules, and combinations thereof, including but not limited to circuits, processors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).

[0181] Referring to FIG. 1, a computing system (100) implementing an artificial neural network of the present disclosure may include a transceiver (110), memory (110), a database (130), and a processor (140). However, not all components shown in FIG. 1 are essential components of the computing system device (100). The computing system device (100) may be implemented with more components than those shown in FIG. 1, or with fewer components than those shown in FIG. 1. Furthermore, the transceiver (110), memory (110), and processor (140) may be implemented in the form of a single chip.

[0182] In one embodiment, the transceiver (110) can communicate with a terminal or other electronic device connected to the computing system device (100) via wired or wireless connection. For example, the transceiver (110) can obtain protein amino acid sequence information, protein structure information, or protein expressions generated using an artificial neural network from another electronic device.

[0183] Various types of data, such as programs like applications and files, can be installed and stored in the memory (110). The processor (140) may access and use the data stored in the memory (110) or store new data in the memory (110). Additionally, one or more instructions may be stored in the memory (110). The processor (140) may execute one or more instructions stored in the memory.

[0184] The processor (140) controls the overall operation of the computing system device (100). Here, the processor (140) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), an application integrated circuit, a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array, controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.

[0185] The processor (140) may control other components included in the computing system device (100) to perform operations for operating the computing system device (100). For example, the processor (140) may acquire protein amino acid sequence information, acquire protein expressions using the neural network, obtain a contrast loss function from the protein expressions, and improve one or more numerical values ​​or parameters of one or more neural networks among the encoder neural network and the projection neural network based on the contrast loss function.

[0186] The database (130) may store various training data for training a learning model. Additionally, the database (130) may store protein amino acid sequence information, protein structure information, simulation result information, etc., and in various embodiments, output data produced by the learning model may be stored. Although FIG. 1 is illustrated as including a database (130) in a computing system device (100), the database (130) may be provided outside the device. In this case, the database (130) may be connected to the computing system device (100) via a wired or wireless connection.

[0187] Additionally, the learning model of the present disclosure may be implemented outside the computing system device (100) (e.g., cloud-based) or included inside the computing system device (100).

[0188] One embodiment of the present disclosure may also be implemented in the form of a recording medium comprising computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include both computer storage media and communication media. A computer storage medium includes both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. A communication medium typically includes computer-readable instructions, data structures, or program modules and includes any information transmission medium.

[0189] Throughout the entire disclosure, the device may include, but is not limited to, a server, smartphone, tablet PC, PC, TV, smart TV, mobile phone, PDA (personal digital assistant), speaker, laptop, media player, micro server, e-book object recognition device, digital broadcasting object recognition device, kiosk, MP3 player, digital camera, robot vacuum cleaner, home appliance, other mobile or non-mobile computing device, a watch, glasses, hair band and ring, etc. equipped with communication functions and data processing functions.

[0190] 6. Applications of Phenotype Prediction Systems

[0191] The phenotypic prediction system of the present disclosure may be used for various purposes, such as diagnosing a disease in an individual, providing information for assessing the risk of a disease in an individual, or searching for biomolecules associated with a disease.

[0192] FIG. 6 is an exemplary diagram illustrating the step of predicting a disease-associated phenotype using the phenotype prediction system of the present disclosure. The phenotype prediction system, which receives omics data (601) of an individual, can generate disease-associated phenotype information (602) of the individual through a reconstruction model and a phenotype prediction model, and the disease-associated phenotype information (602) of the individual generated by the phenotype prediction system of the present disclosure can be used to accurately diagnose the individual's disease or to evaluate the risk of the individual's disease.

[0193] FIG. 7 is an exemplary diagram illustrating the steps of searching for disease-associated biomolecules using the phenotype prediction system of the present disclosure. The phenotype prediction system, which receives omics data (301) as input, generates a latent vector (315) through a reconstruction model, and can obtain information on disease-associated biomarkers (701) from the latent vector. This information on disease-associated biomarkers (701) can be utilized for the discovery of new targets for new drug development.

[0194] The foregoing description of the present disclosure is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present invention. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0195] The scope of the present disclosure is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present disclosure.

[0196] The present invention will be described in more detail below with reference to examples. The embodiments of the present disclosure are exemplary in all respects, and the present invention is not limited to these embodiments.

[0197] Example 1

[0198] Omics features were generated from the omics data of the AD-BXD panel. After the reconstruction model completed training using the generated omics features, latent vectors were generated from the omics features using the reconstruction model and stored in a storage device. Subsequently, a phenotypic prediction model was trained to predict contextual fear acquisition (CFA) and contextual fear memory (CFM) using the stored latent vectors as input values.

[0199] Example 2

[0200] After alternately training the reconstruction model and the phenotypic prediction model, latent vectors were generated from omics features using the reconstruction model and stored in storage. The phenotypic prediction model was additionally trained to predict fear memory acquisition and fear memory recall using the stored latent vectors as input.

[0201] Example 3

[0202] The reconstruction model and the phenotype prediction model were trained to predict fear memory acquisition and fear memory recall by alternating between the reconstruction model and the phenotype prediction model.

[0203] Comparative Example 1

[0204] Omics features were generated from the omics data of the AD-BXD panel, and a phenotype prediction system was trained using the generated omics features as input values.

[0205] Performance evaluation of phenotype prediction systems

[0206] The performance of the phenotype prediction systems of the embodiments and comparative examples of the present disclosure was evaluated using the Root Mean Squared Error (RMSE) and cross-validated using the Pearson correlation coefficient. A lower Root Mean Squared Error indicates superior performance of the phenotype prediction system, and a higher Pearson correlation coefficient indicates superior performance of the phenotype prediction system.

[0207] Fear Memory Acquisition Prediction Performance Root Mean Square Error Pearson Correlation Coefficient Example 10.14 200 0.4404 Example 20.13 76 0.4648 Example 30.13 62 0.4713 Comparative Example 0.14 52 0.4205

[0208] Fear memory recall prediction performance Root mean square error Pearson correlation coefficient Example 10.149 20.5788 Example 20.136 40.6404 Example 30.135 60.6431 Comparative Example 10.15 140.5709

[0209] The phenotype prediction systems of Examples 1 to 3 had lower root mean square error values ​​and higher Pierce correlation coefficient values ​​in predicting fear memory acquisition and fear memory recall compared to the phenotype prediction system that did not use a reconstruction model (Comparative Example) (Tables 1 and 2). The above results show that the phenotype prediction system of the present disclosure exhibits superior phenotype prediction performance compared to the phenotype prediction system of the Comparative Example.

Claims

1. In a phenotype prediction system, Memory for storing one or more instructions; and At least one that executes the one or more instructions stored in the memory. Includes processors, The operation performed by the above one or more instructions is The method includes a step of learning phenotype predictions from omics data, wherein the step of learning phenotype predictions is A step of acquiring at least one piece of data including omics data, A step of generating omics features from the above omics data, A step in which a reconstruction model generates a latent vector from the above omics features, and A phenotype prediction system in which a phenotype prediction model includes the step of predicting a phenotype from the above latent vector.

2. In claim 1, the reconstruction model comprises a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector, and The step of learning the above phenotype prediction is, The step of the first reconstruction sub-neural network generating the latent vector from the omics features, The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector, A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and A phenotype prediction system comprising a step in which the reconstruction model learns based on the above evaluation.

3. In paragraph 2, the step of learning the phenotype prediction above (1) The above reconstruction model and the above phenotype prediction model each learn, or (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or (3) A phenotype prediction system in which the above reconstruction model and the above phenotype prediction model are trained alternately, and then the above phenotype prediction model is further trained.

4. A phenotype prediction system according to claim 1, wherein the phenotype prediction model additionally receives one or more individual features as input to predict a phenotype.

5. A phenotype prediction system according to claim 1, wherein the phenotype is associated with a disease.

6. A phenotype prediction system according to claim 1, wherein the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

7. A phenotype prediction system according to claim 1, further comprising the step of obtaining disease-associated biomarker information or the step of obtaining disease-associated phenotype information of an individual.

8. A phenotype prediction method performed by at least one processor, The above method includes a step of learning phenotype predictions from omics data, and the step of learning phenotype predictions is A step of acquiring at least one piece of data including omics data, A step of generating omics features from the above omics data, A step in which a reconstruction model generates a latent vector from the above omics features, and A phenotype prediction method comprising the step of a phenotype prediction model predicting a phenotype from the above latent vector.

9. In claim 8, the reconstruction model comprises a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector, and The step of learning the above phenotype prediction is, The step of the first reconstruction sub-neural network generating the latent vector from the omics features, The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector, A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and A phenotype prediction method comprising a step of learning the reconstruction model based on the above evaluation.

10. In paragraph 9, the step of learning the phenotype prediction above (1) The above reconstruction model and the above phenotype prediction model each learn, or (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or (3) A phenotype prediction method in which the reconstruction model and the phenotype prediction model are trained alternately, and then the phenotype prediction model is further trained.

11. A phenotype prediction method according to claim 8, wherein the phenotype prediction model additionally receives one or more individual features as input to predict the phenotype.

12. A method for predicting a phenotype in which the phenotype in paragraph 8 is associated with a disease.

13. A phenotype prediction method according to claim 8, wherein the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

14. A phenotype prediction method according to claim 8, further comprising the step of obtaining disease-associated biomarker information or the step of obtaining disease-associated phenotype information of an individual.

15. A program stored on a computer-readable recording medium to execute the method of any one of paragraphs 8 through 14 on a computer.

16. As a computerized method for predicting phenotypes, The above method includes a step of learning phenotype predictions from omics data, and the step of learning phenotype predictions is A step of acquiring at least one piece of data including omics data, A step of generating omics features from the above omics data, A step in which a reconstruction model generates a latent vector from the above omics features, and A phenotype prediction method comprising the step of a phenotype prediction model predicting a phenotype from the above latent vector.

17. In paragraph 16, the reconstruction model comprises a first reconstruction sub-neural network and a second reconstruction sub-neural network, wherein the first reconstruction sub-neural network generates the latent vector from the omics feature, and the second reconstruction sub-neural network generates the reconstructed omics feature from the latent vector, and The step of learning the above phenotype prediction is, The step of the first reconstruction sub-neural network generating the latent vector from the omics features, The step of the second reconstruction sub-neural network generating the reconstructed omics features from the latent vector, A step of evaluating the performance of the reconstruction model by comparing the reconstructed omics features and the input omics features, and A phenotype prediction method comprising a step of learning the reconstruction model based on the above evaluation.

18. In paragraph 17, the step of learning the phenotype prediction above (1) The above reconstruction model and the above phenotype prediction model each learn, or (2) The above reconstruction model and the above phenotype prediction model are trained alternately, or (3) A phenotype prediction method in which the reconstruction model and the phenotype prediction model are trained alternately, and then the phenotype prediction model is further trained.

19. A phenotype prediction method according to claim 16, wherein the phenotype prediction model additionally receives one or more individual features as input to predict the phenotype.

20. A method for predicting a phenotype in which the above phenotype is associated with a disease, in paragraph 16.

21. A phenotype prediction method according to claim 16, wherein the omics data comprises one or more selected from genomic data, transcriptomic data, proteomic data, metabolomic data, lipidomic data, and combinations thereof.

22. A phenotype prediction method according to claim 16, further comprising the step of obtaining disease-associated biomarker information or the step of obtaining disease-associated phenotype information of an individual.