Information processing system, information processing method, and learning model creation method
The information processing system addresses the computational complexity of genome analysis by tokenizing and compressing genetic information, reducing the computational load and enhancing analysis accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KEIO UNIV
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
The computational complexity of genome analysis increases with the number of base pairs, making it difficult to analyze genetic information across the entire genome, particularly for complex diseases involving numerous and widespread factors.
An information processing system comprising a token conversion unit, a compression unit, and an analysis unit, which converts genetic information into tokens, compresses them into a single vector using a first learning model, and analyzes the genetic information using a second learning model to reduce computational load.
The system reduces the computational load of genome analysis, enabling comprehensive analysis of the entire genome and improving the accuracy of genetic information analysis.
Smart Images

Figure 2026068241000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing system, an information processing method, and a method for creating a learning model. [Background technology]
[0002] Information processing systems for analyzing genetic information are known (for example, Non-Patent Document 1). For example, Non-Patent Document 1 describes the use of deep learning models for analyzing genetic information. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Special Publication No. 2022-504916 [Non-patent literature]
[0004] [Non-Patent Document 1] ZigaAvsec et al. “Effective gene expression prediction from sequence byintegrating long-range interactions”, Nature Methods 18, 1196-1203 (2021) [Non-Patent Document 2] YanrongJi et al. “DNABERT: pre-trained Bidirectional EncoderRepresentations from Transformers model for DNA-language in genome”, Bioinformatics,Volume 37, Issue 15, August 2021, Pages 2112-2120 [Overview of the project] [Problems that the invention aims to solve]
[0005] In genome analysis, the computational complexity increases with the number of base pairs. When using deep learning models for genome analysis, the computational complexity required for training also increases with the number of base pairs. The human genome consists of approximately 3.2 billion base pairs. Therefore, analyzing genetic information across the entire genome has been extremely difficult. However, the demand for whole-genome analysis is increasing. For example, complex diseases involving numerous and widespread factors require analysis of the entire genome.
[0006] One aspect of the present invention aims to provide an information processing system and an information processing method that can reduce the computational load of genome analysis. Another aspect of the present invention aims to provide a method for creating a learning model that can reduce the computational load of genome analysis. [Means for solving the problem]
[0007] An information processing system in one aspect of the present invention comprises a token conversion unit, a compression unit, and an analysis unit. The token conversion unit converts genetic information into multiple tokens. The compression unit compresses the multiple tokens into a single vector using a first learning model. The analysis unit analyzes the genetic information using a second learning model that outputs the analysis result of the genetic information based on the input of the above vector. [Effects of the Invention]
[0008] One aspect of the present invention provides an information processing system and an information processing method that can reduce the computational load of genome analysis. Another aspect of the present invention provides a method for creating a learning model that can reduce the computational load of genome analysis. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram of the information processing system in this embodiment. [Figure 2] This is a diagram illustrating the conversion to tokens. [Figure 3] This is a diagram illustrating compression using the first learning model. [Figure 4] This is a diagram for explaining the analysis by the second learning model. [Figure 5] This is a diagram showing an example of the hardware configuration of an information processing system. [Figure 6] This is a flowchart showing an example of the information processing method in the learning phase. [Figure 7] This is a flowchart showing an example of the information processing method in the estimation phase. [Figure 8] This is a diagram showing the performance of the second learning model. [Figure 9] This is a diagram showing the performance of the second learning model. [Figure 10] This is a diagram showing the attention score.
Embodiments for Carrying Out the Invention
[0010] [Explanation of Embodiments of the Present Disclosure] First, the embodiments of the present disclosure will be listed and explained.
[0011] [1] The information processing system in the embodiment of the present disclosure includes a token conversion unit, a compression unit, and an analysis unit. The token conversion unit converts genetic information into a plurality of tokens. The compression unit compresses the plurality of tokens into one vector by the first learning model. The analysis unit analyzes the genetic information by a second learning model that outputs an analysis result of the genetic information based on the input of the above vector.
[0012] In the information processing system in [1] above, a plurality of tokens into which genetic information has been converted are compressed into one vector by the first learning model, and the genetic information is analyzed by the second learning model. The second learning model outputs an analysis result of the genetic information based on the input of the compressed vector. In this case, the computational amount of genome analysis can be reduced. Therefore, for example, comprehensive analysis of the entire genome can be realized.
[0013] [2] In the information processing system of [1] above, the token conversion unit may convert the genetic information into a character array information and then convert it into the plurality of tokens. In this case, the computational amount of genome analysis can be further reduced.
[0014] [3] In the information processing system of [1] or [2] above, the first learning model may have a transformer architecture. In this case, the genetic information can be further compressed.
[0015] [4] In any one of the information processing systems of [1] to [3] above, the second learning model may include a cross-attention mechanism. In this case, the computational amount of genome analysis can be further reduced and the accuracy of the analysis result of genetic information can be further improved.
[0016] [5] In the information processing system of [4] above, the second learning model may further include a self-attention mechanism for analyzing genetic information. The cross-attention mechanism may output diagnostic information of a subject having genetic information according to the input of the output result from the self-attention mechanism. In this case, the variation in the form of the diagnostic information output from the second learning model can be improved.
[0017] [6] In the information processing systems of [1] to [5] above, the second learning model may output an analysis result of genetic information according to the vector and non-genetic information by inputting the vector and non-genetic information. In this case, the accuracy of the analysis result of genetic information output from the second learning model can be further improved.
[0018] [7] In the information processing system of [6] above, the second learning model may be learned based on the first learning information indicating the genetic information and the second learning information indicating non-genetic information. In this case, the accuracy of the analysis result of genetic information output from the second learning model can be further improved.
[0019] [8] In any one of the information processing systems described in [1] to [2] above, the second learning model may be trained based on the vector relating to the genetic information and the results of the analysis of the genetic information by the analysis unit. In this case, the accuracy of the analysis results of the genetic information output from the second learning model can be further improved.
[0020] [9] Another form of the information processing method of the present disclosure comprises converting genetic information into a plurality of tokens, compressing the plurality of tokens into a single vector by a first learning model, and analyzing the genetic information by a second learning model that outputs the results of the analysis of the genetic information based on the input of the compressed vector. In this case, the computational load of genome analysis can be reduced. For this reason, for example, comprehensive analysis of the entire genome can be realized.
[0021]
[10] Another form of the present disclosure is a method for creating a learning model in which a learning model is created by compressing a number of learning tokens converted from learning genetic information obtained from a number of training subjects, and using this compressed vector as input, the learning model outputs the results of genetic information analysis. In this case, a learning model can be created in which the computational cost of genome analysis is reduced. [Details of the embodiments of this disclosure]
[0022] Hereinafter, an example of an embodiment of the information processing system according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same or corresponding parts are denoted by the same reference numerals, and redundant descriptions are omitted.
[0023] First, the schematic configuration of the information processing system in the embodiment of this disclosure will be described with reference to Figures 1 to 4. Figure 1 is a block diagram of the information processing system 1 in this embodiment.
[0024] Information processing system 1 enables, for example, comprehensive analysis of inter-mutation interactions. Information processing system 1 analyzes genetic information. Genetic information includes genetic mutation information. In one example shown in this embodiment, information processing system 1 analyzes genetic mutation information. Information processing system 1 analyzes, for example, the entire human genome. In the example shown in this embodiment, information processing system 1 outputs a disease diagnosis result based on the genome analysis results. For example, information processing system 1 outputs a diagnosis result for a multifactorial disease that cannot be explained by a single mutation. Multifactorial diseases include, for example, amyotrophic lateral sclerosis (ALS). It has been suggested that the onset of ALS is determined by the genome. Information processing system 1 comprehensively analyzes inter-mutation interactions, for example, targeting the entire genome region. As a modification of this embodiment, inter-mutation interactions may be analyzed targeting only a portion of the genome region.
[0025] In this embodiment, the information processing system 1 performs a learning phase and an estimation phase. The information processing system 1 performs the estimation phase based on at least two deep learning models. In the estimation phase, the information processing system 1 analyzes the input genetic information. The information processing system 1 may further perform a deep learning phase based on the information obtained in the estimation phase.
[0026] In the example shown in this embodiment, the information processing system 1 performs an estimation phase based on a first learning model and a second learning model. The first learning model is, for example, a deep learning model that vectorizes genetic information. The second learning model is a deep learning model that analyzes genetic information. The first and second learning models have, for example, a transformer architecture. The first and second learning models include, for example, a cross-attention mechanism and a self-attention mechanism. The first and second learning models correspond to, for example, perceivers. At least one of the first and second learning models is learned in the learning phase. For example, only the second learning model is learned in the learning phase.
[0027] For example, Information Processing System 1 utilizes machine learning in the creation of a deep learning model during the learning phase. Machine learning is a method that autonomously discovers laws or rules by iteratively learning based on given information. For example, in the machine learning performed in Information Processing System 1, model parameters such as activation functions and weights are optimized through learning using a training dataset. This creates a deep learning model.
[0028] The machine learning performed in Information Processing System 1 is, for example, deep learning. This machine learning is supervised learning composed of multiple deep learning models. Information Processing System 1 uses machine learning configured to include neural networks. The machine learning performed in Information Processing System 1 is not limited to supervised learning.
[0029] For example, the neural network used in Information Processing System 1 is a deep learning model with a transformer architecture. This deep learning model includes a model that efficiently vectorizes genetic mutation information and generates contextual vectors, and a model that uses the generated contextual vectors to predict the risk of disease onset. The type of neural network used in Information Processing System 1 is not limited to this.
[0030] The information processing system 1 comprises an estimation unit 10 that operates in the estimation phase and a learning unit 20 that operates in the learning phase. The information processing system 1 includes, as the estimation unit 10, a genetic information acquisition unit 31, a token conversion unit 32, a compression unit 33, an analysis unit 34, a storage unit 35, and an output unit 36. The information processing system 1 includes, as the learning unit 20, a learning information acquisition unit 41, a token conversion unit 32, a compression unit 33, a learning model creation unit 44, and a storage unit 35. In this embodiment, the token conversion unit 32, the compression unit 33, the analysis unit 34, and the storage unit 35 are included in both the estimation unit 10 and the learning unit 20.
[0031] As a variation of this embodiment, the estimation unit 10 and the learning unit 20 may be physically separated from each other. In this case, the estimation unit 10 and the learning unit 20 may each comprise different token conversion units 32, compression units 33, analysis units 34, and storage units 35. As a variation of this embodiment, the information processing system 1 may consist only of the estimation unit 10, without including the learning information acquisition unit 41 and the learning model creation unit 44. The information processing system 1 may also be a learning system consisting only of the learning unit 20.
[0032] Next, we will explain in more detail each functional part of the information processing system 1 in the estimation phase. In the estimation phase, the estimation unit 10 operates. The estimation unit 10 estimates and outputs the results of the genetic information analysis. The estimation unit 10 estimates the results of the genetic information analysis based on the genetic information.
[0033] The genetic information acquisition unit 31 acquires the genetic information of the subject. Genetic information is, for example, DNA information extracted from the subject's blood, saliva, or various tissues. The genetic information acquisition unit 31 may also acquire genetic variation information from short read data generated by NGS technology, as shown in Figure 2. Figure 2 is a diagram illustrating the conversion to tokens. Hereinafter, genetic variation information will also be simply referred to as variation information. Short read data is data in which DNA is fragmented into units of several hundred bases.
[0034] The genetic information acquisition unit 31 maps each short-read data to a reference genome sequence and aligns the short-read data according to the reference genome sequence. For example, GRCh38.p14 provided by the Genome Reference Consortium is used as the human reference genome sequence. For aligning the short-read data, for example, Burrows-Wheeler Aligner (BWA) is used. BWA is a software package designed to perform efficient alignment on large reference genome sequences such as the human genome. BWA is based on Burrows-Wheeler Transform (BWT). BWT is a method for efficiently rearranging and compressing strings and is considered particularly effective for data with many repetitions. Therefore, BWA can rapidly match short-read data to large reference genome sequences.
[0035] The genetic information acquisition unit 31, for example, after aligning the short-read data, uses samtools to sort the BAM files and create an index. A BAM file is, for example, a SAM file compressed into binary format and is used to represent the aligned short-read sequence. Sorting and indexing these BAM files improves the accuracy of the data and the efficiency of the analysis, thereby improving the quality of the dataset creation.
[0036] The genetic information acquisition unit 31, for example, sorts BAM files by genome coordinates or sequence identifiers, and then generates an index file. The index file corresponds to, for example, a BAI file. This allows for the rapid acquisition of data overlapping specific locations, streamlining data access in subsequent analysis. The genetic information acquisition unit 31 extracts mutation information for a reference genome sequence from the BAM files and saves it in a VCF file. The genetic information acquisition unit 31 generates a whole-genome-wide VCF file. In the conversion process from BAM files to VCF files, for example, bcftools is used.
[0037] bcftools is suitable for detecting mutation information from BAM files and converting them to VCF files. For example, the extraction of mutation information from a BAM file is performed in two stages. First, the genetic information acquisition unit 31 generates a VCF file from the sample's BAM file using, for example, the mpileup command of bcftools. "Sample" corresponds to, for example, the genetic information of the subject.
[0038] The mpileup command can process a large number of read data simultaneously. Therefore, the stack of bases at each location is analyzed, and mutation information is detected. During this process, each chromosomal region is transformed individually, and the entire process is parallelized. As a result, the efficiency of the analysis is significantly improved.
[0039] The genetic information acquisition unit 31, for example, integrates the obtained VCF files for each chromosome using the concat command of bcftools. The concat command of bcftools integrates multiple VCF files into a single file while avoiding duplication between files. bcftools ensures the reproducibility and reliability of data transformation and improves the quality of data analysis.
[0040] The genetic information acquisition unit 31, for example, generates a whole-genome-wide VCF file using two shell script files. For example, the genetic information acquisition unit 31 generates a whole-genome-wide VCF file using a pipeline that converts short-read data into chromosome-specific mutation information and a pipeline that integrates the chromosome-specific mutation information into a single file. The processing of these two pipelines can be completed in a few hours per subject. Therefore, a relatively high-quality mutation dataset with unified conversion processing is generated.
[0041] The token conversion unit 32 converts the genetic information obtained by the genetic information acquisition unit 31 into multiple tokens. For example, the token conversion unit 32 converts the genetic information into multiple mutation tokens as character sequence information. For example, the token conversion unit 32 converts all mutation information of a human into 7.5 million mutation tokens. In this specification, "mutation token" is an integer value that represents a mutation.
[0042] For example, if the number of mutation tokens converted from the genetic information is less than 7.5 million, the token conversion unit 32 will fill in the difference with padding tokens. The padding tokens correspond to, for example, the integer value 0.
[0043] For example, as shown in Figure 2, the token conversion unit 32 extracts mutation information obtained from the genetic information acquisition unit 31 and converts them into corresponding integer representations. For example, the token conversion unit 32 converts each mutation information contained in the VCF file into a mutation token. As a result, the genetic mutation information is converted into a format that is easier for machines to handle, improving the ease of applying machine learning models.
[0044] For example, the token conversion unit 32 indicates a base mutation by a mutation token relating to a string of length 4 consisting of five letters: N, A, T, G, and C. 4 That is, mutation tokens are represented by 625 unique integer values, expressing changes in bases relative to the reference genome sequence. Homologous chromosomes are a pair of chromosomes, one inherited from the father and the other from the mother. Whether the mutation in a gene on these homologous chromosomes is homozygous or heterozygous can have different effects on an individual's phenotype. Homozygous means that the same mutation occurs in both chromosomes. Heterozygous means that different mutations occur in the two chromosomes.
[0045] The first two letters of the four-letter genetic code indicate a mutation in one chromosome, and the last two letters indicate a mutation in the other chromosome. For example, the string "AGAG" indicates that a homozygous single nucleotide substitution from A to G has occurred in both homologous chromosomes. The string "AGNN" indicates that a heterozygous single nucleotide substitution from A to G has occurred in one chromosome, and the other chromosome is unmutated. These strings uniquely identify specific mutation types, enabling accurate representation and analysis of gene mutations. Homozygous A insertions are represented by "NANA," and heterozygous A deletions by "NNAN."
[0046] For example, the token conversion unit 32 analyzes each mutation information in the VCF file and converts it into a corresponding string of length 4. This conversion is performed based on the reference base and mutation base described in the VCF file, and the conjugation of the mutation. Next, the token conversion unit 32 determines that the generated string is a base 5 number and converts it into a mutation token represented by a decimal integer value.
[0047] For example, the letters N, A, T, G, and C are assigned to integer values of 0, 1, 2, 3, and 4 respectively, and a string of length 4 is treated as a 4-digit base-5 number. In this case, for example, the string "AGAG" is judged as the base-5 number "1313" and converted to a mutation token represented by the decimal number "208". Therefore, mutation tokens are represented by integer values from 0 to 624 and function as integer codes to uniquely identify each mutation. This allows mutation information to be represented concisely and efficiently. Consequently, computational costs are reduced, data readability is improved, and compatibility with machine learning models is enhanced. In particular, the representation of mutations using integer values is a format that machine learning algorithms can easily handle, contributing to the efficiency of genome data processing and the accuracy of predictive modeling. Furthermore, the conversion process to mutation tokens can be completed in a few minutes per sample. In this specification, "one sample" corresponds, for example, to the genetic information of one subject.
[0048] The compression unit 33 compresses the multiple tokens converted by the token conversion unit 32 into multiple vectors. For example, the compression unit 33 divides the multiple tokens converted by the token conversion unit 32 into multiple groups and converts the multiple tokens contained in each group into a single vector. The compression unit 33 compresses the multiple tokens into a single vector using the first learning model. Figure 3 is a diagram illustrating the compression by the first learning model. The multiple variant tokens converted by the token conversion unit 32 are divided into multiple groups, and the first learning model, taking the multiple variant tokens from each group as input, outputs a single vector in which the multiple variant tokens input for each group are compressed. As a modification of this embodiment, the first learning model may output a single vector compressed for each group in response to the input of multiple variant tokens from multiple groups.
[0049] The compression unit 33 adds a context token to the beginning of multiple variant tokens input to the first learning model. The context token is a special token that always has a fixed value regardless of other variant tokens input. For example, the context token is represented by the integer value 626. For example, 1001 tokens including the context token are all represented by integer values. For example, the context token is obtained by converting 1000 variant tokens into a single vector through the processing of the first learning model. In the first learning model, during the processing of the self-attention mechanism, the inner product of the context token and all tokens including the context token is calculated. In the natural language processing model, variant tokens represented by integer values are converted into vectors with the same number of dimensions as words. The series of processes using context tokens is called embedding. Through the embedding process, multiple tokens are converted into vectors with the same number of dimensions and can be easily processed by the first learning model. In addition, embedding into multidimensional vectors increases the amount of information that can be handled, and the semantic and contextual information of each token is represented in a multidimensional space. Furthermore, by using representations within a unified vector space, the first learning model can capture relationships between tokens, such as similarity and dissimilarity. In other words, embedding is effective for learning complex concepts and contexts that cannot be represented by simple integer values, and embedding into vectors allows different input data to be handled in the same format.
[0050] For example, in the first learning model, the torch.nn.Embedding layer of PyTorch is used as the embedding layer to perform the embedding process. The embedding layer generates a dense vector of tokens by mapping integer values to a fixed-size vector. For example, the vocabulary size is set to 627, and the dimension of the embedding vector for each token is set to 1024. The embedding layer optimizes its uniquely defined token transformation method during the learning process of the first learning model. As a result, unknown properties such as similarity and oppositeness between variants are represented in the vector space.
[0051] If 1001 vectors (each with 1024 dimensions) are generated, these vectors cannot distinguish between vectors corresponding to the same type of mutation. For example, the 10th "AGAG" (a homozygous single nucleotide substitution from A to G) and the 100th "AGAG" are indistinguishable. Therefore, the compression unit 33 generates vectors corresponding to the absolute positions (0-1000) of all tokens, including the context token, and adds them to each token vector. The compression unit 33 distinguishes vectors corresponding to the same type of mutation by their absolute position relative to the context token. In other words, the compression unit 33 generates 1001 vectors with added positional information. Each of the generated vectors has a relationship with other mutation tokens.
[0052] The first learning model includes, for example, a self-attention mechanism stacked in 24 layers. For example, the self-attention mechanism is designed by SelfAttention(X)=MultiHead(X,X,X). For example, the first learning model includes Multi-Head Self Attention with 16 heads. In this case, the number of dimensions of the vector in each head is 64, which is obtained by dividing a 1024-dimensional vector into 16 parts.
[0053] Layer Normalization is effective when the size of a single data point is very large relative to the storage capacity. Layer Normalization performs normalization based on the features within each data point, independent of the batch size. The first learning model, for example, has a feed-forward neural network (FFNN) that operates independently at each location after each self-attention mechanism. The FFNN introduces nonlinearity into the processing of the first learning model, and can improve expressiveness and learning ability by capturing advanced data features that cannot be captured by simple linear transformations or attention mechanisms alone.
[0054] In an FFNN, for example, two linear transformations and two nonlinear transformations are performed. In an FFNN, for example, the first linear transformation, the first nonlinear transformation, the second linear transformation, and the second nonlinear transformation are performed in order. First, in the first linear transformation, for example, the input vector is projected onto a vector space expanded fourfold. Next, for example, in the first nonlinear transformation, the vector on which the first linear transformation was performed is subjected to a nonlinear transformation through the GEGLU function. Next, in the second linear transformation, projection onto the original dimensional space is performed again. Finally, for example, in the second nonlinear transformation, a nonlinear transformation using the ReLU function is performed once more.
[0055] In the first learning model, for example, cross-entropy is used as the loss function, and Adaptive momentestimation (Adam) is used as the optimization algorithm. Such a context-dependent approach more accurately models the effect of each mutation, for example, in cases where diverse mutations combine in complex ways to cause disease. The trained first learning model converts 1,000 consecutive mutation tokens into a single context vector. In other words, a context vector is a single vector containing information on 1,000 mutations. In the first learning model, for example, 7.5 million mutation tokens representing all mutation information for a human are converted into 7,500 context vectors. In other words, the compression unit 33 compresses, for example, 7.5 million tokens into 7,500 context vectors and outputs 7,500 context vectors. This conversion is completed in about 15 minutes per sample.
[0056] For training the first learning model, Masked Language Modeling (MLM) is used, for example. MLM is a task in which a portion of the tokens received as input are randomly masked, and the original tokens are predicted. Masking is done by replacing mutated tokens with masked tokens. Masked tokens are, for example, special tokens derived from the integer value 625. This training helps the model understand the context and learn the relationships between tokens. This allows the first learning model to predict not only a single mutated token, but also the context in which the mutated token exists.
[0057] During the learning process of the first learning model, the transformed vector is subjected to regularization, for example, by dropout. The dropout rate is set to, for example, 10%. By setting the dropout rate, excessive reliance on the output of some neurons is suppressed. For example, the output of the hidden layer is normalized by a torch.nn.LayerNorm layer before each of the 24 layers of self-attention mechanisms included in the first learning model. This suppresses the risk of gradient vanishing and gradient explosion during the learning process. This is extremely important for stabilizing and improving the efficiency of learning models with deep networks. In addition, within the FFNN during the learning process, for example, two dropout processes are performed after each nonlinear transformation, with probabilities set to 10% and 40%.
[0058] The analysis unit 34 analyzes the genetic information using a second learning model. The second learning model outputs the analysis results of the genetic information based on the input of the vector compressed by the compression unit 33. For example, the second learning model may output the analysis results of the genetic information corresponding to the input information based on the input of information indicating genetic information other than the vector compressed by the compression unit 33, and non-genetic information, in addition to the vector compressed by the compression unit 33. The second learning model is trained based on the vector relating to the genetic information and the analysis results of the genetic information. The second learning model may also be trained based on training information indicating genetic information and non-genetic training information. The training information indicating genetic information includes the vector compressed by the compression unit 33.
[0059] For example, the analysis unit 34 includes a second learning model that models the relationship between 7,500 context vectors (each 1024 dimensions) obtained from the compression unit 33 and diseases. The second learning model, for example, relates the 7,500 context vectors to diseases. The second learning model outputs information about diseases in response to the input of context vectors. The second learning model, for example, is designed based on Perceiver IO and can predict the risk of developing ALS from 7,500 context vectors.
[0060] Figure 4 shows the structure of the second learning model. The second learning model includes, for example, two cross-attention mechanisms and one self-attention mechanism. For example, in the second learning model, 7500 context vectors are in the initial latent array ∈ R 512×1024 It is projected onto the output, and contextual semantic information is added by the self-attention mechanism. For example, the shape of the output is the output query array ∈R 1×2 It depends on [the specific factor]. Therefore, it may be adapted to a multi-class classification system to suit the disease and traits to be modeled.
[0061] In the second learning model, the first cross-attention mechanism receives the input array and the latent array as input. The first cross-attention mechanism extracts features. For example, in the first cross-attention mechanism, the input array X ∈ R 7500×1024 However, the first learnable latent array (initial latent array ∈ R) 512×1024 It is projected and encoded. The input array is key "K" and value "V". For example, the latent array is query "Q". The latent array ∈ R represents the amount of information that the self-attention mechanism can directly handle, based on the input of key "K", value "V", and query "Q" to the cross-attention mechanism. 512×1024 This can be obtained. The input array and latent array may be subjected to embedding, for example, by a torch.nn.Embedding layer.
[0062] The potential array projected in the initial cross-attention mechanism is processed, for example, by a self-attention mechanism stacked in 24 layers. The self-attention mechanism analyzes genetic information. The settings of the self-attention mechanism are the same as those of the first learning model except for the following points. In the self-attention mechanism of the second learning model, for example, the number of vectors received as input is 512 instead of 1000, and the dropout rate in the learning process is all 40%.
[0063] In the second cross-attention mechanism in the second learning model, an input array and a potential array are also input. The second cross-attention mechanism outputs an analysis result of genetic information in response to the input of the output result from the self-attention mechanism. For example, the second cross-attention mechanism outputs a diagnostic result of a subject having a target format as an analysis result of genetic information. For example, in the second cross-attention mechanism, the latent array ∈R 512×1024 processed by the self-attention mechanism is projected onto a query array ∈R 1×2 for learnable output, and an output ∈R 1×2 is generated.
[0064] For example, before and after the three attention mechanisms in the second learning model, normalization by the torch.nn.LayerNorm layer and non-linear transformation by FFNN are performed. These attention mechanisms greatly improve the expressiveness when the second learning model predicts the onset risk of diseases from various mutant data. For example, in the second learning model, by using the cross-attention mechanism for 7500 context vectors, information over a wide range of the input array is efficiently integrated into the potential array, and through the subsequent self-attention mechanism, it becomes possible to capture the complex relationships between the information of the input array.
[0065] In the second learning model, by adding position information, the continuity of the context vectors and the relative position information are retained. Therefore, the second learning model can make predictions considering the spatial relationships between mutations, which is effective when multiple mutations far apart affect the onset risk of diseases.
[0066] The storage unit 35 pre-stores information used by each functional unit. The storage unit 35 stores the output from each functional unit. The storage unit 35 stores, for example, information acquired by the genetic information acquisition unit 31. The storage unit 35 stores, for example, tokens converted by the token conversion unit 32. The storage unit 35 stores, for example, vectors compressed by the compression unit 33. The storage unit 35 stores, for example, the analysis results from the analysis unit 34. The storage unit 35 stores the first learning model and the second learning model. The storage unit 35 stores, for example, information to be output from the output unit 36.
[0067] The output unit 36 outputs diagnostic information of a subject who has genetic information. The output unit 36 may also output information analyzed by the analysis unit 34.
[0068] Next, we will describe in more detail each functional part of the information processing system 1 in the learning phase. Some explanations that overlap with the estimation phase will be omitted. In the learning phase, the learning unit 20 operates. In the example shown in this embodiment, the learning unit 20 creates a second learning model using the learning information acquisition unit 41, the token conversion unit 32, the compression unit 33, and the learning model creation unit 44. The learning unit 20 creates a second learning model that has been trained using a learning dataset containing various types of information. In a modified example of this embodiment, the learning unit 20 may create the second learning model using the learning information acquisition unit 41 and the learning model creation unit 44.
[0069] In the example shown in this embodiment, the learning unit 20 acquires a learning dataset for training a second learning model by at least one of the learning information acquisition unit 41, the token conversion unit 32, the compression unit 33, and the learning model creation unit 44. As a modification of this embodiment, the learning unit 20 may acquire a learning dataset that has been previously stored in the storage unit 35, or it may acquire a learning dataset from outside the information processing system 1 via a communication network or the like.
[0070] When the learning unit 20 is operating, in the example shown in this embodiment, the learning information acquisition unit 41 acquires learning genetic information obtained from multiple learning subjects. The learning information acquisition unit 41 has the same function as, for example, the genetic information acquisition unit 31. The learning information acquisition unit 41 and the genetic information acquisition unit 31 may be physically the same part.
[0071] As a modified example of this embodiment, the learning information acquisition unit 41 acquires a learning dataset used for training the second learning data. The learning dataset acquired by the learning information acquisition unit 41 includes learning genetic information acquired from multiple training subjects. In other words, the learning dataset acquired by the learning information acquisition unit 41 includes multiple learning tokens acquired from multiple training subjects.
[0072] The token conversion unit 32 converts the learning genetic information obtained by the learning information acquisition unit 41 into multiple learning tokens. For example, the token conversion unit 32 converts the learning genetic information into multiple learning mutation tokens as character sequence information. For example, if the number of learning mutation tokens converted from the learning genetic information is less than 7.5 million, the token conversion unit 32 fills the gap with padding tokens. The token conversion unit 32 extracts the mutation information obtained from the learning information acquisition unit 41 and converts them into corresponding integer values.
[0073] The compression unit 33 compresses the multiple tokens converted by the token conversion unit 32 into multiple vectors. The compression unit 33, for example, divides the multiple tokens converted by the token conversion unit 32 into multiple groups and converts the multiple tokens in each group into a single vector. The compression unit 33 compresses the multiple training tokens into a single training context vector using the first learning model. The multiple training variant tokens converted by the token conversion unit 32 are divided into multiple groups, and the first learning model, taking the input of multiple training variant tokens for each group, outputs a single training context vector in which the input training variant tokens for each group have been compressed. As a modification of this embodiment, the first learning model may output a single training context vector compressed for each group in response to the input of multiple training variant tokens for multiple groups. For example, the compression unit 33 compresses 7.5 million training variant tokens into 7,500 training context vectors and outputs 7,500 training context vectors. For example, the compression unit 33 creates a training dataset containing 7,500 training context vectors.
[0074] In the example shown in this embodiment, the first learning model is a pre-trained model and is obtained from the storage unit 35. As a variation of this embodiment, the first learning model may be trained using a plurality of training mutation tokens obtained from the token conversion unit 32.
[0075] The learning model creation unit 44 creates a second learning model using a learning dataset containing multiple learning context vectors. In other words, the learning model creation unit 44 creates a second learning model using learning context vectors corresponding to multiple learning genetic information. The second learning model is trained using learning context vectors corresponding to multiple learning genetic information.
[0076] Similar to CEU, cross-entropy is used as the loss function and Adam is used as the optimization algorithm for training the second learning model. This can minimize the difference between the output of the second learning model and the actual labels.
[0077] The learning model creation unit 44 may recreate or update the second learning model based on the information acquired in the estimation phase. For example, the second learning model is repeatedly learned based on a vector relating to genetic information and the results of the analysis of genetic information by the analysis unit 34. The second learning model is learned based on first learning information representing genetic information and non-genetic second learning information. The first learning information includes a vector compressed by the compression unit 33. The first learning information may also include information representing genetic information other than the vector compressed by the compression unit 33. The second learning information is additional information other than the first learning information. The second learning information is non-genetic information representing information other than genetic information, such as acquired information such as lifestyle habits. Acquired information includes, for example, age, height, weight, diet, and exercise habits. At least one of the learning information acquisition unit 41 and the learning model creation unit 44 may be excluded from the information processing system 1 after the completion of the learning phase.
[0078] The storage unit 35 stores the first learning model and the second learning model created by the learning model creation unit 44. The storage unit 35 stores, for example, the learning dataset used for training the second learning model. The storage unit 35 may have already stored the learning dataset used for training the second learning model. In a modified version of this embodiment, the storage unit 35 may store the learning dataset used for training the first learning model.
[0079] Next, the hardware configuration of the information processing system 1 will be described with reference to Figure 5. Figure 5 is a diagram showing an example of the hardware configuration of the information processing system 1.
[0080] Information processing system 1 comprises a processor 101, memory 102, storage 103, communication device 104, output device 105, and input device 106. Information processing system 1 includes one or more computers composed of this hardware and software such as programs. Information processing system 1 may consist of one computer or multiple computers. Information processing system 1 is realized in cooperation with the hardware. In this specification, "computer" refers to an electronic device for processing, storing, and communicating data. Specifically, it includes, but is not limited to, a central processing unit (CPU), memory (RAM, ROM, etc.), storage devices (hard disk, SSD, etc.), input devices (keyboard, mouse, etc.), output devices (display, printer, etc.), and communication interfaces (network card, USB port, etc.). "Computer" also includes an application-specific integrated circuit (ASIC) designed to perform a specific task. A computer can perform a specific task by executing a software program.
[0081] If the information processing system 1 is composed of multiple computers, these computers may be connected locally or via a communication network such as the Internet or an intranet. This connection logically constructs a single information processing system 1.
[0082] The processor 101 executes the operating system and application programs. The memory 102 consists of ROM (Read Only Memory) and RAM (Random Access Memory). For example, at least some of the various functional units of the genetic information acquisition unit 31, token conversion unit 32, compression unit 33, analysis unit 34, storage unit 35, output unit 36, learning information acquisition unit 41, and learning model creation unit 44 can be implemented by the processor 101 and the memory 102.
[0083] The storage 103 is a storage medium composed of a hard disk and flash memory, among other things. Generally, the storage 103 stores a larger amount of data than the memory 102. For example, at least a portion of the storage unit 35 can be realized by the storage 103.
[0084] The communication device 104 is comprised of a network card or a wireless communication module. For example, at least a portion of the genetic information acquisition unit 31, the output unit 36, and the learning information acquisition unit 41 can be implemented by the communication device 104. The output device 105 is comprised of a printer and a display, etc. For example, at least a portion of the output unit 36 can be implemented by the output device 105.
[0085] The input device 106 consists of a keyboard, mouse, and touch panel, among other things. For example, at least a part of the genetic information acquisition unit 31, the output unit 36, and the learning information acquisition unit 41 can be implemented by the input device 106.
[0086] Storage 103 pre-stores the program and data necessary for processing. This program causes the computer to execute each functional element of the information processing system 1. This program allows the computer to execute, for example, each process in the state identification method described later. This program may be provided on a tangible recording medium such as a CD-ROM, DVD-ROM, or semiconductor memory. This program may also be provided as a data signal via a communication network.
[0087] Next, an example of how to use the analysis system in this embodiment will be described with reference to Figures 6 and 7. For example, the method of using the analysis system includes an information processing method for analyzing genetic information and a method for creating a second learning model. In the method for creating the second learning model, a learning phase is performed to create the second learning model. In the information processing method for analyzing genetic information, an estimation phase is performed to analyze the genetic information using the created second learning model. First, an example of how to create the second learning model in the learning phase will be described with reference to Figure 6. Figure 6 is a flowchart showing an example of how to create the second learning model.
[0088] First, genetic information for learning is obtained from multiple training subjects (process S1). In process S1, for example, the learning information acquisition unit 41 acquires genetic information for learning from multiple training subjects.
[0089] Next, the learning genetic information obtained in process S1 is converted into multiple learning tokens (process S2). In process S2, for example, the token conversion unit 32 converts the learning genetic information obtained from multiple training subjects into multiple learning mutation tokens.
[0090] Next, the first learning model compresses multiple training mutation tokens into a single training context vector (process S3). In process S3, for example, the compression unit 33 compresses multiple training tokens into a single training context vector using the first learning model.
[0091] Next, a second learning model is created using a learning dataset containing multiple learning context vectors (process S4). In process S4, for example, the learning model creation unit 44 creates a second learning model using multiple learning context vectors corresponding to the genetic information of multiple training subjects. Next, the second learning model created in process S4 is stored in the storage unit 35 (process S5).
[0092] When process S5 is completed, the series of processes for creating the second learning model is finished. The above describes one example of a method for creating a second learning model, but the order of each process is not limited to this. Instead of processes S1 to S3, the learning dataset may be obtained from outside the information processing system 1, or it may be stored in the storage unit 35 in advance.
[0093] Next, with reference to Figure 7, an example of an information processing method in the estimation phase will be described. Figure 8 is a flowchart of an example of an information processing method.
[0094] First, genetic information obtained from the subject is acquired (process S21). In process S21, for example, the genetic information acquisition unit 31 acquires the subject's genetic information.
[0095] Next, the genetic information obtained in process S21 is converted into multiple tokens (process S22). In process S22, for example, the token conversion unit 32 converts the genetic information obtained from multiple subjects into multiple mutation tokens.
[0096] Next, the first learning model compresses multiple mutation tokens into a single context vector (process S23). In process S23, for example, the compression unit 33 compresses multiple mutation tokens into a single context vector using the first learning model.
[0097] Next, the subject's genetic information is analyzed by the second learning model (process S24). In process S24, for example, the analysis unit 34 analyzes the subject's genetic information using the second learning model, which outputs the analysis results of genetic information based on the input of a context vector.
[0098] Next, the analysis results from process S24 are output (process S25). The output unit 36 converts the information analyzed by the analysis unit 34 and outputs diagnostic information of the subject who has genetic information.
[0099] Once process S25 is completed, the series of processes of the analysis method are finished. The above is an example of an analysis method, but the order of each process is not limited to this.
[0100] Next, the effects of the information processing system 1, analysis method, and learning model creation method in the above-described embodiment will be explained.
[0101] In information processing system 1, multiple tokens into which genetic information has been transformed are compressed into a single vector by the first learning model, and the genetic information is then analyzed by the second learning model. The second learning model outputs the analysis result of the genetic information based on the input of the compressed vector. In this case, the computational load of genome analysis can be reduced. Therefore, for example, comprehensive analysis of the entire genome can be realized.
[0102] In the information processing system 1, the token conversion unit 32 may convert genetic information into multiple tokens as character sequence information. In this case, the computational load of genome analysis can be further reduced.
[0103] In information processing system 1, the first learning model may have a transformer architecture. In this case, the genetic information can be further compressed.
[0104] In information processing system 1, the second learning model includes a cross-attention mechanism. In this case, the computational load of genome analysis can be further reduced, and the accuracy of the genetic information analysis results can be further improved.
[0105] In the information processing system 1, the second learning model may further include a self-attention mechanism for analyzing genetic information. The cross-attention mechanism may output diagnostic information of a subject with genetic information in response to the input of the output result from the self-attention mechanism. In this case, the variety of formats of the diagnostic information output from the second learning model may be improved.
[0106] In the information processing system 1, the second learning model may, upon input of compressed vectors and non-genetic information, output the results of genetic information analysis corresponding to the compressed vectors and non-genetic information. In this case, the accuracy of the genetic information analysis results output from the second learning model can be further improved.
[0107] In the information processing system 1, the second learning model may be trained based on first learning information representing genetic information and second learning information representing non-genetic information. In this case, the accuracy of the analysis results of the genetic information output from the second learning model can be further improved.
[0108] In the information processing system 1, the second learning model may be trained based on a vector relating to genetic information and the results of the analysis of genetic information by the analysis unit 34. In this case, the accuracy of the analysis results of genetic information output from the second learning model can be further improved.
[0109] Next, with reference to Figures 8 to 10, an example of verification of the analysis method using the information processing system 1 and processes S1 to S25 described above will be explained. Hereinafter, the analysis method using the information processing system 1 and the method using processes S1 to S25 described above will be referred to as the "proposed method."
[0110] In this validation, the R2522 and KALS data sets are used as disease group datasets. The SM, TMGH, and HEV data sets are used as control group datasets. The R2522 group consists of data collected from the JaCALS (Japanese Consortium for ALS Research) cohort and is derived from sporadic ALS patients collected from all over Japan. The R2522 group contains genomic data from a total of 99 patients. The SM group is a genetic disease cohort collected from various facilities in Japan, and recruits a variety of diseases, including familial ALS. The SM group contains genomic data from a total of 10 patients. The TMGH group consists of data collected from the Tokyo Metropolitan Institute of Gerontology cohort and includes healthy individuals and a variety of diseases other than ALS. The TMGH group contains genomic data from a total of 31 patients. The KALS group consists of data collected from the Keio University Hospital cohort and is derived from sporadic ALS patients. The KALS group contains genome data from a total of 37 individuals. The HEV group was provided by the RIKEN BioResource Research Center (RIKEN BRC) and is a healthy control cohort. The HEV group contains genome data from a total of 28 individuals.
[0111] For each sample, the maximum number of mutations was 7,316,289, the minimum was 6,090,338, the mean was 6,656,020.80, and the standard deviation was 204,307. Note that the mutation counts shown here represent the number of mutation tokens excluding padding tokens, not the number of mutations contained in the VCF file.
[0112] Figure 9 shows the performance of the trained second-trained model on training data, test data, and unknown cohort data as a confusion matrix. In Figure 9, the labels in the confusion matrix are 1 for the case group and 0 for the control group. The inference results are arranged along the horizontal axis (X axis), and the actual labels are arranged along the vertical axis (Y axis). In Figure 9, a darker color indicates a higher density of data. [A] shows the performance on the training data, [B] shows the performance on the test data, and [C] shows the performance on data belonging to the unknown cohort.
[0113] The performance of the second trained model on the training dataset was very high. Accuracy, precision, and recall all reached 100%, indicating that the second trained model was perfectly fitted to the training dataset. Performance on the test data was slightly lower than on the training dataset. However, extremely high results were still observed, with accuracy at 92%, precision at 100%, and recall at 83%. These results may suggest overfitting to the training data, and may also be influenced by the size and diversity of the dataset. Furthermore, for data from an unknown cohort (consisting only of cases), recall was 64%, as shown by the confusion matrix results.
[0114] Next, we will explain the effect of additional training to adapt the second learned model to an unknown cohort. In Figure 10, the labels in the mixture matrix are 1 for the case group and 0 for the control group. The inference results are arranged along the horizontal axis (X axis), and the actual labels are arranged along the vertical axis (Y axis). Note that darker colors indicate a higher density of data. (A)(B)(C)(D)(E) show the performance when newly trained with 2, 4, 6, 8, and 10 data points, respectively. [1] and [2] show the performance for the new cohort and the performance for the original training data, respectively. In Figure 10, [A-1] to [E-1] show how accurately the second learned model after additional training can predict the samples belonging to the new cohort targeted for additional training. In Figure 10, [A-2] to [E-2] show the impact of additional training on the performance for the original training data. As shown in [A-1] through [C-1], when samples from new cohorts of 2, 4, and 6 were used for additional training, the recall rate of the second trained model was 99%, demonstrating a relatively high recall rate.
[0115] In [A-2] through [C-2], the performance on the original training data is shown to be 89%, 89%, and 91%, respectively. However, as shown in [D-1] and [E-1], when 8 and 10 samples were used for additional training, the recall of the second trained model was 89% and 82%, respectively. As shown in [D-2] through [E-2], the performance on the training data was 98% in all cases, which is an extremely high value.
[0116] The visualized data file includes attention scores for each individual sample, attention scores for the disease group and the control group, and the attention score for all samples combined from both groups. This data allows for the assessment of the impact of specific chromosomal regions on the disease. Therefore, this data can be used for disease prediction.
[0117] When the attention scores of two different samples from a randomly selected group are visualized, it can be confirmed that different chromosomal regions have a significant impact on the disease. For example, in the control group, the region between positions 12 million and 13 million on chromosome 19 was found to have a significant impact on the disease, while in the diseased group, the region between positions 32 million and 33 million on chromosome 20 was found to have a significant impact.
[0118] When the attention scores for both the disease group and the control group were visualized, it was confirmed that the Y chromosome and the region of chromosome 1 from position 221 million to 222 million had a significant impact on the disease in both groups. Furthermore, it was confirmed that chromosomes 8 through 12 had a significant impact on the disease in both groups.
[0119] When the attention scores for all samples from both groups were visualized, it was confirmed that the Y chromosome, the latter half of the X chromosome, the region of chromosome 1 from position 221 million to 222 million, and chromosomes 8 to 12 had a significant impact on the disease. Figure 10 shows the attention scores for all samples from both groups. In Figure 10, darker colors indicate higher attention scores.
[0120] In this validation, the fill-in-the-blank accuracy of CEUs trained using MLM on the training data converged to approximately 46%. This result is sufficiently high considering the complexity and diversity of genetic variation. This accuracy is extremely important for the application of MLM in variation analysis for the following reasons:
[0121] When tokenizing genetic variations, each variation token represents a single-nucleotide genetic variation and its conjugation. With 625 variation token patterns, a prediction accuracy of 46% for a specific masked variation token indicates that CEU effectively captures genetic information and accurately encodes it using variation tokens.
[0122] Many genetic variations, even if they are different from each other, do not affect biological function. For example, synonymous variations do not directly affect the function of the proteins they synthesize. Synonymous variations are variations that code for the same amino acid even if the codon changes. Therefore, an accuracy of 46% is considered a practical level for predicting genetic variations.
[0123] When genetic information is analyzed using context vectors obtained from the first learning model, extremely high accuracy is achieved. This result indicates that the token-based representation of genetic information obtained from the first learning model accurately captures actual biological phenomena.
[0124] The classification performance of the second learning model on the test data is extremely high. Therefore, for the following reasons, the second learning model is effective as a clinical diagnostic support tool and a tool for investigating the causes of genetic diseases.
[0125] The cause of sporadic ALS is not yet clear at the genetic level. For example, the explanatory power of known causative genes for sporadic ALS is currently only about 15%. The high classification performance of the second learning model on the test data means that the information processing system 1 using the second learning model can effectively capture disease-specific patterns within an unclear pathological mechanism.
[0126] For 13 unknown samples, the second learning model demonstrated extremely high results, achieving accuracy of 92%, precision of 100%, and recall of 83%. This indicates that the second learning model possesses strong predictive capabilities even for unknown data. In particular, a precision of 100% means that healthy individuals were never misdiagnosed as having the disease.
[0127] All data used in this validation were from a sample of Japanese individuals. Approximately half of the data used in this validation was from ALS patients. The fact that the second learning model demonstrated high classification performance even with limited data diversity indicates that the second learning model is independent of specific populations and disease states.
[0128] Regarding unknown cohorts, the adaptability of the second-learned model improved with increasing sample size for additional training. The larger the sample size used for additional training, the more flexibly the second-learned model could adapt to unknown cohorts while retaining knowledge of a certain amount of training data. In additional training using a relatively small number of samples, the recall rate of the second-learned model for the new cohort remained high. This suggests that the second-learned model retained knowledge of the original training data while learning the characteristics of the new cohort. In fact, in training with a small number of samples that showed high recall for the new cohort, the judgment performance of the control group included in the original training data dropped drastically.
[0129] Next, we examine the case where 7500 context tokens are encoded by the second learning model. In Figure 9, the attention score matrix of the first cross-attention mechanism is visualized. Concentrations of attention scores were observed in the Y chromosome, the latter half of the X chromosome, the region of chromosome 1 from position 221 million to 222 million, and chromosomes 8 to 12. In Figure 9, the attention scores of all samples are visualized for known ALS causative gene regions. The locations of four major ALS-related genes—C9orf72, SOD1, TARDBP, and FUS—were investigated. The correspondence between the location information of the GRCh38.p14 reference genome sequence used in this validation and the location information of each gene is publicly available from GENCODE. The Comprehensive gene annotation dataset (Release 44), which covers the entire region, was used. This revealed that C9orf72, SOD1, TARDBP, and FUS are located at positions 27,535,640 to 27,573,866 on chromosome 9, positions 31,659,666 to 31,668,931 on chromosome 21, positions 11,012,344 to 11,030,528 on chromosome 1, and positions 31,180,138 to 31,191,605 on chromosome 16, respectively.
[0130] When examining the attention score matrix at the locations of the four genes investigated, it was confirmed that at the location of C9orf72, attention scores were concentrated in the region behind the gene, including the gene region itself. Mutations in C9orf72 were the most common genetic cause, accounting for 25% to 40% of familial ALS cases and approximately 6% of sporadic ALS cases. The concentration of attention scores in the C9orf72 region suggests that appropriate genetic cause identification has been achieved. Furthermore, in Figure 10, a broad concentration of attention scores was observed in the region surrounding C9orf72. This surrounding region is thought to correspond to the region related to transcription factors that regulate C9orf72 expression, or to the enhancer region.
[0131] Furthermore, a concentration of attention scores was observed around SOD1. This SOD1 mutation is associated with approximately 20% of ALS cases and is the second most common gene in familial ALS. This further confirms the validity of the analysis method using the second learning model.
[0132] In Figure 10, attention scores are concentrated across the entire Y chromosome region. Since the Y chromosome plays a role in determining maleity, this suggests that sex differences are important in the risk of developing ALS. In fact, the male-to-female ratio of ALS incidence is 3:2, indicating a slight tendency for men to be affected.
[0133] Information Processing System 1 makes it possible to comprehensively model the genetic background of complex multifactorial diseases that cannot be captured by single gene mutations, across the entire genome. Information Processing System 1 is expected to greatly advance our understanding of hereditary and intractable diseases such as ALS and make a significant contribution to elucidating the genetic factors involved.
[0134] Furthermore, Information Processing System 1 is applicable not only to ALS but to all hereditary diseases and traits. Information Processing System 1 enables a comprehensive and contextual understanding of the impact of genetic diversity on disease onset and trait expression, potentially establishing a new standard for genetic research.
[0135] Information Processing System 1 is considered a powerful tool, particularly for analyzing complex traits involving multiple genes, and the interactions between environmental and genetic factors. The realization of whole-variant analysis using Information Processing System 1 will revolutionize methodologies in genetic research, enabling the exploration of genetic factors with unprecedented precision and resolution. The versatility and adaptability of this technology will lead to the elucidation of unknown genetic diseases, the identification of new therapeutic targets, and the realization of personalized medicine.
[0136] Information processing system 1 applies natural language processing technology to extract the context of variant sequences, for example, converting 1,000 consecutive variants into a single context vector. Information processing system 1 can represent all variant information of a single subject with 7,500 context vectors, enabling the introduction of advanced analysis using a cross-attention mechanism.
[0137] For example, Information Processing System 1 achieved an accuracy of 92%, precision of 100%, and recall of 83% for 13 unknown samples collected from the same cohort as the training data, and achieved a recall of 82% for 88 samples collected from an unknown cohort with an additional 10 samples. In this way, the genetic background of complex multifactorial diseases that cannot be captured by single gene mutations was comprehensively modeled across the entire genome.
[0138] Furthermore, in information processing system 1, it was confirmed that the attention scores of the second learning model were concentrated on specific gene regions, particularly C9orf72 and SOD1. At the location of C9orf72, the attention scores were concentrated in the posterior region containing this gene region. Mutations in C9orf72 are a major genetic cause accounting for many cases of familial ALS, and the second learning model appropriately paid attention to this important gene region, effectively functioning in the search for the genetic cause. Therefore, the second learning model is useful for effectively analyzing the context of genetic variation and identifying the genetic causes of hereditary diseases.
[0139] Because information processing system 1 can handle all mutations, the quantity and quality of information obtained from genomic data have been significantly improved. As a result, the accuracy of analysis, particularly for non-coding regions and complex multifactorial diseases, has been improved.
[0140] Information processing system 1 can also contribute to personalized medicine. For example, information processing system 1 promotes a deeper understanding of the relationship between genetic variations and diseases, enabling more precise diagnosis, the formulation of more effective treatment strategies, and applications in preventive medicine. For instance, integration with protein structure prediction models such as AlphaFold enables the design of drugs tailored to an individual's genetic characteristics on a computer. Personalized drug discovery makes it possible to propose more practical medical technologies.
[0141] The genome data analysis technology provided by Information Processing System 1 is expected to be applied to the analysis of various unexplained genetic diseases, including ALS. For example, it is thought that unknown genetic causes may be revealed, leading to the identification of new therapeutic targets and a better understanding of disease mechanisms.
[0142] While embodiments and modifications of the present invention have been described above, the present invention is not necessarily limited to the embodiments and modifications described above, and various modifications are possible without departing from the spirit of the invention. The numerical values shown in the embodiments and modifications of the present invention are merely examples. The information processing system of the present invention can be applied not only to the medical field but also to other fields such as agriculture, environmental science, and forensic medicine. Even if an element described in the claims of the present invention is described in the singular form, it shall be interpreted as including the plural form. Furthermore, each element described in the claims is included in the scope of the present invention whether it is used in combination with other elements or used alone. [Explanation of Symbols]
[0143] 1... Information processing system, 32... Token conversion unit, 33... Compression unit, 34... Analysis unit.
Claims
1. A token conversion unit that converts genetic information into multiple tokens, A compression unit that compresses the plurality of tokens into a single vector using a first learning model, An information processing system comprising: an analysis unit that analyzes the genetic information using a second learning model that outputs the results of analyzing the genetic information based on the input of the aforementioned vector; and an analysis unit that analyzes the genetic information.
2. The information processing system according to claim 1, wherein the token conversion unit converts the genetic information into character sequence information and into the plurality of tokens.
3. The information processing system according to claim 1, wherein the first learning model has a transformer architecture.
4. The information processing system according to claim 1, wherein the second learning model includes a cross-attention mechanism.
5. The second learning model further includes a self-attention mechanism for analyzing the genetic information, The information processing system according to claim 4, wherein the cross-attention mechanism outputs diagnostic information of a subject having the genetic information in response to the input of the output result from the self-attention mechanism.
6. The information processing system according to claim 1, wherein the second learning model outputs the results of analyzing the genetic information corresponding to the vector and the non-genetic information, based on the input of the vector and the non-genetic information.
7. The information processing system according to claim 6, wherein the second learning model is learned based on first learning information representing genetic information and second learning information representing non-genetic information.
8. The information processing system according to claim 1, wherein the second learning model is learned based on the vector relating to the genetic information and the results of the analysis of the genetic information by the analysis unit.
9. Converting genetic information into multiple tokens, The first learning model compresses the multiple tokens into a single vector, An information processing method comprising: analyzing the genetic information using a second learning model that outputs the results of analyzing the genetic information based on the input of the aforementioned vector.
10. A method for creating a learning model, which involves creating a learning vector obtained by compressing multiple learning tokens converted from learning genetic information acquired from multiple training subjects, and using this vector as input to create a learning model that outputs the results of analyzing the genetic information.
Citation Information
Patent Citations
A multi-omics search engine for integrated analysis of cancer genetic and clinical data
JP2022504916A