Information processing program, information processing method, and information processing device
Patent Information
- Application Number
- JP2025560440
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-01
AI Technical Summary
Conventional virus mutation prediction methods fail to accurately reflect the influence between structurally distant amino acids and the differences in properties at positions with the same amino acid name, leading to decreased prediction accuracy.
An information processing program that determines the weight of input feature amounts in a regression model predicting the amino acid sequence of a virus after mutation, using feature amounts related to the three-dimensional structure of the virus protein and the contribution degree to the prediction of a machine learning model.
Improves the accuracy of virus mutation prediction by considering the three-dimensional structure and property contributions of amino acids, enabling more precise forecasting of amino acid sequence changes.
Abstract
Description
Information processing program, information processing method, and information processing device
[0001] The present invention relates to an information processing program, an information processing method, and an information processing device.
[0002] Because viruses mutate repeatedly, predicting mutations is important in developing vaccines for coronaviruses and other viruses.
[0003] Conventionally, viral proteins have been treated as amino acid sequences, and time series analysis has been performed to correlate these with epidemic periods, or LSTM (Long Short-Term Memory) has been used to predict the amino acid sequences of mutated viruses.
[0004] International Publication No. 2022 / 019331 Patent Publication No. 2022-521686 U.S. Patent Application Publication No. 2012 / 0265513 Patent Publication No. 2022-527381 U.S. Patent Application Publication No. 2019 / 0266493
[0005] However, such conventional methods for predicting viral mutations cannot reflect the influence of structurally distant amino acids or the differences in properties of amino acids with the same name at different positions within the virus.
[0006] For example, even if a virus has the same chemical formula, there are often cases where it has different shapes or properties, such as isomers. Conventional virus mutation prediction methods have difficulty tracking these variations, which leads to a problem of reduced accuracy in predicting virus mutations.
[0007] In one aspect, the present invention aims to improve the accuracy of predicting viral mutations.
[0008] Therefore, this information processing program causes a computer to execute a process of determining the weight of an input feature in a regression model that uses the amino acid sequence of a virus as an input feature and predicts the amino acid sequence of the virus after mutation, based on a first feature related to the three-dimensional structure of the virus's protein and a second feature related to the contribution to the prediction of the machine learning model obtained based on the first feature.
[0009] According to one embodiment, the accuracy of virus mutation prediction can be improved.
[0010] 1 is a diagram schematically showing the configuration of an information processing device according to an embodiment; FIG. 2 is a diagram illustrating an example of amino acid sequence and antigen cluster name information used in an information processing device according to an embodiment; FIG. 3 is a block diagram illustrating an example of the hardware configuration of a computer that realizes the functions of an information processing device according to an embodiment; FIG. 4 is a diagram illustrating an example of amino acid three-dimensional structure information output by a three-dimensional structure calculation processing unit in an information processing device according to an embodiment; FIG. 5 is a diagram illustrating an example of chemical parameter information created by a chemical parameter calculation processing unit in an information processing device according to an embodiment; FIG. 6 is a diagram illustrating graph information in an information processing device according to an embodiment; FIG. 7 is a diagram illustrating the processing of a graph data shaping processing unit in an information processing device according to an embodiment; FIG. 8 is a diagram illustrating graph AI input information in an information processing device according to an embodiment; FIG. 9 is a diagram illustrating statistical information in an information processing device according to an embodiment; FIG. 10 is a diagram illustrating the processing of a graph AI calculation processing unit in an information processing device according to an embodiment; FIG. 11 is a diagram illustrating weight vector information in an information processing device according to an embodiment; FIG. 12 is a flowchart illustrating the processing in a training phase in an information processing device according to an embodiment; FIG. 13 is a flowchart illustrating the processing of a graph AI calculation processing unit in an information processing device according to an embodiment.
[0011] Hereinafter, embodiments of the present information processing program, information processing method, and information processing device will be described with reference to the drawings. However, the embodiments shown below are merely examples, and are not intended to exclude various modifications or application of techniques not explicitly stated in the embodiments. In other words, the present embodiment can be implemented with various modifications within the scope of its purpose. Furthermore, each figure does not intend to include only the components shown in the figure, but may include other functions, etc.
[0012] (A) Configuration FIG. 1 is a diagram schematically showing the configuration of an information processing device 1 according to an embodiment.
[0013] The information processing device 1 performs training (machine learning) of a regression model 110 that predicts the amino acid sequence of a viral protein after mutation (training phase).
[0014] In the training phase, the information processing device 1 receives input of the amino acid sequence of a virus at a certain time in the past and the name of an antigen cluster, and the amino acid sequence of the virus after mutation is used as the correct answer data.
[0015] A virus from a certain point in the past may simply be called a past virus. Furthermore, the amino acids contained in this past virus may also be called past amino acids. An antigen cluster name may simply be called a cluster name. Furthermore, the amino acid sequence and antigen cluster name of a past virus may also be called the past amino acid sequence and antigen cluster name.
[0016] Furthermore, in the information processing device 1, the inference unit 106 uses the trained regression model 110 to predict (infer) the amino acid sequence of the virus protein after mutation (prediction phase).
[0017] In the prediction phase, the current (latest) amino acid sequence of the virus is input to the information processing device 1, and the regression model 110 predicts the post-mutation amino acid sequence of the virus. In the prediction phase, the post-mutation amino acid sequence predicted by the regression model 110 based on the input current (latest) amino acid sequence of the virus and the antigen cluster name may be referred to as the future amino acid sequence.
[0018] FIG. 2 is a diagram illustrating an example of amino acid sequence and antigen cluster name information used in the information processing device 1 according to an embodiment.
[0019] 2, the amino acid sequence and antigen cluster name information are shown in the form of a data table. Hereinafter, the amino acid sequence and antigen cluster name information may be represented by adding the symbol T1.
[0020] The amino acid sequence and antigen cluster name information T1 shown in Fig. 2 shows a correspondence between a number, a cluster name, a date, and an amino acid name.
[0021] In the amino acid sequence and antigen cluster name information T1 illustrated in Fig. 2, each piece of data is shown as a character string for convenience, but in practice it may be a uniquely linked integer value or the like. By expressing the data as an integer value, it can be used efficiently in various calculations and is highly convenient. The same applies to other information described below.
[0022] No. is information that identifies the virus. The cluster name is the antigen cluster name of the virus. The date may be the date and time when the virus appeared or was discovered. The amino acid name indicates the type of amino acid contained in the virus, and represents one of 20 types of amino acids. In Figure 2, for convenience, the amino acid names (types of amino acids) are represented using letters such as D and N.
[0023] If a virus contains multiple amino acids, the names of the multiple amino acids may be listed in the amino acid sequence and antigen cluster name information T1 in accordance with the virus. The amino acid names may be listed, for example, in the order of peptide bonds from beginning to end.
[0024] The multiple amino acids contained in the virus may be represented by numbers. The numbers representing the amino acids contained in the virus may be called amino acid numbers. In the example shown in Figure 2, the amino acid number 0 is added to the amino acid name to represent the 0th amino acid of the multiple amino acids contained in the virus.
[0025] The amino acid sequence and antigen cluster name information T1 may be prepared by, for example, a user. Alternatively, for example, a processing unit (not shown) may generate the amino acid sequence and antigen cluster name information T1 by extracting information on amino acids and antigen clusters from information on known viruses.
[0026] (A-1) Hardware Configuration Example The functions of the information processing device 1 according to an embodiment may be implemented by one computer or two or more computers. Furthermore, at least some of the functions of the information processing device 1 may be implemented using hardware (HW) resources and network (NW) resources provided by a cloud environment.
[0027] 3 is a block diagram showing an example of the hardware (HW) configuration of a computer 10 that realizes the functions of the information processing device 1 according to an embodiment. When multiple computers are used as HW resources that realize the functions of the information processing device 1, each computer may have the HW configuration shown in FIG.
[0028] As shown in FIG. 3, the computer 10 may, as its HW configuration, illustratively include a processor 10a, a graphics processing unit 10b, a memory 10c, a storage unit 10d, an IF (Interface) unit 10e, an IO (Input / Output) unit 10f, and a reading unit 10g.
[0029] The processor 10a is an example of a processing unit that performs various controls and calculations, and is a control unit that executes various processes. The processor 10a may be connected to each block in the computer 10 via a bus 10j so that they can communicate with each other. The processor 10a may be a multiprocessor including multiple processors, a multi-core processor having multiple processor cores, or a configuration having multiple multi-core processors.
[0030] Examples of the processor 10a include integrated circuits (ICs) such as a CPU, MPU, APU, DSP, ASIC, and FPGA. Note that the processor 10a may be a combination of two or more of these integrated circuits. CPU is an abbreviation for Central Processing Unit, MPU is an abbreviation for Micro Processing Unit, APU is an abbreviation for Accelerated Processing Unit, DSP is an abbreviation for Digital Signal Processor, ASIC is an abbreviation for Application Specific IC, and FPGA is an abbreviation for Field-Programmable Gate Array.
[0031] The graphics processing device 10b controls screen display for an output device such as a monitor in the IO unit 10f. The graphics processing device 10b may also be configured as an accelerator that executes machine learning processing and prediction processing using a machine learning model. Examples of the graphics processing device 10b include various types of arithmetic processing devices, such as a graphics processing unit (GPU), an APU, a DSP, an ASIC, an FPGA, or other integrated circuits (ICs).
[0032] The memory 10c is an example of HW that stores various types of data, programs, and other information. The memory 10c may be, for example, a volatile memory such as a dynamic random access memory (DRAM) or a non-volatile memory such as a persistent memory (PM), or both.
[0033] The storage unit 10d is an example of HW that stores various types of data, programs, and other information. Examples of the storage unit 10d include various storage devices such as a magnetic disk device such as a hard disk drive (HDD), a semiconductor drive device such as a solid state drive (SSD), and a nonvolatile memory. Examples of nonvolatile memory include a flash memory, a storage class memory (SCM), and a read-only memory (ROM).
[0034] The storage unit 10d may store a program 10h (information processing program) that realizes all or part of the various functions of the computer 10.
[0035] For example, the processor 10a of the information processing device 1 can implement functions in a training phase and a prediction phase, which will be described later, by expanding the program 10h stored in the storage unit 10d into the memory 10c and executing it.
[0036] The IF unit 10e is an example of a communication IF that controls the connection and communication between the computer 10 and other computers. For example, the IF unit 10e may include an adapter that complies with a LAN (Local Area Network) such as Ethernet (registered trademark) or optical communication such as FC (Fibre Channel). The adapter may support either or both wireless and wired communication methods.
[0037] For example, the computer 10 may be connected to other computers and databases (not shown) via the IF unit 10e and a network so as to be able to communicate with each other. The program 10h may be downloaded from the network to the computer 10 via the communication IF and stored in the storage unit 10d.
[0038] The IO unit 10f may include one or both of an input device and an output device. Examples of input devices include a keyboard, a mouse, and a touch panel. Examples of output devices include a monitor, a projector, and a printer. The IO unit 10f may also include a touch panel that combines an input device and an output device. The output device may be connected to the graphics processing device 10b.
[0039] The reading unit 10g is an example of a reader that reads data and program information recorded on the recording medium 10i. The reading unit 10g may include a connection terminal or device to which the recording medium 10i can be connected or inserted. Examples of the reading unit 10g include an adapter compliant with USB (Universal Serial Bus) or the like, a drive device that accesses a recording disk, and a card reader that accesses a flash memory such as an SD card. Note that the recording medium 10i may store the program 10h, and the reading unit 10g may read the program 10h from the recording medium 10i and store it in the memory unit 10d.
[0040] Examples of the recording medium 10i include non-transitory computer-readable recording media such as magnetic / optical disks and flash memories. Examples of magnetic / optical disks include flexible disks, CDs (Compact Discs), DVDs (Digital Versatile Discs), Blu-ray Discs, and HVDs (Holographic Versatile Discs). Examples of flash memories include semiconductor memories such as USB memories and SD cards.
[0041] The above-described HW configuration of the computer 10 is an example. Therefore, the HW in the computer 10 may be increased or decreased (for example, adding or deleting any block), divided, integrated in any combination, or the HW may be added or deleted as needed.
[0042] (A-2) Example of Functional Configuration As shown in Fig. 1, the information processing device 11 may exemplarily include functions as a three-dimensional structure calculation processing unit 101, a graph AI calculation processing unit 102, a graph AI 103, a weight vector calculation processing unit 104, a chemical parameter calculation processing unit 105, an inference unit 106, a graph data shaping processing unit 107, and a regression model 110. These functions may be realized by the hardware of the computer 10 (see Fig. 3). AI is an abbreviation for Artificial Intelligence.
[0043] The 3D structure calculation processor 101 analyzes the 3D structure of a viral protein. When the 3D structure calculation processor 101 receives an amino acid sequence of a virus, it analyzes the 3D structure of the amino acids. The 3D structure calculation processor 101 outputs 3D structure information of the amino acids as the analysis result. The 3D structure information of the amino acids may include, for example, the coordinates of each atom.
[0044] The function of the three-dimensional structure calculation processing unit 101 may be realized by using a known protein structure calculation tool, such as AlphaFold2.
[0045] FIG. 4 is a diagram illustrating an example of amino acid three-dimensional structure information output by the three-dimensional structure calculation processing unit 101 in the information processing device 1 according to an embodiment.
[0046] 4, the amino acid tertiary structure information is shown in the form of a data table. Hereinafter, the amino acid tertiary structure information may be referred to by adding the symbol T2.
[0047] The amino acid three-dimensional structure information T2 shown in Fig. 4 shows the coordinate values of each amino acid in association with a number that identifies the virus.
[0048] The coordinate values of each amino acid include the coordinate values of x, y, and z. In Fig. 4, for example, the coordinates of the amino acid with amino acid number 0 are represented by assigning amino acid number 0 to each of amino acid x, amino acid y, and amino acid z.
[0049] In the amino acid three-dimensional structure information T2, the amino acid names may also be arranged, for example, in the order of peptide bonds from the beginning to the end.
[0050] The amino acid three-dimensional structure information T2 corresponds to a first feature amount related to the three-dimensional structure of a viral protein. The amino acid three-dimensional structure information T2 output by the three-dimensional structure calculation processing unit 101 may be stored in a predetermined storage area of the memory 10c or the storage unit 10d.
[0051] The chemical parameter calculation processor 105 generates chemical parameters for each amino acid contained in the virus based on the amino acid three-dimensional structure information created by the three-dimensional structure calculation processor 101. The chemical parameters may be, for example, electric charge, etc. The chemical parameter calculation processor 105 may calculate the electric charge, etc. for each amino acid.
[0052] The chemical parameter calculation processing unit 105 may generate chemical parameters by using various known methods. For example, the chemical parameter calculation processing unit 105 may calculate feature quantities such as charge by using a known molecular dynamics simulator.
[0053] FIG. 5 is a diagram illustrating chemical parameter information created by the chemical parameter calculation processing unit 105 in the information processing device 1 according to an embodiment.
[0054] 5, the chemical parameter information is represented in the form of a data table including a plurality of chemical parameters. Hereinafter, the chemical parameter information may be represented by adding the symbol T3.
[0055] The chemical parameter information T3 shown in Fig. 5 indicates the chemical parameter values of a plurality of amino acids in association with a number that identifies a virus.
[0056] In FIG. 5, for example, the chemical parameters of the amino acid with the amino acid number 0 are represented by assigning the amino acid number 0 to the amino acid chemical parameters.
[0057] In the chemical parameter information T3, the amino acid names may also be arranged in the order of peptide bonds, for example, from the beginning to the end.
[0058] Furthermore, the chemical parameter calculation processing unit 105 may generate multiple types of chemical parameters for each amino acid.
[0059] The chemical parameter information T3 corresponds to a third feature amount related to the properties resulting from the three-dimensional structure. The chemical parameter information T3 generated by the chemical parameter calculation processing unit 105 may be stored in a predetermined storage area of the memory 10c or the storage unit 10d. Note that the first feature amount related to the three-dimensional structure of the viral protein may include a third feature amount related to the properties resulting from the three-dimensional structure.
[0060] The graph data shaping processor 107 creates graph information based on the amino acid three-dimensional structure information T2 created by the three-dimensional structure calculation processor 101 and the chemical parameter information T3 created by the chemical parameter calculation processor 105. The graph information may be referred to as graph data.
[0061] FIG. 6 is a diagram illustrating graph information in the information processing device 1 according to an embodiment.
[0062] 6, the graph information is expressed in the form of a data table. Hereinafter, the graph information may be denoted by the symbol T4.
[0063] FIG. 7 is a diagram for explaining the processing of the graph data shaping processing unit 107 in the information processing device 1 according to an embodiment.
[0064] The graph data formatting processor 107 combines (synthesizes) the amino acid sequence and antigen cluster name information T1, the amino acid three-dimensional structure information code T2, and the chemical parameter information T3 to generate graph information T4.
[0065] When generating the graph information T4, the graph data shaping processor 107 may combine the amino acid sequence and antigen cluster name information T1, the amino acid three-dimensional structure information with the code T2, and the chemical parameter information T3 based on the number that identifies the virus.
[0066] The graph AI calculation processing unit 102 creates (shapes) data to be input to the graph AI 103 (graph AI input information T5: input data) based on the graph information T4 generated by the graph data shaping processing unit 107.
[0067] The graph AI calculation processing unit 102 generates graph AI input information T5 by converting information about multiple viruses included in the graph information T4 into data in a format that can be processed by the graph AI 103.
[0068] In addition, in the training phase, the graph AI calculation processing unit 102 performs training (machine learning) of the graph AI 103 using the graph AI input information T5.
[0069] Here, Graph AI 103 is a machine learning model that performs graph-based relationship learning and realizes graph classification (class classification).
[0070] A graph is a mathematical model that is composed of a set of nodes and a set of edges between the nodes.
[0071] When applying a graph to a virus, amino acids correspond to nodes, and bonds between amino acids correspond to edges. The bonds between amino acids may be, for example, peptide bonds, or other bonds such as those formed by electrostatic forces.
[0072] The graph AI 103 performs graph classification based on the graph and edge information, using the amino acid conformation as an explanatory variable and the antigen cluster name as a response variable.
[0073] In graph classification, parameters for each node or edge may be used as node attributes or edge attributes to aid in classification.
[0074] To have the graph AI 103 perform graph classification, it is necessary to explicitly assign edges to the graph AI 103. Therefore, the graph AI calculation processing unit 102 determines that adjacent amino acids have an edge based on the amino acid sequence. In addition, amino acids within a certain distance due to electrostatic force or the like may also be determined to have an edge.
[0075] The function of the graph AI 103 can be realized using a known method. For example, the function of the graph AI 103 may be realized by Deep Tensor (registered trademark).
[0076] Graph AI103 corresponds to a machine learning model that uses amino acid three-dimensional structure information T2 (first feature related to the three-dimensional structure of the viral protein) and chemical parameter information T3 (third feature related to properties resulting from the three-dimensional structure) as input data.
[0077] The graph AI calculation processing unit 102 performs classification of virus antigen clusters based on the three-dimensional structure using the graph AI 103 once, and calculates the contribution of each amino acid after this classification. The classification of virus antigen clusters based on the three-dimensional structure using the graph AI 103 is an example of prediction using a machine learning model that uses the first feature amount and the third feature amount as input data.
[0078] Based on the graph information T4, the graph AI calculation processing unit 102 creates graph AI input information T5 by listing the attributes of the two amino acids connected by each edge in the amino acid sequence that constitutes the virus, in units of bonds. Hereinafter, the two amino acids connected by an edge may be referred to as an amino acid pair. The amino acid that is the starting point of the edge in the amino acid pair may be referred to as the start node, and the amino acid that is the end point of the edge may be referred to as the end node.
[0079] FIG. 8 is a diagram for explaining graph AI input information T5 in the information processing device 1 according to an embodiment.
[0080] 8 shows the graph information T4 exemplified in FIG. 6 and graph AI input information T5 created by the graph AI calculation processing unit 102 based on this graph information T4.
[0081] In the graph AI input information T5 illustrated in Fig. 8, the number specifying an edge is associated with information on the amino acid pair to which the edge is connected.
[0082] The amino acid pair information includes a virus identification number, a cluster name, and the amino acid name, amino acid sequence number, chemical parameters, and amino acid coordinate values (x, y, z) of each of the start and end nodes. In the example shown in Figure 8, the start node information is suffixed with "s" and the end node information is suffixed with "e."
[0083] Therefore, for example, amino acid name s represents the start node, and amino acid name e represents the end node. Furthermore, amino acid sequence number s, chemical parameter s, amino acid xs, amino acid name ys, and amino acid zs represent the attribute information of the start node (start node attribute). Similarly, amino acid sequence number e, chemical parameter e, amino acid xe, amino acid name ye, and amino acid ze represent the attribute information of the end node (end node attribute).
[0084] In the training phase, the graph AI calculation processing unit 102 uses the graph AI input information T5 as training information to train the graph AI 103.
[0085] 8 , the cluster name is used as a response variable in the training phase of the graph AI 103. In addition, the amino acid name s, the amino acid name e, the start node attribute, and the end node attribute are used as explanatory variables in the training phase of the graph AI 103.
[0086] Graph AI 103 may be a deep neural network (DNN) that includes multiple hidden layers between an input layer and an output layer.
[0087] A NN, for example, inputs input data to an input layer and sequentially executes predetermined calculations in a hidden layer composed of a convolutional layer, a pooling layer, etc., thereby executing forward processing (forward propagation processing) in which information obtained by the calculations is sequentially transmitted from the input side to the output side. After executing the forward processing, backward processing (backpropagation processing) is executed to determine parameters to be used in the forward processing in order to reduce the value of an error function obtained from the output data (graph classification result) output from the output layer and the correct answer data (cluster name). Then, an update processing is executed to update variables such as weights based on the results of the backpropagation processing. For example, gradient descent may be used as an algorithm to determine the update width of the weights used in the backpropagation calculation.
[0088] In addition, in the training phase, the graph AI calculation processing unit 102 inputs graph AI input information T5 to the graph AI 103, causes graph classification (class classification) to be performed, and then causes statistical information to be calculated.
[0089] The statistical information may be, for example, the contribution (contribution score, node contribution) used to obtain a prediction result when the graph AI 103 performs graph classification. The statistical information may also be referred to as a statistical quantity. The graph AI calculation processing unit 102 obtains a statistical quantity for each amino acid contained in the virus. The graph AI calculation processing unit 102 generates a statistical quantity (node contribution) for each amino acid based on the statistical information.
[0090] The statistics for each amino acid contained in the virus correspond to a second feature (statistical feature) based on the contribution (statistical information) of each amino acid contained in the protein to the prediction. Therefore, the graph AI calculation processing unit 102 obtains the second feature (statistical feature) based on the contribution (statistical information) of each amino acid contained in the protein to the prediction by making a prediction using the graph AI 103.
[0091] FIG. 9 is a diagram illustrating statistical information in the information processing device 1 according to an embodiment.
[0092] 9, a plurality of pieces of statistical information are represented in the form of a data table. Hereinafter, statistical information may be represented by adding the symbol T6.
[0093] 9 shows statistical information T6 in which values of statistical information for a plurality of amino acids are associated with virus-specific numbers.
[0094] In FIG. 9, for example, the statistical information of the amino acid with amino acid number 0 is represented by assigning amino acid number 0 to the amino acid statistics.
[0095] In the statistical information T6, the amino acid names may also be arranged in the order of peptide bonds, for example, from the beginning to the end.
[0096] The statistical information generated by the graph AI calculation processing unit 102 may be stored in a predetermined storage area of the memory 10c or the storage unit 10d.
[0097] In the graph AI (graph AI 103), the contribution is obtained for each three-dimensional structure or each amino acid. Therefore, the graph AI calculation processing unit 102 may calculate a sample average of the contribution in a predetermined unit such as a cluster, a year, or an amino acid, and use the average as statistical information.
[0098] The prediction results performed by the graph AI calculation processing unit 102 in the graph AI 103 and the values of the statistical information calculated by the graph AI 103 may be stored in a predetermined memory area of the memory 10c or the storage unit 10d.
[0099] FIG. 10 is a diagram for explaining the processing of the graph AI calculation processing unit 102 of the information processing device 1 according to an embodiment.
[0100] As described above, in the training phase, the graph AI calculation processing unit 102 inputs the graph AI input information T5 to the graph AI 103 to perform graph classification (see symbol P1). The graph AI calculation processing unit 102 also acquires statistical information (contribution rate) calculated by the graph AI 103 (see symbol P2).
[0101] The graph AI calculation processing unit 102 may change the values contained in the graph AI input information T5, check how the inference result changes (see symbol P3), and if the inference result improves, may perform processing such as reflecting the changes in the graph AI input information T5.
[0102] The weight vector calculation processing unit 104 receives as input the amino acid three-dimensional structure information (amino acid three-dimensional structure information T2) generated by the three-dimensional structure calculation processing unit 101, the chemical parameters (chemical parameter information T3) generated by the chemical parameter calculation processing unit 105, and the statistics for each amino acid (node contribution: statistical information T6) generated by the graph AI calculation processing unit 102.
[0103] The weight vector calculation unit 104 uses this information to create a fixed-length vector (weight vector of input feature values) for each amino acid sequence. The feature value weight vector is used as a weight for the feature values (input feature values) input to a regression model 110 (NN: Neural Network) described later.
[0104] The weight vector calculation processing unit 104 determines weights (weight vectors) of input features in the regression model 110 based on the amino acid three-dimensional structure information T2 (first feature), the chemical parameter information T3 (third feature), and the statistical feature (second feature).
[0105] The weight vector calculation unit 104 may set weights for amino acid sequences by embedding graph data into fixed-length vectors using a function such as a transformer, which is a known machine learning model. That is, the weight vector calculation unit 104 sets numerical values of regularity such as importance corresponding to amino acid sequences.
[0106] FIG. 11 is a diagram illustrating weight vector information in the information processing device 1 according to an embodiment.
[0107] 11, a plurality of pieces of weight vector information are represented in the form of a data table. Hereinafter, weight vector information may be represented by adding the symbol T7.
[0108] In FIG. 11, for example, the weight vector of the amino acid with amino acid number 0 is represented by assigning amino acid number 0 to the weight vector.
[0109] In the weight vector information T7, the amino acid names may also be arranged in the order of peptide bonds, for example, from the beginning to the end.
[0110] The weight vector calculation processing unit 104 determines hyperparameters such as the input / output variables and dimensions of the latent variables of the model based on the contribution of each amino acid.
[0111] The weight vector information T7 is not limited to the example shown in Fig. 11 and can be modified as appropriate. For example, a single virus may be assigned multiple weight vectors. These multiple weight vectors may be managed in chronological order, etc.
[0112] The weight vector information generated by the weight vector calculation processing unit 104 may be stored in a predetermined storage area of the memory 10c or the storage unit 10d.
[0113] The inference unit 106 predicts (infers) the amino acid sequence of the virus after mutation.
[0114] For example, the inference unit 106 uses a regression model 110 to predict the amino acid sequence of the virus after mutation.
[0115] The inference unit 106 trains the regression model 110 in the training phase, and causes the regression model 110 to predict the amino acid sequence after mutation in the prediction phase.
[0116] The inference unit 106 uses the regression model 110 to predict the amino acid sequence at time t+Δt (Δt>0) based on the amino acid sequence at time t.
[0117] The regression model 110 may achieve regression using techniques such as SVR, NN, GA (Genetic Algorithms), and time series analysis. In this regression, for example, an amino acid sequence may first be associated with a vector of numbers, and a sequence of numbers corresponding to amino acid names (20 types, such as proline) may be input and output as a vector. For example, an amino acid sequence may be represented by numbers 0 to 19, and a regression problem may be solved as to the order in which these numbers are output.
[0118] The regression model 110 may be a deep neural network (DNN) that includes multiple hidden layers between an input layer and an output layer.
[0119] The regression model 110 corresponds to a regression model that predicts the amino acid sequence of a virus after mutation using the amino acid sequence of the virus as an input feature (explanatory variable).
[0120] In the training phase, the inference unit 106 trains a regression model 110 that predicts the amino acid sequence of the virus after mutation, using the amino acid sequence of the virus and the weight vector of the features (weights of the input features) generated by the weight vector calculation processing unit 104 as input features (explanatory variables).
[0121] The inference unit 106 trains the machine learning model 100 using the amino acid sequence at the previous time (time t) and the weight vector of the features generated by the weight vector calculation processing unit 104 as learning data, and the amino acid sequence at the next time t + Δt (Δt > 0) as correct answer data.
[0122] In this way, by using the weight vector of the feature quantities generated by the weight vector calculation processing unit 104 in training the regression model 110, the three-dimensional structure of the virus protein is reflected in the regression model 110.
[0123] Here, the regression calculation assumes that the input and output data lengths (dimensions when vectorized) are fixed. However, the length of the amino acid sequence for each virus is not constant. Therefore, a fixed-length amino acid sequence may be created and used by extracting a portion of a predetermined length from the amino acid sequence. To create a fixed-length amino acid sequence, for example, the beginning and end of the amino acid sequence may be removed to extract the predetermined-length portion. The method for making an amino acid sequence of a fixed length is not limited to this, and can be modified as appropriate.
[0124] In the prediction phase, only the amino acid sequence is input to the inference unit 106. The inference unit 106 inputs this amino acid sequence into the regression model 110 to obtain the amino acid sequence after mutation. The inference unit 106 may also output statistics (such as contributions) that can be used to explain the prediction.
[0125] (B) Operation The processing in the training phase in the information processing device 1 according to the embodiment configured as described above will be described with reference to the flowchart (steps A1 to A6) shown in FIG.
[0126] When the amino acid sequence of a virus at time t is input to the 3D structure calculation processor 101, the 3D structure calculation processor 101 performs 3D structure analysis of amino acids in step A1. The 3D structure calculation processor 101 generates amino acid 3D structure information T2.
[0127] The amino acid three-dimensional structure information T2 is input to the chemical parameter calculation processing unit 105. In step A2, the chemical parameter calculation processing unit 105 generates chemical parameters for each amino acid contained in the virus based on the amino acid three-dimensional structure information T2, and generates chemical parameter information T3.
[0128] The amino acid three-dimensional structure information T2 created by the three-dimensional structure calculation processing unit 101 and the chemical parameter information T3 created by the chemical parameter calculation processing unit 105 are input to the graph data shaping processing unit 107. In step A3, the graph data shaping processing unit 107 generates graph information T4 based on the amino acid three-dimensional structure information T2 and the chemical parameter information T3.
[0129] The graph information T4 generated by the graph data shaping processor 107 is input to the graph AI 103. The graph AI calculation processor 102 generates graph AI input information T5 based on the graph information T4 by arranging the attributes of each of the two amino acids connected by each edge constituting the virus in units of bond.
[0130] The graph AI calculation processing unit 102 uses the graph AI input information T5 as training information to train the graph AI 103. In step A4, the graph AI calculation processing unit 102 causes the graph AI 103 to calculate statistical information (contribution degree) and generates statistical information T6.
[0131] The contribution of each amino acid (statistical information T6) generated by the graph AI calculation processing unit 102 and the chemical parameter information T3 generated by the chemical parameter calculation processing unit 105 are input to the weight vector calculation processing unit 104.
[0132] In step A5, the weight vector calculation processor 104 creates a weight vector (weight vector information T7) of the NN feature quantities using the contribution of each amino acid, the chemical parameter information T3, and the amino acid three-dimensional structure information T2. That is, the weight vector calculation processor 104 determines hyperparameters such as the dimensions of the input / output variables and latent variables of the model based on the contribution of each amino acid.
[0133] The weight vector information T7 generated by the weight vector calculation processing unit 104 and the amino acid sequence of the virus at time t are input to the inference unit 106.
[0134] In step A6, the inference unit 106 trains the machine learning model 100 using the amino acid sequence at the previous time (time t) and the weight vector of the features generated by the weight vector calculation processing unit 104 as input features (explanatory variables), and the amino acid sequence at the next time t + Δt (Δt > 0) as correct data.
[0135] For example, the inference unit 106 converts the amino acid sequence into a fixed dimension, and then inputs it into the regression model 110 to predict the amino acid sequence.
[0136] The inference unit 106 compares the predicted amino acid sequence with the correct data (the amino acid sequence after mutation). As a result of this comparison, the inference unit 106 performs a backward process (backpropagation process) to determine parameters to be used in the forward process in order to reduce the value of the error function obtained. The inference unit 106 then performs an update process to update variables such as weights based on the results of the backpropagation process.
[0137] In the prediction phase in the information processing device 1 according to one embodiment configured as described above, the current amino acid sequence of the virus is input to the inference unit 106. In step A6, the inference unit 106 converts the amino acid sequence to a fixed length, and then inputs it to the regression model 110, which predicts the amino acid sequence after mutation.
[0138] The amino acid sequences output by the regression model 110 in the prediction phase may be used as training data in the subsequent training phase.
[0139] Next, the processing of the graph AI calculation processing unit 102 of the information processing device 1 according to an embodiment will be described with reference to the flowchart (steps B1 to B3) shown in FIG.
[0140] In step B1, the graph AI calculation processing unit 102 formats the graph information T4 generated by the graph data formatting processing unit 107 to create graph AI input information T5.
[0141] The graph AI calculation processing unit 102 trains the graph AI 103 using the created graph AI input information T5 (step B2).
[0142] At this time, the graph AI calculation processing unit 102 uses the information other than the cluster name in the graph AI input information T5 as explanatory variables, and uses the cluster name as the objective variable.
[0143] In step B3, the graph AI calculation processing unit 102 inputs the graph AI input information T5 to the graph AI 103 to perform graph classification and predict (infer) cluster names. At this time, the graph AI calculation processing unit 102 uses information other than the cluster names in the graph AI input information T5 as explanatory variables.
[0144] Then, the graph AI calculation processing unit 102 causes the graph AI 103 to calculate statistical information (contribution of each amino acid), and then ends the process.
[0145] (C) Effect Thus, according to the information processing device 1 of one embodiment, in the training phase of training the regression model 110 that predicts the amino acid sequence of a virus after mutation, the graph AI calculation processing unit 102 inputs amino acid three-dimensional structure information T2 (first feature) related to the three-dimensional structure of the virus protein and chemical parameter information T3 (third feature) related to the properties resulting from the three-dimensional structure, and causes the graph AI 103 to perform graph classification (prediction).
[0146] In addition, the graph AI calculation processing unit 102 inputs the graph AI input information T5 into the graph AI 103, performs graph classification (class classification), and then calculates statistical information to generate statistics (node contribution) for each amino acid.
[0147] The weight vector calculation processing unit 104 creates a weight vector (weight of input feature) for each amino acid sequence using the amino acid three-dimensional structure information T2 generated by the three-dimensional structure calculation processing unit 101, the chemical parameter information T3 generated by the chemical parameter calculation processing unit 105, and the statistics (node contribution) for each amino acid generated by the graph AI calculation processing unit 102.
[0148] Then, in the training phase, the inference unit 106 trains a regression model 110 that predicts the amino acid sequence of the virus after mutation, using the amino acid sequence and the weight vector of the features generated by the weight vector calculation processing unit 104 as input features (explanatory variables).
[0149] As a result, the three-dimensional structure of the viral protein is reflected in the regression model 110. Therefore, in the prediction phase, the regression model 110 can predict viral mutations taking into account the properties specific to the three-dimensional structure of the viral protein, thereby improving prediction accuracy.
[0150] Proteins are made up of multiple amino acids linked by peptide bonds, and an amino acid sequence is an arrangement of amino acid names in the order of these bonds. However, even amino acids that are far apart in an amino acid sequence can be bound together by electrostatic forces, etc., and can have unique shapes and properties. In other words, due to their unique shapes and properties, different features can be obtained even if the amino acid sequence is the same.
[0151] In the information processing device 1, prediction accuracy can be improved by predicting virus mutations using feature amounts based on the three-dimensional structure of proteins as clues.
[0152] (D) Others The disclosed technology is not limited to the above-described embodiment, and various modifications can be made without departing from the spirit of the present embodiment. The configurations and processes of the present embodiment can be selected or combined as needed.
[0153] For example, in the above-described embodiment, an example is shown in which the chemical parameter information T3 is generated based on the amino acid three-dimensional structure information T2, but the present invention is not limited to this, and the chemical parameter information T3 does not have to be generated based on the amino acid three-dimensional structure information T2. Moreover, the information processing device 1 does not have to generate the chemical parameter information T3, and the chemical parameter information T3 may be acquired from an external source.
[0154] For example, in the above-described embodiment, the amino acid 3D structure information T2 (first feature amount) and the chemical parameter information T3 (third feature amount) are input to the graph AI 103 to perform graph classification (prediction). On the other hand, the graph AI 103 may also be input with only the amino acid 3D structure information T2 (first feature amount) to perform graph classification (prediction).
[0155] For example, in the above-described embodiment, the graph AI calculation processing unit 102 generates statistics (node contributions), and the weight vector calculation processing unit 104 creates a weight vector (weights of input features) for each amino acid sequence, but these may be performed by the same processing unit.
[0156] For example, in the above-described embodiment, an example is shown in which the contribution degree is used as the statistical information, but the present invention is not limited to this, and information other than the contribution degree may be used as the statistical information.
[0157] Furthermore, the above disclosure enables those skilled in the art to implement and manufacture the present embodiment.
[0158] 1 Information processing device 10 Computer 10a Processor 10b Graphic processing device 10c Memory 10d Storage unit 10e IF unit 10f IO unit 10g Reading unit 10h Program 10i Recording medium 10j Bus 101 Three-dimensional structure calculation processing unit 102 Graph AI calculation processing unit 103 Graph AI 104 Weight vector calculation processing unit 105 Chemical parameter calculation processing unit 106 Inference unit 106 107 Graph data shaping processing unit 110 Regression model T1 Amino acid sequence and antigen cluster name information T2 Amino acid three-dimensional structure information (first feature amount) T3 Chemical parameter information (third feature amount) T4 Graph information T5 Graph AI input information T6 Statistical information T7 Weight vector information
Claims
1. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of a mutated virus using the amino acid sequence of the virus as the input feature. An information processing program characterized by having a computer perform the processing.
2. The first feature includes a third feature relating to properties resulting from the three-dimensional structure. The information processing program according to claim 1, characterized in that...
3. By using the machine learning model with the first and third features as input data, a second feature is obtained based on the contribution of each amino acid contained in the protein corresponding to the input data to the prediction. Based on the first, second, and third features, the weights of the input features are determined in the regression model that predicts the amino acid sequence of the mutated virus using the amino acid sequence of the virus as input features. The information processing program according to claim 2, characterized in that it causes the computer to perform the processing.
4. In the process of training the aforementioned regression model, The input features are defined as the weights of the aforementioned input features and the amino acid sequence of the aforementioned virus. The information processing program according to claim 3, characterized in that...
5. The first feature is generated by performing a three-dimensional structural analysis of the amino acids of the aforementioned protein. The information processing program according to claim 3 or 4, characterized in that it causes the computer to perform the processing.
6. Based on the first feature described above, the third feature described above is generated for each amino acid contained in the virus. The information processing program according to claim 5, characterized in that it causes the computer to perform the processing.
7. The information processing program according to claim 3, characterized in that the weights of the input features are weights set for each amino acid sequence of the virus.
8. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of a mutated virus using the amino acid sequence of the virus as the input feature. An information processing method characterized in that the processing is performed by a computer.
9. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of a mutated virus using the amino acid sequence of the virus as the input feature. An information processing apparatus characterized by including a control unit that performs processing.
10. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of the mutated virus using the amino acid sequence of the virus as an input feature. The regression model is trained using the weights of the aforementioned input features and the amino acid sequence of the virus as input features. An information processing program characterized by having a computer perform the processing.
11. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of the mutated virus using the amino acid sequence of the virus as an input feature. The regression model is trained using the weights of the aforementioned input features and the amino acid sequence of the virus as input features. An information processing method characterized in that the processing is performed by a computer.
12. Based on a first feature relating to the three-dimensional structure of the viral protein and a second feature relating to the contribution of the machine learning model to the prediction obtained based on the first feature, the weights of the input features are determined in a regression model that predicts the amino acid sequence of the mutated virus using the amino acid sequence of the virus as an input feature. The regression model is trained using the weights of the aforementioned input features and the amino acid sequence of the virus as input features. An information processing apparatus characterized by including a control unit that performs processing.