Device for constructing sequence transformation neural network for transforming input sequence and learning method using the same

A sequence-transformation neural network enhances neoantigen prediction accuracy, enabling effective immune anti-cancer vaccine therapy by accurately identifying and synthesizing cancer-specific peptides for T cell activation.

JP2025528901APending Publication Date: 2025-09-02LG MANAGEMENT DEV INST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025511633
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-23
Filing Date
2023-08-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing artificial intelligence-based neural networks struggle to provide accurate data prediction, particularly in the context of neoantigen prediction for immune anti-cancer vaccine therapy, necessitating improved learning algorithms.

Method used

A sequence-transformation neural network is constructed using a computer device with a processor and memory, performing attention operations to train the network based on labeled input data, determining key and query values, and generating output data to predict neoantigens accurately.

Benefits of technology

The solution enables more accurate neoantigen prediction, facilitating immune anti-cancer vaccine therapy by identifying and synthesizing neoantigens that can activate T cells, thereby accelerating cancer treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528901000001_ABST
    Figure 2025528901000001_ABST
Patent Text Reader

Abstract

A neural network construction device according to one embodiment of the present invention includes at least one memory and at least one processor communicating with the at least one memory, wherein the at least one processor is configured to: receive first input data and second input data corresponding to the first input data; train a sequence-transformation neural network by performing an attention operation based on predetermined label information and the labeled first input data and the labeled second input data; and determine output data to be output by the sequence-transformation neural network trained based on the first input data, the second input data, and the label information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus and method for predicting neoantigens, and more particularly to an apparatus for constructing a sequence transformation neural network for transforming an input sequence, and a learning method using the same. [Background technology]

[0002] In recent years, various concepts and models (neural networks) have been developed in the field of artificial intelligence, and research into data prediction using these concepts and models is actively progressing.

[0003] However, when predicting data using an artificial intelligence-based neural network, in order to derive more accurate results, i.e., results with a high prediction probability, it was necessary to develop learning algorithms and prediction algorithms that were appropriate for the model. Summary of the Invention [Problem to be solved by the invention]

[0004] An embodiment of the present invention aims to provide an algorithm that can derive more accurate results in data prediction using an artificial intelligence-based neural network. Furthermore, to facilitate neoantigen prediction by utilizing the algorithm and enable the provision of immune anti-cancer vaccine therapy to patients, an apparatus and method for extracting neoantigen candidates using an artificial intelligence-based neoantigen prediction model are provided.

[0005] The problems to be solved by the present invention are not limited to those described above, and other problems not mentioned will also be clearly understood by those skilled in the art through the following description. [Means for solving the problem]

[0006] To achieve the above technical objectives, at least one computer according to the present invention comprises a neural network construction device for constructing a sequence-transformation neural network for transforming an input sequence having each network input. The neural network construction device includes at least one memory and at least one processor in communication with the at least one memory. The at least one processor receives first input data and corresponding second input data, and performs attention operations based on predetermined label information and the labeled first input data and second input data to train the sequence-transformation neural network, and determines output data to be output by the trained sequence-transformation neural network based on the first input data, the second input data, and the label information.

[0007] Furthermore, a learning method for transforming an input sequence, which is executed by a sequence transformation neural network on a computer device according to the present invention, can include the steps of: determining a first key value and a first value value by performing a predetermined pre-learning operation based on first input data; inputting second input data corresponding to the first input data and matching position information to each sequence in the second input data; generating a first query value by performing a predetermined self-attention operation based on the second key value, second value value, and second query value corresponding to the second input data; and determining output data corresponding to the first input data and the second input data by performing the predetermined attention operation based on the first key value, the first value, and the first query value.

[0008] Additionally, a computer program stored on a computer-readable recording medium for implementing the present invention may also be provided.

[0009] Furthermore, a computer-readable recording medium storing a computer program for implementing the present invention may also be provided. [Effects of the Invention]

[0010] According to the means for solving the above-mentioned problems, the present invention provides an algorithm that can derive more accurate results when predicting data using an artificial intelligence-based neural network. Furthermore, by using the algorithm of the present invention, it is possible to easily predict neoantigens, which can have the effect of enabling the provision of immune anti-cancer vaccine therapy to patients.

[0011] The effects of the present invention are not limited to those described above, and other effects not mentioned above can also be easily understood by those skilled in the art from the following description. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating a learning process for constructing a sequence-to-sequence neural network in accordance with an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram illustrating the attention algorithm of a sequence-transformation neural network constructed according to one embodiment of the present invention. [Figure 3] FIG. 3 is a diagram illustrating an example of performing a scaled dot-product attention operation in the attention algorithm of a sequence-transformation neural network constructed according to an embodiment of the present invention. [Figure 4] FIG. 4 is a diagram illustrating an example of processing second input data in the attention algorithm of the sequence transformation neural network constructed according to an embodiment of the present invention. [Figure 5] FIG. 5 is a diagram illustrating an example of processing first input data in the attention algorithm of a sequence-to-sequence neural network constructed according to an embodiment of the present invention. [Figure 6]FIG. 6 is a diagram showing an example of generating output data using attention scores calculated by the attention algorithm in a sequence transformation neural network constructed according to an embodiment of the present invention. [Figure 7] FIG. 7 is a diagram illustrating the configuration of a computer device for converting an input sequence in one embodiment of the present invention. [Figure 8] FIG. 8 is a diagram showing a detailed configuration of a neural network constructed by a processor in the computer device of FIG. [Figure 9] FIG. 9 is a chart illustrating an example of processing first input data based on the attention algorithm of a sequence-to-sequence neural network constructed according to an embodiment of the present invention. [Figure 10] FIG. 10 is a chart illustrating an example of processing second input data based on the attention algorithm of a sequence-to-sequence neural network constructed according to an embodiment of the present invention. [Figure 11] FIG. 11 is a chart showing an example of generating output data using attention scores calculated by the attention algorithm of a sequence-transformation neural network constructed according to an embodiment of the present invention. [Figure 12-13] 12 and 13 are charts showing an example of predicting the degree of binding between a peptide sequence and MHC, as well as T cell activation, based on the attention algorithm of a sequence transformation neural network constructed according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0013] The advantages and features of the present invention, as well as methods for achieving them, will be described in detail in the following examples with reference to the accompanying drawings. However, the present invention is not limited to the following examples and can be implemented in various forms. The embodiments described in this specification are provided to enable those skilled in the art to which the present invention pertains to properly understand the technical scope of the present invention, and the scope of the present invention is defined by the claims.

[0014] The terms used herein are for the purpose of describing the examples and are not intended to limit the scope of the present invention. Terms used in the singular herein are understood to include the plural unless otherwise specified. Furthermore, the terms "comprises" and / or "comprising" used herein do not exclude the presence or addition of other elements in addition to the elements listed. The same reference numerals throughout the specification refer to the same elements, and "and / or" encompasses any one of the listed elements or any combination thereof. Furthermore, reference numerals such as "first" and "second" are used merely to distinguish between different elements, and these terms do not necessarily limit the specific order or role of elements. Therefore, a "first element" described herein may also be treated as a "second element" within the scope of the technical concept.

[0015] Unless otherwise defined, all terms (including technical and scientific terms) described herein shall be interpreted in the sense commonly understood by those skilled in the art to which the present invention pertains. Furthermore, commonly used terms defined in dictionaries shall be interpreted within their ordinary scope and not in an excessively restrictive or expansive manner unless otherwise clearly defined.

[0016] Throughout this specification, the same reference numerals indicate the same components. This specification does not comprehensively describe all components related to the examples, and general content that can be easily understood by a person skilled in the art and redundant explanations are omitted as appropriate. Note that the terms "unit, module, component, block" used in this specification refer to components that can be implemented as software or hardware, and depending on the specific embodiment, multiple "units, modules, components, blocks" may be provided as a single component, or one "unit, module, component, block" may include multiple components.

[0017] In this specification, the term "connected" between one part and another part refers not only to a direct connection between the two parts but also to an indirect connection between the two parts, where an indirect connection also includes a connection via a wireless communication network.

[0018] Furthermore, when a part "comprises" one component, it does not mean that it is limited to only that component, but may also include other components, unless otherwise specified to the contrary.

[0019] In this specification, when one component is positioned "on" another component, it includes not only the case where one component is in direct contact with the other component, but also the case where another component is interposed between the two components.

[0020] Terms such as "first" and "second" are used merely to distinguish one component from another, and these components are not limited by these terms.

[0021] References to the singular are to be construed as including the plural unless the context clearly indicates otherwise.

[0022] The identification numbers assigned to each step are used for convenience of explanation and do not limit the order in which the steps are performed. Unless a specific order is clearly dictated by the context, the steps may be performed in a form different from the order described.

[0023] In the detailed description of the invention below, the terms used are defined as follows:

[0024] In this specification, the term "a sequence transformation neural network construction apparatus according to the present invention" refers to a concept that includes various devices that can execute arithmetic processing and provide results to users. For example, the sequence transformation neural network construction apparatus according to the present invention includes a computer device, a server device, a portable terminal, etc., and may refer to all of these or any one of them.

[0025] Examples of the computer device include laptops, desktops, tablet PCs, slate PCs, and the like equipped with a web browser.

[0026] The server device is a server that communicates with external devices and processes information, and examples thereof include an application server, computing server, database server, file server, game server, mail server, proxy server, web server, etc.

[0027] The portable terminal is a wireless communication device with portability and mobility, and examples thereof include all handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminals, smartphones, etc. Furthermore, examples of the portable terminal include wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, and head-mounted devices (HMDs).

[0028] As used herein, the term "antigen" refers to a substance that induces an immune response.

[0029] Neoantigens are proteins that are newly produced in cancer cells when specific mutations occur in tumor DNA.

[0030] Neoantigens arise from genetic mutations and are characterized by their specific expression only in cancer cells.

[0031] Neoantigens include polypeptide or nucleotide sequences. Genetic mutations include frameshift or non-shift indels, missense or nonsense mutations, splice site mutations, genomic rearrangements or gene fusions, and any genomic or expression alteration that results in the generation of de novo ORFs. Splice variants are also an example of genetic mutations. Specific post-translational modifications in tumor cells include aberrant phosphorylation and splice antigens generated by the proteasome.

[0032] As used herein, the term "epitope" refers to the particular portion of an antigen to which an antibody or T-cell receptor specifically binds.

[0033] As used herein, "MHC" refers to a protein that presents "peptides" synthesized in specific cells on the cell surface, enabling T cells to recognize those cells.

[0034] As used herein, "peptide" refers to a polymer of amino acids. For simplicity, "peptide" as used below refers to an amino acid polymer or amino acid sequence that is displayed on the surface of cancer cells.

[0035] As used herein, the term "MHC-peptide complex" refers to a complex structure of MHC and peptides presented on the surface of cancer cells. T cells recognize this MHC-peptide complex and induce an immune response.

[0036] The principles of operation and embodiments of the present invention will be described below with reference to the accompanying drawings.

[0037] Cancer is a disease in which normal cells mutate and grow indefinitely. Traditionally, the main treatments used have been surgery, radiation, and chemotherapy, but in recent years, active research has been conducted into immune anti-cancer vaccine therapy by utilizing neural network construction devices related to the human immune system.

[0038] Cancer cells produce neoantigens, whose epitopes are presented on the surface of cancer cells by major histocompatibility complexes (MHC). T cells recognize these MHC-epitope complexes and induce an immune response.

[0039] Therefore, in order to identify neoantigens produced by cancer cells, it is necessary to predict the binding between MHC and peptides.

[0040] Specifically, cancer cells fragment abnormal peptides generated by genetic mutations and present them on the cell surface via MHC. When T cells recognize the abnormal peptide fragments, an immune response is induced that destroys the cancer cells.

[0041] Therefore, by identifying abnormal peptide fragments that are presented on the cell surface among the abnormal substances produced by cancer cells and that can bind to T cells, and then directly synthesizing and administering these to patients, the recognition of neoantigens by antigen-presenting cells can be promoted, thereby accelerating the maturation and proliferation of immature T cells and enabling cancer treatment.

[0042] The present invention relates to a technique for constructing a sequence-transformation neural network by learning and predicting the degree of binding between MHC and peptide sequences and T cell activation based on the constructed sequence-transformation neural network. The predetermined operations and algorithms for this can be executed by a computer, and the detailed configuration of the computer will be described below with reference to Figures 7 and 8. In this specification, the term "computer" refers to a neural network construction device for constructing a neural network.

[0043] FIG. 1 is a diagram illustrating a learning process for constructing a sequence-to-sequence neural network in accordance with an embodiment of the present invention.

[0044] Referring to FIG. 1, in one embodiment of the present invention, when training a sequence-to-sequence neural network NN, the training is performed using previously stored open data.

[0045] Specifically, MHC features 110 are input as the first input data, and peptide features 120 are input as the second input data, and then processed as each input data. These are used to learn the degree of MHC binding and the degree of T cell activation.

[0046] This constructs a sequence transformation neural network NN for transforming the input sequence.

[0047] The open data already stored for use in learning includes data on the degree of MHC binding according to the MHC type corresponding to each peptide and the degree of T cell activation, which will be described in detail later.

[0048] FIG. 2 is a diagram showing an overview of the attention algorithm in a sequence-to-sequence neural network constructed according to one embodiment of the present invention.

[0049] 2, a computer device receives first input data, performs predetermined pre-learning, and then determines a first key value and a first value corresponding to the first input data (210). Next, a second input data corresponding to the first input data is received, and generates a first query value corresponding to the first input data by performing multi-head self-attention (220).

[0050] Next, a scaled dot product attention operation is performed using the first key value, the first value value, and the first query value, and each attention head is concatenated to output a matrix in which each sequence is converted into a vector (230). Next, a convolutional neural network (CNN) merge operation is performed according to the peptide length, and after embedding processing, output data is generated and output through a linear layer after vector conversion (240).

[0051] Here, the output data relates to data that has been extracted (filtered) to include only at least one or more MHC candidates that bind to and activate T cells.

[0052] 3 shows an example of implementing scaled dot product attention in the sequence-transformation neural network attention algorithm constructed according to one embodiment of the present invention. The processing operation shown in 230 in FIG. 2 will now be described in more detail.

[0053] Referring to Figure 3, self-attention is performed based on the first key value, the first value value, and the first query value corresponding to the first input data, and the resulting attention score is output (231). This may be performed by concatenating all of these attention heads (multiple attention heads) (232).

[0054] Here, self-attention is a method for extracting relationships between three elements: key, value, and query, and refers to a scaled dot product attention operation.

[0055] In this case, the computer device calculates the first query value for each sequence of the first input data and attention scores corresponding to all first key values, and applies a softmax function to obtain a probability distribution whose sum is 1.

[0056] This is called an "attention distribution," and each value is called an "attention weight."

[0057] The computer device can calculate an attention value by weighting the attention weight corresponding to each sequence with the hidden state, and then concatenate the attention value with the hidden state at time t to generate a single vector.

[0058] That is, the computer device can perform a predetermined operation, i.e., an operation of combining multiple attention heads generated by applying a softmax function, based on the first key value, the first value value, and the first query value corresponding to each sequence of the first input data and the second input data.

[0059] Here, the attention weight corresponds to the attention energy related to the second input data corresponding to the first input data.

[0060] This allows a matrix of attention values ​​to be calculated that includes all of the attention values ​​corresponding to each sequence of the first input data and the second input data. The result is expressed by the following equation (1).

[0061]

number

[0062] The number 1 represents the attention operation, where Q denotes the query on which the attention operation is performed, K denotes the key, and V denotes the value corresponding to the query and the key.

[0063] The softmax function is an activation function used in multi-class classification. When the number of classes is N, an N-dimensional vector is input and the function estimates the probability of belonging to each class. The softmax function is expressed by the following equation 2.

[0064]

number

[0065] The formula 2 represents the softmax function, n represents the number of neurons in the output layer, and k represents the order of the classes.

[0066] Meanwhile, the first key value and the first value corresponding to the first input data are determined based on the first input data, but the first query value may be generated based on the second input data, as will be described in detail later with reference to Figures 4 and 5.

[0067] On the other hand, the output data can be configured as data indicating the degree of matching between sequences based on attention scores calculated by performing self-attention in the attention algorithm of the sequence-transformation neural network.

[0068] Specifically, the output data is obtained by generating output data in which T-cell immunogenicity is matched for each sequence corresponding to the first input data, and then performing a normalization operation on the T-cell immunogenicity corresponding to the sequence of the output data to generate label information corresponding to the second input data.

[0069] 4 is a diagram showing an example of processing second input data in the attention algorithm of the sequence-transformation neural network constructed according to an embodiment of the present invention. The processing operation shown in 220 of FIG. 2 will be specifically described below.

[0070] Referring to Figure 4, the computer device receives second input data corresponding to the first input data in the constructed sequence-transformation neural network, and matches position information to each sequence in the second input data. Here, the second input data is peptide features (sequences), and either physicochemical / biochemical properties (AA index) or amino acid substitution matrices (BLOSUM) can be used as peptide features. As an example, the second input data can be composed of 9 to 10 sequences.

[0071] Next, the input data is subjected to an embedding process, and attention weights are calculated to generate second key values, second value values, and second query values ​​corresponding to the second input data, i.e., corresponding to each sequence of the second input data. Next, a predetermined self-attention calculation is performed based on the second key values, second value values, and second query values ​​(221). At this time, the attention weights can be calculated through pre-training.

[0072] Here, self-attention is multi-head attention, and d model This can be done by performing num_heads parallel attentions on the second key value, second value value, and second query value, each with dimensions equal to the dimension of the first key value divided by num_heads, and then concatenating all these attention heads (matrices for each attention value).

[0073] Here, each attention head has a different attention weight (W Q , W K , W V ) is applied. Then, another matrix of weights (W O ) to obtain the final output value of multi-head attention.

[0074] These output values ​​are then passed through a linear layer to generate a first query value corresponding to the first input data.

[0075] 5 is a diagram showing an example of processing the first input data in the attention algorithm of the sequence transformation neural network constructed according to an embodiment of the present invention. The processing operation shown in 210 of FIG. 2 will be specifically described below.

[0076] 5, the computer device inputs first input data into the constructed sequence transformation neural network and performs a predetermined pre-learning operation (211). Next, each of these output values ​​is passed through a linear layer to determine a first key value and a first value value (212, 213).

[0077] Here, the first input data may include at least one of information on the type and structure of the MHC.

[0078] Specifically, the first input data is an MHC feature (three-dimensional structure), and this MHC feature can use 181 sequences close to the binding site. As an example, the MHC type can be changed to 360 sequences. On the other hand, the first input data can be composed of sequences corresponding to a predetermined range based on the binding origin between the peptide sequence and the MHC sequence.

[0079] 6 shows an example of generating output data using attention scores calculated by the attention algorithm of a sequence-transformation neural network constructed according to an embodiment of the present invention. The processing operation shown in 240 of FIG. 2 will now be described in detail.

[0080] Referring to Figure 6, the computer device outputs a matrix in which each sequence is converted into a vector, as shown in 230 of Figure 2, and then performs a CNN merging operation according to the peptide length (241) and performs an embedding process to convert the data into a vector (242). That is, an embedding vector corresponding to the first input data and the second input data is generated. Next, output data is generated and output via a linear layer (243). Here, output data can also be generated by adding and combining information about the experimental method to the embedding vector using one-hot encoding.

[0081] FIG. 7 is a diagram showing the configuration of a computer device for converting an input sequence in one embodiment of the present invention.

[0082] 7, a computer device 800 may include a memory 810, a processor 820, a communication interface 830, an input / output interface 840, and an input / output device 850. Although only one of each component is shown in FIG. 7, this is for the sake of convenience, and the computer device may include at least one or more of these components, if necessary.

[0083] The memory 810 stores data and / or various information to support various functions of the computer device 800. Specifically, the memory 810 can store a number of application programs (applications) executed by the computer device 800, as well as data and instructions required for the operation of the computer device 800. At least some of these application programs can be downloaded from an external server via wireless communication. Meanwhile, the application programs can be loaded into the processor 820, set up on the computer device 800, and executed by the processor 820 to realize various operations (or functions).

[0084] In particular, the memory 810 stores at least one machine learning model and at least one process to construct a sequence transformation neural network for transforming an input sequence.

[0085] Meanwhile, the memory 810 may include at least one type of recording medium such as a flash memory type, a hard disk type, a multimedia card micro type, a card-type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. The memory may store information temporarily, permanently, or semi-permanently, and may be internal or removable.

[0086] The processor 820 controls all components within the computer device 800 and processes input and output signals, data, information, etc. It also executes instructions, algorithms, and application programs stored in the memory 810 and executes various processes to build a sequence transformation neural network for transforming input sequences, thereby providing appropriate information and functions to each user or processing them.

[0087] The processor 820 determines a first key value and a first value by performing a predetermined pre-learning operation based on the first input data. Second input data corresponding to the first input data is also input, and position information is matched to each sequence in the second input data. The processor 820 generates a first query value by performing a predetermined self-attention operation based on the second key value, second value, and second query value corresponding to the second input data. Then, the processor 820 can determine output data corresponding to the first input data by performing a predetermined attention operation based on the first key value, first value, and first query value.

[0088] The processor can also perform a predetermined operation based on the first key value, the first value value, and the first query value corresponding to each sequence of the first input data and the second input data to generate multiple attention heads and perform an operation to combine them.

[0089] The operation of combining attention heads can be defined by the concatenate operation.

[0090] A concatenation operation refers to a function that concatenates multiple text strings into a single text string.

[0091] The processor may be configured to output output data including an attention energy associated with second input data corresponding to the first input data.

[0092] The output data may be configured to be capable of generating output data in which T-cell immunogenicity is matched to each sequence of the output data corresponding to the first input data.

[0093] Meanwhile, the neural network may be configured to generate label information corresponding to the second input data by performing a normalization operation related to T cell immunogenicity corresponding to the sequence of output data.

[0094] When the neural network predicts binding information between a peptide sequence and an MHC, the output data may include binding information corresponding to the first input data and the second input data.

[0095] The binding information refers to any of the probability of binding, the degree of binding, the strength of binding, the dissociation constant, and IC50.

[0096] Also, if the neural network predicts T cell activation corresponding to a peptide sequence, the output data can include a T cell activation probability corresponding to the first input data and the second input data.

[0097] Furthermore, the first input data according to an embodiment of the present invention may include at least one of information on the type and structure of MHC.

[0098] The second input data may also include sequences of multiple peptides.

[0099] On the other hand, when the second input data includes at least one peptide sequence that does not bind to an MHC, the sequence conversion neural network constructed in the neural network construction device can predict binding information between the MHC and the peptide sequence.

[0100] On the other hand, the sequence transformation neural network can be configured such that when the second input data includes at least one peptide sequence that does not activate T cells, the neural network can predict the probability of T cell activation.

[0101] Specifically, referring to Table 1 below, when the second input data includes a sequence that is negative for MHC binding, the sequence conversion neural network can predict binding information between the peptide sequence and MHC.

[0102] In addition, in the same table, when the second input data includes a peptide sequence that is negative for T cell activation, the sequence conversion neural network can predict the probability of T cell activation.

[0103] [Table 1]

[0104] Table 1 above shows an example of data used for neural network training according to one embodiment of the present invention. The data used for training may include peptide sequences matched with Kd. Kd refers to the dissociation constant of an antigen-binding site; a lower Kd value indicates a higher affinity of the antibody for the antigen, and conversely, a higher Kd value indicates a lower affinity.

[0105] On the other hand, if the binding information between the peptide sequence and MHC included in the output data corresponding to the first input data is less than a predetermined value, the processor may change the variables of the sequence-transformation neural network using a loss function determined based on the binding information.

[0106] Here, the predicted binding information can be determined based on the above-mentioned Kd.

[0107] Also, Kd can be matched in the form of a specific value, but can also be matched in a specific range (inequality).

[0108] Each Kd can be used as label information when the neural network operates on a model that predicts binding information between MHC and peptide sequences. Furthermore, the neural network can reflect the range of the label information in the loss function.

[0109] In particular, when training is performed using a peptide sequence with Kd within a specific range, if the Kd of the output data corresponding to the input data falls within that range, the model determines that there is no loss and training can be performed.

[0110] For example, when training with the peptide sequence "GQIVTMFEA" and MHC type "A*0101," if the Kd output by the neural network is 97 nM, this Kd falls within the label information "<100 nM," so the neural network determines the loss to be 0 and can proceed with training.

[0111] In another example, when training is performed using the peptide sequence "MGQIVTMFE" and the MHC type "B*0101," if the Kd output by the neural network is 20003 nM, this Kd is included in the label information ">20000 nM," so the neural network determines the loss to be 0 and can perform training.

[0112] That is, in the present invention, the label information of the training data used for training can include a specific value and a specific range. If the label information is in a specific range and the value predicted by the neural network is within that range, the neural network determines that the loss is 0 and can perform training.

[0113] Meanwhile, the above-described operation is merely one embodiment of the present invention, and there is no particular limitation on the form of pre-processed data used for learning by the neural network.

[0114] The second input data may also include at least one of a substitution matrix (BLOSUM) and physicochemical / biochemical properties (AAindex) of amino acids corresponding to the peptide sequence.

[0115] The first input data may also be configured with sequences corresponding to a predetermined range based on the binding origin between the peptide sequence and the MHC sequence.

[0116] The at least one processor may also generate output data in which T-cell immunogenicity is matched for each of the output data sequences corresponding to the first input data, and generate label information corresponding to the second input data by performing a normalization operation related to the T-cell immunogenicity corresponding to the output data sequence.

[0117] The processor may also generate output data by adding and combining information corresponding to experimental results to an embedding vector obtained by performing a predetermined attention operation based on the first key value, the first value value, and the first query value.

[0118] On the other hand, the computer device 800 may include one or more components that enable communication with external devices. For example, the computer device 800 may include a communication interface 830 for wireless communication and an input / output interface 840 for wired communication.

[0119] Specifically, the communication interface 830 can transmit and receive signals based on wireless communication with external devices via the network 600. To this end, the communication interface 830 can include at least one wireless communication module, a short-range communication module, or the like.

[0120] First, the wireless communication module may include a Wi-Fi module, a WiBro (Wireless Broadband) module, or a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), WLAN (Wireless LAN), DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), 4G, 5G, and 6G.

[0121] The short-range communication module is for short-range communication and is compatible with Bluetooth. TM Near-field communication can be supported using any of the following technologies: RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless Fidelity), Wi-Fi Direct, and Wireless Universal Serial Bus (Wireless USB).

[0122] On the other hand, the input / output interface 240 is connected to the input / output device 250 by a wire, i.e., can transmit and receive signals based on wired communication. For this purpose, the input / output interface 240 can include at least one wired communication module, and the wired communication module can include various wired communication modules such as a local area network (LAN) module, a wide area network (WAN) module, a value added network (VAN) module, etc., as well as various cable communication modules such as a universal serial bus (USB), a high definition multimedia interface (HDMI), a digital visual interface (DVI), a recommended standard 232 (RS-232), power line communication, or a plain old telephone service (POTS).

[0123] 7 are not essential components for the configuration of the computer device according to the present invention, and the construction device described herein may have more or fewer components than those listed above. In other words, there are no particular limitations on the number, names, operations, etc. of these components.

[0124] FIG. 8 is a diagram showing a detailed configuration of the processor in the computer device of FIG.

[0125] 8, a processor of a computer device 800 constructs a neural network NN. The neural network NN may include a first input layer 821, a second input layer 822, an attention layer 823, and an output layer 824.

[0126] Referring to Figure 8, a first key value and a first value value are determined and output from first input data input via a first input layer 821, and a second input layer 822 generates a first query value from second input data corresponding to the first input data.

[0127] Next, the attention layer 823 performs self-attention based on the first key value, the first value value, and the first query value, and the output layer 824 uses the output value to generate and output output data.

[0128] When the neural network NN operates on a model that predicts binding between an MHC and a peptide sequence, the output layer 824 can output output data including MHC binding information corresponding to the first input data and the second input data.

[0129] That is, the output layer 824 can include binding information between the MHC input as the first input data and the peptide sequence input as the second input data.

[0130] The combined information included in the output data can determine a loss function based on the difference between each label corresponding to the particular first input data and second input data.

[0131] Based on this, the neural network construction device can change each variable that constitutes the neural network NN and perform learning to predict peptide sequences that can bind to MHC.

[0132] On the other hand, when the neural network NN operates using a model that predicts the activity of T-cells corresponding to MHC and peptide sequences, the output data output by the output layer 824 includes the degree of activity of T-cells corresponding to the first input data and the second input data.

[0133] Specifically, the output layer 824 can output output data in the form of positive when the T cells are activated, and negative when the T cells are deactivated.

[0134] According to one embodiment of the present invention, when a neural network NN operates on a model that predicts T-cell activity corresponding to an MHC and peptide sequence, label information for the data can be generated using a Better distribution.

[0135] To label the data required for learning, the processor can perform multiple tests that output the degree of T cell activation corresponding to the first input data and the second input data.

[0136] Furthermore, the processor according to one embodiment of the present invention can perform multiple tests based on first input data and second input data that are similar to each other.

[0137] In the neural network NN, the degree of activation of T cells corresponding to the first input data and the second input data can be output.

[0138] The processor in which the neural network is constructed can determine the number of multiple tests and the number of T cell activations corresponding to the first input data and the second input data.

[0139] The processor that constructs the neural network can generate a better distribution corresponding to the first input data and the second input data based on the number of T cell activations relative to the number of times of the multiple tests.

[0140] That is, the processor can output the T cell activity related to the first input data and the second input data in a better distribution.

[0141] The beta distribution is a continuous probability distribution defined in the interval [0, 1] by two parameters α and β. The parameters are the parameters that determine the shape of the distribution.

[0142] The probability density function of the Better distribution is expressed by the following equation 3.

[0143]

number

[0144] Referring to Equation 3, f denotes a Better distribution, where B(a, b) denotes a Better function, which is expressed by the following Equation 4.

[0145]

number

[0146] The processor driving the neural network can also determine the mean and variance of the Better distribution.

[0147] Furthermore, the processor may determine label information corresponding to the first input data and the second input data based on the mean and variance of the Better distribution. Specifically, the processor may determine the difference between the mean and variance of the label information corresponding to the first input data and the second input data.

[0148] Typically, with such a data distribution, a large number of tests will result in a smaller variance, while a small number of tests may result in a larger variance.

[0149] Based on this, the labels corresponding to the first input data and the second input data, i.e., the MHC and peptide sequences, are expressed by the following equation (5):

[0150]

number

[0151] Referring to Equation 5, L denotes label information determined by the processor, E denotes the mean of the better distribution, and V denotes the variance of the better distribution.

[0152] On the other hand, the operation of each layer constituting the neural network described in Fig. 8 is merely one embodiment of the present invention, and there is no limitation to the operation. Fig. 9 is a chart showing an example of processing first input data based on the attention algorithm of a sequence transformation neural network constructed according to an embodiment of the present invention.

[0153] 9, the computer device 800 inputs first input data to the constructed sequence-to-sequence neural network (S1010), executes predetermined pre-learning operations (S1020), and then determines a first key value and a first value value from each of the output values ​​via a linear layer (S1030).

[0154] FIG. 10 is a chart illustrating an example of processing second input data based on the attention algorithm of a sequence-to-sequence neural network constructed according to an embodiment of the present invention.

[0155] Referring to FIG. 10, the computer device 800 inputs second input data corresponding to the first input data into the constructed sequence transformation neural network (S1110), and matches position information to each sequence in the second input data (S1120).

[0156] Next, after the embedding process for the input data, attention weights are calculated to generate second key values, second value values, and second query values ​​corresponding to the second input data, i.e., corresponding to each sequence of the second input data, and a predetermined self-attention calculation is performed based on these second key values, second value values, and second query values ​​(S1130).

[0157] Next, the output value is passed through a linear layer to generate a first query value corresponding to the first input data (S1140).

[0158] FIG. 11 is a chart showing an example of generating output data using attention scores calculated based on the attention algorithm of a sequence-transformation neural network constructed according to an embodiment of the present invention.

[0159] Referring to FIG. 11, the first key value, first value value, and first query value of the first input data are input (S1210), self-attention is performed based on these first key value, first value value, and first query value, and an attention score is output as the output value (S1220), and this can be performed by concatenating all these attention heads (S1230).

[0160] Next, a CNN merging operation is performed according to the length of the second input data, i.e., the length of the peptide sequence (S1240), and an embedding vector relating to the first input data and the second input data is generated (S1250).

[0161] Next, output data is generated and output via the linear layer (S1260).

[0162] As mentioned above, the output data can include binding information between MHC and peptide sequences, or the probability of T cell activation.

[0163] 12 and 13 are charts showing the operation of predicting binding information between MHC and peptide sequences and the degree of T cell activity corresponding to the peptide sequences, based on the attention algorithm of a sequence-transformation neural network constructed according to one embodiment of the present invention.

[0164] FIG. 12 shows the operation of a neural network that predicts binding information between MHC and peptide sequences.

[0165] 12, first and second input data are input to the neural network. In FIG. 12, the first input data is information on the type and structure of MHC (S1310).

[0166] The second input data may also include a peptide sequence. As shown in Figure 12, when the neural network predicts binding information between an MHC and a peptide sequence, the second input data may include at least one peptide sequence that is not bound to an MHC (S1320).

[0167] In the case of FIG. 12, information on the type and structure of MHC may be input as first input data to the neural network, and information on multiple peptide sequences may be input as second input data.

[0168] Next, the first input data and the second input data are input, and the neural network performs an attention operation using the query, key, and value derived from each data, thereby predicting binding information between the MHC and the peptide sequence (S1330).

[0169] FIG. 13 shows the operation of a neural network to predict the degree of T cell activation corresponding to a peptide sequence.

[0170] Referring to FIG. 13, first input data and second input data are input to the neural network.

[0171] In the case of FIG. 13, information on the type and structure of MHC may be input as first input data to the neural network, and information on multiple peptide sequences may be input as second input data (S1410, S1420).

[0172] Here, the second input data can include either a peptide sequence that activates T cells or a peptide sequence that does not activate T cells (S1420).

[0173] The neural network to which the first input data and the second input data are input can predict the probability of T cell activation by performing an attention operation using a query, a key, and a value derived from each data (S1430).

[0174] Meanwhile, this embodiment may be configured in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, which, when executed by a processor, generates a program module to perform the operations described in this embodiment. The recording medium may be configured to be read by a computer.

[0175] Examples of recording media that can be read by a computer include various recording media that store instructions that can be deciphered by a computer, such as ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, and optical data storage devices.

[0176] The embodiments described herein have been described with reference to the accompanying drawings. Those skilled in the art in the technical field to which the present invention pertains can implement the present invention in forms different from the examples described herein without changing the technical idea or essential features. The embodiments in the present specification are merely examples and should not be interpreted as limiting.

Claims

1. 1. A neural network construction apparatus, comprising: at least one computer constructing a sequence transformation neural network for transforming an input sequence having respective network inputs; at least one memory; at least one processor in communication with the at least one memory; The at least one processor First input data is input; second input data corresponding to the first input data is input; training the sequence-transformation neural network by performing an attention operation based on predetermined label information and the first input data and the second input data that have been labeled; a neural network construction device that determines output data to be output by the sequence transformation neural network that has been trained based on the first input data, the second input data, and the label information;

2. The at least one processor determining a first key value and a first value by performing a predetermined pre-learning operation based on the first input data; Matching position information to each sequence in the second input data; generating a first query value by performing a predetermined self-attention operation based on a second key value, a second value value, and a second query value corresponding to the second input data; 2. The neural network construction device according to claim 1, wherein the output data corresponding to the first input data and the second input data is determined by executing the predetermined attention operation based on the first key value, the first value value, and the first query value.

3. The at least one processor 3. The neural network construction device according to claim 2, further comprising: a calculation for combining a plurality of attention heads generated by a predetermined calculation based on the first key value, the first value value, and the first query value corresponding to each sequence of the first input data and the second input data.

4. The at least one processor 3. The neural network construction device according to claim 2, wherein the output data includes attention energy related to the second input data corresponding to the first input data.

5. The at least one processor 5. The neural network construction device of claim 4, wherein the output data is generated by adding and combining information corresponding to an experimental method to an embedding vector determined by executing the predetermined attention operation based on the first key value, the first value value, and the first query value.

6. the first input data includes at least one of information on the type and structure of MHC; 3. The neural network construction device according to claim 2, wherein the second input data includes information on a plurality of peptide sequences.

7. The sequence transformation neural network 7. The neural network construction device according to claim 6, wherein the neural network construction device is constructed so as to predict the degree of binding between the MHC and the peptide sequence when the second input data includes at least one peptide sequence that is not bound to the MHC.

8. The at least one processor 7. The neural network construction device according to claim 6, wherein variables of the sequence-transformation neural network are changed using a loss function determined by binding information between the peptide sequence and the MHC corresponding to the first input data and the second input data included in the output data.

9. The second input data is 9. The neural network construction device according to claim 8, further comprising at least one of a substitution matrix (BLOSUM) of amino acids corresponding to the peptide sequence and physicochemical / biochemical properties (AAindex).

10. The first input data is 8. The neural network construction device according to claim 7, wherein the neural network construction device is obtained by a sequence corresponding to a predetermined range based on the binding origin between the peptide sequence and the MHC.

11. The sequence transformation neural network 7. The neural network construction device according to claim 6, wherein the neural network construction device is constructed so as to predict the degree of activation of the T cells matching the peptide sequence when the second input data includes at least one peptide sequence that does not activate the T cells.

12. The at least one processor performing a plurality of tests that output a degree of activation of the T cells corresponding to the first input data and the second input data; determining the number of tests and the number of activations of the T cells; generating a better distribution corresponding to the first input data and the second input data based on the number of times of activation of the T cells relative to the number of times of the plurality of tests; 12. The neural network construction device according to claim 11, wherein label information corresponding to the first input data and the second input data is determined based on a mean and a variance of the Better distribution.

13. 1. A learning method for transforming an input sequence, performed by a sequence-transformation neural network on a computer device, comprising: determining a first key value and a first value by performing a predetermined pre-learning operation based on first input data; inputting second input data corresponding to the first input data, and matching position information to each sequence in the second input data; generating a first query value by performing a predetermined self-attention operation based on a second key value, a second value value, and a second query value corresponding to the second input data; determining output data corresponding to the first input data and the second input data by performing the predetermined attention operation based on the first key value, the first value value, and the first query value; A sequence transformation neural network construction device and a learning method using the same, including:

14. A computer-readable recording medium, which is coupled to a computer that is hardware, and stores a program for executing the artificial intelligence-based sequence-transformation neural network of claim 13.

Citation Information

Patent Citations

  • Machine learning system for digital assistant

    JP2022003520A

  • Enhancing hybrid self-attention structure with relative-position-aware bias for speech synthesis

    US20200258496A1

  • Sequence model processing method and apparatus

    US20210174170A1

  • Systems and methods for evaluation of structure and property of polynucleotides

    WO2022159635A1