Data expansion method, device, and program for predicting MHC class II binding and immunogenicity models
The data augmentation method enhances the reliability of MHC class II binding and immunogenicity predictions by expanding and transforming training data, addressing the challenge of insufficient data in neural networks.
Patent Information
- Application Number
- JP2025533457
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-09
- Filing Date
- 2023-12-08
- Publication Date
- 2025-12-16
AI Technical Summary
Existing artificial intelligence-based neural networks face challenges in achieving highly accurate data predictions for MHC class II binding and immunogenicity due to insufficient training data, necessitating methods to augment and enhance the input data for improved reliability.
A data augmentation method that selects and extends original data based on specific conditions, applies sequence transformations, and adjusts labeling to generate expanded data, including adding or removing amino acids and normalizing labels, specifically for peptide features binding to MHC class II.
The expanded data improves the quality and reliability of predictions by enhancing the learning model's ability to predict MHC class II binding and immunogenicity, addressing the limitations of insufficient training data.
Smart Images

Figure 2025540817000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a data augmentation method and apparatus for expanding training data. More particularly, the present disclosure relates to a data augmentation method, apparatus, and program for a model that predicts MHC class II binding and immunogenicity. [Background technology]
[0002] In recent years, various concepts and learning models have been developed in the field of artificial intelligence technology, and research into data prediction using these concepts and models has been actively conducted.
[0003] However, when predicting data based on an artificial intelligence-based neural network, it is necessary to develop a learning model or estimation algorithm to obtain highly accurate results.
[0004] Furthermore, in order to improve the reliability of data prediction results, measures are being sought to input a larger number of data. Summary of the Invention [Problem to be solved by the invention]
[0005] The embodiments disclosed in the present disclosure aim to provide a data augmentation method, device, and program for a model predicting MHC class II binding and immunogenicity, which augments data input during training of artificial intelligence-based predictions.
[0006] The problems to be solved by the present disclosure are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0007] A data extension device according to one aspect of the present disclosure for solving the above-mentioned problems may include a memory and a processor in communication with the memory and configured to perform extension of original data to be learned, wherein the processor may be configured to perform the following steps: selecting, from the original data, a plurality of extension target data including first type data and second type data to be extended according to a predetermined selection condition; extending the selected plurality of extension target data according to the predetermined extension condition, wherein the first type data and the second type data are extended according to the extension condition of the first type data and the extension condition of the second type data, respectively, to generate a plurality of extended data; and changing labeling of the plurality of extended data, wherein the labels of the first type data and the second type data are changed according to different labeling conditions, respectively; wherein the original data may be peptide features for binding to MHC (Major Histocompatibility Complex) class II features.
[0008] In addition, when selecting multiple data to be extended, the processor may select a first type of data from the original data that includes at least one positive data that meets a first selection condition, and the first selection condition may be that the IC50 label is less than a predetermined concentration value and the peptide length is less than a predetermined number.
[0009] In addition, when selecting multiple data to be extended, the processor selects a second type of data from the original data that includes at least one negative data that meets a second selection condition, and the second selection condition may be that the IC50 label is greater than a predetermined concentration value and the peptide length is equal to or greater than a predetermined number.
[0010] Furthermore, when generating multiple pieces of extended data, the processor may randomly add any amino acid to each of the multiple pieces of data to be extended, where the randomly selected amino acid sequence may be added as a single sequence to the N-terminus of the original peptide sequence of the first type data, or the randomly selected amino acid sequence may be added as a single sequence to the C-terminus of the original peptide sequence of the first type data, or the randomly selected amino acid sequence may be added as a single sequence to each of the N-terminus and C-terminus of the original peptide sequence of the first type data, thereby extending the first type data among the multiple pieces of extended data.
[0011] Furthermore, when generating multiple pieces of extended data, the processor may add a sequence to each of the multiple pieces of extended data using an amino acid sequence pattern of a human protein, where one sequence may be added to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, one sequence may be added to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, or one sequence may be added to each of the N-terminus and C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, thereby extending the first type data among the multiple pieces of extended data.
[0012] In addition, when generating multiple extended data, the processor may remove sequences at both ends of the original peptide sequence until the length of the original peptide sequence of the second type data among the multiple extended data matches a predetermined number of sequences.
[0013] In addition, when changing the labeling of the multiple extended data, the processor may normalize the labels of the original data to which each of the multiple extended data corresponds, and obtain a final pseudo label according to the labeling conditions for each of the first type data and the second type data, where the final pseudo label may be calculated using a predetermined label constant value based on the normalized pseudo label of the original data and the binding affinity of the peptide to the MHC class II molecule.
[0014] The processor may compare the multiple augmented data with the original data to remove duplicate data.
[0015] In addition, a data extension method according to another aspect of the present disclosure may be a method executed by a computer device, and may include the steps of selecting, from original data, a plurality of extension target data including first type data and second type data to be extended according to predetermined selection conditions; extending the selected plurality of extension target data according to predetermined extension conditions, in which the first type data and the second type data are extended according to the extension conditions of the first type data and the extension conditions of the second type data, respectively, to generate a plurality of extended data; and changing the labeling of the plurality of extended data, in which the labels of the first type data and the second type data are changed according to different labeling conditions, respectively, and the original data may be peptide features for binding to MHC (Major Histocompatibility Complex) class II features.
[0016] In addition, when selecting multiple data to be extended, the data extension method may select a first type of data from the original data that includes at least one positive data that matches a first selection condition, and the first selection condition may be that the IC50 label is less than a predetermined concentration value and the peptide length is less than a predetermined number.
[0017] In addition, when selecting multiple data to be extended, the data extension method may select a second type of data from the original data that includes at least one negative data that meets a second selection condition, and the second selection condition may be that the IC50 label is greater than a predetermined concentration value and the peptide length is greater than or equal to a predetermined number.
[0018] Furthermore, when generating multiple pieces of extended data, the data extension method may randomly add any amino acids to each of the multiple pieces of extension target data, where a randomly selected amino acid sequence may be added as a single sequence to the N-terminus of the original peptide sequence of the first type data, a randomly selected amino acid sequence may be added as a single sequence to the C-terminus of the original peptide sequence of the first type data, or a randomly selected amino acid sequence may be added as a single sequence to each of the N-terminus and C-terminus of the original peptide sequence of the first type data, thereby extending the first type data among the multiple pieces of extension data.
[0019] Furthermore, when generating a plurality of pieces of extended data, the data extension method may add a sequence to each of the plurality of pieces of extension target data using an amino acid sequence pattern of a human protein, wherein one sequence may be added to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, one sequence may be added to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, or one sequence may be added to each of the N-terminus and C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein, thereby extending the first type data among the plurality of pieces of extension target data.
[0020] In addition, when generating multiple extended data, the data extension method can remove sequences from both ends of the original peptide sequence until the length of the original peptide sequence of the second type data among the multiple extended data matches a predetermined number of sequences.
[0021] In addition, when changing the labeling of the multiple extended data, the data extension method may normalize the labels of the original data to which the multiple extended data respectively correspond, and obtain final pseudo labels according to the labeling conditions for the first type data and the second type data, where the final pseudo labels may be calculated using a predetermined label constant value based on the normalized pseudo labels of the original data and the binding affinity of the peptide to the MHC class II molecule.
[0022] After changing the labeling of the plurality of augmented data, the data augmentation method may compare the plurality of augmented data with the original data to remove redundant data.
[0023] In addition, a computer program stored on a computer-readable recording medium for executing a method for implementing the present disclosure may further be provided.
[0024] In addition, a computer-readable recording medium having a computer program for executing the method for implementing the present disclosure recorded thereon can be further provided. [Effects of the Invention]
[0025] According to the solution of the present disclosure described above, the input data of the learning model for predicting binding and immunogenicity to MHC class II is expanded. This expansion is performed by selecting data to be expanded based on various conditions including IC50 label, which can improve the quality of the expanded data and also improve the reliability of the predicted results of binding and immunogenicity learned based on the expanded data.
[0026] The effects of the present disclosure are not limited to the effects described above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]
[0027] [Figure 1A] FIG. 1A is a diagram showing the structure of MHC class II according to the present disclosure.
[0028] [Figure 1B] FIG. 1B is an exemplary diagram for briefly explaining the data augmentation method according to the present disclosure.
[0029] [Figure 1C] FIG. 1C is a diagram showing the overall structure of a learning model according to the present disclosure.
[0030] [Figure 2] FIG. 2 is a diagram illustrating the configuration of a computer device according to the present disclosure.
[0031] [Figure 3] FIG. 3 is an exemplary diagram illustrating a method for selecting data to be extended according to the present disclosure.
[0032] [Figure 4] 4 to 6 are exemplary diagrams for explaining a method for extending data to be extended according to the present disclosure. [Figure 5] 4 to 6 are exemplary diagrams for explaining a method for extending data to be extended according to the present disclosure. [Figure 6] 4 to 6 are exemplary diagrams for explaining a method for extending data to be extended according to the present disclosure.
[0033] [Figure 7] FIG. 7 is an exemplary diagram illustrating the pseudo-labeling method according to the present disclosure.
[0034] [Figure 8] FIG. 8 is a flowchart illustrating the data extension method according to the present disclosure.
[0035] [Figure 9]9 and 10 are flowcharts for explaining the data extension method of FIG. 8 in detail. [Figure 10] 9 and 10 are flowcharts for explaining the data extension method of FIG. 8 in detail. DETAILED DESCRIPTION OF THE INVENTION
[0036] Throughout this disclosure, the same reference numerals refer to the same elements. This disclosure does not describe all elements of the embodiments, and content that is well known in the technical field to which the disclosure belongs or content that is duplicated between embodiments will be omitted. The terms "unit," "module," "member," and "block" used in this specification can be implemented in software or hardware, and depending on the embodiment, multiple "units," "modules," "members," and "blocks" may be implemented as a single component, or one "unit," "module," "member," or "block" may include multiple components.
[0037] Throughout this specification, when a first component is described as being "connected" to a second component, this includes not only when the first component is directly connected to the second component, but also when the first component is indirectly connected to the second component, where indirect connection includes when connected via a wireless communication network.
[0038] Also, when a part is described as "including" certain elements, unless otherwise specified, it means that it can further include other elements, rather than excluding other elements.
[0039] Throughout this specification, when a first member is described as being "on" a second member, this includes both when the first member is in contact with the second member, and when there is a third member between the two members.
[0040] Terms such as "first" and "second" are used to distinguish one component from another, and the components are not limited to the terms mentioned above.
[0041] Any reference to the singular includes any reference to the plural, unless the context clearly indicates otherwise.
[0042] The identification numbers in each process are used for convenience of explanation, and the identification numbers do not describe the order of each process, and each process may be performed in an order different from the order described unless the context clearly dictates a specific order.
[0043] Hereinafter, the principles of operation and embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0044] In this specification, a "data augmentation device according to the present disclosure" includes all of various devices that can perform computational processing and provide the results to a user. For example, the data augmentation device according to the present disclosure may include all of a computer, a server device, and a mobile terminal, or may take any one of these forms.
[0045] Here, the computer may include, for example, a notebook computer, a desktop computer, a laptop computer, a tablet PC, a slate PC, or the like equipped with a web browser.
[0046] The server device is a server that communicates with external devices and processes information, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, a web server, and the like.
[0047] The mobile terminal may include, for example, any kind of handheld-based wireless communication device such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, smartphone, etc., as a wireless communication device that ensures portability and mobility, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD), etc.
[0048] An "antigen" in this disclosure can be a substance that induces an immune response.
[0049] Neoantigens can refer to novel proteins formed in cancer cells when specific mutations occur in tumor DNA. Neoantigens arise through mutations and are characterized by their expression only in cancer cells. Neoantigens can include polypeptide or nucleotide sequences. Mutations can include frameshift or non-lattice shift indels, missense or nonsense substitutions, splice site alterations, genomic rearrangements or gene fusions, or any genomic or expression changes that result in novel open reading frames (ORFs). Mutations can also include splice variants. Tumor cell-specific post-translational modifications can include aberrant phosphorylation. Tumor cell-specific post-translational modifications can include spliced antigens generated by the proteasome.
[0050] "Epitope" in this disclosure may refer to the specific site on an antigen to which an antibody or T-cell receptor normally binds.
[0051] In the present disclosure, "major histocompatibility complex (MHC)" may refer to proteins that enable T cells to recognize specific cells by presenting "peptides" synthesized by those cells on their cell surface.
[0052] In the present disclosure, "peptide" refers to a polymer of amino acids. For convenience of explanation, in the following description, "peptide" can refer to an amino acid polymer or amino acid sequence expressed on the surface of cancer cells.
[0053] In the present disclosure, "MHC class II (major histocompatibility complex class II)" can be a protein that is expressed in antigen-presenting cells and regulates various immune responses by activating helper T cells.
[0054] The "MHC class II-peptide complex" in the present disclosure is expressed on the surface of antigen-presenting cells or cancer cells and may be a complex structure of MHC class II and a peptide. Helper T cells can recognize the MHC class II-peptide complex and initiate an immune response.
[0055] Cancer cells can produce neoantigens. MHC class II is mainly expressed in antigen-presenting cells. Antigen-presenting cells can degrade neoantigens produced by cancer, and epitopes derived from the neoantigens can be presented on the surface by MHC class II. Helper T cells recognize MHC class II-epitopes and induce immune responses. Therefore, to identify neoantigens produced by cancer cells, it is necessary to predict MHC-peptide binding affinity.
[0056] The present disclosure relates to expanding data input to a learning model that predicts whether MHC class II binds to a peptide sequence and T cell activation based on a sequence transformation neural network implemented through learning. A series of operations or algorithms for this can be executed by a computer device, and the detailed configuration of the computer device will be described below with reference to FIG. 2. In the present disclosure, the computer device may refer to a data expansion device.
[0057] FIG. 1A is a diagram showing the structure of MHC class II according to the present disclosure, and FIG. 1B is an exemplary diagram for briefly explaining the data expansion method according to the present disclosure.
[0058] MHC is a group of cell surface molecules that function as biochemical markers to distinguish individuals and can function as a mediator that allows substances that are the target of immune responses to be recognized as antigens.
[0059] MHC may be classified into MHC class I and MHC class II based on molecular structure, and the degree of difficulty in prediction may vary depending on differences in antigen-binding sites.
[0060] MHC class II may have binding sites composed of different substances and can bind to peptides of 13 to 17 amino acids.
[0061] As shown in Figure 1A, MHC class II, unlike MHC class I, can be composed of two chains, an α chain and a β chain, and can be formed in a chain-like structure with both ends open.
[0062] Because MHC class II has an open structure at both ends, peptides containing a core binding sequence can bind to MHC class II even if other sequences are added to both ends, whereas peptides lacking the core binding sequence will not bind to MHC class II even if the sequences at both ends are removed.
[0063]
[0064] 1B, the present disclosure may process data augmentation in a way that adds or removes sequences from the original peptide sequence based on the open structure of MHC class II. Specifically, in positive augmentation, both ends of the original sequence, P 13 and P1, one sequence (P r ) can be added. In negative augmentation, both ends of the original sequence are 13 The data may be expanded by removing n arrays D1, D2, D3-1, and D3-2 from P1 and P2, respectively, in which case the length of the final array of the expanded data may be limited to 10 or more.
[0065]
[0066] FIG. 1C is a diagram showing the overall structure of a learning model according to the present disclosure.
[0067] Referring to FIG. 1C, a learning model based on a sequence transformation neural network (NN) according to the present disclosure can accept MHC class II α chain features and MHC class II β chain features as first input data, and peptide features and extended peptide features as second input data.
[0068] The first input data may be subjected to predetermined pre-learning based on an alpha chain feature of MHC class II and a beta chain feature of MHC class II, thereby determining a first key and a first value for the first input data, and a first query for the first input data may be generated through multi-head self-attention processing based on second input data corresponding to the first input data.
[0069] Based on the first key, first value, and first query described above, a scaled dot product attention process may be performed, and by concatenating each attention head, a matrix in which each array is converted into a vector may be output as the first input data for learning.
[0070] The second input data may include peptide features (sequences) and extended peptide features, which may be amino acid features using both an amino acid substitution matrix (BLOSUM) and physicochemical properties (AA index).
[0071] Upon receiving the first input data and the second input data, the learning model can learn MHC class II binding affinity and immunogenicity according to the first input data and the second input data. In this case, MHC class II binding affinity may refer to the possibility of binding between a peptide sequence and MHC class II, and immunogenicity may refer to the presence or absence of T cell activation.
[0072] By the above-described process, an array transformation neural network (NN) for transforming an input array can be implemented.
[0073]
[0074] FIG. 2 is a diagram illustrating the configuration of a computer device according to the present disclosure.
[0075] The following description will be made with reference to FIG. 3, which is an exemplary diagram for explaining a method for selecting data to be extended according to the present disclosure, FIGS. 4 to 6, which are exemplary diagrams for explaining a method for extending data to be extended according to the present disclosure, and FIG. 7, which is an exemplary diagram for explaining a method for pseudo-labeling according to the present disclosure.
[0076] 2, a computer device 100 may include a memory 110, a processor 120, a communication interface 130, an input / output interface 140, and an input / output device 150. Although each component is shown as one in FIG. 2, this is for convenience of explanation, and each component may be provided as at least one, if necessary.
[0077] The memory 110 may store a computer program for providing a data expansion method, and the stored computer program may be read and executed by the processor 120. The memory 110 may store any type of information generated or determined by the processor 120 and any type of information received by the communication interface 130.
[0078] The memory 110 can store data supporting various functions of the computer device 100 and programs for the operation of the processor 120, can store input and output data (e.g., original data, data to be extended, extended data, etc.), can store multiple application programs or applications executed by the computer device 100, and data and commands for the operation of the computer device 100. At least some of the application programs can be downloaded from an external server via wireless communication.
[0079] Such memory 110 may include at least one type of storage medium among flash memory type, hard disk type, solid state disk type (SSD type), silicon disk drive type (SDD type), multimedia card micro type, card type memory (such as SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. The memory may also be a database separate from the device but connected by wire or wireless means.
[0080] 2, the processor 120 controls all components within the computer device 100, processes input or output signals, data, information, etc., and executes commands, algorithms, and application programs stored in the memory 110 to perform various processes and execute data expansion procedures, thereby providing or processing appropriate information or functions to each user. The components shown in FIG. 2 are not essential for implementing the computer device 100 according to the present disclosure, and the computer device 100 described above in this disclosure may include more or fewer components than those listed above. In this case, the computer device 100 may refer to a data expansion device.
[0081] The processor 120 may be in communication with the memory 110 and may perform augmentation of the original data to be learned.
[0082] The processor 120 may select a plurality of extension target data, including first type data and second type data to be extended, from the original data according to a predetermined selection condition, where the first type data may mean positive data, and the second type data may mean negative data.
[0083] The original data may be peptide features for binding with MHC (Major Histocompatibility Complex) class II features. In this case, the original data may include whether each peptide feature binds to MHC class II and its immunogenicity. Based on this information, processor 120 selects data to be augmented according to predetermined selection criteria.
[0084] The selection conditions may include a predetermined IC50 value, ie, a concentration value, and a predetermined number of peptide lengths.
[0085] As an example, referring to FIG. 3, when selecting a plurality of data to be extended, the processor 120 can select type 1 data including at least one or more positive data that meet the first selection condition from the original data.
[0086] The first selection condition can be a condition that the IC50 label is less than a predetermined concentration value and the peptide length is less than a predetermined number. For example, the first selection condition can be a condition that the IC50 label is greater than 0.01 nM and less than 500 nM (0.01 nM < IC50 label < 500 nM), and the peptide length is 10 or more and less than 20 (10 ≤ peptide length < 20).
[0087] In this case, the half maximal inhibitory concentration (IC) 50 can mean the maximum concentration when the activity of the cells (enzyme / protein activity) is halved when the drug is administered. In this case, the index indicating the activity of the cells can be a protein. The smaller the IC50 value, the higher the affinity can be meant.
[0088] That is, the selected type 1 data is a positive subset consisting of a plurality of positive data, the qualitative label can be positive high, and the quantitative label can be greater than 0.01 nM and less than 500 nM for IC50 and the peptide length can be 10 or more and less than 20. If the peptide length is too long, structural deformation may occur, so the peptide length can be restricted.
[0089] As another example, referring to FIG. 3, when selecting a plurality of data to be extended, the processor 120 can select type 2 data including at least one negative data that meets the second selection condition from the original data.
[0090] The second selection condition can be a condition where the IC50 label exceeds a predetermined concentration value and the peptide length is equal to or more than a predetermined number. For example, the second selection condition can be a condition where the IC50 label exceeds 50,000 nM and is less than 5,000,000 nM (50,000 nM < IC50 label < 5,000,000 nM), and the peptide length exceeds 11 and is equal to or less than 30 (11 < peptide length ≤ 30).
[0091] That is, the selected second type of data is a negative subset composed of a plurality of negative data, the qualitative label is negative, and the quantitative label can exceed 50,000 nM and be less than 5,000,000 nM for IC50, and the peptide length exceeds 11 and is equal to or less than 30.
[0092] The processor 120 can expand the selected plurality of data to be expanded according to the expansion conditions of the first type of data and the expansion conditions of the second type of data, thereby expanding the first type of data and the second type of data according to the predetermined expansion conditions to generate a plurality of expanded data.
[0093] The expansion conditions of the first type of data can include conditions where random expansion is applied to the positive subset and conditions where a human protein pattern is applied. Also, the expansion conditions of the second type of data can be conditions for removing the peptide sequence of the negative subset to shorten the length.
[0094] As an example, when generating a plurality of expanded data, the processor 120 can randomly add an arbitrary amino acid to each of the plurality of data to be expanded, and can do so as follows. In this case, the processor 120 can randomly add an arbitrary amino acid other than cysteine as a sequence.
[0095] Specifically, referring to Figure 4, the processor 120 may add a randomly selected amino acid sequence to the N-terminus of the original peptide sequence of the first type data as a single sequence A1. The processor 120 may also add a randomly selected amino acid sequence to the C-terminus of the original peptide sequence of the first type data as a single sequence A2. The processor 120 may also add randomly selected amino acid sequences to the N-terminus and C-terminus of the original peptide sequence of the first type data as a single sequence A3-1 and A3-2, respectively. The processor 120 can extend the first type data among the multiple extension target data using the method described above.
[0096] In this case, the number of original peptide sequences shown in Figure 4 is only an illustrative example and can be considered to include the binding core sequence.
[0097] As another example, when generating multiple extension data, the processor 120 adds a sequence to each of the multiple extension target data using an amino acid sequence pattern (4-mer) of a human protein, which can be done as follows.
[0098] Specifically, referring to FIG. 5 , the processor 120 may add one sequence (a) to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein (a, b, c, d, e, f). The processor 120 may add one sequence (f) to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein. The processor 120 may add one sequence (a, f) to each of the N-terminus and C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein. The processor 120 can extend the first type data of multiple extension target data using the above-described method.
[0099] For example, the processor 120 may add sequences to the N-terminus and C-terminus using a sequence pattern (4-mer) that appears more than 100 times in human proteins. According to this embodiment, it is possible to expect the effect of enabling data expansion similar to that in reality. In this case, the number of original peptide sequences shown in FIG. 5 is merely an example for explanation purposes and can be considered to include a binding core sequence.
[0100] 6, when generating a plurality of extended data, the processor 120 can remove sequences D6-1 and D6-2 from both ends of the original peptide sequence until the length of the original peptide sequence of the second type data among the plurality of data to be extended matches a predetermined number of sequences (N). In this case, the processor 120 can generate additional extended data each time it removes one sequence from both ends of the original peptide sequence.
[0101] The processor 120 may be implemented to change the labeling of the plurality of augmented data, but to change the labels according to different labeling conditions for each of the first type data and the second type data.
[0102] When changing the labeling of the plurality of augmented data, the processor 120 may normalize the labels of the original data corresponding to each of the plurality of augmented data, and obtain a final pseudo label according to the labeling conditions of the first type data and the second type data. In this case, changing the pseudo label is a technique for assigning the most probable label in the form of a virtual label, and is a method that can overcome the limitations of data with insufficient label values.
[0103] The final pseudo-label can be calculated based on the normalized pseudo-label of the original data and the binding affinity of the peptide to the MHC class II molecule using a predetermined label constant value.
[0104] Specifically, referring to FIG. 7, the processor 120 can, for convenience, normalize the original label of the original data corresponding to the extended data in the range from 0 to 1 and then change the normalized label to the final pseudo label. In this case, the label of the original data corresponding to the extended data can mean the label of the original data (data to be extended) before the extension of the extended data.
[0105] As described above, 0 can mean low affinity or low immunogenicity, and 1 can mean high affinity or high immunogenicity.
[0106] When the extended data is a positive subset of the first type of data, the processor 120 can calculate the final pseudo label (pseudo label > x - k) with a value exceeding the value obtained by subtracting a predetermined label constant value (k) from the normalized pseudo label (x) of the original data. In this case, the predetermined label constant value k may be determined as the constant with the highest peptide binding affinity performance obtained through the learning of the model for MHC class II molecules and may vary for each task. For example, in the case of binding affinity, the predetermined label constant value k may be 0.25, and in the case of immunogenicity, the predetermined label constant value k may be 0.15.
[0107] When the extended data is a negative subset of the second type of data, the processor 120 can calculate the final pseudo label (pseudo label < x - k) with a value less than the value obtained by subtracting a predetermined label constant (k) from the normalized pseudo label (x) of the original data.
[0108] That is, processor 120 can use 1-log(IC50) / log50000, where 1 represents high affinity, to change the label of the original data to a range between 0 and 1, and then proceed with the subsequent steps. In the present disclosure, a smaller IC50 value indicates higher affinity, while a larger normalized label indicates higher affinity.
[0109] The processor 120 may compare the multiple pieces of augmented data with the original data and remove duplicate data.
[0110] For example, if both the original data and the augmented data are A, the processor 120 may delete one of them, thereby obviating noise and unnecessary training procedures.
[0111] The processor 120 may perform a validation procedure by generating a validation set for validating a fair learning model, where the validation set may consist of only the original data without any augmentation.
[0112] Meanwhile, the computing device 100 may include one or more components that enable communication with external devices, and may include, for example, a communication interface 130 for wireless communication and an input / output interface 140 for wired communication.
[0113] Specifically, the communication interface 130 may transmit and receive signals based on wireless communication with external devices via the network 200. To this end, the communication interface 130 may include at least one wireless communication module, a short-range communication module, or the like.
[0114] First, the wireless communication module may include a Wi-fi module, a WiBro (Wireless broadband) module, as well as a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), WLAN (Wireless LAN), DLNA (registered trademark) (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), 4G, 5G, and 6G.
[0115] The short-range communication module may be configured for short-range communication and may support short-range communication using at least one of Bluetooth (registered trademark), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless Fidelity), Wi-Fi Direct, and Wireless Universal Serial Bus (Wireless USB) technologies.
[0116] The input / output interface 140 may be connected to the input / output device 150 in a wired manner, i.e., may transmit and receive signals based on wired communication. To this end, the input / output interface 140 may include at least one wired communication module, which may include various wired communication modules such as a LAN (Local Area Network) module, a WAN (Wide Area Network) module, or a VAN (Value Added Network) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), a recommended standard 232 (RS-232), power line communication, or a plain old telephone service (POTS).
[0117] Although not shown in the drawings, the data expansion device 100 of the present disclosure may further include an output unit and an input unit.
[0118] The output unit may display a user interface (UI) for providing the data augmentation results, etc. The output unit may output information in any format generated or determined by the processor 120 and information in any format received by the communication interface 130.
[0119] The output unit may include at least one of a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, and a 3D display. Some of these display modules may be transparent or light-transmitting so that the outside can be seen through them. This is sometimes called a transparent display module, and a representative example of this transparent display module is a TOLED (Transparent OLED).
[0120] The input unit may receive information input by a user. The input unit may include keys and / or buttons on a user interface for receiving the information input by the user, or physical keys and / or buttons. A computer program for controlling a display according to an embodiment of the present disclosure may be executed based on user input via the input unit.
[0121]
[0122] FIG. 8 is a flowchart illustrating the data extension method according to the present disclosure.
[0123] The processor 120 of the computer device 100 may select 1100 a plurality of extension target data, including first type data and second type data to be extended, from the original data according to a predetermined selection condition. In this case, the first type data may mean positive data, and the second type data may mean negative data. The original data may be peptide features for binding to MHC (Major Histocompatibility Complex) class II features.
[0124] Next, the processor 120 may extend the selected plurality of extension target data in accordance with the predetermined extension conditions by extending the first type data and the second type data, respectively, in accordance with the extension conditions of the first type data and the extension conditions of the second type data, thereby generating a plurality of extension data (1200).
[0125] The processor 120 may then relabel the plurality of augmented data, where the labels may be changed according to different labeling conditions for each of the first type data and the second type data (1300).
[0126] When changing the labeling of the plurality of extended data, the processor 120 may normalize the label of the original data corresponding to each of the plurality of extended data, and obtain a final pseudo label according to the labeling conditions of the first type data and the second type data. The final pseudo label may be calculated using a predetermined label constant value based on the normalized pseudo label of the original data and the binding affinity of the peptide to the MHC class II molecule.
[0127] The processor 120 may then compare the multiple augmented data with the original data to remove duplicate data (1400).
[0128] The processor 120 may generate a validation set for validating a fair learning model and perform a validation procedure (1500). In this case, the validation set may consist of only the original data without any augmentation.
[0129]
[0130] FIG. 9 is a flowchart for explaining in detail the data reinforcement method of FIG. 8, and will be explained by taking as an example a method for extending the extension target data of the first type data.
[0131] First, when selecting a plurality of data to be extended, the processor 120 can select type 1 data including at least one or more positive data that meet the first selection condition from the original data (2100).
[0132] The first selection condition can be a condition where the IC50 label is less than a predetermined concentration value and the peptide length is less than a predetermined number. For example, the first selection condition can be a condition where the IC50 label is greater than 0.01 nM and less than 500 nM (0.01 nM < IC50 label < 500 nM), and the peptide length is 10 or more and less than 20 (10 ≤ peptide length < 20).
[0133] Next, when generating a plurality of extended data, the processor 120 can randomly add an arbitrary amino acid to each of the plurality of data to be extended, as follows (2200). In this case, the processor 120 can randomly add an arbitrary amino acid other than cysteine as a sequence.
[0134] Specifically, referring to FIG. 4, the processor 120 can add a randomly selected amino acid sequence as one sequence A1 to the N-terminal of the original peptide sequence of the type 1 data. Also, the processor 120 can add a randomly selected amino acid sequence as one sequence to the C-terminal of the original peptide sequence of the type 1 data. Further, the processor 120 can add a randomly selected amino acid sequence as one sequence to each of the N-terminal and C-terminal of the original peptide sequence of the type 1 data. The processor 120 can extend the type 1 data among the plurality of data to be extended by the method described above.
[0135] In this case, the number of the original peptide sequences shown in FIG. 4 is only an example for explanation, and can also be regarded as including a binding core sequence.
[0136] As another example, when generating multiple extension data, the processor 120 adds a sequence to each of the multiple extension target data using an amino acid sequence pattern (4-mer) of a human protein, which can be done as follows (2300).
[0137] Specifically, referring to FIG. 5, the processor 120 may add one sequence (a) to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein. The processor 120 may add one sequence (f) to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein. The processor 120 may add one sequence (a, f) to each of the N-terminus and C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of a human protein. The processor 120 can extend the first type data of multiple extension target data using the above-mentioned method.
[0138] For example, the processor 120 may add sequences to the N-terminus and C-terminus using a sequence pattern (4-mer) that appears more than 100 times in human proteins. According to this embodiment, it is possible to expect the effect of enabling data expansion similar to that in reality. In this case, the number of original peptide sequences shown in FIG. 5 is merely an example for explanation purposes and can be considered to include a binding core sequence.
[0139] Processor 120 can then execute steps 1300 and onward in FIG.
[0140]
[0141] FIG. 10 is a flowchart for explaining in detail the data extension method of FIG. 8, and will be explained by taking as an example a method for extending extension target data of the second type data.
[0142] Referring to FIG. 10, when selecting a plurality of data to be extended, the processor 120 can select type 2 data including at least one negative data that meets the second selection condition from the original data (3100).
[0143] The second selection condition can be a condition that the IC50 label exceeds a predetermined concentration value and the peptide length is more than a predetermined number. For example, the second selection condition can be a condition that the IC50 label is more than 50000 nM and less than 5000000 nM (50000 nM < IC50 label <5000000> nM), and the peptide length is more than 11 and less than or equal to 30 (11 < peptide length ≤ 30).
[0144] Next, referring to FIG. 6, when generating a plurality of extended data, the processor 120 can remove sequences from both ends of the original peptide sequence until the length of the original peptide sequence of the type 2 data among the plurality of data to be extended matches a predetermined number of sequences (N) (3200). In this case, each time the processor 120 removes one sequence from both ends of the original peptide sequence, additional extended data can be generated.
[0145] Next, the processor 1 can execute the steps after step 1300 in FIG. 8.
[0146] <000×0482> On the other hand, since the method according to the present disclosure described above is executed by hardware such as a server, it can be implemented as a program (or application) and stored in a medium.
[0148] The disclosed embodiment can be implemented in the form of a recording medium storing computer-executable commands. The commands can be stored in the form of program code, and when executed by a processor, program modules can be generated to execute the processing of the disclosed embodiment. The recording medium can be implemented as a computer-readable recording medium.
[0149] The computer-readable recording medium includes any type of recording medium that stores computer-interpretable instructions, such as a read-only memory (ROM), a random access memory (RAM), a magnetic tape, a magnetic disk, a flash memory, an optical data storage device, etc.
[0150] As described above, the disclosed embodiments have been described with reference to the accompanying drawings. Those skilled in the art in the technical field to which the present disclosure pertains will understand that the present disclosure can be implemented in a manner different from the disclosed embodiments without changing the technical idea or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be interpreted as limiting.
Claims
1. A data expansion device, Memory and a processor in communication with the memory and configured to perform augmentation of the original data to be learned; The processor: selecting a plurality of extension target data including first type data and second type data to be extended from the original data in accordance with a predetermined selection condition; a step of extending the selected plurality of extension target data in accordance with a predetermined extension condition, in which the first type data and the second type data are extended in accordance with the extension condition of the first type data and the extension condition of the second type data, respectively, to generate a plurality of extension data; changing the labeling of the plurality of extension data, wherein the labels of the first type data and the second type data are changed according to different labeling conditions; configured to run The data expansion device, wherein the original data is peptide features for binding to MHC (Major Histocompatibility Complex) class II features.
2. 2. The data expansion device according to claim 1, When selecting the plurality of extension target data, The processor selects the first type data including at least one positive data item that meets a first selection condition from the original data; The first selection condition is that the IC50 label is less than a predetermined concentration value and the peptide length is equal to or less than a predetermined number.
3. 2. The data expansion device according to claim 1, When selecting the plurality of extension target data, The processor selects the second type data from the original data that includes at least one negative data that meets a second selection condition; The second selection condition is that the IC50 label is greater than a predetermined concentration value and the peptide length is equal to or greater than a predetermined number.
4. 2. The data expansion device according to claim 1, When generating the plurality of extension data, The processor randomly adds any amino acid to each of the plurality of extension target data, wherein: a data extension device that extends the first type data among the plurality of extended data by adding a randomly selected amino acid sequence as a single sequence to the N-terminus of the original peptide sequence of the first type data, adding a randomly selected amino acid sequence as a single sequence to the C-terminus of the original peptide sequence of the first type data, and adding a randomly selected amino acid sequence as a single sequence to each of the N-terminus and the C-terminus of the original peptide sequence of the first type data.
5. 2. The data expansion device according to claim 1, When generating the plurality of extension data, The processor adds a sequence to each of the plurality of pieces of extension target data using an amino acid sequence pattern of a human protein, wherein: a data extension device that adds one sequence to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein, adds one sequence to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein, and adds one sequence to each of the N-terminus and the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein, thereby extending the first type data among the plurality of data to be extended.
6. 2. The data expansion device according to claim 1, When generating the plurality of extension data, The processor removes sequences from both ends of the original peptide sequence of the second type data among the plurality of extension target data until the length of the original peptide sequence matches a predetermined number of sequences.
7. 2. The data expansion device according to claim 1, When changing the labeling of the plurality of augmented data, The processor normalizes labels of the original data corresponding to each of the plurality of augmented data, and obtains a final pseudo label according to the labeling conditions for each of the first type data and the second type data, wherein: a data augmentation device that calculates the final pseudo-label based on the normalized pseudo-label of the original data and the binding affinity of the peptide to the MHC class II molecule using a predetermined label constant value.
8. 2. The data expansion device according to claim 1, The processor compares the plurality of pieces of extended data with the original data and deletes duplicate data.
9. 1. A data augmentation method executed by a computer device, comprising: selecting a plurality of extension target data including first type data and second type data to be extended from the original data in accordance with a predetermined selection condition; a step of extending the selected plurality of extension target data in accordance with a predetermined extension condition, in which the first type data and the second type data are extended in accordance with the extension condition of the first type data and the extension condition of the second type data, respectively, to generate a plurality of extension data; changing the labeling of the plurality of extension data, wherein the labels of the first type data and the second type data are changed according to different labeling conditions; and The data augmentation method, wherein the original data is peptide features for binding to MHC (Major Histocompatibility Complex) class II features.
10. 10. The data extension method according to claim 9, When selecting the plurality of extension target data, The data extension method includes selecting, from the original data, the first type data including at least one positive data item that meets a first selection condition; A data augmentation method, wherein the first selection condition is that the IC50 label is less than a predetermined concentration value and the peptide length is equal to or less than a predetermined number.
11. 10. The data extension method according to claim 9, When selecting the plurality of extension target data, The data extension method includes selecting, from the original data, the second type data including at least one negative data that meets a second selection condition; The second selection condition is that the IC50 label is greater than a predetermined concentration value and the peptide length is equal to or greater than a predetermined number.
12. 10. The data extension method according to claim 9, When generating the plurality of extension data, The data extension method randomly adds any amino acid to each of the plurality of extension target data, wherein: A data extension method for extending the first type data among the plurality of extended data by adding a randomly selected amino acid sequence as a single sequence to the N-terminus of the original peptide sequence of the first type data, adding a randomly selected amino acid sequence as a single sequence to the C-terminus of the original peptide sequence of the first type data, and adding a randomly selected amino acid sequence as a single sequence to each of the N-terminus and the C-terminus of the original peptide sequence of the first type data.
13. 10. The data extension method according to claim 9, When generating the plurality of extension data, The data augmentation method includes adding a sequence to each of the plurality of augmentation target data using an amino acid sequence pattern of a human protein, wherein: A data extension method for extending the first type data among the plurality of extension target data by adding one sequence to the N-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein, adding one sequence to the C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein, and adding one sequence to each of the N-terminus and C-terminus of the original peptide sequence of the first type data using the amino acid sequence pattern of the human protein.
14. 10. The data extension method according to claim 9, When generating the plurality of extension data, The data extension method includes removing sequences at both ends of the original peptide sequence until the length of the original peptide sequence of the second type data among the plurality of extension target data matches a predetermined number of sequences.
15. 10. The data extension method according to claim 9, When changing the labeling of the plurality of augmented data, The data augmentation method normalizes labels of the original data corresponding to the plurality of augmented data, and obtains final pseudo labels according to the labeling conditions for the first type data and the second type data, wherein: The data augmentation method includes calculating the final pseudo-labels based on the normalized pseudo-labels of the original data and the binding affinity of peptides to MHC class II molecules using a predetermined label constant value.
16. 10. The data extension method according to claim 9, After changing the labeling of the plurality of augmented data, The data extension method includes comparing the plurality of extended data with the original data and deleting duplicate data.
17. A program stored on a computer-readable recording medium for causing a computer to execute the data extension method according to any one of claims 9 to 16.