Data augmentation method, data augmentation device, learning method, and cell type classification system
The data augmentation method addresses the issue of preserving local structures in non-sequence data by using neighborhood relationships and graph-based techniques, enhancing learning model accuracy by generating relevant data.
Patent Information
- Application Number
- PCT/JP2024/043755
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing data augmentation methods, such as CutMix, fail to adequately consider the relationships between variables in non-sequence data, leading to the generation of pseudo-data that deviates from actual data, particularly in tabular data where the local structure is not positively defined, which can adversely affect learning performance.
A data augmentation method that extracts and maintains relationships between variables by selecting and replacing values based on neighborhood relationships within datasets, using techniques like graph construction and clustering to ensure that the local structure is preserved, allowing for data augmentation across different classes and datasets.
This approach effectively generates new data that maintains the local structure of the original data, reducing the likelihood of pseudo-data generation and improving the accuracy of learning models, especially in non-sequence data like tabular data.
Smart Images

Figure JP2024043755_24072025_PF_FP_ABST
Abstract
Description
Data augmentation method, data augmentation device, learning method, and cell type classification system
[0001] The present invention relates to a data augmentation method and device, a learning method, and a cell type classification system.
[0002] Currently, the development of AI (artificial intelligence) using machine learning models and statistical models is gaining momentum. In such AI development, the amount of training data used to train the model is an important factor that greatly influences the final AI performance. Generally, the more data used for training, the better. However, securing large amounts of target data in the real world is considered a difficult challenge in terms of cost and time. Therefore, a technique called "data augmentation" is widely used, in which existing data is processed in some way to increase the amount of training data.
[0003] For example, Patent Document 1 describes a method or system called CutMix, which generates new training data from two sample images. CutMix is a method for expanding training data by cutting and pasting regions cut out at a given ratio from two training data images.
[0004] Non-Patent Document 1 also describes a method of applying the above-mentioned CutMix to ratio data, i.e., data (especially bacterial flora data) in which the sum of the values of the variables of each sample is 1. The differences from CutMix are that (1) a process of selecting multiple variables instead of selecting pixel regions, and (2) a process of normalizing the variables of the generated data so that the sum is 1 have been newly added.
[0005] Furthermore, Patent Document 2 describes a teacher data extension device, a program, and a method, in which the teacher data extension device includes a relationship acquisition unit that determines the relationship between multiple features included in each of multiple teacher data, a feature selection unit that selects one or more of the multiple features based on the relationship, and a teacher data extension unit that generates new teacher data by replacing the feature values selected by the feature selection unit for one or more teacher data with the feature values of other teacher data classified into the same class.In particular, the document describes selecting two or more uncorrelated or independent features that can be combined in actual teacher data with the same label information, and replacing the selected features with other teacher data.
[0006] Patent Literature 3 describes a data extension method for time series data to which first labels representing characteristics in the time series are assigned, in which a regularity of the first labels is extracted based on a correlation coefficient between minimum constituent unit data and other minimum constituent unit data, and the time series data is extended so as to maintain the regularity. Patent Literature 3 also describes a data extension method in which second labels having a smaller number of labels than the first labels are assigned, and minimum constituent data are selected and combined based on the regularity of the second labels so as to maintain the regularity of the second labels.
[0007] Japanese Patent Publication No. 2020-187736 Japanese Patent No. 7063393 Japanese Patent No. 7310908
[0008] "Data Augmentation for Compositional Data: Advancing Predictive Models of the Microbiome", NeurIPS, 2022, Gordon-Rodriguez E et al., [Retrieved December 11, 2023], Internet (https: / / arxiv.org / abs / 2205.09906)
[0009] CutMix, a representative data augmentation method, replaces specific regions of an image with the same parts of another image, which is said to improve local structure recognition performance. The "local structure of data" here refers to the existence of subsets with particularly strong correlations when considering the correlations within the variable sets that make up a given piece of data. For example, if a full-body image of a cat is used as a single piece of data, the variable set corresponds to the pixels contained in the image, and local structures in this case would be the cat's eyes, nose, etc.
[0010] CutMix contributes to improving the recognition performance of local structures, but to achieve this, it is important that the local structure be preserved during the data augmentation process. For data in which the arrangement of variables is explicitly defined temporally, spatially, or sequentially, such as image data, time series, or sequence data including character strings, maintaining the sequence is sufficient to maintain the local structure (see, for example, Patent Documents 1 and 3 and Non-Patent Document 1 mentioned above). However, typical CutMix techniques encounter problems when augmenting non-sequential data, such as table data, where the local structure is not explicitly defined. In other words, replacing variables with those from different samples can destroy the local structure, potentially resulting in features that differ from the actual data.
[0011] Figure 17 shows an example of how the local structure of data collapses when there is a one-to-one relationship. The example in Figure 17 shows a state in which there is a positive correlation between marker 1 and marker 2, and the ellipse in the figure represents the approximate range in which the data is distributed. In the example in the figure, when considering a CutMix of the data indicated by the white circles (marker 1 = 0.8, marker 2 = 0.8) and the data indicated by the black circles (marker 1 = 0.2, marker 2 = 0.2), pseudo-data that deviates from the actual data, such as the data indicated by the white crosses (marker 1 = 0.8, marker 2 = 0.2) and the data indicated by the black crosses (marker 1 = 0.2, marker 2 = 0.2), may be generated. Such data that deviates from the actual data does not contribute to learning and may even have a negative impact, so it is desirable not to use it for learning.
[0012] However, in actual data that requires data augmentation, hundreds to thousands of variables may be correlated with each other, making it difficult to maintain the local structure of the data through simple combinations.
[0013] Furthermore, in the technology described in Patent Document 2, only uncorrelated or independent features are selected or replaced, and when replacing variables, two pieces of data are selected only from data pairs classified into the same class.
[0014] As described above, the conventional techniques disclosed in the above-mentioned Patent Documents 1 to 3 and Non-Patent Document 1 do not appropriately consider the relationship between variables in data augmentation.
[0015] The present invention has been made in consideration of the above circumstances, and one embodiment of the present invention aims to provide a data augmentation method and data augmentation device that can perform data augmentation based on the relationships between variables, as well as a learning method and cell type classification system that use such technology.
[0016] A data expansion method according to a first aspect of the present invention is a data expansion method executed by a data expansion device including a processor. The processor acquires a training dataset, extracts relationships between multiple variables contained in the entire training dataset, selects one or more sets of variables from the multiple variables based on the relationships, determines subsets consisting of two or more teacher data sets in the training dataset, and replaces the values of the selected one or more sets of variables with the values of the same variables in source teacher data, which is teacher data other than the target teacher data set in the subset, to generate new teacher data. According to the first aspect, data expansion is performed based on the relationships between multiple variables, making it difficult to generate pseudo data that deviates from the actual data. Note that in the first aspect, it is preferable to select one or more sets of variables from multiple variables that are highly (or strongly) related.
[0017] In the data augmentation method according to the second aspect, in the first aspect, a processor extracts a relationship using a second dataset different from the training dataset or a combined dataset combining the training dataset and the second dataset, and selects and generates new training data based on the relationship using the second dataset or the combined dataset. In the second aspect, the training dataset and the second dataset being "different" includes cases where at least one of the measurement means, device, environment, and measurement conditions is different, even when the same variables or samples are measured. According to this second aspect, it is possible to use different datasets, multiple datasets, or combined datasets, thereby generating diverse data.
[0018] In the data augmentation method according to the third aspect, in the first aspect, a processor extracts, as a relationship, a distance or a similarity between any two pairs of variables contained in a training dataset. The third aspect specifies a specific aspect of the relationship between the variables.
[0019] In the data augmentation method according to the fourth aspect, in the second aspect, the processor extracts, as a relationship, a distance or a similarity between any two pairs of variables contained in the second dataset or the combined dataset. The fourth variable defines another specific aspect of the relationship between the variables.
[0020] A data augmentation method according to a fifth aspect is the third or fourth aspect, in which the processor extracts variables whose distance is less than a first threshold and / or variables whose similarity is greater than a second threshold as variables that satisfy a condition for the relationship. The fifth aspect specifically defines the condition for extracting variables.
[0021] In the data augmentation method according to the sixth aspect, in the first or third aspect, a processor extracts neighborhood relationships in a local structure constituted by one or more variables included in a training dataset based on the relationships, and selects one or more sets of variables based on the neighborhood relationships. In the sixth aspect, one or more sets of variables that satisfy a condition for the neighborhood relationships can be selected. The condition for the neighborhood relationships can be, for example, that the strength of the relationship between the variables is equal to or greater than a predetermined level. Note that when selecting "one or more sets of variables that satisfy the condition for the neighborhood relationships," it is not necessary to select all variables that satisfy the condition.
[0022] In the data augmentation method according to the seventh aspect, in the second or fourth aspect, the processor extracts a neighborhood relationship in a local structure constituted by one or more variables included in the second dataset or the combined dataset based on the relationship, and selects one or more sets of variables based on the neighborhood relationship. In the seventh aspect, as in the sixth aspect, one or more sets of variables that satisfy a condition for the neighborhood relationship can be selected. Also, as in the sixth aspect, it is not necessary to select all variables that satisfy the condition for the neighborhood relationship.
[0023] A data augmentation method according to an eighth aspect is the sixth or seventh aspect, in which the processor extracts neighborhood relationships by clustering a plurality of variables based on the relationships. The eighth aspect specifically defines one aspect of a neighborhood relationship extraction technique.
[0024] A data expansion method according to a ninth aspect is any one of the sixth to eighth aspects, in which a processor constructs a first graph in which variables belonging to a plurality of variables are represented as points and relationships are represented as edges, and extracts, as a neighborhood relationship, a set of points or a part of the set that can be reached by tracing the points connected by the edges from each point in the first graph. The ninth aspect specifically defines another aspect of the neighborhood relationship extraction method. In the ninth aspect, a set of points (a set of points that can be reached by tracing the points connected by the edges) or a part of the set of points can be extracted as points (variables) that satisfy a condition for the neighborhood relationship.
[0025] In a data expansion method according to a tenth aspect, in the ninth aspect, a processor constructs a second graph in which the edges of a first graph are weighted by the magnitude of the relationship, and extracts, as a neighborhood relationship, a set of points or a part of the set that can be reached by tracing the points connected by the edges from each point in the second graph while taking the weights into consideration. The tenth aspect further specifies the neighborhood relationship extraction method in the ninth aspect. Note that in the tenth aspect, a set of points (a set of points that can be reached by tracing the points connected by the edges while taking the weights into consideration) or a part of the set of points can also be extracted as points (variables) that satisfy conditions for neighborhood relationships.
[0026] In a data augmentation method according to an eleventh aspect, in any one of the sixth to tenth aspects, a processor selects one or more variables based on a neighborhood relationship. In the eleventh aspect, variables that satisfy a condition for the neighborhood relationship can be selected.
[0027] In a data augmentation method according to a twelfth aspect, in any one of the first to eleventh aspects, the processor retains a first label for each training data set included in the training dataset, replaces values of selected variables in one or more training data sets in a subset consisting of two or more training data sets with values of the same variables in other training data sets in the subset, and generates new training data sets. For each training data set, the processor mixes the first label of the training data set with the second label of the original training data set with a weight corresponding to the number of variables replaced from each of the original training data sets. As described in the second aspect, the processor may generate labels for the new training data sets using a second dataset different from the training dataset or a combined dataset combining the training dataset and the second dataset.
[0028] A data augmentation method according to a thirteenth aspect is any one of the first to twelfth aspects, in which a processor generates new teacher data using replacement teacher data and replacement source teacher data that belongs to a class different from the class to which the replacement teacher data belongs. Conventional techniques such as those described in Patent Document 2 above select or replace only uncorrelated or independent features, and two pieces of data when replacing variables are selected only from data pairs classified into the same class. However, the present invention also allows data augmentation to be performed using data that belong to different classes.
[0029] A data augmentation method according to a fourteenth aspect is any one of the first to thirteenth aspects, wherein the training dataset is a dataset composed of non-sequential data. Non-sequential data is data expressed in a tabular format in which the arrangement of variables is not explicitly defined in terms of time, space, or order, and the order of the variables is essentially meaningless. While conventional techniques such as those described in Patent Documents 1 and 3 and Non-Patent Document 1 are difficult to apply to non-sequential data, the present invention provides advantageous effects not found in conventional techniques in data augmentation of non-sequential data (e.g., no pseudo-data deviating from the actual data is generated).
[0030] A fifteenth aspect of the data augmentation method is any one of the second, fourth, and seventh aspects, wherein the second dataset and the combined dataset are datasets composed of non-sequential data. In the fifteenth aspect, the meaning of non-sequential data is the same as in the fourteenth aspect.
[0031] A data extension method according to a sixteenth aspect is any one of the first to fifteenth aspects, wherein the processor displays information indicating the relationships on a display device, thereby allowing a user to easily understand the relationships between variables.
[0032] A data augmentation method according to a seventeenth aspect is any one of the sixth to eleventh aspects, wherein the processor displays information indicating the neighborhood relationships on a display device, thereby enabling a user to easily understand the neighborhood relationships of the variables.
[0033] In a data augmentation method according to an eighteenth aspect, in the tenth aspect, the processor displays the first graph and / or the second graph on a display device, allowing a user to easily understand the neighborhood relationships of the variables.
[0034] A data augmentation method according to a nineteenth aspect is any one of the first to eighteenth aspects, wherein the plurality of variables are variables measurable in a biological sample. Variables measurable in a biological sample are often non-sequential data, and the data augmentation method of the present invention is effective.
[0035] The data augmentation method according to the twentieth aspect is the 19th aspect, wherein the measurable variables in a biological sample are any one or a combination of methylation status, mutations, and gene expression levels at one or more locations in nucleic acid extracted from a living organism.
[0036] In addition, a program that causes a computer to execute a data extension method relating to any one of aspects 1 to 20, and a non-transitory, tangible recording medium that records computer-readable code of such a program can also be cited as aspects of the present invention.
[0037] A training method according to a 21st aspect is a training method for a cell type classification system including a processor, wherein the processor acquires a training dataset consisting of training data obtained by measuring the methylation state and / or nucleic acid sequence mutations of one or more samples of cell-free free DNA, each training data being assigned a first label indicating a cell type, and augments the training dataset using the data augmentation method according to any one of aspects 1 to 20 to generate an augmented training dataset including the training dataset and a second training dataset newly generated by the data augmentation, each training data being assigned a first label, and trains a classifier that classifies and predicts the first label using the augmented training dataset. According to the 21st aspect, the classifier is trained using a dataset that has been data augmented based on the relationship between multiple variables and is less likely to include pseudo data that deviates from the actual data, thereby enabling the construction of a classifier that is efficient and has high classification prediction accuracy.
[0038] In addition, a program that causes a computer to execute the learning method according to the 21st aspect, and a non-transitory, tangible recording medium on which computer-readable code of such a program is recorded, can also be cited as aspects of the present invention.
[0039] A cell type classification system according to a 22nd aspect includes a classifier constructed by the learning method according to the 21st aspect, thereby enabling highly accurate classification prediction of cell type data.
[0040] A cell type classification system according to a 23rd aspect is the 22nd aspect, wherein the classifier predicts a classification of cancer type or healthy as the first label.
[0041] A data extension device according to a 24th aspect is a data extension device including a processor, in which the processor acquires a training dataset, extracts relationships between multiple variables contained in the entire training dataset, selects one or more sets of variables from the multiple variables based on the relationships, determines a subset consisting of two or more teacher data in the training dataset, and replaces the values of the selected one or more sets of variables for replacement teacher data, which is one or more teacher data in the subset, with the values of the same variables in source teacher data, which is teacher data other than the replacement teacher data in the subset, to generate new teacher data.
[0042] The data extension device according to the twenty-fourth aspect may have a configuration for executing the same processing as the data extension methods according to the second to twentieth aspects.
[0043] As described above, according to the data augmentation method and data augmentation device of the present invention, and the learning method and cell type classification system using such technology, data augmentation can be performed based on the relationships between variables.
[0044] FIG. 1 is a diagram showing the overall configuration of a data extension device according to a first embodiment. FIG. 2 is a diagram showing the configuration of a data processing unit. FIG. 3 is a diagram showing information recorded in a recording device. FIG. 4 is a conceptual diagram showing data extension according to the present invention. FIG. 5 is a diagram showing the relationship between variables and the process of extracting neighborhoods. FIG. 6 is a diagram showing the process of extracting neighborhoods using a graph. FIG. 7 is a diagram showing the process of selecting variables. FIG. 8 is a diagram showing the process of generating new training data. FIG. 9 is a diagram showing sample values used in the examples. FIG. 10 is a diagram showing the correlation between markers. FIG. 11 is a diagram showing the results of data extension according to a conventional method (comparative example). FIG. 12 is a diagram showing the results of data extension according to the present invention (part 1). FIG. 13 is a diagram showing the results of data extension according to the present invention (part 2). FIG. 14 is a partially enlarged view of the results of data extension according to a conventional method (comparative example). FIG. 15 is a diagram showing the configuration of a cell type classification system according to a second embodiment. FIG. 16 is a diagram showing the configuration of a data processing unit in the cell type classification system. FIG. 17 is a diagram for explaining the issues with data extension according to the conventional method.
[0045] Hereinafter, embodiments of a data augmentation method, a data augmentation device, a learning method, and a cell type classification system according to the present invention will be described. In the description, the accompanying drawings will be referred to as necessary.
[0046] [First Embodiment] [Data Expansion Device] FIG. 1 is a block diagram showing the configuration of a data expansion device 10 (data expansion device) according to a first embodiment. The data expansion device 10 is a device that expands training data (a data expansion method) and can be implemented using a computer. As shown in FIG. 1, the data expansion device 10 includes a data processing unit 100, a recording device 200, a display device 300 (output device), and an operation unit 400 (input device), which are interconnected to transmit and receive necessary information. These components can be installed in various configurations, and each component may be installed in a single location (e.g., within a single enclosure or room) or in separate locations and connected via a network. The data expansion device 10 is also connected to an external server 500 and an external database 510 via a network NW such as a LAN or the Internet, and can acquire training data, etc., used for data expansion as needed.
[0047] 2 is a diagram showing the configuration of the data processing unit 100 (processor). The data processing unit 100 includes a processor 110, a ROM 140 (Read Only Memory), and a RAM 150 (Random Access Memory). The processor 110 has functions as a relationship extraction unit 112, a neighborhood extraction unit 114, a variable selection unit 116, a data generation unit 118, and an input / output control unit 120.
[0048] [Overview of Functions of the Data Processing Unit] The relationship extraction unit 112 extracts relationships between multiple variables included in the entire dataset (e.g., the training dataset 202, the second dataset 204, the combination dataset 206, etc.). The neighborhood extraction unit 114 extracts neighborhood relationships in a local structure formed by one or more variables included in a dataset such as the training dataset 202 based on the relationships between the multiple variables. The variable selection unit 116 selects one or more sets of variables from multiple variables based on the extracted relationships (relationships between multiple variables). The data generation unit 118 generates new training data for the dataset. The input / output control unit 120 controls input / output (acquisition from an external device, transmission to an external device, input / output to the recording device 200, display on the display device 300, etc.) of datasets, datasets and data expansion results, variable relationships and neighborhoods, data expansion conditions, etc. Details of processing using these functions will be described later.
[0049] The functions of each part of the data processing unit 100 (processor 110) described above can be realized using various processors. The various processors include, for example, a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) to realize various functions. The various processors also include a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacturing. Furthermore, the various processors also include dedicated electrical circuits, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically to execute specific processing.
[0050] The functions of each unit may be realized by a single processor or by a combination of multiple processors. Furthermore, multiple functions may be realized by a single processor. Examples of multiple functions configured by a single processor include: a first configuration, as typified by client and server computers, in which a single processor is configured by combining one or more central processing units (CPUs) and software, and this processor realizes multiple functions; a second configuration, as typified by system-on-chip (SoC), in which a processor is used to realize the functions of the entire system on a single integrated circuit (IC) chip; and various functions are thus configured as hardware structures using one or more of the various processors described above. Furthermore, the hardware structures of these various processors are, more specifically, electrical circuits (circuitry) that combine circuit elements such as semiconductor devices.
[0051] When the above-described processor or electrical circuit executes software (programs), processor-readable code for the software to be executed is stored in a non-transitory, tangible recording medium such as ROM 140, and the processor references the software. The software stored in the non-transitory, tangible recording medium includes a program (compound search program) for executing the compound search method of the present invention. Instead of ROM 140, the code may be recorded in a non-transitory, tangible recording medium such as various magneto-optical recording devices or semiconductor memories. When processing using the software, for example, RAM 150 is used as a temporary storage area, and programs and / or data stored in non-volatile memory (non-transitory, tangible recording medium) such as an EEPROM (Electrically Erasable and Programmable Read Only Memory) or a flash memory (not shown) may be referenced.
[0052] It should be noted that the above-mentioned "non-transitory and tangible recording medium" does not include non-tangible recording media such as the carrier signal itself and the propagated signal itself.
[0053] The processing using these units of the data processing unit 100 will be described in detail later.
[0054] [Configuration of Recording Device] The recording device 200 is composed of a non-transitory and tangible recording medium such as various types of magneto-optical recording media and semiconductor memory, and its input / output control unit. In the recording device 200 of this embodiment, as shown in Fig. 3, for example, a training dataset 202 (training dataset), a second dataset 204 (second dataset), a combination dataset 206 (combination dataset), relationship information 208, neighborhood relationship information 210, and a data expansion condition 212 are associated and recorded. The data expansion device 10 can perform data expansion using the method of the present invention using the training dataset 202, the second dataset 204, and the combination dataset 206.
[0055] The second data set 204 is a data set different from the training data set 202, and the combined data set 206 is a data set obtained by combining the training data set 202 and the second data set 204. Each data set may be assigned a label.
[0056] Furthermore, the recording device 200 may record a second training data set and an extended training data set in addition to the training data set 202. The second training data set is a training data set that is newly generated by data augmenting the training data set 202 using the data augmentation method according to the present invention, and each piece of teacher data is assigned the same first label as that of the training data set 202. The extended training data set is a data set that includes the training data set 202 and the second training data set.
[0057] The data expansion conditions 212 may include information that identifies the dataset to be subjected to data expansion, a method for extracting variable relationships and neighborhood relationships, and values of various parameters used in data expansion (e.g., the number of new training data to be generated by data expansion, thresholds for similarity and path length, the number of variables to be selected, and the ratio of the number of variables to be selected).
[0058] Furthermore, the data processing unit 100 (processor 110) can display the data recorded in the recording device 200 on the display device 300 in response to a user operation or automatically without a user operation.
[0059] [Configuration of Display Device and Operation Unit] The display device 300 is equipped with a monitor 310 (display device) and can display input information, information recorded in the recording device 200, results of processing by the data processing unit 100, etc. The operation unit 400 (input device) includes a keyboard 410 and a mouse 420 as input devices and / or pointing devices, and a user can perform operations necessary to execute the data expansion method according to the present invention via these devices and the screen of the monitor 310. Note that the display device 300 may be configured as a touch panel display and the display screen may be used as the operation unit 400, or a microphone may be provided in the operation unit 400 to allow voice input.
[0060] [Outline of Data Augmentation Method] Figure 4 is a diagram for explaining an outline of a data augmentation method according to one aspect of the present invention. In the example of Figure 4, each sample (sample S1, S2) of the dataset has values for markers m1 to m6. Also, as shown in part (a) of Figure 4, markers m1, m2, and m3 have a positive correlation (see part (b) of the same figure), and markers m4 and m5 have a negative correlation. The other markers are not correlated with each other (shown by the approximately circular dotted line in part (a) of the same figure).
[0061] In such cases, a data augmentation method according to one aspect of the present invention organizes markers (multiple variables) into a local structure. Specifically, markers m1 to m3 are grouped into one group G1, and markers m4 and m5 are grouped into another group G2. Marker m6 has no correlation with other markers, so it can be determined to independently constitute group G3. Markers (variables) belonging to the same group are then swapped simultaneously or collectively. Specifically, as shown in part (c) of FIG. 4, a new sample S3 is generated using markers m1 to m3 of group G1 in sample S1 and markers m4 and m5 of groups G2 and G3 in sample S2. According to the present invention, data augmentation that takes into account the relationships between variables in this way maintains the relationships between variables, making it less likely that pseudo-data that deviates from the actual data will be generated.
[0062] [Data Augmentation Method Processing] [Overall Processing] The present invention relates to a data augmentation method that maintains the aforementioned local structure even for data in which the local structure is not explicitly defined. One aspect of the present invention includes four steps: (1) a relationship extraction step, (2) a variable selection step, and (3) a generation step. In the (1) relationship extraction step, relationships between multiple variables are extracted. The relationship between the variables can be, for example, a distance or similarity that can be calculated between any two variables. In the (2) variable selection step, one or more variable sets are selected based on the relationships between the variables. In the (3) generation step, new training data is generated by replacing the values of variables in a certain training data set with the values of corresponding variables in another training data set. In this case, the variables to be replaced are the variable set selected in (2).
[0063] The data augmentation method of the present invention may include a neighborhood extraction step. In the neighborhood extraction step, neighborhood relationships in the variable space can be extracted based on the relationships determined in the relationship extraction step (1) above. For example, in one form, when a graph is constructed based on the distance between variables, variables connected to each other by edges on the graph are considered to be neighbors. When neighborhood extraction is performed, in one form, the above-mentioned variable selection step selects a variable set consisting of multiple variables that are in a neighborhood relationship with each other from the neighbor variables corresponding to each variable extracted in the neighborhood extraction step, with the aim of preserving local structure.
[0064] According to the present invention configured in this way, it is possible to perform data expansion taking into account the relationships between variables, and when performing neighborhood extraction, it is possible to generate training data that maintains local structure.
[0065] [Steps of Data Augmentation Method] Each step of the data augmentation method will be described below. Note that the following describes an aspect in which the above-mentioned neighborhood extraction step is performed based on the relationships extracted in the relationship extraction step, and a variable set is selected based on the results.
[0066] [Relationship Extraction Step] In the relationship extraction step, the relationship extraction unit 112 (processor) assigns a value that can be calculated between any two variable series (any pair of two variables; a set having two or more observed values for a certain variable) included in the training dataset 202 (training dataset). In the relationship extraction step, the relationship extraction unit 112 may perform calculations using a second dataset 204 (second dataset) other than the given training dataset 202, or using a combined dataset 206 (combined dataset) that combines the training dataset 202 and the second dataset 204.
[0067] The second dataset 204 may be in any format. However, the relationship extraction unit 112 may use, for example, data having the same set of variables for measurement samples not included in the training dataset 202. Alternatively, the relationship extraction unit 112 may use a dataset in which different aspects of the same sample are measured as the second dataset 204. For example, if the training dataset 202 is data in which gene expression levels are measured, gene mutations can also be measured for the same sample. Therefore, the relationship extraction unit 112 can perform the relationship extraction process by using a gene mutation dataset as the second dataset 204 or by combining the second dataset 204 and the training dataset 202. In addition, a dataset in which the same variables and / or the same sample are measured but with at least one of the measurement means, measurement device, measurement environment, and measurement conditions changed can generally be used as the second dataset 204. The relationship extraction unit 112 may determine which dataset to perform relationship extraction on (an example of the data extension condition 212) in response to a user operation via the operation unit 400, or may automatically determine the data without a user operation.
[0068] The relationship extraction unit 112 preferably uses, as a value that can be calculated between two variable series (two variable pairs) (i.e., the relationship between multiple variables included in the entire dataset in the present invention), either a quantitative value (similarity) that increases the stronger the relationship between the variables, or a quantitative value (distance) that decreases the stronger the relationship between the variables. The relationship extraction unit 112 may calculate such distance or similarity for the training dataset 202, the second dataset 204, or the combined dataset 206.
[0069] The distance or similarity may be, for example, one or a combination of Euclidean distance, Manhattan distance, cosine similarity, correlation coefficient, mutual information, etc., but is not limited to these examples. The relationship extraction unit 112 may determine whether to extract distance or similarity as the "relationship between multiple variables" and what index to use as the distance or similarity (an example of the data extension condition 212) in response to a user operation via the operation unit 400, or may automatically determine the same without a user operation.
[0070] [Neighborhood Extraction Process] In the neighborhood extraction process, the neighborhood extraction unit 114 (processor) identifies a set of variables (a set of variables with strong relationships) that are in a neighborhood relationship with a certain variable based on the relationships between the variables extracted in the relationship extraction process. In particular, neighborhood relationships between variables can be extracted by clustering multiple variables. In clustering, the user must determine criteria such as the number of clusters, and the relationships extracted in the relationship extraction process, such as the distance or similarity between variables (examples of quantitative values indicating the strength of the relationship), are used. The algorithm used for neighborhood extraction is not particularly limited, and hierarchical clustering or non-hierarchical clustering may be used. The neighborhood extraction unit 114 may determine which algorithm to use (an example of the data expansion condition 212) in response to a user operation via the operation unit 400, or may automatically determine the algorithm without user operation.
[0071] In the case of clustering, a variable does not belong to multiple clusters, so if there are complex correlations, it is difficult to sufficiently extract neighborhood relationships. In such cases, as another embodiment of the neighborhood extraction process, the neighborhood extraction unit 114 preferably constructs a graph (first graph) based on the results of the relationship extraction process, with variables as points and relationships as edges. It is also preferable to construct a second graph in which the edges of the first graph are weighted by the magnitude of the extracted relationships. When using a graph, a variable may have neighborhood relationships with multiple other variables. When focusing on a certain variable in the constructed graph, the set of points that can be reached by tracing the edges corresponds to the extracted set of neighborhood variables (a set of variables that satisfy the conditions for neighborhood relationships). Whether or not to connect certain points (variables) with edges depends on a threshold set by the user and the value of the relationship between variables output by the relationship extraction process.
[0072] FIG. 5 is a diagram illustrating an example of the relationship extraction process and the neighborhood extraction process. This diagram illustrates a case where the relationship extraction unit 112 extracts relationships between three variables A, B, and C. As shown in part (a) of FIG. 5, the relationship (here, similarity; the higher the value, the stronger the relationship) between variables A and B is 0.7, the relationship between variables A and C is 0.4, and the relationship between variables B and C is 0.1. If the threshold (an example of a "second threshold" for similarity) is set to 0.5, as shown in part (b) of FIG. 5, the neighborhood extraction unit 114 connects only variables A and B with an edge, but does not connect variables A and C, or variables B and C. The neighborhood extraction unit 114 determines that variables A and B are a neighborhood variable set (variables that satisfy the neighborhood relationship condition of "similarity greater than the second threshold"), and determines that variable C is a neighborhood variable set consisting of only variable C. Furthermore, when the threshold value (another example of the "second threshold value") is set to 0.2, as shown in part (c) of Figure 5, the neighborhood extraction unit 114 connects variables A and B, and variables A and C with edges, but does not connect variables B and C.
[0073] In such a graph, not all connected points need to be included in the neighborhood variable set. For example, the neighborhood extraction unit 114 can designate a certain reference variable X and, when extracting its neighborhood relationship, determine that only points whose path length from X is less than a threshold are neighborhoods (variables that satisfy the conditions for neighborhood relationships) (the threshold is set by the user; it is different from the threshold for whether or not to connect edges). For example, in FIG. 6 (first graph, second graph), if the path length threshold is set to 3 (an example of a "first threshold" for distance), the neighborhood extraction unit 114 determines that "variables A, B, C, D, and E whose distance from variable X is less than the first threshold are neighborhoods of variable X." Note that in FIG. 6, the numbers enclosed in squares above the edges represent the path weights, and the numbers in parentheses near each variable, such as (1, 0.1), represent the sum of the path lengths and weights from variable X, respectively. Also, in FIG. 6, solid lines indicate variables in the neighborhood of variable X, and dashed lines indicate relationships that are "connected by edges in graph construction but not determined to be neighborhoods."
[0074] In the relation extraction and neighborhood extraction exemplified in FIGS. 5 and 6, variables for which either the distance or the similarity is less than a threshold value may be extracted, or variables for which both are less than a threshold value may be extracted.
[0075] Furthermore, when constructing a graph, the neighborhood extraction unit 114 may assign weights to each edge based on the results of the relationship extraction process. If a smaller weight is assigned as the relationship between variables becomes stronger, the neighborhood extraction unit 114 determines that only points where the sum of the weights of paths from variable A is less than a threshold are neighbors (variables that satisfy the conditions for a neighborhood relationship). For example, if the path weight threshold in FIG. 6 is set to 0.8, the neighborhood extraction unit 114 determines that "variables A, B, C, D, and E are neighbors of variable X." Furthermore, if a larger weight is assigned as the relationship between variables becomes stronger, the neighborhood extraction unit 114 can convert the weights into their inverses or into values obtained by subtracting each weight from the maximum weight, and evaluate whether or not the variables are neighbors based on the sum of these values.
[0076] The relationship extraction unit 112, the neighborhood extraction unit 114, and the input / output control unit 120 (processor 110; processor) can record the information indicating the relationship and the information indicating the neighborhood relationship, such as the graphs illustrated in FIGS. 5 and 6, in the recording device 200 as relationship information 208 and neighborhood relationship information 210, or display them on the display device 300. Furthermore, the relationship extraction unit 112, the neighborhood extraction unit 114, and the input / output control unit 120 may distinguish variables in a neighborhood relationship (variables that satisfy the conditions for the neighborhood relationship) by changing the type, thickness, color, etc. of the line, or by surrounding them with a line. The manner of distinguishing display may be changed depending on the strength of the relationship. Furthermore, the relationship extraction unit 112, the neighborhood extraction unit 114, and the input / output control unit 120 may determine what information and how to display it in response to a user operation via the operation unit 400, or may automatically determine the manner without relying on a user operation.
[0077] [Necessity and Effects of the Neighborhood Extraction Process] In the above section "Neighborhood Extraction Process," one aspect of neighborhood extraction was described. However, in the present invention, the neighborhood extraction process can be omitted by sorting candidates by weights (reflecting the results of relationship extraction) using the method described in the section "[Example of Execution of Variable Selection Algorithm]" below. However, as will be explained below, it is preferable to perform the neighborhood extraction process because it is effective in terms of calculation efficiency.
[0078] For example, if neighborhood extraction is not performed in advance and candidates are sorted using the weights described above, the order of relationships between a variable and all other variables must be determined each time during the variable selection process, which is computationally expensive. On the other hand, if neighborhood extraction is performed once, the number of neighbors of a variable will be far fewer than the total number of variables, making it possible to reduce the amount of calculation required for variable selection. Variable selection process 5 (see the "Variable Selection Process" section below) is typically repeated the same number of times as the amount of data to be generated. Since the amount of data to be generated is expected to be, for example, several hundred to several thousand, the amount of calculation required for each loop significantly affects the overall amount of calculation.
[0079] [Variable Selection Step] In the variable selection step, the variable selection unit 116 (processor 110; processor) selects one or more variables based on neighborhood relationships. Value replacement is performed on the selected variables (set) in the generation step, which will be described later. In the variable selection step, first, the variable selection unit 116 sets the number of variables to be selected. Next, the variable selection unit 116 selects variables according to the results of the neighborhood extraction step so as to satisfy the set number of variables. As one form of the variable selection step, the variable selection unit 116 can select variables according to an algorithm having the following steps 1 to 5. (Step 1) Set the ratio λ of the number of variables to be selected. (Step 2) Define the number of variables to be selected as n_target = round(λN). (Step 3) Define a list List_selected to store the selected variables. (Step 4) Define a variable n_selected to record the number of selected variables. (Step 5) Repeat the following steps (A) to (D) while n_target > n_selected is satisfied. (a) Select one variable x and add it to List_selected. (b) n_seletect += 1 (c) Obtain x's neighbor variables x_neighbors from the neighbor extraction process. (d) Repeat (i)-(iii) for variable i using breadth-first search from x_neighbors. (i) Add variable i to List_selected. (ii) n_selected += 1 (iii) If n_selected >= n_target, end the repetition.
[0080] [Example of execution of variable selection algorithm] Figure 7 shows an example of execution of the above algorithm. When x1 is selected in 5-(A) on the first round, A and B are stored in List_selected and n_selected = 2. When x2 is selected in 5-(A) on the second round, one of B / E / F is stored in List_selected and n_selected = 3, and the selection ends (F is selected in Figure 7). The order of B / E / F can be random, or in the case of a weighted graph, they can be rearranged in the order according to the weight of the edge connecting to x2. If x2 is selected first, B, E, and F are stored in List_selected, and the remaining one is selected from A and B, which are neighbors of x1. In addition, the threshold in the neighborhood extraction process is set based on the data and the reference variable x in the variable selection process. i It may be changed depending on the situation.
[0081] Figure 7 shows an example where N = 7, λ = 0.9, and n_target = 6. The double circle indicates the selected variable. The solid line encloses the neighborhood of x1, and the dashed line encloses the neighborhood of x2.
[0082] In the above algorithm, a set of neighborhood variables is selected from one or more variables that serve as the basis for neighborhood extraction. Therefore, not all of the selected variable sets are necessarily neighbors. In the example shown in Figure 7, A and B are neighbors with respect to x1, and F and B are neighbors with respect to x2, but the neighborhood relationship between A and F is not necessary. In addition, λ in the above algorithm can be a random number or a fixed value determined by the user. When using random numbers, it is preferable to use a beta distribution or uniform distribution, but other random number distributions can also be used.
[0083] [Generation Step] The generation step is a step of replacing the value of a corresponding variable with the value of other training data based on the variables (set) selected in the variable selection step. Hereinafter, the generation step will be described with reference to FIG. 8 .
[0084] For example, let us consider four training data {x1, x2, x3, x4} included in the training data as subsets, and the selected variable set is 60% of all variables (M variables), three subsets (20% each): [u, v, w] (the value (vector) of variable set u of the ith data is u i In this case, select one or more training data (replacement training data), x1 in this case, from the subset, and calculate the value (vector) [u 1 ,v 1 ,w 1 ] is the value (vector) of other teacher data included in the subset (the replacement teacher data that is the teacher data other than the replacement teacher data in the subset) [u k ,v k ,w k ],k ∈ Replace with {2,3,4}. Here, if [u4,v3,w2] is selected, then the new x new =[u4,v3,w 2, ...] is generated. Which value of the source training data to use is determined by a random number, and it is allowed to use the same training data multiple times (e.g., [u2,v2,w4]). If there are multiple replacement training data, the same operation as in the example above is performed on all replacement training data. For simplicity, other variables that are not selected are omitted, but x new For other variables in , the value of x1, which is the original value, is used.
[0085] [Correct Labels for Training Data] Furthermore, as described above as the twelfth aspect of the present invention, the data generation unit 118 (processor 110; processor) can generate correct labels for the new training data along with the new training data. In the above example, when the correct labels for the four training data are {G(x1)=A, G(x2)=A, G(x3)=B, G(x4)=C} (function G(·) outputs the correct labels for the input training data), new) is preferably a mixture of {C, B, A} (a mixture of the first label and the second label), i.e., the first label is preferably retained. The data generator 118 encodes the correct label (one-hot encoding) and assigns the correct label according to the weight of the original to be replaced. That is, the data generator 118 sets A=[0,0,1], B=[0,1,0], C=[1,0,0], and then calculates G(x new ) = [0.2, 0.2, 0.6], and a label is assigned to the new training data. As another example, if the correct label (second label) of the original data is A, and half of its variables are replaced with training data whose correct label (first label) is B, the label for the new training data will be [0, 0.5, 0.5]. It is also possible to assign weights during training without encoding the correct label.
[0086] According to the first embodiment, by repeating the steps described above, a desired number of new training data can be generated, and correct labels for the new training data can also be generated.
[0087] [Data Augmentation Using Data Belonging to Different Classes] In conventional technology (see, for example, the aforementioned Patent Document 2), when replacing variables, two pieces of data are selected only from data pairs classified into the same class. In contrast, in this embodiment, when selecting samples, it is possible to select training data classified into other classes, and in neighborhood extraction, it is possible to select a set of variables based on the neighborhood relationships in a local set of variables based on the extracted relationships (for example). This makes it possible to generate new training data (perform data augmentation) using the replacement training data and the replacement source training data that belongs to a class different from the class to which the replacement training data belongs.
[0088] Regarding such a technique, the following Non-Patent Document 2 states that "learning by mixing data belonging to different classes improves discrimination performance because the decision boundary performance of the two classes in the classifier becomes smoother." Non-Patent Document 2 also states that mixing not only inputs but also labels is effective.
[0089] [Non-Patent Document 2] "mixup: BEYOND EMPIRICAL RISK MINIMIZATION", published as a conference paper at ICLR 2018, Hongyi Zhang et al., [searched January 12, 2024], Internet (https: / / arxiv.org / abs / 1710.09412). Non-Patent Document 2 is a document about data augmentation using a method called mixup, but this method also has the same effect in CutMix-like data augmentation as the present invention.
[0090] [Data Augmentation for Non-Sequential Data] When performing CutMix on an image, the techniques described in Patent Document 1 and Non-Patent Document 1 do not guarantee that the local structure of data, which is trivially preserved by selecting adjacent pixels in an image, can be maintained for non-sequential data. Furthermore, the technique described in Patent Document 3 limits data augmentation to time-series data (the method described in Patent Document 3 cannot define regularity for non-sequential data). Thus, conventional techniques pose a problem when augmenting non-sequential data, such as table data, where the local structure is not explicitly defined. In other words, replacing variables with those of different samples can destroy the local structure, potentially resulting in features that differ from the actual data. In contrast, this embodiment enables data augmentation while maintaining the local structure, even for non-sequential data, such as table data.
[0091] [Data Augmentation for Measurement Data of Biological Samples] Among non-series data, correlations between variables are often significant for measurement data of biological samples. In this regard, as described above in the nineteenth and twentieth aspects of the present invention, the multiple variables in data augmentation may be variables measurable in a biological sample. Specifically, these "variables measurable in a biological sample" may be any one or a combination of methylation states, mutations, and gene expression levels at one or more sites in nucleic acids extracted from a living organism. Taking methylation states as an example, the methylation state at a certain site may correlate with another methylation state as a result of underlying biological cellular regulation, often resulting in a many-to-many high-order correlation. Furthermore, this correlation itself may be described as a characteristic of a cell and is an important factor in determining the properties of the cell. The present invention is particularly effective when using such training datasets.
[0092] [Example] This example shows the results of data expansion of omics data using the large-scale cancer database TCGA [K. Tomczak et al., 2015]. (1) Data: Eight cancer types (64 rectal cancer, 204 colon cancer, 18 bile duct cancer, 246 liver cancer, 314 lung adenocarcinoma, 245 lung squamous cell carcinoma, 6 ovarian cancer, and 118 pancreatic cancer samples) obtained from TCGA were used. Approximately 4.5 million CpG sites were measured for these samples using DNA methylation microarray analysis (Illumina, Infinium Human Methylation 450K BeadChip).
[0093] (2) Processing For simplicity, 20 CpG sites are extracted from the above data and named marker_0 to marker_19. These correspond to the variables in the present invention. Next, missing values are imputed by the average for each CpG site, and the average measurement values for each cancer type are calculated, resulting in the table shown in part (a) of Figure 9. The columns of the table shown in this part represent each cancer type, and each row represents the measurement value for each marker. The meanings of the abbreviations in each column are as shown in part (b) of Figure 9.
[0094] First, the relationship extraction process is performed by calculating the Pearson correlation coefficient between each marker from the data in the table above (see the third and fourth aspects). The results are shown in Figure 10. In Figure 10, the vertical and horizontal columns are labeled marker0 to marker20. The intersections represent the correlation coefficients between corresponding variables (the larger the value, the stronger the relationship). The color scale corresponds to the correlation coefficient, with the color closer to white (= +1.0) or black (= -1.0) indicating a larger absolute value of the correlation coefficient. In Figure 10, for example, if we focus on marker0 and marker1, or marker18 and marker19, we can see that there is a strong correlation and a local structure. It can also be seen that marker9 and marker14 have a negative correlation.
[0095] Figure 11 shows the results of applying the CutMix technology to omics data described in Non-Patent Document 1. Figures 12 and 13 show the results of applying the data augmentation method for preserving local structure according to the present invention. Figure 12 shows the results of identifying clusters with a correlation coefficient threshold of 0.9 or higher (clusters that satisfy the conditions for neighborhood relationships) from the results of the relationship extraction step using the data augmentation method according to the eighth aspect of the present invention, and extracting neighborhoods. Figure 13 shows the results of using the data augmentation method according to the tenth aspect of the present invention, constructing a graph with markers (variables) as points and weighted by correlation coefficients from the results of the relationship extraction step, and extracting points that can be traced by taking into account the weight of edges as neighborhoods. In both Figures 12 and 13, the variable selection and generation steps were performed from the extracted neighborhoods.
[0096] (3) Results For example, markers 0 and 1, markers 18 and 19, etc. show strong correlation coefficients (close to white in Figure 10, indicating a strong positive correlation), suggesting the formation of a local structure. Conventional CutMix generates uncorrelated data, whereas the present invention maintains this correlation. Conversely, markers 9, 10, and 14 show a negative correlation (close to black in Figure 10). However, conventional CutMix generates data with high measurement values in some cases, failing to maintain the local structure. However, the present invention maintains the negative correlation and enables data expansion.
[0097] The above results will be explained in more detail. In Figures 11 to 13, the rows represent variables, the columns represent samples, and the colors represent the degree of methylation (measured values). The variables of a given sample are arranged vertically. The tournament-like lines represent a tree diagram, showing the similarity between samples or variables. In other words, items placed close to each other have a high degree of similarity, and the lines connect nearby pairs hierarchically in descending order.
[0098] Figures 11 to 13 are compared, focusing on the above variables. In Figure 11 (conventional method), when marker0 and marker1 are viewed vertically (for each sample), samples with large discrepancies in methylation degree are observed. Specifically, as shown in parts (a) and (b) of Figure 14 (Figure 14 is an excerpt from Figure 11), the methylation degrees (corresponding to the color intensity) of marker0 and marker1 are significantly different (indicated by the arrows). In contrast, when marker0 and marker1 are viewed vertically in Figures 12 and 13 (present invention), although there is some discrepancy, there is little data generated with large discrepancies in methylation degree, suggesting that a correlation is maintained. Conversely, for marker9 and 14, data with a tendency for the methylation degree to always be inversely correlated is expected to be generated. From the same perspective, it can be qualitatively seen that Figures 12 and 13 generate the expected data compared to Figure 11.
[0099] However, it is not ideal to have no discrepancy at all. It is known that mixing a moderate amount of noise into the data can be effective in expanding the data.
[0100] [Examples of Data to Which the Present Invention Can Be Applied] The data augmentation device and data augmentation method of the present invention are particularly advantageous for data in which the local structure is not explicitly defined. "Data in which the local structure is not explicitly defined" refers to data expressed in a tabular format in which the order of variables is essentially meaningless. The variables may be independent or dependent of each other, and may be values or categorical values measured, observed, or recorded by different measuring devices or means. The data is a compilation of variables that are thought to be potentially related to the labels assigned to each sample. For example, the datasets described in "4.2 Datasets" in the following non-patent document 3 are also tabular data.
[0101] [Non-Patent Document 3] "Revisiting Deep Learning Models for Tabular Data", NeurIPS 2021 camera-ready, Yury Gorishniy et al. [Retrieved December 22, 2023], Internet (https: / / arxiv.org / abs / 2106.11959)
[0102] Below are examples of data other than omics data to which this invention can be applied. (1) Housing-related data: Variables include household income, age of the house, number of rooms, number of people in the household, and area, with the price of the house being the label. (2) Financial income data: Variables include a person's age, occupation, educational history, marital history, job position, race, and gender, with the person's income being the label. (3) Measurement data of physical phenomena: Variables include mechanical observations made by multiple detection devices, with the label indicating whether the observation captured the target phenomenon. (4) Vegetation data: Variables include the elevation, orientation, slope, soil type, and distance to a water source at a given point, with the label indicating the forest cover level at that point. (5) Web advertising data: Variables include the display coordinates, additional text, and URL of an image on the web, with the label indicating whether the image is an advertisement (the image itself is not used).
[0103] [Second Embodiment] [Learning Method and Cell Type Classification System] Next, a second embodiment of the present invention will be described. In the second embodiment, a learning method will be described in which a training dataset is expanded using the data expansion device and data expansion method of the present invention to generate an expanded training dataset, and a classifier that classifies and predicts the label (first label) of training data is trained using the expanded training dataset. Also, a classification device (cell type classification system) equipped with the classifier constructed using this learning method will be described.
[0104] In the second embodiment, the training data may be composed of sequential data or non-sequential data. In the following, however, a case will be described in which the training data set is composed of training data obtained by measuring the methylation state and / or mutations in nucleic acid sequences of one or more samples of cell-free free DNA, and each training data set is assigned a first label indicating the cell type (non-sequential data).
[0105] 15 is a diagram showing the configuration of a cell type classification system 1 according to the second embodiment. In the figure, the same reference numerals are used to designate components common to the first embodiment, and detailed descriptions thereof will be omitted.
[0106] The cell type classification system 1 comprises a data expansion device 10 and a cell type classification device 20 (cell type classification device), which are connected via a network NW. The data expansion device 10 and the cell type classification device 20 may be configured as separate devices as shown in FIG. 15 , or may be configured as a single device. The cell type classification device 20 comprises a data processing device 600, a recording device 700, a display device 800, and an operation unit 900. The operation unit 900 comprises a keyboard 910 and a mouse 920. The data expansion device 10 can have the same configuration as in the first embodiment.
[0107] Furthermore, an external server 500 and an external database 510 are connected via a network NW to the cell type classification system 1. The external server 500 and the external database 510 can have the same configuration as in the first embodiment.
[0108] 16 is a diagram showing the configuration of the data processing unit 600. The data processing unit 600 includes a processor 610 (processor), a ROM 640, and a RAM 500. The processor 610 may include the same processor 110 as the data processing unit 100 of the data extension device 10 or its functions.
[0109] The processor 610 includes a learning control unit 612, a discriminator 614, and an input / output control unit 616. The functions of each unit of the processor 610 can be realized using various processors, similar to the data processing unit 100 (processor 110) in the first embodiment.
[0110] [Learning Method of Cell Type Classification Device] In the cell type classification system 1 configured as described above, the input / output control unit 616 and the learning control unit 612 (processor 610) acquire a training dataset. This training dataset is composed of training data obtained by measuring the methylation status and / or nucleic acid sequence mutations of cell-free free DNA in more than one sample, and each training data is assigned a first label indicating the cell type. The input / output control unit 616 and the learning control unit 612 may acquire the training dataset from the recording device 700 or the recording device 200 (data expansion device 10), or from another recording device or recording medium such as the external database 510. Furthermore, as described above for the first aspect, training may be performed using a second dataset different from the training dataset (e.g., the second dataset 204 shown in FIG. 2 ) or a combined dataset combining the training dataset and the second dataset (e.g., the combined dataset 206 shown in FIG. 2 ).
[0111] The learning control unit 612 (processor 610) uses the data extension device 10 to perform data extension using the data extension method according to any one of the first to twentieth aspects of the present invention, thereby generating an extended learning data set including a learning data set and a second learning data set newly generated by the data extension and in which each training data set is assigned a first label. Note that if the processor 610 has the same functions as the processor 110, the extended learning data set may be generated by the processor 610 without using the data extension device 10.
[0112] The learning control unit 612 trains a classifier 614 that classifies and predicts the first label using this extended learning dataset (labeled training data). The data extension described in the first embodiment provides an extended learning dataset in which the generation of pseudo data that deviates from the actual data is suppressed. This allows for efficient learning, thereby enabling the construction of a classifier 614 that can perform accurate classification predictions. The classifier 614 may be a classifier that uses a machine learning algorithm such as deep learning. While the classifier 614 may be a linear classifier including logistic regression and its derivatives, or a nonlinear classifier using a support vector machine, a decision tree, or a neural network, the latter type of nonlinear classifier is expected to be more effective when mixing data belonging to different classes.
[0113] Note that a classifier that has been separately trained using the above-described method may be transplanted (e.g., layer structure, number of channels, filter size, weight parameter values, etc.) and used in the cell type classification device 20. In this case, the learning control unit 612 may be omitted.
[0114] [Cell Type Classification] In the cell type classification system 1, the classifier 614 constructed by the above-described learning method can classify and predict cancer type or healthy as the first label.
[0115] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described aspects and various modifications are possible.
[0116] 1 Cell type classification system 10 Data expansion device 20 Cell type classification device 100 Data processing unit 110 Processor 112 Relationship extraction unit 114 Neighborhood extraction unit 116 Variable selection unit 118 Data generation unit 120 Input / output control unit 200 Recording device 202 Learning dataset 204 Second dataset 206 Combination dataset 208 Relationship information 210 Neighborhood relationship information 300 Display device 310 Monitor 400 Operation unit 410 Keyboard 420 Mouse 500 External server 510 External database 600 Data processing unit 610 Processor 612 Learning control unit 614 Classifier 616 Input / output control unit 700 Recording device 800 Display device 900 Operation unit 910 Keyboard 920 Mouse
Claims
1. A data augmentation method executed by a data augmentation device including a processor, wherein the processor: obtains a learning dataset; extracts relationships between a plurality of variables included in the entire learning dataset; selects one or more sets of variables from the plurality of variables based on the relationships; determines a subset consisting of two or more teacher data in the learning dataset, and for replacement destination teacher data that is one or more of the teacher data in the subset, replaces the values of the selected one or more sets of variables with the values of the same variables in replacement source teacher data that is teacher data other than the replacement destination teacher data in the subset to generate new teacher data.
2. The data augmentation method according to claim 1, wherein the processor extracts the relationships using a second dataset different from the learning dataset or using a combined dataset obtained by combining the learning dataset and the second dataset, and based on the relationships, performs the selection and generation of the new teacher data using the second dataset or the combined dataset.
3. The data augmentation method according to claim 1, wherein the processor extracts, as the relationships, the distances or similarities of any two-variable pairs in the learning dataset.
4. The data augmentation method according to claim 2, wherein the processor extracts, as the relationships, the distances or similarities of any two-variable pairs in the second dataset or the combined dataset.
5. The data augmentation method according to claim 3 or 4, wherein the processor extracts, as variables satisfying the conditions for the relationships, variables for which the distance is less than a first threshold value and / or variables for which the similarity is greater than a second threshold value.
6. The data augmentation method according to claim 1, wherein the processor extracts a neighborhood relationship in a local structure formed by one or more variables included in the learning dataset based on the relationships, and selects the one or more sets of variables based on the neighborhood relationship.
7. The data augmentation method according to claim 2, wherein the processor extracts a neighborhood relationship in a local structure formed by one or more variables included in the second dataset or the combined dataset based on the relationships, and selects the one or more sets of variables based on the neighborhood relationship.
8. The data augmentation method according to claim 6 or 7, wherein the processor extracts the proximity relationship by clustering the plurality of variables based on the relationship.
9. The data augmentation method according to claim 6 or 7, wherein the processor constructs a first graph with variables belonging to the plurality of variables as points and the relationship as edges, and extracts, as the proximity relationship, a set of points reachable by traversing points connected by edges from each point of the first graph or a part of the set.
10. The data augmentation method according to claim 9, wherein the processor constructs a second graph in which the edges of the first graph are weighted by the magnitude of the relationship, and extracts, as the proximity relationship, a set of points reachable by traversing points connected by the edges in consideration of the weights from each point of the second graph or a part of the set.
11. The data augmentation method according to claim 6 or 7, wherein the processor selects one or more variables based on the proximity relationship.
12. The data augmentation method according to any one of claims 1 to 4, wherein the processor holds a first label of each piece of teacher data included in the learning data set, replaces the value of the selected variable of replacement destination teacher data, which is one or more pieces of teacher data within a subset composed of two or more pieces of teacher data, with the value of the same variable in replacement source teacher data, which is other teacher data within the subset, to generate the new teacher data, and for each replacement destination teacher data, mixes the first label of the replacement destination teacher data and the second label of the replacement source teacher data with weights corresponding to the number of variables replaced from each replacement source teacher data among all variables, and generates a label for the new teacher data together with the new teacher data.
13. The data augmentation method according to any one of claims 1 to 4, wherein the processor generates the new teacher data using the replacement destination teacher data and the replacement source teacher data belonging to a class different from the class to which the replacement destination teacher data belongs.
14. The data augmentation method according to any one of claims 1 to 4, wherein the learning data set is a data set composed of non-sequential data.
15. The data augmentation method according to any one of claims 2, 4, and 7, wherein the second dataset and the combined dataset are datasets composed of non-sequential data.
16. The data augmentation method according to any one of claims 1 to 4, wherein the processor causes a display device to display information indicating the relationship.
17. The data augmentation method according to claim 6 or 7, wherein the processor causes a display device to display information indicating the neighborhood relationship.
18. The data augmentation method according to claim 10, wherein the processor causes a display device to display the first graph and / or the second graph.
19. The data augmentation method according to any one of claims 1 to 4, wherein the plurality of variables are variables measurable in a biological sample.
20. The data augmentation method according to claim 19, wherein the variable measurable in the biological sample is any one or a combination of a methylation state, a mutation, and a gene expression level at one or more locations of a nucleic acid extracted from a living body.
21. A learning method for a cell type classification system including a processor, wherein the processor obtains a learning dataset composed of teacher data in which the methylation state and / or the mutation of the nucleic acid sequence of one or more cell-free free DNAs are measured, and each teacher data is provided with a first label indicating the cell type, expands the learning dataset by the data augmentation method according to any one of claims 1 to 4, and generates an extended learning dataset including the learning dataset and a second learning dataset newly generated by the data augmentation and provided with the first label for each teacher data, and learns a discriminator that classifies and predicts the first label using the extended learning dataset.
22. A cell type classification system including a discriminator constructed by the learning method according to claim 21.
23. The cell type classification system according to claim 22, wherein the discriminator classifies and predicts a cancer type or health as the first label.
24. A data augmentation device including a processor, wherein the processor: obtains a learning data set; extracts relationships among a plurality of variables included in the entirety of the learning data set; selects one or more sets of variables from the plurality of variables based on the relationships; determines a subset consisting of two or more teacher data in the learning data set; and for replacement target teacher data which is one or more of the teacher data in the subset, replaces values of the selected one or more sets of variables with values of the same variables in replacement source teacher data which is teacher data other than the replacement target teacher data in the subset, to generate new teacher data.
Citation Information
Patent Citations
Method, apparatus and computer program for training artificial intelligence model that judges danger in work site
KR1020220092361A
Medium composition for Streptomyces sp.
KR1020250105961A
Teacher data extending device, teacher data extending method, and program
WO2020070876A1