Data augmentation device and data augmentation method

The data augmentation device addresses the limitations of existing methods by diffusing and reversing data to generate synthetic samples that enhance deep learning model training, improving accuracy by incorporating boundary and outer samples.

WO2025109700A1PCT designated stage expired Publication Date: 2025-05-30NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/041883
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing data augmentation techniques, such as diffusion models, struggle to generate learning data that effectively supports the learning of deep learning models, as they tend to produce samples concentrated around the data distribution's center, lacking boundary and outer samples crucial for model learning.

Method used

A data augmentation device and method that involves diffusing input data to a predetermined stage in the noise-creation process and then reversing this diffusion to reconstruct synthetic samples, thereby retaining important information from the input samples while adding new information.

Benefits of technology

This approach enables the generation of synthetic samples that are useful for training deep learning models, improving learning accuracy by providing a richer dataset that includes boundary and outer samples typically missed by conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023041883_30052025_PF_FP_ABST
    Figure JP2023041883_30052025_PF_FP_ABST
Patent Text Reader

Abstract

In a data augmentation device (10), an acquisition unit (15a) acquires data to be processed. A diffusion unit (15b) diffuses the acquired data to a prescribed stage in the process of becoming noise. An inverse diffusion unit (15c) removes noise from the diffused data and inversely diffuses the resulting data.
Need to check novelty before this filing date? Find Prior Art

Description

Data extension device and data extension method

[0001] The present invention relates to a data expansion device and a data expansion method.

[0002] In recent years, advances in deep learning technology have led to the widespread application of AI in various industrial fields. However, deep learning models require large amounts of training data. Therefore, data augmentation techniques that transform training data samples to maximize the information obtained from a training dataset are expected. For example, a technique for augmenting training samples using a diffusion model, a type of generative model, has been proposed (see Non-Patent Document 1). A diffusion model is a generative model that learns the inverse diffusion process to restore data from noise.

[0003] Ho, Jonathan, Ajay Jain, Pieter Abbeel, “Denoising Diffusion Probabilistic Models”, 34th conference on neural information processing systems, 2020

[0004] However, conventional techniques can sometimes make it difficult to generate training data useful for training deep learning models. For example, a diffusion model generates new data points from complete noise, which allows it to generate unknown samples. However, the generated data are typical samples that are unevenly distributed around the center of the data distribution. As a result, the amount of information is small, and boundary samples and outer samples that are important for model training are not generated, making it difficult to say that the data is useful for model training.

[0005] The present invention has been made in consideration of the above, and aims to generate training data that is useful for training a deep learning model.

[0006] In order to solve the above-mentioned problems and achieve the object, the data expansion device of the present invention is characterized by having an acquisition unit that acquires data to be processed, a diffusion unit that diffuses the acquired data to a predetermined stage in the process of becoming noise, and a de-diffusion unit that removes the noise from the diffused data and de-diffuses it.

[0007] According to the present invention, it is possible to generate training data that is useful for training a deep learning model.

[0008] FIG. 1 is a diagram for explaining an overview of a data extension device of this embodiment. FIG. 2 is a diagram for explaining an overview of a data extension device of this embodiment. FIG. 3 is a diagram for explaining an overview of a data extension device of this embodiment. FIG. 4 is a schematic diagram illustrating a general configuration of a data extension device of this embodiment. FIG. 5 is a diagram for explaining data extension processing. FIG. 6 is a diagram for explaining data extension processing. FIG. 7 is a flowchart showing a data extension processing procedure. FIG. 8 is a diagram for explaining an example. FIG. 9 is a diagram for explaining an example. FIG. 10 is a diagram showing an example of a computer that executes a data extension program.

[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.

[0010] [Outline of the Data Expansion Device] Figures 1 to 3 are diagrams for explaining the outline of the data expansion device of this embodiment. In general, data expansion increases the variation of input training data samples by predefined transformations such as rotation. Using these training data to train a deep learning model makes the model more robust, and improvement in accuracy can be expected. However, since essential information of the input sample, such as the posture of a cat, does not change, there is a limit to the improvement in accuracy.

[0011] In contrast, a diffusion model is a type of generative model that learns the inverse diffusion process to restore data from noise. As shown in Figure 1, a conventional diffusion model generates new data points (synthetic samples) from complete noise such as Gaussian noise.

[0012] In this case, as shown in Fig. 2, the generated data is a typical sample that is unevenly distributed in the center of the data distribution, and boundary samples and outer samples that are important for model training are difficult to generate. Therefore, the amount of information is even less than that of real data, and it is difficult to say that it is useful for model training. In the example shown in Fig. 2, "Real" represents a real sample, and "Synthetic" represents a synthetic sample.

[0013] Therefore, the data extension device of this embodiment diffuses (destroys) the input real samples halfway through the diffusion process, and then reconstructs the samples in the de-diffusion process to output synthetic samples, as shown in Fig. 3. In the example shown in Fig. 3, the real samples are diffused up to 40% of the stage before becoming complete noise, and then synthetic samples are reconstructed in which 40% of the information of the real samples remains.

[0014] In this way, the input real samples are used to preserve important information from the real samples, while the diffusion model adds new information. Therefore, the output synthetic samples have the properties of both the real samples and the synthetic samples, making them useful training data for deep learning models.

[0015] [Configuration of the Data Expansion Device] Fig. 4 is a schematic diagram illustrating the overall configuration of the data expansion device of this embodiment. As illustrated in Fig. 4, the data expansion device 10 of this embodiment is realized by a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.

[0016] The input unit 11 is realized using input devices such as a keyboard and a mouse, and inputs various instruction information such as a command to start processing to the control unit 15 in response to input operations by an operator. The output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, etc. For example, the output unit 12 displays the results of the data expansion process described below.

[0017] The communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication between the control unit 15 and external devices via telecommunication lines such as a LAN (Local Area Network) or the Internet. For example, the communication control unit 13 controls communication between the control unit 15 and a management device or the like that manages various types of information.

[0018] The storage unit 14 is realized by a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 14 stores in advance the processing program that operates the data extension device 10, data used during execution of the processing program, and the like, or temporarily stores the data each time processing is performed. In particular, in this embodiment, the storage unit 14 stores, for example, parameters of a diffusion model 14a for the data extension process and parameters of a model 14b to be trained. The storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13.

[0019] The control unit 15 is implemented using a CPU (Central Processing Unit), an NP (Network Processor), an FPGA (Field Programmable Gate Array), or the like, and executes a processing program stored in memory. As a result, the control unit 15 functions as an acquisition unit 15a, a diffusion unit 15b, a de-diffusion unit 15c, and a learning unit 15d, as illustrated in FIG. 4, to perform data augmentation processing. Note that each or some of these functional units may be implemented in different hardware. For example, the learning unit 15d may be implemented as a device separate from the other functional units. The control unit 15 may also include other functional units.

[0020] The acquiring unit 15a acquires a training dataset D to be processed from a management device that manages various types of data, etc., via the input unit 11 or the communication control unit 13. The training dataset D is composed of, for example, a set of an input sample x and a class label y. The acquiring unit 15a may store the acquired training dataset D in the storage unit 14.

[0021] The diffusion unit 15b diffuses the acquired data to a predetermined stage in the process of becoming noise. For example, the diffusion unit 15b diffuses the acquired data to a predetermined percentage of the process of becoming complete noise, or to a predetermined number of steps t re The noise generation process is continued until it reaches

[0022] The despreading unit 15c removes noise from the spread data and despreads the data. For example, the despreading unit 15c removes noise from the spread data to reconstruct the acquired data.

[0023] Specifically, the diffusion unit 15b and the despreading unit 15c use a diffusion model 14a that, upon inputting acquired data, diffuses the data to a predetermined stage and then despreads the data to output reconstructed data. That is, the diffusion unit 15b diffuses the acquired data to a predetermined stage using the diffusion model 14a that, upon inputting acquired data, diffuses the data to a predetermined stage and then outputs despread data. Furthermore, the despreading unit 15c despreads the diffused data using the diffusion model 14a. That is, the diffusion unit 15b and the despreading unit 15c generate reconstructed data x^ from the acquired data x using an arbitrary trained diffusion model 14a.

[0024] 5 and 6 are diagrams for explaining the data extension process. Fig. 5 illustrates an example of the data extension process. Fig. 6 illustrates a pseudo code for the data extension process.

[0025] Specifically, the diffusion unit 15b performs a predetermined step t re In other words, the diffusion unit 15b diffuses (destroys) the input sample x in t steps according to the following equation (1): t and spread it as

[0026]

[0027] Then, the despreading unit 15c performs step t reThe despreading unit 15c generates the composite sample x̂ by applying the despreading process to reconstruct the real samples from step 1 back to step 1 (lines 6 to 10 in FIG. 6). That is, the despreading unit 15c generates the composite sample x̂ according to the following equation (2):

[0028]

[0029] The learning unit 15d uses the despread data as training data to train the model 14b that outputs a predetermined feature from the data. For example, the learning unit 15d uses a synthetic sample x̂ reconstructed from a data sample x of the training data D to train the model f θ The supervised loss Lsup(f θ Then, the learning unit 15d calculates the model f θ The process of updating the parameter θ of the model f is repeated until a predetermined maximum number of learning steps is reached. θ Learn about the following.

[0030] This generates a synthetic sample x̂ that retains important information from the input sample x while adding new information. By using this as training data for the model 14b, it becomes possible to train the model 14b with high accuracy.

[0031] The learning unit 15d may use the acquired data, that is, the real sample x, together with the synthetic sample x^ as learning data.

[0032] [Data Extension Processing] Next, the data extension processing by the data extension device 10 according to this embodiment will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the procedure of the data extension processing. The flowchart in Fig. 7 starts, for example, when the user performs an operation input to instruct the start of the processing.

[0033] First, the acquisition unit 15a acquires a training dataset D to be processed (step S1). The acquisition unit 15a also acquires (x, y) from the dataset D, where x is a data sample and y is a class label.

[0034] Next, the spreading unit 15b spreads (destroys) the input sample x in the spreading process according to the above formula (1) (step S2), and the despreading unit 15c reconstructs x^ in the despreading process according to the above formula (2) (step S3).

[0035] Then, the learning unit 15d uses the synthetic sample x̂ reconstructed from the data sample x of the learning data D to generate a model f θ The supervised loss Lsup(f θ (x^), y) is calculated (step S4).

[0036] The learning unit 15d also learns the model f by backpropagation. θ The parameter θ is updated (step S5). If the number of learning steps has not reached the predetermined maximum number of learning steps (step S6, True), the process returns to step S1. On the other hand, if the number of learning steps has reached the predetermined maximum number of learning steps (step S6, False), the series of data expansion processes is terminated.

[0037] [Effects] As described above, in the data expansion device 10 of this embodiment, the acquisition unit 15a acquires data to be processed. The diffusion unit 15b diffuses the acquired data to a predetermined stage in the process of becoming noise. The dediffusion unit 15c removes noise from the diffused data and dediffuses it.

[0038] Specifically, the diffusion unit 15b receives the acquired data, diffuses the data to a predetermined stage using the diffusion model 14a, which diffuses the data to a predetermined stage and then outputs de-diffused data. The de-diffusion unit 15c de-diffuses the diffused data using the diffusion model 14a.

[0039] This generates synthetic samples that retain important information from the input samples while adding new information. Therefore, by using the synthetic samples generated by the data augmentation process of this embodiment as training data for the deep learning model 14b, it is possible to train the model 14b with high accuracy.

[0040] The learning unit 15d uses the de-diffused data as learning data to learn the model 14b that outputs a predetermined feature from the data. This makes it possible to learn the model 14b with high accuracy using synthetic samples that are useful for learning.

[0041] 8 and 9 are diagrams for explaining an example. In this example, CIRAR-10, which classifies general images into 10 classes, was used as the image dataset. The ratio of the training dataset to the validation dataset was 9:1. As neural networks, ResNet-18 was used as the classifier (training model 14b), and EDM was used as the diffusion model 14a.

[0042] The generated data was used as training data to train Model 14b for 100 epochs, and the accuracy with the highest Top-1 accuracy in the validation data set was adopted. Each experimental pattern shown below was performed five times, and the average value and standard deviation were reported.

[0043] FIG. 8 illustrates the results of this example. The experimental patterns illustrated in FIG. 8 are as follows: "Base Model" is data that has undergone data extension using normal transformations such as rotation, left-right flip, and random cropping. "GDA" is a synthetic sample generated from conventional complete noise using a diffusion model. "RDA" is a data extension process (t re = 50).

[0044] As shown in FIG. 8, it has been found that the data augmentation process (RDA) of the above embodiment can improve the learning accuracy of the model 14b.

[0045] 9 illustrates the distribution of input real sample (Real), GDA, and RDA data. As shown in FIG. 9, the data augmentation process (RDA) according to the above embodiment generates synthetic samples in areas not generated by conventional GDA, and further generates synthetic samples in boundary and outer areas that differ from the real samples, thereby confirming that the real samples can be complemented.

[0046] [Program] A program written in a computer-executable language may be created to execute the processes executed by the data expansion device 10 according to the above embodiment. In one embodiment, the data expansion device 10 can be implemented by installing a data expansion program that executes the data expansion process as package software or online software on a desired computer. For example, by executing the data expansion program on an information processing device, the information processing device can function as the data expansion device 10. The information processing device referred to here includes desktop and notebook personal computers. Other examples of information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System) phones, as well as slate terminals such as PDAs (Personal Digital Assistants). The functions of the data expansion device 10 may also be implemented on a cloud server.

[0047] 10 is a diagram showing an example of a computer that executes a data expansion program. The computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0048] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1031. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to a mouse 1051 and a keyboard 1052, for example. The video adapter 1060 is connected to a display 1061, for example.

[0049] Here, the hard disk drive 1031 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. The various pieces of information described in the above embodiments are stored in the hard disk drive 1031 or the memory 1010, for example.

[0050] The data extension program is stored in the hard disk drive 1031 as a program module 1093 in which instructions to be executed by the computer 1000 are written. Specifically, the program module 1093 in which each process executed by the data extension device 10 described in the above embodiment is written is stored in the hard disk drive 1031.

[0051] Furthermore, data used for information processing by the data extension program is stored as program data 1094, for example, in the hard disk drive 1031. Then, the CPU 1020 reads the program module 1093 and program data 1094 stored in the hard disk drive 1031 into the RAM 1012 as necessary, and executes each of the above-described procedures.

[0052] The program module 1093 and program data 1094 related to the data expansion program are not limited to being stored in the hard disk drive 1031, but may be stored in a removable storage medium, for example, and read by the CPU 1020 via the disk drive 1041. Alternatively, the program module 1093 and program data 1094 related to the data expansion program may be stored in another computer connected via a network such as a LAN or a WAN (Wide Area Network), and read by the CPU 1020 via the network interface 1070.

[0053] Although the present invention has been described above as an embodiment, the present invention is not limited to the description and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention.

[0054] REFERENCE SIGNS LIST 10 Data expansion device 11 Input unit 12 Output unit 13 Communication control unit 14 Storage unit 14a Diffusion model 14b Model 15 Control unit 15a Acquisition unit 15b Diffusion unit 15c Dediffusion unit 15d Learning unit

Claims

1. An acquisition unit that acquires data to be processed, a diffusion unit that diffuses the acquired data to a predetermined stage in the process of becoming noise, and an inverse diffusion unit that removes the noise from the diffused data and performs inverse diffusion. A data augmentation device, characterized by comprising:

2. The diffusion unit, when the acquired data is input, uses a diffusion model that outputs the data after diffusion to the predetermined stage and then inverse diffusion to diffuse the acquired data to the predetermined stage. The inverse diffusion unit uses the diffusion model to perform inverse diffusion on the diffused data. The data augmentation device according to claim 1, characterized in that:

3. The data augmentation device according to claim 1, further comprising a learning unit that uses the inversely diffused data as learning data and performs learning of a model that outputs a predetermined feature amount from the data.

4. A data augmentation method executed by a data augmentation device, comprising: an acquisition step of acquiring data to be processed; a diffusion step of diffusing the acquired data to a predetermined stage in the process of becoming noise; and an inverse diffusion step of removing the noise from the diffused data and performing inverse diffusion. A data augmentation method, characterized by including: