Information processing program, information processing method, and information processing device

By generating additional data with noise and adjusting its ratio with original data, the method enhances white-box attack resistance in machine learning models without significantly affecting accuracy.

JP7679630B2Active Publication Date: 2025-05-20FUJITSU LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021012143
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-28
Publication Date
2025-05-20
Estimated Expiration
2041-01-28

AI Technical Summary

Technical Problem

Conventional defense methods against white-box attacks in machine learning models suffer from a trade-off between attack resistance and accuracy, and defenses against black-box attacks are insufficient, leaving models vulnerable to training data inference.

Method used

A method involving a first machine learning model trained with noise or constant values to generate additional data, followed by upsampling or downsampling to create second training data, which is used to train a machine learning model resistant to white-box attacks.

Benefits of technology

Generates a machine learning model that effectively resists white-box attacks on training data estimation while maintaining model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679630000001
    Figure 0007679630000001
  • Figure 0007679630000002
    Figure 0007679630000002
  • Figure 0007679630000003
    Figure 0007679630000003
Patent Text Reader

Abstract

To generate a machine learning model that is resistant to white-box attacks on training data estimation.SOLUTION: A first machine learning model A trained with first training data is used to mechanically generate meaningless data as an initial value to create additional data. The first training data and the additional data are combined to create second training data. The second training data is used to train the machine learning model.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing program, an information processing method, and an information processing device. [Background technology]

[0002] In recent years, the development and use of systems using machine learning has progressed rapidly. On the other hand, various security issues specific to systems using machine learning have been found. For example, there is a training data inference attack, which infers and steals the training data used in machine learning.

[0003] As a training data inference attack, for example, a face recognition edge device used in a face recognition system is analyzed to extract a machine learning model, and a training data inference attack is performed on this machine learning model to infer the face images used as training data.

[0004] Training data inference attacks are attacks that target trained models (machine learning models) that have completed the training phase, and are classified into black-box attacks and white-box attacks. A black-box attack infers training data from input data and output data in the inference phase.

[0005] As a defense against black-box attacks, for example, a method is known that simply reduces the information of the output of a trained model by adding noise or deleting confidence levels. In addition, a method is known to counter attacks by faking gradients and directing the attacker to a decoy dataset prepared in advance.

[0006] White-box attacks infer training data from a trained machine learning model itself. A known defense against white-box attacks is to generate a trained machine learning model that is resistant to training data estimation by adding appropriate noise to the parameters when updating the parameters of the machine learning model. One example of a defense against such white-box attacks is DP-SGD (Differential Private - Stochastic Gradient Descent). [Prior art documents] [Patent documents]

[0007] [Patent Document 1] JP 2020-115312 A [Patent Document 2] JP 2020-119044 A Summary of the Invention [Problem to be solved by the invention]

[0008] In general, there is always the risk that an attacker may obtain the machine learning model itself, and defenses against black-box attacks alone are insufficient.

[0009] On the other hand, in conventional defense methods against white-box attacks, the estimation accuracy decreases by adding noise to the parameters of the machine learning model, so there is a trade-off between the strength of the attack resistance and the accuracy. Therefore, there is an issue that it cannot be introduced in systems that require the accuracy of the machine learning model. In one aspect, the present invention aims to enable the generation of machine learning models that are resistant to white-box attacks on training data estimation. [Means for solving the problem]

[0010] For this purpose, the information processing program uses a first machine learning model trained by the first training data to mechanically generate additional data by using noise or a constant value as an initial value; Based on a result of comparing the number of the additional data with the number of the first training data to confirm the ratio of the additional data to the first training data, upsampling or downsampling is performed on at least one of the first training data and the additional data so that the ratio of the additional data to the first training data becomes a predetermined value, and the upsampling or downsampling is performed. The computer is caused to perform a process of combining the first training data and the additional data to create second training data, and training a machine learning model using the second training data. Effect of the Invention

[0011] According to one embodiment, a machine learning model can be generated that is resistant to white-box attacks of training data estimation. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating a hardware configuration of an information processing apparatus according to an embodiment; [Diagram 2] FIG. 1 is a diagram illustrating a functional configuration of an information processing apparatus as an example of an embodiment. [Diagram 3] FIG. 2 is a diagram for explaining processing of a mini-batch creation unit in an information processing device as an example of an embodiment. [Figure 4] FIG. 1 is a diagram illustrating an overview of a training method for a machine learning model in an information processing device as an example of an embodiment. [Diagram 5] 1 is a flowchart illustrating a method for training a machine learning model in an information processing device as an example of an embodiment. [Figure 6] 11 is a diagram for explaining the results of a white-box attack on training data estimation for a machine learning model generated by an information processing device as an example of an embodiment. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, the embodiments of the present information processing program, information processing method, and information processing device will be described with reference to the drawings. However, the embodiments shown below are merely examples, and there is no intention to exclude the application of various modified examples and techniques not explicitly stated in the embodiments. In other words, this embodiment can be implemented with various modifications within the scope of its purpose. Furthermore, each figure is not intended to include only the components shown in the figure, but may include other functions, etc.

[0014] (A) Configuration FIG. 1 is a diagram illustrating a hardware configuration of an information processing device 1 as an example of an embodiment. 1, the information processing device 1 includes, as components, a processor 11, a memory 12, a storage device 13, a graphics processing device 14, an input interface 15, an optical drive device 16, a device connection interface 17, and a network interface 18. These components 11 to 18 are configured to be able to communicate with each other via a bus 19.

[0015] The processor (control unit) 11 controls the entire information processing device 1. The processor 11 may be a multiprocessor. The processor 11 may be, for example, any one of a CPU, an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array). The processor 11 may also be a combination of two or more types of elements among the CPU, MPU, DSP, ASIC, PLD, and FPGA.

[0016] Then, when the processor 11 executes a control program (information processing program: not shown), the functions of a training processing unit 100 (first training execution unit 101, additional training data creation unit 102, and second training execution unit 105) illustrated in FIG. 2 are realized.

[0017] The information processing device 1 realizes the function of the training processing unit 100 by executing a program {an information processing program or an OS (Operating System) program} recorded on, for example, a computer-readable non-transitory recording medium.

[0018] The programs describing the processing contents to be executed by the information processing device 1 can be recorded on various recording media. For example, the programs to be executed by the information processing device 1 can be stored in the storage device 13. The processor 11 loads at least a part of the programs in the storage device 13 into the memory 12 and executes the loaded programs.

[0019] The programs to be executed by the information processing device 1 (processor 11) may also be recorded in a non-transitory portable recording medium such as the optical disk 16a, the memory device 17a, or the memory card 17c. The programs stored in the portable recording medium are executable after being installed in the storage device 13, for example, under the control of the processor 11. The processor 11 may also read and execute the programs directly from the portable recording medium.

[0020] The memory 12 is a storage memory including a ROM (Read Only Memory) and a RAM (Random Access Memory). The RAM of the memory 12 is used as a main storage device of the information processing device 1. The RAM temporarily stores at least a part of the OS program and the control program to be executed by the processor 11. The memory 12 also stores various data necessary for processing by the processor 11.

[0021] The storage device 13 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a storage class memory (SCM), and stores various data. The storage device 13 is used as an auxiliary storage device for the information processing device 1. The storage device 13 stores an OS program, a control program, and various data. The control program includes an information processing program.

[0022] As the auxiliary storage device, a semiconductor storage device such as an SCM or a flash memory can also be used. Furthermore, a plurality of storage devices 13 may be used to configure a RAID (Redundant Arrays of Inexpensive Disks).

[0023] In addition, the memory device 13 may store various data generated when the first training execution unit 101, the additional training data creation unit 102 (additional data creation unit 103, mini-batch creation unit 104), and the second training execution unit 105 described below execute each process.

[0024] A monitor 14a is connected to the graphics processing device 14. The graphics processing device 14 displays an image on the screen of the monitor 14a in accordance with an instruction from the processor 11. Examples of the monitor 14a include a display device using a CRT (Cathode Ray Tube) and a liquid crystal display device.

[0025] A keyboard 15a and a mouse 15b are connected to the input interface 15. The input interface 15 transmits signals sent from the keyboard 15a and the mouse 15b to the processor 11. The mouse 15b is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include a touch panel, a tablet, a touch pad, and a trackball.

[0026] The optical drive device 16 uses a laser beam or the like to read data recorded on an optical disk 16a. The optical disk 16a is a portable, non-transient recording medium on which data is recorded so that it can be read by the reflection of light. Examples of the optical disk 16a include a DVD (Digital Versatile Disk), a DVD-RAM, a CD-ROM (Compact Disc Read Only Memory), and a CD-R (Recordable) / RW (ReWritable).

[0027] The device connection interface 17 is a communication interface for connecting peripheral devices to the information processing device 1. For example, a memory device 17a or a memory reader / writer 17b can be connected to the device connection interface 17. The memory device 17a is a non-transitory recording medium equipped with a communication function with the device connection interface 17, such as a Universal Serial Bus (USB) memory. The memory reader / writer 17b writes data to a memory card 17c or reads data from the memory card 17c. The memory card 17c is a card-type non-transitory recording medium.

[0028] The network interface 18 is connected to a network (not shown). The network interface 18 may be connected to other information processing devices, communication devices, etc. via the network. For example, an input image or an input text may be input via the network.

[0029] 2 is a diagram illustrating a functional configuration of the information processing device 1 as an example of an embodiment. The information processing device 1 has a function as a training processing unit 100 as shown in FIG. In the information processing device 1, the processor 11 executes a control program (information processing program), thereby realizing the function of a training processing unit 100.

[0030] The training processing unit 100 realizes learning processing (training processing) in machine learning by using training data. That is, the information processing device 1 functions as a training device that trains a machine learning model by the training processing unit 100.

[0031] The training processing unit 100 realizes a learning process (training process) in machine learning, for example, by using training data (teacher data) to which a correct answer label is assigned. The training processing unit 100 trains a machine learning model by using the training data, and generates a trained machine learning model that is resistant to training data estimation.

[0032] The machine learning model may be, for example, a deep learning model (deep neural network). The neural network may be a hardware circuit, or may be a virtual network by software that connects layers virtually constructed on a computer program by the processor 11 or the like. As shown in FIG. 2, the training processing unit 100 includes a first training execution unit 101, an additional data creation unit 103, and a second training execution unit 105.

[0033] The first training execution unit 101 trains the machine learning model using the training data and generates a trained machine learning model. The training data is configured, for example, as a combination of input data x and correct output data y.

[0034] The training of the machine learning model performed by the first training execution unit 101 using training data may be referred to as the first training. Also, the machine learning model before training by the first training execution unit 101 may be referred to as the first machine learning model. Since the first machine learning model is a machine learning model before training, it may be called an empty machine learning model. Also, the machine learning model may simply be called a model. In addition, hereinafter, the training data used by the first training execution unit 101 for the first training may be referred to as first training data or training data A.

[0035] Furthermore, the trained machine learning model generated by the first training execution unit 101 may be referred to as a second machine learning model or a machine learning model A. Through the first training by the first training execution unit 101, model parameters of the machine learning model A are set.

[0036] The first training execution unit 101 can use a known method to train the first machine learning model using training data A to generate a second machine learning model (machine learning model A), and detailed explanations of this are omitted.

[0037] The additional training data creation unit 102 creates training data that is used when the second training execution unit 105, which will be described later, performs additional training on the second machine learning model (machine learning model A) generated by the first training execution unit 101. Hereinafter, the training data used when performing additional training on the second machine learning model may be referred to as second training data or training data B. Training data B may also be referred to as additional training data. The additional training data creation unit 102 includes an additional data creation unit 103 and a mini-batch creation unit 104 .

[0038] The additional data creation unit 103 creates a plurality of additional data. The additional data is data that is not input to the machine learning model A in a general machine learning model operation, and is artificial data that is classified into a specific label in a classifier. The additional data creation unit 103 creates the additional data by, for example, a gradient descent method in which the gradient of the machine learning model A is obtained and the input is updated in a direction in which the confidence level increases.

[0039] Below, we explain a method (steps 1 to 4) for generating additional data using a simple gradient descent method. (Step 1) First, the additional data generating unit 103 sets an objective function. Input for machine learning model A: X Output of machine learning model A: f(X) = (f 1 (X), …, f n (X) Target label: t In this case, the objective function can be expressed, for example, by the following equation (1). L(X) = (1-f t (X) 2 (1) When the value of L(X) above is minimized, X is classified as label t with confidence level 1. Because of this dependency on label t, step 1 must be performed for all labels.

[0040] (Step 2) As the initial value, prepare the input of meaningless data (e.g., noise or a constant value) for the machine learning model A (hereinafter, the initial value X 0 (represented as ). Initial value X 0 The additional data may be prepared or set in advance by an operator or the like, or may be generated by the additional data creating unit 103. (Step 3) The additional data creation unit 103 0 The differential value of L(X) around 0 ) to get the (Step 4) The additional data creation unit 103 0 -λL′(X 0 ) is the additional data. λ is a hyperparameter.

[0041] The method of creating the additional data is not limited to the above, and can be modified as appropriate. For example, other objective functions may be used. Furthermore, the above step 4 may be repeated a certain number of times. Furthermore, the formula in the above step 4 may be modified.

[0042] The additional data creation unit 103 generates meaningless data (X 0 ) is mechanically generated as the initial value to create additional data. In addition, the additional data may be generated using an optimization method other than the gradient descent method, such as an evolutionary algorithm, and may be implemented in various modified forms.

[0043] When the input data is image data, for example, a fooling image may be used as the additional data. Note that the fooling image can be generated by a known method, and the description thereof will be omitted.

[0044] The mini-batch creation unit 104 adds the additional data created by the additional data creation unit 103 to the training data A, thereby creating second training data (training data B, additional training data).

[0045] The mini-batch creation unit 104 upsamples the training data A or downsamples the additional data so that the number of samples of the additional data is sufficiently smaller than the number of samples of the training data A. For example, the mini-batch creation unit 104 adjusts the amount of training data A and additional data so that the ratio of the additional data to the training data A becomes a predetermined value (α).

[0046] That is, when the additional data is less than a predetermined ratio α with respect to the training data A, the mini-batch creation unit 104 performs at least one of downsampling the training data A and upsampling the additional data so that the ratio of the additional data with respect to the training data A becomes α. On the other hand, when the additional data is equal to or greater than the predetermined ratio α with respect to the training data A, the mini-batch creation unit 104 performs at least one of upsampling the training data A and downsampling the additional data so that the ratio of the additional data with respect to the training data A becomes α. Note that a method such as adding noise may be used for upsampling.

[0047] In addition, by increasing the ratio of the additional data to the training data A, it is possible to improve the resistance to white-box attacks of the machine learning model (machine learning model B) generated by the second training execution unit 105 described later using the second training data (training data B). On the other hand, by increasing the ratio of the additional data to the training data A, there is a risk that the accuracy of the machine learning model (machine learning model B) will decrease. Therefore, it is desirable to set the threshold value (α) representing the ratio of the additional data to the training data A to as large a value as possible within a range in which the accuracy of the machine learning model (machine learning model B) is maintained. The mini-batch generation unit 104 generates a plurality of mini-batches using the training data A and the additional data.

[0048] FIG. 3 is a diagram for explaining the processing of the mini-batch creation unit 104 in the information processing device 1 as an example of an embodiment. The mini-batch creation unit 104 shuffles the data so that a certain proportion of additional data is included in each mini-batch in order to stabilize training (machine learning) by the second training execution unit 105 described later.

[0049] That is, the mini-batch creation unit 104 randomly rearranges (shuffles) the training data A and the additional data separately, and divides each into N equal parts (N is a natural number equal to or greater than 2). Hereinafter, 1 / N of the training data A generated by dividing the training data into N equal parts may be referred to as divided training data A. Also, 1 / N of the additional data generated by dividing the additional data into N equal parts may be referred to as divided additional data.

[0050] The mini-batch creation unit 104 creates one mini-batch by combining one piece of divided training data A extracted from training data A divided into N pieces (N division) and divided additional data extracted from additional data divided into N. The mini-batch is used for training the machine learning model by the second training execution unit 105 described later.

[0051] That is, the mini-batch creation unit 104 extracts a certain number of pieces from each of the shuffled data (training data A and additional data) and combines them to create one mini-batch. A collection of these multiple mini-batches may be referred to as training data B.

[0052] The mini-batch creation unit 104 corresponds to a second training data creation unit that combines training data A (first training data) with additional data to create training data B (second training data). In addition, the mini-batch creation unit 104 performs upsampling or downsampling on at least one of the training data A and the additional data in the training data B so that the ratio of the additional data to the training data A (first training data) becomes a predetermined value (α).

[0053] The size of the mini-batch can be set appropriately based on the know-how of machine learning. The mini-batch creation unit 104 shuffles the training data A and the additional data, respectively, to prevent bias in the gradient of the parameters set by training.

[0054] The second training execution unit 105 trains the machine learning model using the training data B created by the additional training data creation unit 102, thereby creating a machine learning model that is resistant to training data estimation attacks.

[0055] In the present information processing device 1, the second training execution unit 105 performs training (additional training) using training data B on the machine learning model A trained by the first training execution unit 101. Hereinafter, the training of the machine learning model that the second training execution unit 105 performs using the training data B may be referred to as second training.

[0056] Furthermore, the trained machine learning model generated by the second training execution unit 105 may be referred to as the machine learning model B. The machine learning model B may also be referred to as a third machine learning model.

[0057] The second training execution unit 105 and the first training execution unit 101 can use a known method to train the second machine learning model using training data B to generate a third machine learning model (machine learning model B), the specific explanation of which is omitted.

[0058] The second training execution unit 105 performs further training (additional training) on ​​the trained machine learning model A using mini-batches generated by dividing the training data B by N, which are created by the additional training data creation unit 102, to generate an additionally trained machine learning model B. Through the second training (additional training) performed by the second training execution unit 105, model parameters of the machine learning model B are set.

[0059] The second training execution unit 105 trains the machine learning model using training data B (second training data), and performs retraining of the machine learning model A (first machine learning model) using the training data B (second training data).

[0060] The machine learning model B generated by the second training (additional training) performed by the second training execution unit 105 is resistant to a white-box attack that estimates the training data A. By performing further training (additional training) on ​​a trained machine learning model A, the time required to train the machine learning model can be shortened.

[0061] (B) Operation A training method for a machine learning model in the information processing device 1 as an example of an embodiment configured as described above will be described according to the flowchart (steps S1 to S10) shown in Fig. 5 with reference to Fig. 4. Fig. 4 is a diagram showing an overview of the training method for a machine learning model in the information processing device 1.

[0062] In step S1, an operator prepares an empty machine learning model (first machine learning model) and training data A. Information constituting the empty machine learning model and training data A is stored in a predetermined storage area such as the storage device 13.

[0063] In step S2, the first training execution unit 101 performs training (first training) on ​​an empty machine learning model (first machine learning model) using training data A to generate a trained machine learning model A (see symbol A1 in Figure 4). In step S3, the additional data creation unit 103 generates additional data by applying an optimization method to the machine learning model A (see reference symbol A2 in FIG. 4).

[0064] In step S4, the mini-batch creation unit 104 compares the amount of additional data with the amount of training data A, and checks whether the amount of additional data relative to the training data A is less than a predetermined ratio α.

[0065] If the result of the check is that the additional data is less than a predetermined ratio α with respect to the training data A (see the YES route in step S4), the process proceeds to step S6. In step S6, the mini-batch creation unit 104 adjusts the ratio of the additional data to the training data A to α by performing at least one of downsampling of the training data A and upsampling of the additional data.

[0066] On the other hand, if the confirmation result indicates that the ratio of the additional data to the training data A is equal to or greater than a predetermined ratio α (see NO route in step S4), the process proceeds to step S5. In step S5, the mini-batch creation unit 104 adjusts the ratio of the additional data to the training data A to α by performing at least one of upsampling of the training data A and downsampling of the additional data.

[0067] After that, in step S7, the mini-batch creation unit 104 randomly rearranges the training data A and the additional data separately. Furthermore, the mini-batch creation unit 104 divides the training data A and the additional data into N equal parts.

[0068] In step S8, the mini-batch creation unit 104 creates N-partitioned training data B by combining the training data A divided into N pieces (N-partitioned) and the additional data divided into N parts (see symbol A3 in FIG. 4).

[0069] In step S9, the second training execution unit 105 performs further training (additional training) on ​​the trained machine learning model A using each mini-batch of the N-divided training data B created by the additional training data creation unit 102, thereby generating an additionally trained machine learning model B (see symbol A4 in Figure 4).

[0070] In step S10, the second training execution unit 105 outputs the generated machine learning model B. Information constituting the machine learning model B is stored in a predetermined storage area such as the storage device 13.

[0071] (C) Effects Thus, according to the information processing device 1 as an example of an embodiment, the additional training data creation unit 102 creates training data B including additional data, and the second training execution unit 105 uses this training data B to perform further training (additional training) on ​​the trained machine learning model A, thereby generating an additionally trained machine learning model B.

[0072] The additional data is data that is not input in general machine learning model operations and is meaningless data for machine learning model A (X 0 ) as the initial value. Therefore, even if a white-box attack to infer training data is performed on machine learning model B that has undergone additional training, the influence of the additional data can prevent inference of training data A. When a white-box attack to infer training data is performed on machine learning model B, the additional data functions as a decoy and can prevent inference of training data A.

[0073] FIG. 6 is a diagram for explaining the result of a white-box attack on training data estimation for a machine learning model generated by the information processing device 1 as an example of an embodiment.

[0074] Fig. 6 shows an example of a training data inference attack performed on a machine learning model that infers (classifies) a number represented by an input number image based on the input number image. Fig. 6 also shows the result of a training data inference attack performed based on a machine learning model trained by a conventional method of adding noise to the parameters of the machine learning model, and the result of a training data inference attack performed based on a trained machine learning model created by the information processing device 1.

[0075] 6, "Model performance (Accuracy)" indicates the performance (accuracy) of the machine learning model trained by the conventional method and the machine learning model trained by the present information processing device 1. It can be seen that the performance (0.9863) of the machine learning model trained by the present information processing device 1 is equivalent to the performance (0.9888) of the machine learning model trained by the conventional method.

[0076] The section "Resistance to training data estimation (attack results)" shows images (number images) generated by performing a training data estimation attack on each machine learning model, along with the numerical values ​​that are the original correct data for those number images.

[0077] In the results of a training data inference attack based on a machine learning model trained by a conventional method, the number images in the training data are reproduced by a white-box attack. In contrast, in the results of a training data inference attack based on a trained machine learning model by the present information processing device 1, the number images in the training data are not reproduced except for a few, and it can be seen that the reproduction rate of the number images in the training data by the white-box attack is low. In other words, this shows that the trained machine learning model by the present information processing device 1 is resistant to training data inference attacks.

[0078] In addition, in a conventional defense method against white-box attacks that adds noise to the parameters of a machine learning model, the noise also significantly affects the inference ability of the model, resulting in a significant decrease in accuracy. In contrast, in a machine learning model trained by the information processing device 1, the additional data is less likely to affect the inference ability of normal inputs, so the decrease in accuracy can be relatively suppressed.

[0079] (D) Other The disclosed technology is not limited to the above-described embodiment, and can be implemented in various modifications without departing from the spirit of the embodiment. For example, the configurations and processes of the present embodiment may be selected or removed as necessary, or may be combined as appropriate.

[0080] In the above embodiment, the second training execution unit 105 performs further training (additional training) on ​​the machine learning model A that has been trained by the first training execution unit 101, but the present invention is not limited to this. The second training execution unit 105 may train an empty machine learning model using the second training data. Furthermore, the above disclosure enables a person skilled in the art to implement and manufacture the present embodiment.

[0081] (E) Notes Regarding the above embodiment, the following supplementary notes are further disclosed. (Appendix 1) Using a first machine learning model trained with the first training data, create additional data by mechanically generating meaningless data as initial values; combining the first training data with the additional data to generate second training data; Training a machine learning model using the second training data. An information processing program that causes a computer to execute a process.

[0082] (Appendix 2) Training the machine learning model using the second training data is retraining the first machine learning model using the second training data. 2. An information processing program according to claim 1.

[0083] (Appendix 3) The additional data is generated using an optimization technique. 3. The information processing program according to claim 1 or 2, characterized in that the program causes the computer to execute the process.

[0084] (Appendix 4) upsampling or downsampling is performed on at least one of the first training data and the additional data so that a ratio of the additional data to the first training data is a predetermined value; 4. The information processing program according to any one of claims 1 to 3, which causes the computer to execute a process.

[0085] (Appendix 5) Using a first machine learning model trained with the first training data, create additional data by mechanically generating meaningless data as initial values; combining the first training data with the additional data to generate second training data; Training a machine learning model using the second training data. An information processing method characterized in that the processing is executed by a computer.

[0086] (Appendix 6) Training the machine learning model using the second training data is retraining the first machine learning model using the second training data. 6. The information processing method according to claim 5,

[0087] (Appendix 7) The additional data is generated using an optimization technique. 7. The information processing method according to claim 5 or 6, characterized in that the processing is executed by the computer.

[0088] (Appendix 8) upsampling or downsampling is performed on at least one of the first training data and the additional data so that a ratio of the additional data to the first training data is a predetermined value; 8. The information processing method according to any one of claims 5 to 7, wherein the processing is executed by the computer.

[0089] (Appendix 9) an additional data creation unit that uses a first machine learning model trained with the first training data to mechanically generate meaningless data as initial values ​​to create additional data; a second training data creation unit that creates second training data by combining the first training data and the additional data; a second training execution unit that trains a machine learning model using the second training data; An information processing device comprising:

[0090] (Appendix 10) The second training execution unit executes retraining of the first machine learning model using the second training data. 10. The information processing device according to claim 9.

[0091] (Appendix 11) The additional data generating unit generates the additional data by using an optimization method. 11. The information processing device according to claim 9 or 10.

[0092] (Appendix 12) The second training data creation unit performs upsampling or downsampling on at least one of the first training data and the additional data so that a ratio of the additional data to the first training data becomes a predetermined value. 12. The information processing device according to any one of claims 9 to 11. [Explanation of symbols]

[0093] 1. Information processing device 11 Processors 12. Memory 13 Storage device 14 Graphics Processing Unit 14a Monitor 15 Input Interface 15a Keyboard 15b Mouse 16 Optical drive device 16a Optical disc 17 Device connection interface 17a Memory Device 17b Memory Reader / Writer 17c Memory Card 18 Network Interface 18a Network 19 Bus 100 Training Processing Unit 101 1st Training Execution Department 102 Additional training data creation department 103 Additional Data Creation Department 104 Minibatch Creation Department 105 2nd Training Execution Department

Claims

1. Using a first machine learning model trained with the first training data, generate additional data by mechanically generating noise or a constant value as an initial value; based on a result of comparing the number of the additional data with the number of the first training data to confirm a ratio of the additional data to the first training data, upsampling or downsampling at least one of the first training data and the additional data so that the ratio of the additional data to the first training data becomes a predetermined value; combining the upsampled or downsampled first training data with the additional data to generate second training data; Training a machine learning model using the second training data. Making a computer execute a process An information processing program comprising:

2. Training the machine learning model using the second training data is retraining the first machine learning model using the second training data.

2. The information processing program according to claim 1.

3. The additional data is generated using an optimization technique.

3. The information processing program according to claim 1, which causes the computer to execute a process.

4. Using a first machine learning model trained with the first training data, generate additional data by mechanically generating noise or a constant value as an initial value; based on a result of comparing the number of the additional data with the number of the first training data to confirm a ratio of the additional data to the first training data, upsampling or downsampling at least one of the first training data and the additional data so that the ratio of the additional data to the first training data becomes a predetermined value; combining the upsampled or downsampled first training data with the additional data to generate second training data; Training a machine learning model using the second training data. An information processing method characterized in that the processing is executed by a computer.

5. an additional data creation unit that creates additional data by mechanically generating additional data using a first machine learning model trained by the first training data as an initial value, the additional data being generated using noise or a constant value; a second training data creation unit that, based on a result of comparing the number of the additional data with the number of the first training data to confirm a ratio of the additional data to the first training data, upsamples or downsamples at least one of the first training data and the additional data so that the ratio of the additional data to the first training data becomes a predetermined value, and creates second training data by combining the first training data that has been upsampled or downsampled with the additional data; a second training execution unit that trains a machine learning model using the second training data; An information processing device comprising:

Citation Information

Patent Citations

  • Intermediate process state estimation method

    JP2020003906A

  • Model generation device, model generation method, model generation program, model generation system, inspection system, and monitoring system

    JP2020115312A

  • Learning method, learning program and learning apparatus

    JP2020119044A

  • Protecting cognitive systems from gradient-based attacks via the use of deceptive gradients

    JP2021501414A

  • Learned model update device, learned model update method, and program

    WO2019207770A1