Information processing method

By dividing images into patches and using separate servers for learning, the method ensures privacy protection and maintains recognition performance in neural network learning, addressing data restoration and adjustment challenges.

JP7708035B2Active Publication Date: 2025-07-15TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022133601
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-07-15
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Existing neural network learning methods fail to provide sufficient privacy protection, as data can be potentially restored from learning results, and data adjustment is difficult when improving model performance, especially in supervised learning with sensitive information.

Method used

The method divides images into small patches, storing them separately on individual servers, using a two-part learning model structure with Upper and Lower models operating on separate servers, ensuring privacy by preventing restoration of the original image from calculation results.

Benefits of technology

This approach enables privacy-protected learning while maintaining recognition performance, allowing data analysis with privacy-protected patches and accelerating learning through individual loss calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708035000002
    Figure 0007708035000002
  • Figure 0007708035000003
    Figure 0007708035000003
  • Figure 0007708035000004
    Figure 0007708035000004
Patent Text Reader

Abstract

To enable privacy protection in learning of models.SOLUTION: In an information processing method for generating a learned model for recognizing an image, a learning model consists of a plurality of first models and a second model different from the first models. The information processing method divides an image to be used for learning for each patch, inputs and calculates each of a plurality of divided patches into each of the plurality of predetermined first models for each patch, integrates the output of calculation results of each of the first models in the second model, learns the learning model, and generates the learned model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing method.

Background Art

[0002] Patent Document 1 discloses a technique of dividing a neural network into a plurality of parts and specifying parameters based on input / output characteristics.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the prior art, by dividing a neural network, it is possible to consider privacy protection by giving each neural network the learning results up to a certain point, respectively. However, there is also a method of restoring the original data from the data during the calculation. Therefore, it is assumed that sufficient privacy protection cannot be achieved only by dividing the neural network, and there is room for improvement.

[0006] An object of the present disclosure is to provide an information processing method that enables privacy protection in the learning of a learning model.

Means for Solving the Problems

[0007] The information processing method according to claim 1 is an information processing method for generating a learned model for recognizing an image, wherein the learning model is composed of a plurality of first models and a second model different from the first model. The image used for learning is divided patch by patch, and each of the divided Each of the plurality of patches is stored independently, and each of the plurality of first models operates independently on a predetermined server. Each of the plurality of patches is input and calculated to each of the predetermined plurality of first models patch by patch, and the outputs of the calculation results of each of the first models are integrated in the second model, and the learned model is generated by learning the learning model.

[0008] The information processing method according to claim 1 is configured to divide each of the patches and divide the learning model corresponding to the patches. Thereby, privacy protection in the learning of the learning model can be enabled. In addition, since the original image cannot be restored from a single server, privacy protection is possible.

[0009] The information processing method according to claim 2 is the information processing method according to claim 1, wherein in the learning model, the number of the plurality of first models corresponding to the plurality of patches is configured to be equal to the number of the plurality of patches. According to the information processing method according to claim 2, learning data can be handled in a manner that enables privacy protection.

[0011] Claim 3The information processing method described in [reference] is, in the information processing method described in claim 1, the second model receives and integrates the outputs of the plurality of first models, calculates the loss of the second model, and trains the learning model. Claim 3 According to the information processing method described in [reference], the calculation of patches can be handled independently, and the model can be configured in a manner that enables privacy protection.

[0012] Claim 4 The information processing method described in [reference] is, in the information processing method described in claim [reference], 3 in the information processing method described in [reference], the learning model further includes a third model, and the third model receives the output of each of the plurality of first models, calculates an individual loss for each output, and calculates an overall loss based on the individual loss and the loss of the second model, and trains the learning model. Claim 4 According to the information processing method described in [reference], by considering individual losses and providing feedback in learning as the overall loss, the learning of the learning model can be accelerated.

Advantages of the Invention

[0013] According to the technology of the present disclosure, privacy protection in the learning of a learning model is enabled.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiment for Carrying Out the Invention

[0015] The outline of the embodiment of the present invention will be described. Image recognition has been improved by deep learning, especially in supervised learning that requires correct data, and has become widely used in general. On the other hand, the number of applications that utilize models learned using sensitive information such as personal information, such as face recognition and emotion recognition, is also increasing. In addition, laws related to privacy protection, such as GDPR (General Data Protection Regulation) and CCPA (California Consumer Privacy Act), have been enacted in various countries, and privacy protection has become important in data collection and model learning.

[0016] Regarding the existing learning methods as the premise of this embodiment, there are the above - mentioned problems. As learning methods related to the problems, there are federated learning such as Non - Patent Document 1 and the method of split neural network such as Non - Patent Document 2.

[0017] In the method of Non - Patent Document 1, without collecting data on the central server, only the parameters of the results learned at the edge (the terminal on the data - providing side) are collected, and the server uses the collected parameters to learn the model. However, in Non - Patent Document 1, as the model grows, the number of parameters increases, the communication volume increases accordingly, and the computational amount at the edge increases. In addition, it is necessary to arrange the latest model to be learned on each edge.

[0018] In the method of Non-Patent Document 2, by splitting a CNN (Convolutional Neural Network) in the middle and holding and training it on separate user terminals, the user who provides the data only needs to send the learning result of the neural network up to the middle to the terminal of the user who processes the learning, rather than the data itself. However, in order to increase the amount of data, it is necessary to increase the data providers (clients), and for this purpose, the communication during learning becomes complicated. Also, since data is not collected, although privacy protection is possible, data adjustment such as data confirmation in case of low accuracy or provision of correct data cannot be performed.

[0019] Although each method realizes privacy protection by not collecting data centrally, there is a problem that data adjustment cannot be performed when it is desired to improve the performance of the model. Also, methods such as masking parts containing privacy information such as faces are conceivable, but are likely to affect the recognition performance.

[0020] When training a neural network in deep learning for image recognition, a large amount of images are required, and in the case of supervised learning, it is necessary to label the correct data after collection. The collected images contain information related to privacy such as, for example, faces and vehicle license plates. Even if security is ensured, when storing in one place, it contains privacy information and careful handling is required. Also, collection of privacy information requires consent from the target person, and large-scale collection is difficult.

[0021] In this embodiment, the images used for learning are divided into small-sized patches and stored in individual servers. Also, based on the split neural network of Non-Patent Document 2, a patch split neural network (Patch split NN) is used as the configuration of the learning model adapted to the split patches. An example of the configuration of the learning model will be described later.

[0022] By making patches of a small size, even if the original image contains privacy information, it can be divided into sizes that cannot be discriminated, and each patch can be made into non-privacy information. Also, by storing the patches separately, privacy information cannot be restored unless a certain number of the patches that make up the original image are taken out at the same time. However, some users may be given the authority to restore the original image to enable data analysis.

[0023] The learning model divides the CNN into a two-part structure of an Upper model and a Lower model. During learning, the Upper model and the Lower model are operated on separate servers so that the data (patches) used for learning do not gather in one place as they are. By outputting to the Lower through the calculation of the Upper model, it becomes impossible to restore from the calculation results. Also, by preparing the number of Upper models corresponding to the number of patches, the patches do not gather in one server. Note that when the number of Upper models is less than the number of patches, they are appropriately assigned to the Upper models. The Upper model is an example of the first model of the present disclosure, and the Lower model is an example of the second model of the present disclosure.

[0024] According to the method of the present embodiment, a learned model having sufficient recognition performance can be generated even when learning is performed by dividing patches, as compared with the case where the original image is used as it is.

[0025] FIG. 1 is a diagram showing the configuration of the information processing system 10. As shown in FIG. 1, the information processing system 10 includes a user terminal 102, a plurality of patch servers 110 (110 1~N ) that store patches, and a plurality of Upper servers 112 (112 that store each of the plurality of Upper models 1~N) and a Lower server 114 that stores the Lower model are connected via the network N. The user terminal 102 is a terminal for inputting images used for learning. The patch server 110 is a storage server that stores the divided patches. The Upper server 112 and the Lower server 114 are servers that store and execute the learning model.

[0026] FIG. 2 is a block diagram showing the hardware configuration of the computer (CM) of the user terminal 102 and each server (110, 112, 114). As shown in FIG. 2, the computer (CM) includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17. Each configuration is connected to be communicable with each other via a bus 19.

[0027] The CPU 11 is a central processing unit that executes various programs and controls each unit. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 performs control of each of the above configurations and various arithmetic processes according to a program stored in the ROM 12 or the storage 14. In the present embodiment, an information processing program is stored in the ROM 12 or the storage 14. Since other hardware configurations may use a general configuration as a computer, the description is omitted.

[0028] FIG. 3 is a diagram schematically showing the data flow of the information processing system 10 and a configuration example of the learning model. A patch is each of a plurality of patches obtained by dividing an image used for learning. The learning model learned in the information processing system 10 is divided into each of a plurality of upper models and a Lower model. This will be described later. Note that the division into patches is assumed to be performed by the user terminal 102, but the division process into batches may be performed by one patch server 110 and distributed to other patch servers 110.

[0029] In (1) of FIG. 3, the image input by the user is divided into a plurality of patches and uploaded to each of the patch servers 110. In (2), the patches are stored in separate patch servers 110. In (3), a plurality of Upper models are prepared and the patches are input to the respective Upper models. Individual learning by the Upper models among the learning models is performed. Here, each individual Upper model is started and operates on an individual and independent Upper server 112. In (4), the output results of each Upper model are transferred to the Lower model, and the output results of each Upper model are integrated by the Lower model among the learning models to calculate the loss, and the learning model is learned. By calculating the integrated loss, the recognition result of the learning model is obtained. In this way, the information processing system 10 generates and outputs a learned model. (5) is an arbitrary process according to the learning status of the learning model. In (5), according to the learning status of the learned learning model, necessary patches are collected, the image is restored and analyzed from the collected patches, and the correct data is appropriately labeled on the patches.

[0030] An example of the patch division pattern will be described. As the division pattern, either (A) simple division or (B) overlap may be used. (A) Simple division divides, for example, the original size of the image without excess or deficiency. In (A) simple division, for an image of 32x32 size, if the division size is set to 16x16 size, it is divided into 4 patches. The division size may be set to 8x8 size and divided into 8 patches. In the case of (B) overlap, it is divided with a fixed length and overlapped so that there are no breaks. For an image of 32x32 size, if the division size is 16x16 size and the slide is 8, it is divided into 9 patches. If the division size is 8x8 size and the slide is 4, it is divided into 49 patches. The division pattern may be appropriately used according to the complexity of the image.

[0031] FIG. 4 is a diagram showing a configuration example of a neural network of a learning model. The upper part shows the learning model as a Patch split NN, and the lower part shows ResNet18, which is a type of base CNN, for comparison. The Upper model assumes a case where it is configured in the same manner as the number of patches. Each patch indicated by numbers 1 to N is input to the corresponding Upper model (Upper in the figure). The Lower model is provided with a concatenate layer that integrates the outputs from the individual Upper models, and after integration, it is input to the Lower layer. Note that even if the Upper model is not equal to the number of patches, as long as it is at the patch level where privacy information can be protected even if there is a patch model that inputs several patches. The Upper model corresponds to the Input Layer and Residual Block1 of ResNet18. The Lower model corresponds to Residual Blocks 2 to 4 and the Output Layer of ResNet18.

[0032] (Flow of control) FIG. 5 is a sequence for explaining the processing flow as an information processing method executed by the information processing system 10 of the present embodiment. In the information processing method, a plurality of computers (each of the user terminal 102, the patch server 110, the Upper server 112, and the Lower server 114) are combined to execute the processing.

[0033] In step S100, the user terminal 102 divides an image into patches and transmits each patch to each of the patch servers 110.

[0034] In step S102, each of the patch servers 110 outputs to the Upper model (Upper server 112 where there is an Upper model) corresponding to the stored patch.

[0035] In step S104, each of the Upper servers 112 performs the calculation of the Upper model and outputs the calculation result to the Lower model (Lower server 114 where there is a Lower model).

[0036] In step S106, the Lower server 114 integrates the calculation results of each of the received Lower models, calculates the loss, and trains the learning model. As for the learning method, a method of adjusting the weight parameters of the Upper model and the Lower model from the calculated loss may be used, similar to the learning of the CN.

[0037] As described above, the information processing system 10 of the present embodiment is configured to be divided into each of the patches and the learning model is divided corresponding to the patches, thereby enabling privacy protection in the learning of the learning model.

[0038] [Modification Example] FIG. 6 is a configuration example in the case of further providing a multi-calculation model that improves the above-described embodiment and calculates individual losses of the Upper. Note that the multi-calculation model may be provided in the Lower server 114 or may be provided in an individual server. The multi-calculation model is an example of the third model of the present disclosure.

[0039] In the multi-calculation model, the calculation results of the individual Upper models are received, and individual losses (upper_loss) are calculated for the individual calculation results in the Adaptation Net. By introducing such a configuration during learning, the learning of the Upper model can be accelerated. Note that such a configuration is not necessary during inference.

[0040] The Lower server 114 receives the calculation results of the individual losses calculated for the Upper model, and calculates the overall loss according to the following formula (1). For the learning of the learning model, the overall loss may be used instead of the loss of the Lower model. [Equation] ···(1)

[0041] Full_Loss is the loss of the entire neural network, Lower_Loss is the loss calculated by the Lower model, and Upper_Loss is the individual loss calculated for each Upper model. Also, Num_Of_Model is the number of Upper models, Num_Of_Path is the number of divided patches, and α is the coefficient of the entire Upper_Loss. Note that each layer in the Adaptation Net may be a neural network layer of AdaptiveAvgPool, Linner, BatchNorn, Linner, BatchNorn, ReLU, Linner in order from the input. In the verification example using the method in this modification example, it has been confirmed that when verifying using CIFAR10 and CIFAR100 for the dataset, recognition accuracy equivalent to or higher than that in the case of learning using the original image as it is can be obtained.

[0042] In addition, in the above embodiment, various processes executed by the CPU 11 by reading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacture, such as an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), and a dedicated electric circuit which is a processor having a circuit configuration dedicated to executing specific processes such as an ASIC (Application Specific Integrated Circuit). Further, each of the above-described processes may be executed by one of these various processors, or may be executed by a combination of two or more processors of the same type or different types (for example, a plurality of FPGAs, a combination of a CPU and an FPGA, etc.). Also, the hardware structure of these various processors is, more specifically, an electric circuit combining circuit elements such as semiconductor elements.

[0043] Also, in the above embodiment, the information processing program has been described in a mode where it is pre-stored (installed) in a computer-readable non-transitory recording medium. For example, the information processing program is pre-stored in the ROM 12 or the storage 14. However, it is not limited to this, and each program may be provided in a form recorded on a non-transitory recording medium such as a CD-ROM (Compact Disc Read Only Memory), a DVD-ROM (Digital Versatile Disc Read Only Memory), and a USB (Universal Serial Bus) memory. Further, the information processing program may be in a form downloaded from an external device via a network.

[0044] The processing flow described in the above embodiment is an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within a range not departing from the gist.

Explanation of Reference Numerals

[0045] 10 Information processing system 102 User terminal 110 Patch server 112 Upper server 114 Lower server

Claims

1. An information processing method for generating a learned model for recognizing an image, comprising: The learning model is composed of a plurality of first models and a second model different from the first models; Dividing the images used for learning into patches; Each of the plurality of divided patches is stored independently, and each of the plurality of first models operates independently on a predetermined server. Each of the plurality of patches is input and calculated to each of the predetermined plurality of first models respectively; Integrating the outputs of the calculation results of each of the first models in the second model, and learning the learning model to generate the learned model; Information processing method.

2. The information processing method according to claim 1, wherein in the learning model, the number of the plurality of first models corresponding to the plurality of patches is configured to be equal to the number of the plurality of patches.

3. The information processing method according to claim 1, wherein the second model receives and integrates the outputs of the calculation results of the plurality of first models, calculates the loss of the second model, and learns the learning model.

4. The learning model further includes a third model; The third model receives the output of each of the plurality of first models, calculates an individual loss for each output; The information processing method according to claim 3, wherein an overall loss is calculated based on the individual losses and the loss of the second model, and the learning model is learned.

Citation Information

Patent Citations

  • Solid-state imaging system, solid-state imaging device, information processing device, image processing method and program

    JP2020047191A

  • Information processing device, program, and information processing method

    JP6784162B2

  • Neural architecture search system and search method

    WO2022102054A1

  • Image processing apparatus, image processing method, and program

    WO2022124067A1