Image conversion apparatus, image processing system, and learning method
The image conversion device and system address the challenge of extracting features from blurred images and preventing subject identification by converting images using an encoder and decoder, and processing them with an extractor and estimator, achieving effective feature extraction and secure image handling.
Patent Information
- Application Number
- JP2023207072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-19
AI Technical Summary
Blurred images have reduced information content, making it difficult to appropriately detect desired information, and existing technologies struggle to convert images while preventing subject identification.
An image conversion device and system that convert a first image into a second image using an encoder and decoder, where the second image is processed by an extractor and estimator to extract features without identifying the subject, and is transmitted over a communication network for further processing.
The system effectively extracts features from images while preventing subject identification, enabling secure handling and processing of image data while maintaining the ability to extract relevant information.
Smart Images

Figure 2025091674000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image conversion device that converts an image representing a subject, an image processing system, and a learning method for learning an image converter for converting an image.
Background Art
[0002] By collecting and analyzing a large number of data generated by many information processing devices existing in society, advanced utilization of the data becomes possible. Image data generated by information processing devices such as smartphones carried by users and drive recorders installed in vehicles ridden by users may include personal information that can identify the user, such as the user's face image. Therefore, it is not easy to receive the provision of image data that may include personal information, and even if it can be collected, prevention of leakage of personal information is required, and the handling becomes complicated.
[0003] Patent Document 1 describes an image processing system that can reduce the risks arising from handling personal information. The image processing system described in Patent Document 1 inputs a second anonymized image that is unclear to the extent that an individual in the shooting target cannot be identified into an inference model that outputs subject information regarding the person who is the subject from a first anonymized image that cannot be identified by a human, and obtains the subject information.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Since a blurred image has a reduced amount of information compared to a clear image, there are cases where desired information to be detected from the image data cannot be appropriately detected from the blurred image.
[0006] An object of the present disclosure is to provide an image conversion device that can convert an image so as to appropriately extract features of a subject while preventing identification of the subject. **Means for Solving the Problems**
[0007] The gist of the present disclosure is as follows.
[0008] (1) An image conversion device including a conversion unit that converts a first image representing a subject into a second image that is different from the first image and that cannot identify the subject, wherein the second image is an image from which features regarding the subject are extracted by an extractor that extracts the features from the first image. An image conversion device.
[0009] (2) An image processing system including a vehicle including an imaging device that generates a first image representing a subject, and a server device communicably connected to the vehicle via a communication network and including an extractor capable of extracting features regarding the subject from the first image, wherein the vehicle further includes an image conversion device that converts the first image into a second image that is different from the first image and that cannot identify the subject, the second image is an image from which the features extracted from the first image by the extractor are extracted when input to the extractor, and the second image is transmitted to the server device via the communication network, and the server device extracts the features by inputting the second image received via the communication network to the extractor. An image processing system.
[0010] (3) The image processing system according to (2) above, wherein the server device transmits the features extracted from the second image to the vehicle via the communication network, and the vehicle controls the operation of the vehicle based on the features received via the communication network. The image processing system according to (2) above.
[0011] An extractor pre-trained to extract features regarding the subject from a first image representing the subject, an image converter having an encoder that extracts feature amounts from the first image and a decoder that generates a second image different from the first image corresponding to the extracted features, and an estimator that estimates, from the second image, the class to which the subject represented by the first image corresponding to the second image belongs among a plurality of classes; The computer Inputs the first image representing a subject belonging to any correct class among the plurality of classes and having correct features to the encoder of the image converter to cause the decoder to generate the second image, Inputs the second image to the extractor to extract the features, Inputs the second image to the estimator to estimate the class, Updates the parameters defining the encoder so that the error between the correct features and the extracted features extracted from the second image corresponding to the first image by the extractor becomes small and the error between the correct class and the estimated class estimated from the first image by the estimator becomes large, Updates the parameters defining the decoder so that the error between the correct features and the extracted features becomes small, Updates the parameters defining the estimator so that the error between the correct class and the estimated class becomes small, A learning method including the above.
[0012] (5) Further prepares a discriminator that discriminates whether or not the subject represented by the input image belongs to a predetermined type, The computer Inputs the second image to the discriminator to discriminate whether or not the subject represented by the image belongs to the predetermined type, Updates the parameters defining the decoder so that it is determined that the subject represented by the second image does not belong to the predetermined type, The learning method according to (4) above.
[0013] According to the image conversion device of the present disclosure, an image can be converted so as to appropriately extract features of a subject while preventing identification of the subject.
Brief Description of Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0015] Hereinafter, with reference to the drawings, an image conversion device capable of converting an image so as to appropriately extract features of a subject while preventing identification of the subject will be described in detail. The image conversion device converts a first image representing a subject into a second image that is different from the first image and that cannot identify the subject. The subject of the first image is, for example, a vehicle occupant including a vehicle driver. The second image is an image from which features regarding the subject are extracted by an extractor that extracts the features.
[0016] FIG. 1 is a schematic configuration diagram of an image conversion system including an image conversion device.
[0017] In this embodiment, the image conversion system 100 includes a plurality of vehicles 1 and a server device 10. Each of the plurality of vehicles 1 is connected to a server device 10 via a wireless base station WBS (not shown) and a communication network NW by accessing the wireless base station WBS connected to the communication network NW via a gateway (not shown). Hereinafter, one of the plurality of vehicles 1 is also referred to as "vehicle 1". In the image conversion system 100, a plurality of wireless base stations WBS may be connected to the communication network NW.
[0018] FIG. 2 is a diagram for explaining the operation of the image conversion system 100.
[0019] An image conversion device 4 is mounted on the vehicle 1. In the vehicle 1, a first image P1 representing a subject is converted by an image converter C included in the image conversion device 4 into a second image P2 that is different from the first image P1 and that cannot identify the subject, and is transmitted to the server device 10 via the communication network NW.
[0020] The server device 10 extracts an extraction feature F related to the subject by inputting the received second image P2 to an extractor A. The extraction feature F may be transmitted to the vehicle 1 via the communication network NW. In that case, the vehicle 1 may execute control of operations of the vehicle 1, such as notification of a message to the driver and change of the upper limit speed, based on the received extraction feature F.
[0021] FIG. 3 is a schematic configuration diagram of the vehicle 1 on which the image conversion device 4 is mounted.
[0022] The vehicle 1 includes a driver monitor camera 2, a data communication module (DCM) 3, and an image conversion device 4. The driver monitor camera 2, the data communication module 3, and the image conversion device 4 are communicably connected via an in-vehicle network compliant with a standard such as a controller area network.
[0023] The driver monitoring camera 2 is an example of a sensor for generating a first image representing the face regions of vehicle occupants including the driver in a time series. The driver monitoring camera 2 has an infrared camera including a two-dimensional detector composed of an array of photoelectric conversion elements sensitive to infrared light, such as a CCD or a C-MOS, and an imaging optical system for forming an image of the region to be photographed on the two-dimensional detector. Further, the driver monitoring camera 2 has a light source that emits infrared light. The driver monitoring camera 2 is attached, for example, in the front of the vehicle interior, facing the face of the driver sitting on the driver's seat. The driver monitoring camera 2 irradiates the driver with infrared light every predetermined time (for example, 1 / 30 second to 1 / 10 second), and generates face images representing the driver's face at predetermined time intervals. The face images generated by the driver monitoring camera 2 are an example of the first images.
[0024] The data communication module 3 is an example of a communication interface, and is a device that executes wireless communication processing compliant with a predetermined wireless communication standard such as so-called 4G (4th Generation) or 5G (5th Generation). The data communication module 3 is communicably connected to the server device 10 via the wireless base station WBS and the communication network NW, for example, by accessing the wireless base station WBS. The data communication module 3 includes the data of the second image received from the image conversion device 4 in a wireless signal, and transmits the wireless signal to the server device 10. Further, the data communication module 3 receives the data included in the wireless signal received from the server device 10, and passes it to a vehicle control device (not shown) that controls the operation of the vehicle. Note that the data communication module 26 may be implemented as a part of the image conversion device 4.
[0025] FIG. 4 is a hardware schematic diagram of the image conversion device 4. The image conversion device 4 includes a communication interface 41, a memory 42, and a processor 43.
[0026] The communication interface 41 is an example of a communication unit and has a communication interface circuit for connecting the image conversion device 4 to the in-vehicle network. The communication interface 41 supplies the received data to the processor 43. Also, the communication interface 41 outputs the data supplied from the processor 43 to the outside.
[0027] The memory 42 is an example of a storage unit and has a volatile semiconductor memory and a non-volatile semiconductor memory. The memory 42 stores various data used for processing by the processor 43, such as, for example, a face image received from the driver monitor camera 2, parameters for defining a neural network used as an image converter for converting the face image, and an image converted from the face image. Also, the memory 42 further stores various application programs, such as, for example, a computer program for image processing executed by the image conversion device 4.
[0028] The processor 43 is an example of a control unit and has one or more processors and its peripheral circuits. The processor 43 may further have other arithmetic circuits such as a logical arithmetic unit, a numerical arithmetic unit, or a graphic processing unit.
[0029] FIG. 5 is a functional block diagram of the processor 43 included in the image conversion device 4.
[0030] The processor 43 of the image conversion device 4 has an image conversion unit 431 as a functional block. The image conversion unit 431 included in the processor 43 is a functional module implemented by a program executed on the processor 43. The computer program for realizing the functions of each part of the processor 43 may be provided in a form recorded on a computer-readable portable recording medium such as a semiconductor memory, a magnetic recording medium, or an optical recording medium. Alternatively, each of these parts included in the processor 43 may be implemented in the image conversion device 4 as an independent integrated circuit, a microprocessor, or firmware.
[0031] The image conversion unit 431 converts the first image P1 into the second image P2 by inputting the first image P1 into the image converter C.
[0032] FIG. 6 is a diagram for explaining the outline of the learning process of the image converter C, and FIG. 7 is a flowchart of the learning process of the image converter C.
[0033] In the learning of the image converter C, a pre-trained extractor A, an image converter C, and an estimator I are prepared in advance (step S1).
[0034] The image converter C includes an encoder N for extracting features from the first image P1 DROP and a decoder N for generating a second image P2 different from the first image P1 corresponding to the features extracted by the encoder N DROP . In FIG. 6, the encoder N RECO is, for example, a convolutional neural network (CNN) having a plurality of convolutional layers, and the decoder is, for example, a neural network having a plurality of transposed convolutional layers. DROP
[0035] The extractor A is, for example, a CNN and is pre-trained to extract features regarding the subject from the first image P1. The extractor A, for example, detects accessories (such as masks, sunglasses, earrings, etc.) worn on the face of the driver of the vehicle 1 from the first image P1 and outputs the confidence corresponding to each detection. By using a large number of images representing the subject as teacher data and performing the learning of the CNN in advance according to a predetermined learning method such as the error backpropagation method, the CNN operates as an extractor A that extracts features regarding the subject from the first image P1.
[0036] The estimator I estimates, from the second image P2, the class to which the subject represented in the first image P1 corresponding to the second image P2 belongs among a plurality of classes. The estimator is, for example, a CNN and estimates, from the second image P2, the class to which the subject represented in the first image P1 corresponding to the second image P2 belongs and outputs the confidence corresponding to each class. The class may be, for example, an identifier for identifying the person who is the subject.
[0037] The first image P1 used for training the image converter C may include images P1X, P1Y, and P1Z representing different subjects X, Y, and Z, respectively. At this time, the estimator I may be configured to estimate from the second image P2 which of X, Y, and Z is the subject represented in the corresponding first image P1.
[0038] In the training of the image converter C, first, a sufficient number of first images P1 are prepared as teacher data. Each of the first images P1 represents, as a subject, one of a plurality of persons wearing zero or more accessories, and an identifier (correct class) representing the person and accessory information (correct feature) representing the accessories worn by the person are assigned as labels.
[0039] In the training of the image converter C, by causing a computer to execute a predetermined program, the computer is caused to execute a process of converting the first image P1 into the second image P2 by the image converter C (step S2). That is, the computer inputs each of the first images P1 into the encoder N DROP of the image converter C and obtains the second image P2 from the decoder N RECO .
[0040] The computer inputs the second image P2 into the extractor A and extracts the extracted feature F. The computer also inputs the second image P2 into the estimator I and estimates the class (estimated class) to which the subject belongs (step S3).
[0041] The computer updates the parameters defining the encoder N DROP , the decoder N RECO , and the estimator I, respectively (step S4).
[0042] The computer updates the parameters defining the encoder N DROP so that the error between the correct feature and the extracted feature F becomes small and the error between the correct class and the estimated class becomes large. The encoder N DROPThe update of the parameters that define [the object] is, for example, according to the loss function L expressed by the following formula (1). DROP It may follow a learning method such as the error backpropagation method using [the loss function L].
[0043] L DROP = CELoss(Z, y A ) - CELoss(Z, y1) … Formula (1) Here, CELoss(Z, y A ) is the cross - entropy error of the extracted feature F, and CELoss(Z, y1) is the cross - entropy error of the estimated class.
[0044] The computer updates the parameters that define the decoder N RECO so that the error between the correct feature and the feature F becomes small. The update of the parameters that define the decoder N RECO may follow a learning method such as the error backpropagation method using the following loss function L RECO .
[0045] L RECO = CELoss(Z, y A ) … Formula (2)
[0046] The computer updates the parameters that define the estimator I so that the error between the correct class and the extracted class becomes small. The update of the parameters that define the estimator I may follow a learning method such as the error backpropagation method using the following loss function L I .
[0047] L I = CELoss(Z, y1) … Formula (3)
[0048] By performing learning in this way, an image converter C can be obtained that converts the first image P1 that can identify the person who is the subject into a second image P2 that cannot identify the person who is the subject and can extract the features extracted from the first image P1.
[0049] Furthermore, a discriminator D for determining whether the subject represented in the second image P2 belongs to a predetermined type may be prepared. The predetermined type may be, for example, an animal species such as a cat or a dog. The discriminator D is, for example, a neural network that outputs "1" when the subject is determined to belong to a predetermined type and "-1" when it is determined not to belong. The neural network constituting the discriminator D may have a fully connected layer.
[0050] When the discriminator D is prepared in the learning of the image converter C, the computer inputs the second image P2 to the discriminator D and obtains each discrimination result.
[0051] The computer updates the parameters defining the decoder N RECO so that the subject represented in the second image P2 is determined not to belong to a predetermined type. When updating the parameters defining the decoder N RECO , for example, the following formula (2)' may be used instead of the above formula (2).
[0052] L RECO = CELoss(Z, y A ) - λN D … Formula (2)' Here, N D is the determination result by the discriminator D, and λ is a hyperparameter set as appropriate. The hyperparameter λ may be changed to a larger value when the ratio of the determination result by the discriminator D being determined to belong to a predetermined type is larger than a predetermined threshold.
[0053] The discriminator D is pre-trained. Also, in addition to the second image P2, the computer may input a third image P3 representing a subject belonging to a predetermined type to the discriminator D and update the parameters defining the discriminator D using each discrimination result. The update of the parameters defining the discriminator D may follow a learning method such as the error backpropagation method using the loss function L D represented by, for example, the following formula (4).
[0054] L D= (1 - N D (P)) + (1 + N D (Z)) … Equation (4) Here, N D (P) is the discrimination result of the third image P3, and N D (Z) is the discrimination result of the second image.
[0055] When using the image converter C learned using the discriminator D in this way, the first image P1 can be converted into a second image P2 representing a subject of a predetermined type different from the subject represented in the first image P1.
[0056] By executing the learning process in this way, it is possible to configure the image conversion unit C that converts an image so as to appropriately extract the features of the subject while preventing the identification of the subject.
[0057] The parameters defining the image conversion unit C configured in this way are stored in the memory 42 of the image conversion device 4 and used for the image conversion process by the image conversion device 4.
[0058] Note that the image conversion unit C can also be configured by a machine learning model other than neural networks such as SVM and AdaBoost. The same applies to the extractor A, estimator I, and discriminator D.
[0059] Those skilled in the art should understand that various changes, substitutions, and modifications can be made to this without departing from the spirit and scope of the present disclosure.
Description of Reference Numerals
[0060] 1 Vehicle 4 Image conversion device 431 Conversion unit 10 Server device 100 Image conversion system C Image converter A Extractor I Estimator D Discriminator
Claims
1. A conversion unit that converts a first image representing a subject into a second image that is different from the first image and that cannot identify the subject, The second image is an image from which features regarding the subject are extracted by an extractor that extracts features regarding the subject from the first image, An image conversion device.
2. A vehicle including an imaging device that generates a first image representing a subject, and a server device communicably connected to the vehicle via a communication network and including an extractor capable of extracting features regarding the subject from the first image, the image processing system comprising: The vehicle further includes an image conversion device that converts the first image into a second image that is different from the first image and that cannot identify the subject, and the second image is an image from which the features extracted from the first image by the extractor are extracted when input to the extractor, and transmits the second image to the server device via the communication network, The server device extracts the features by inputting the second image received via the communication network to the extractor, An image processing system.
3. The server device transmits the features extracted from the second image to the vehicle via the communication network, The vehicle controls the operation of the vehicle based on the features received via the communication network, The image processing system according to claim 2.
4. Preparing an extractor pre-trained to extract features regarding a subject from a first image representing the subject, an image converter having an encoder that extracts feature amounts from the first image and a decoder that generates a second image different from the first image corresponding to the extracted features, and an estimator that estimates, from the second image, the class to which the subject represented by the first image corresponding to the second image belongs among a plurality of classes, The computer is configured to: Input the first image representing a subject that belongs to any correct class among the plurality of classes and has correct features into the encoder of the image converter to generate the second image in the decoder, Input the second image into the extractor to extract the features, Input the second image into the estimator to estimate the class, Update the parameters defining the encoder so that the error between the correct features and the extracted features extracted from the second image corresponding to the first image becomes small, and the error between the correct class and the estimated class estimated from the first image by the estimator becomes large, Update the parameters defining the decoder so that the error between the correct features and the extracted features becomes small, Update the parameters defining the estimator so that the error between the correct class and the estimated class becomes small, A learning method including the above.
5. Further prepare a discriminator for discriminating whether the subject represented in the input image belongs to a predetermined type, The computer, Input the second image into the discriminator to discriminate whether the subject represented in the image belongs to the predetermined type, Update the parameters defining the decoder so that it is determined that the subject represented in the second image does not belong to the predetermined type, The learning method according to claim 4.
Citation Information
Patent Citations
Image processing system, server device, image processing method, and image capturing device
JP2023037362A