Image processing device, image processing method, image processing system, and program
The image processing device ensures effective training data generation for machine learning models in autonomous driving by anonymizing images based on attributes, addressing privacy concerns and maintaining face orientation, enhancing sustainable transportation systems.
Patent Information
- Application Number
- JP2022106676
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Conventional image processing technologies for generating training data for machine learning models in autonomous driving systems often compromise privacy by converting original facial images, leading to a loss of feature information.
An image processing device and method that acquires image attributes, performs anonymization based on these attributes, and ensures that the anonymized images meet predetermined requirements to maintain privacy while being effective for training, involving face replacement or mosaic processing, and ensuring continuity and consistency of face orientations.
Generates training data that is effective for machine learning models while protecting privacy, maintaining face orientation information, thereby contributing to sustainable transportation systems.
Smart Images

Figure 0007805260000001 
Figure 0007805260000002 
Figure 0007805260000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, an image processing system, and a program. [Background technology]
[0002] In recent years, efforts to provide access to sustainable transportation systems that take into consideration vulnerable traffic participants have been gaining momentum. To achieve this, efforts are being made to further improve traffic safety and convenience through research and development of autonomous driving technology. For example, a technology that annotates individual facial images to generate training data used in training machine learning models is known. Patent Document 1 discloses a technology that generates a composite facial image by referencing facial images of multiple people stored in a facial image database and enables annotation operations to be performed on the generated composite facial image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 5930450 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology described in Patent Document 1 protects the privacy of multiple people by having an annotator perform annotation operations on a composite face image synthesized from the facial images of multiple people. However, with conventional technology, the original image is converted to protect privacy, which can result in loss of feature information from the original image. As a result, it has sometimes been impossible to generate training data that is effective for training a machine learning model while protecting the privacy of people depicted in the facial image.
[0005] The present invention has been made in consideration of the above circumstances, and an object of the present invention is to provide an image processing device, an image processing method, an image processing system, and a program that can generate training data effective for training a machine learning model while protecting the privacy of people depicted in facial images, thereby contributing to the development of sustainable transportation systems. [Means for solving the problem]
[0006] The image processing device, image processing method, image processing system, and program according to the present invention employ the following configuration. (1): An image processing device according to one embodiment of the present invention includes an image attribute acquisition unit that acquires image attributes, which are the shooting manner of an input image; an image conversion unit that performs an anonymization process on the input image; and an image determination unit that determines whether the input image that has been anonymized satisfies predetermined requirements. If the image determination unit determines that the input image that has been anonymized satisfies the predetermined requirements, it performs predetermined processing on the input image that has been anonymized. The predetermined requirements are determined according to the acquired image attributes.
[0007] (2) In the above aspect (1), the predetermined process is a process of saving the input image that has been subjected to the anonymization process as a target image for annotation work.
[0008] (3): In the above aspect (1), the predetermined processing is a processing of saving the input image that has been subjected to the anonymization processing as learning information for generating a behavior prediction model that predicts the behavior of the person depicted in the input image.
[0009] (4) In the above aspect (1), the predetermined process is a process of transmitting the input image, which has been subjected to the anonymization process, to an image server via a communication means.
[0010] (5): In the above aspect (1), the image attribute is information that at least indicates whether the input image is an image of the interior of a vehicle equipped with a camera that captured the input image, or an image of the exterior of the vehicle.
[0011] (6): In the above aspect (5), the anonymization process includes a process of changing the face of a person depicted in the input image to the face of another person, and the predetermined requirement includes whether the gaze direction of the face of the person coincides with the gaze direction of the face of the other person when the image attribute indicates that the image is an image of the interior of the vehicle.
[0012] (7): In the aspect of (6) above, the predetermined requirements include, when the image attribute indicates that the image is an image of the interior of the vehicle, whether the gaze direction of the person's face matches the gaze direction of the face of the other person, and whether the facial direction of the person's face matches the facial direction of the face of the other person, and when the image attribute indicates that the image is an image of the exterior of the vehicle, whether the facial direction of the person's face matches the facial direction of the face of the other person.
[0013] (8): In the above aspect (6), when the image attribute indicates that the image is an image of the exterior of the vehicle, the specified requirement does not include whether the gaze direction of the face of the person coincides with the gaze direction of the face of the other person.
[0014] (9): In the above aspect (1), if the image determination unit determines that the input image to which the anonymization process has been applied does not satisfy the specified requirements, the image conversion unit applies the anonymization process to the input image again.
[0015] (10): In the above aspect (1), if the image determination unit determines that the input image that has been anonymized does not satisfy the specified requirements, the image conversion unit does not perform the specified processing on the input image that has been anonymized.
[0016] (11): Another aspect of the image processing system of the present invention includes an image attribute acquisition unit that acquires image attributes, which are the shooting manner of an input image; an image conversion unit that performs an anonymization process on the input image; and an image determination unit that determines whether the input image that has been anonymized satisfies predetermined requirements. If the image determination unit determines that the input image that has been anonymized satisfies the predetermined requirements, it performs predetermined processing on the input image that has been anonymized. The predetermined requirements are determined according to the acquired image attributes.
[0017] (12): Another aspect of the image processing method of the present invention is a method in which a computer acquires image attributes, which are the shooting manner of an input image, performs an anonymization process on the input image, determines whether the anonymized input image satisfies predetermined requirements, and if it determines that the anonymized input image satisfies the predetermined requirements, performs predetermined processing on the anonymized input image, where the predetermined requirements are determined according to the acquired image attributes.
[0018] (13): Another aspect of the present invention provides a program that causes a computer to acquire image attributes, which are the shooting manner of an input image, perform an anonymization process on the input image, determine whether the anonymized input image satisfies predetermined requirements, and, if it is determined that the anonymized input image satisfies the predetermined requirements, perform predetermined processing on the anonymized input image, where the predetermined requirements are determined according to the acquired image attributes. [Effects of the Invention]
[0019] According to (1) to (13), it is possible to generate learning data that is effective for training a machine learning model while protecting the privacy of people depicted in face images. [Brief explanation of the drawings]
[0020] [Figure 1]1 is a diagram showing an overview of a system 1 including an image processing device 100 according to the present embodiment. [Figure 2] 1 is a diagram illustrating an example of a functional configuration of an image processing device 100 according to the present embodiment. [Figure 3] 2A and 2B are diagrams showing examples of an interior image and an exterior image acquired from a vehicle M1. [Figure 4] 10 is a diagram for explaining the processing executed by the image processing unit 130. FIG. [Figure 5] 10 is a diagram for explaining the processing executed by the image conversion unit 140. FIG. [Figure 6] 10A and 10B are diagrams showing an example of time-series images of the interior of a vehicle converted by an image conversion unit 140. FIG. [Figure 7] 10 is a diagram for explaining the processing executed by the image determination unit 150. FIG. [Figure 8] FIG. 10 is a diagram illustrating an example of annotation work performed by an annotator. [Figure 9] FIG. 10 is a diagram illustrating an example of driving assistance using a trained model 180. [Figure 10] FIG. 10 is a diagram showing an example of the flow of processing executed by an image conversion unit 140. [Figure 11] 10 is a diagram showing an example of the flow of processing executed by an image determination unit 150. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0021] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of an image processing device, an image processing method, an image processing system, and a program according to the present invention will be described with reference to the accompanying drawings.
[0022] [overview] Fig. 1 is a diagram showing an overview of a system 1 including an image processing device 100 according to this embodiment. As shown in Fig. 1, the system 1 includes at least one vehicle M1 and one vehicle M2, the image processing device 100, and a terminal device 200. For ease of explanation, the vehicle M1 and the vehicle M2 are illustrated as different vehicles, but these vehicles may be the same.
[0023] The vehicle M1 is, for example, a four-wheel drive vehicle such as a hybrid vehicle or an electric vehicle, and includes at least a camera that captures images of the interior of the vehicle M1 and a camera that captures images of the exterior of the vehicle M1. While the vehicle M1 is traveling, the vehicle M1 transmits images of the interior and exterior of the vehicle captured by these cameras to the image processing device 100 via a network NW such as a cellular network, a Wi-Fi network, or the Internet.
[0024] The image processing device 100 is a server device that, upon receiving captured image data including vehicle interior images and vehicle exterior images from the vehicle M1, performs image conversion, described below, on the received captured image data. This image conversion is a process for protecting the privacy of people appearing in the vehicle interior images and vehicle exterior images. The image processing device 100 transmits the obtained converted image data to the terminal device 200 via the network NW.
[0025] The terminal device 200 is a terminal device such as a desktop personal computer or a smartphone. When the user of the terminal device 200 acquires converted image data from the image processing device 100, the user performs an annotation assignment process (described later) on the acquired converted image data. When the annotation assignment process is completed, the user of the terminal device 200 transmits the annotated image data, in which the annotations have been assigned to the converted image data, to the image processing device 100.
[0026] When the image processing device 100 receives annotated image data from the terminal device 200, it uses the received annotated image data as learning data and generates a trained model (described later) using an arbitrary machine learning model. This trained model is, for example, a behavior prediction model that, in response to an input of an exterior image of a vehicle, outputs the predicted behavior (trajectory) of a person depicted in the exterior image of the vehicle, or, in response to input of an interior image and an exterior image of the vehicle, takes into account the line of sight of the driver depicted in the interior image of the vehicle and calls attention to a pedestrian depicted in the exterior image of the vehicle.
[0027] The image data used as the training data at this time may be annotated image data in which annotations are added to converted image data, or may be annotated image data in which the converted image data is reconverted into captured image data while leaving the annotations intact (i.e., annotated image data in which annotations are added to captured image data). By using annotated image data in which annotations are added to captured image data as training data, it is possible to use training data that is more realistic and in which the effects of image conversion have been removed.
[0028] After generating the trained model, the image processing device 100 distributes the generated trained model to the vehicle M2 via the network NW. Like the vehicle M1, the vehicle M2 is a four-wheel drive vehicle such as a hybrid vehicle or an electric vehicle, and while the vehicle M2 is traveling, at least one of an interior image and an exterior image captured by a camera is input into the trained model to obtain behavior prediction data of people present around the vehicle M2. The driver of the vehicle M2 can refer to the obtained behavior prediction data and use it to drive the vehicle M2. Each process will be described in more detail below.
[0029] [Functional configuration of image processing device] 2 is a diagram illustrating an example of the functional configuration of an image processing device 100 according to this embodiment. The image processing device 100 includes, for example, a communication unit 110, a transmission / reception control unit 120, an image processing unit 130, an image conversion unit 140, an image determination unit 150, a trained model generation unit 160, and a storage unit 170. These components are implemented by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be implemented by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or may be implemented by a combination of software and hardware. The program may be stored in advance in a storage device (a storage device having a non-transitory storage medium) such as a hard disk drive (HDD) or flash memory, or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or CD-ROM and installed by inserting the storage medium into a drive device. The storage unit 170 is, for example, an HDD, flash memory, or random access memory (RAM). The storage unit 170 stores, for example, captured image data 172, converted image data 174, annotation image data 176, annotated image data 178, and a trained model 180. For convenience of explanation, the image processing device 100 includes the trained model generation unit 160 and the storage unit 170 that stores the trained model 180. However, the function of generating a trained model and the generated trained model may be stored in a server device different from the image processing device 100.
[0030] The communication unit 110 is an interface that communicates with the communication device 10 of the vehicle M via the network NW. For example, the communication unit 110 includes a NIC (Network Interface Card), an antenna for wireless communication, and the like.
[0031] The transmission / reception control unit 120 transmits and receives data to and from the vehicles M1 and M2 and the terminal device 200 using the communication unit 110. More specifically, the transmission / reception control unit 120 first acquires from the vehicle M1 a plurality of interior and exterior images captured in time series by a camera mounted on the vehicle M1. The time series in this case refers to images captured at a predetermined interval (e.g., every second) during one driving cycle from when the vehicle M1 starts to when it stops.
[0032] FIG. 3 is a diagram showing an example of an interior image and an exterior image acquired from vehicle M1. The left part of FIG. 3 shows an interior image acquired from vehicle M1, and the right part of FIG. 3 shows an exterior image acquired from vehicle M1. As shown in the left part of FIG. 3, the interior image is captured with a camera installed so as to capture at least the facial area of the driver of vehicle M1, and as shown in the right part of FIG. 3, the exterior image is captured with a camera installed so as to capture at least the area ahead in the traveling direction of vehicle M1. The transmission / reception control unit 120 associates the interior image and the exterior image acquired from vehicle M1 with an image ID and stores them in the storage unit 170 as captured image data 172.
[0033] 4 is a diagram for explaining the processing executed by the image processing unit 130. The image processing unit 130 performs image processing on the captured image data 172, and acquires information such as image attributes, facial attributes, and orientation of each image included in the captured image data 172. More specifically, when an image is input, the image processing unit 130 acquires image attributes indicating whether each image included in the captured image data 172 is an image inside or outside the vehicle, using a trained model that outputs a classification result indicating whether the image is an image inside or outside the vehicle.
[0034] Furthermore, when an image is input, the image processing unit 130 acquires face attributes of each image included in the captured image data 172 using a trained model that outputs, for all faces included in the image, the face area, the size of the face (area of the face area), and the distance from the shooting position of the image to the face. In FIG. 3, as an example, a face area FA1 of person P1 is acquired from the inside-of-vehicle image, and a face area FA2 of person P2, a face area FA3 of person P3, and a face area FA4 of person P4 are acquired from the outside-of-vehicle image. For convenience, the face areas FA1, FA2, FA3, and FA4 are acquired as rectangular areas, but the present invention is not limited to such a configuration. For example, a trained model that acquires face areas along the contours of the faces of people may be used.
[0035] Furthermore, when an image is input, the image processing unit 130 acquires directional information of faces captured in each image included in the captured image data 172 using a trained model that outputs, for example, a vector, at least one of the face direction and the gaze direction for all faces included in the image. More specifically, for an image of the captured image data 172 having an attribute of an in-vehicle image, the image processing unit 130 acquires directional information using a trained model that outputs, when the image is input, the face direction and the gaze direction for all faces included in the image. On the other hand, for an image of the captured image data 172 having an attribute of an outside-vehicle image, the image processing unit 130 acquires directional information using a trained model that outputs, when the image is input, the face direction for all faces included in the image. This is because, compared to an outside-vehicle image, faces captured in an inside-vehicle image are generally closer to the shooting position and tend to be captured large enough to allow the gaze direction to be extracted. In Figure 3, as an example, the facial direction FD1 and gaze direction ED1 of person P1 are obtained from the in-vehicle image, and the facial direction FD2 of person P2, the facial direction FD3 of person P3, and the facial direction FD4 of person P4 are obtained from the outside-vehicle image.
[0036] When the image processing unit 130 acquires image attributes, facial attributes, and direction information for each image of the captured image data 172, it records the image attributes, facial attributes, and direction information in association with the image. Note that, as an example, in the above description, the image processing unit 130 acquires image attributes, facial attributes, and direction information using a trained model, but the present invention is not limited to such a configuration, and the image processing unit 130 may acquire the image attributes, facial attributes, and direction information using any known method.
[0037] The image conversion unit 140 executes a process for replacing the face of a person captured in each image with the face of another person using any software that implements such a function, without changing the direction information of the person, for the captured image data 172 processed by the image processing unit 130. FIG. 5 is a diagram for explaining the process executed by the image conversion unit 140. As shown in FIG. 5, the image conversion unit 140 replaces the faces of persons P1, P2, and P3 shown in FIG. 4 with the faces of other persons without changing the line of sight direction ED1 and facial directions FD1, FD2, and FD3. On the other hand, the face of person P4 is covered with a mosaic MS as a result of the mosaic process performed by the image conversion unit 140.
[0038] That is, the image conversion unit 140 determines whether to replace each face captured in each image of the captured image data 172 with the face of another person or to apply mosaic processing based on the facial attributes of the face. More specifically, for each face captured in each image of the captured image data 172, the image conversion unit 140 determines whether the size of the face is equal to or greater than a first threshold Th1, and if it is determined that the size of the face is equal to or greater than the first threshold Th1, it determines to replace the face with the face of another person. On the other hand, if it is determined that the size of the face is less than the first threshold Th1, the image conversion unit 140 determines to apply mosaic processing to the face. Replacing the face of a person captured in a captured image with the face of another person or applying mosaic processing is an example of an "anonymization process."
[0039] Furthermore, for each face captured in each image of the captured image data 172, the image conversion unit 140 determines whether the distance to the face is equal to or less than a second threshold Th2. If it is determined that the distance to the face is equal to or less than the second threshold Th2, the image conversion unit 140 determines to replace the face with the face of another person. On the other hand, if it is determined that the distance to the face is greater than the second threshold Th2, the image conversion unit 140 determines to apply mosaic processing to the face. The image conversion unit 140 repeatedly performs this determination process the number of times equal to the number of faces captured in the image, and, according to the determination result, replaces each face with the face of another person or applies mosaic processing to the face. The image conversion unit 140 stores image data obtained by performing such processing on the captured image data 172 in the storage unit 170 as converted image data 174. This allows for the selection of data useful as learning data for generating a behavior prediction model, and also allows for the privacy of people captured in each image to be protected when annotators, described later, perform annotation work.
[0040] At least one of the process of determining whether the face size is equal to or larger than the first threshold Th1 and the process of determining whether the face distance is equal to or smaller than the second threshold Th2 may be performed. When both processes are performed, the image conversion unit 140 may determine to replace the face with the face of another person when the face size is equal to or larger than the first threshold Th1 and the face distance is equal to or smaller than the second threshold Th2, or may determine to replace the face with the face of another person when the face size is equal to or larger than the first threshold Th1 or the face distance is equal to or smaller than the second threshold Th2.
[0041] Furthermore, the image conversion unit 140 may select faces to be used as learning data by performing mosaic processing on faces captured in each image of the captured image data 172 for which directional information has not been obtained.
[0042] FIG. 6 shows an example of time-series interior vehicle images converted by the image conversion unit 140. FIG. 6 shows an example of time-series interior vehicle images converted at three time points: t, t+1, and t+2. These time-series interior vehicle images were captured and facially converted from the same person. However, as shown in FIG. 6, depending on the operation of the facial conversion software, the same person's face may be converted into the faces of multiple different people. Even if the same person's face has been converted into the faces of multiple different people, using such converted image data as training data without modification can degrade the accuracy of the behavior prediction model, which is undesirable. Therefore, the image determination unit 150 performs the process described below to determine the continuity of the time-series interior vehicle images and exterior vehicle images.
[0043] FIG. 7 is a diagram illustrating the processing executed by the image determination unit 150. As shown in FIG. 7, the image determination unit 150 first extracts feature points representing a face of a person captured in a converted image. For example, the image determination unit 150 extracts feature points representing the right eye REP, the left eye LEP, the nose NP, the right corner of the mouth RMP, the left corner of the mouth LMP, and the ears EP from the face of the person captured in the converted image. The image determination unit 150 extracts feature points of the face of a person tracked as the same person from each of the converted images in the time series and compares these feature points. Whether or not the person has been "tracked as the same person" can be determined by, for example, associating the same person captured in the captured images before converting the image.
[0044] 7, the image determination unit 150 extracts feature points of a person captured in a converted image at time t and feature points of a person captured in a converted image at time t+1. The image determination unit 150 performs matching by determining whether or not these two sets of extracted feature points substantially match by translation and rotation.
[0045] If the matching result determines that the extracted feature points are substantially identical, the image determination unit 150 determines that the faces of the persons tracked as the same person are still the faces of the same person after the transformation (i.e., there is continuity in the faces). On the other hand, if the matching result determines that the extracted feature points are not substantially identical, the image determination unit 150 determines that the faces of the persons tracked as the same person are not the faces of the same person after the transformation (i.e., there is no continuity in the faces). In this case, the image conversion unit 140 performs a conversion process again on the faces determined to have no continuity. At this time, the image conversion unit 140 may perform a conversion process again only on the faces determined to have no continuity, or may perform a conversion process again on the faces of all persons captured in the time-series converted images. Furthermore, for example, the image conversion unit 140 may perform a mosaic process on the faces determined to have no continuity without performing a conversion process again, and exclude them from the targets to be used as learning data. Furthermore, for example, if the image determination unit 150 determines that the faces of people tracked as the same person are not the faces of the same person after conversion (i.e., there is no continuity in the faces), the image determination unit 150 may restrict the application of a predetermined process to the converted images in time series, i.e., exclude the converted images in time series from being used as learning data. This can prevent discontinuity from occurring due to unintended operations of the face conversion software.
[0046] The image determination unit 150 further inputs the converted image into the trained model that outputs at least one of the face direction and the gaze direction, and acquires the face direction FD or the gaze direction ED in the converted image. The image determination unit 150 determines whether the face direction FD or the gaze direction ED of a person's face captured in the converted image substantially matches the face direction FD or the gaze direction ED of a face captured in the captured image before conversion. As described above, both the face direction FD and the gaze direction ED are acquired for the in-vehicle image, and only the face direction FD is acquired for the outside-vehicle image. Therefore, for the in-vehicle image, the image determination unit 150 determines whether the face direction FD and the gaze direction ED substantially match between the captured image before conversion and the converted image, and for the outside-vehicle image, the image determination unit 150 determines whether the face direction FD and the gaze direction ED substantially match between the captured image before conversion and the converted image. More specifically, for example, the image determination unit 150 calculates the angle difference between a vector representing the face direction FD in the captured image before conversion and a vector representing the face direction FD in the converted image, and determines that the face directions FD approximately match if the calculated angle difference is within a threshold value. The same applies to the gaze direction ED. Satisfying the continuity of the face or the consistency of the directional information is an example of a "predetermined requirement."
[0047] If it is determined that the face direction FD or the gaze direction ED does not substantially match between the captured image before conversion and the converted image, the image conversion unit 140 performs conversion processing again on the captured image for the faces determined to have a face direction FD or gaze direction ED that does not substantially match. In this case, the image conversion unit 140 may perform conversion processing again only on the faces determined to have a substantially mismatch, or may perform conversion processing again on all faces included in the converted image, including the faces determined to have a substantially mismatch. For example, the image conversion unit 140 may perform mosaic processing on the faces determined to have a substantially mismatch without performing conversion processing again, and exclude them from use as learning data. For example, if the image determination unit 150 determines that the faces do not substantially match, the image determination unit 150 may restrict the application of a predetermined process to the time-series converted images, i.e., exclude the time-series converted images from use as learning data. This prevents information degradation due to unintended operation of the face conversion software.
[0048] Note that when there are multiple faces captured in the converted image (or when the number of faces captured in the converted image is equal to or greater than a predetermined value), the above-described determination process regarding the continuity of the converted image and the determination process regarding the consistency of the directional information performed by the image determination unit 150 may be performed only on faces assumed to be of higher importance, rather than on all faces captured in the converted image. As an example of a face assumed to be of higher importance, the image determination unit 150 may perform these determination processes only on faces in the captured image before conversion whose face size is equal to or greater than a third threshold Th3 that is greater than the first threshold Th1, or may perform these determination processes only on faces whose face distance is equal to or less than a fourth threshold Th4 that is smaller than the second threshold Th2. Furthermore, for example, the image determination unit 150 may perform these determination processes on faces of people present ahead in the traveling direction of the vehicle M1 or whose faces are facing ahead in the traveling direction of the vehicle M1 in the captured image before conversion, assuming that these faces are of higher importance. Furthermore, for example, if continuity or consistency is denied for a certain face captured in a converted image, reconversion processing may be performed for that face and faces that are considered to be of high importance.
[0049] When the image determination unit 150 confirms continuity and consistency of the time-series converted images, it stores the converted image data 174, for which continuity and consistency have been confirmed, in the storage unit 170 as annotation image data 176. At this time, the converted image data 174 may be stored in the storage unit 170 as annotation image data 176 together with information indicating the purpose of use, for example, information indicating that the converted image data 174 is annotation image data for generating a behavior prediction model for predicting the behavior of a person depicted in the input image. The transmission / reception control unit 120 transmits the annotation image data 176 to the terminal device 200. An annotator, who is a user of the terminal device 200, generates annotated image data by annotating the annotation image included in the received annotation image data 176 and transmits the annotated image data to the image processing device 100. The image processing device 100 stores the received annotated image data in the storage unit 170 as annotation image data 178.
[0050] It is sufficient that at least one of the determination process regarding the continuity of the converted image performed by the image determination unit 150 described above and the determination process regarding the consistency of the face direction information is performed, and if at least one of the continuity and consistency is established, the converted image data 174 may be stored in the memory unit 170 as image data for annotation 176.
[0051] Furthermore, for example, if there are missing images in the time series of captured images (or their converted images) obtained at a predetermined interval (e.g., every second) during one driving cycle due to a malfunction of the camera or the like, the image determination unit 150 does not need to store all of these time series images in the memory unit 170 as annotation image data 176.
[0052] FIG. 8 is a diagram showing an example of annotation work performed by an annotator. The left part of FIG. 8 shows annotations to a converted image of an interior image, and the right part of FIG. 8 shows annotations to a converted image of an exterior image. The annotator assigns information to the converted image of the interior image, indicating whether the driver's gaze direction ED1 shown in the converted image is appropriate for the situation shown in the converted image of the exterior image at the same time (e.g., 1 if appropriate, 0 if inappropriate). For example, in the case of FIG. 8, the converted image of the exterior image indicates the presence of a pedestrian on the left side of the vehicle's traveling direction, while the converted image of the interior image indicates that the driver is looking leftward. In other words, since it is assumed that the driver is paying appropriate attention to pedestrians, the annotator assigns information indicating that the driver's gaze direction ED1 is appropriate (i.e., 1).
[0053] Furthermore, the annotator specifies a risk area RA for the converted image of the vehicle exterior image, excluding, for example, people who have been pixelated. Because the image conversion unit 140 and the image determination unit 150 convert the face of the person in the original image into the face of another person, the privacy of that person is protected. At the same time, because the face direction and gaze direction of the person are maintained even after conversion, the annotator can accurately specify the risk area RA while referring to the face direction and gaze direction of the other person in the converted image. This makes it possible to generate learning data that is effective for training a machine learning model while protecting the privacy of the people in the face image.
[0054] When the annotated image data 178 is stored in the storage unit 170, the trained model generation unit 160 uses an arbitrary machine learning model and the annotated image data 178 as training data to generate a trained model. As described above, this trained model is, for example, a behavior prediction model that, when an exterior vehicle image is input, outputs the predicted behavior (trajectory) of a person depicted in the exterior vehicle image, or, when an interior vehicle image and an exterior vehicle image are input, takes into account the driver's line of sight depicted in the interior vehicle image to call attention to a pedestrian depicted in the exterior vehicle image. The trained model generation unit 160 stores the generated trained model in the storage unit 170 as trained model 180.
[0055] When the trained model 180 is generated, the transmission / reception control unit 120 distributes the generated trained model 180 to the vehicle M2 via the network NW. When the vehicle M2 receives the trained model 180, the vehicle M2 uses the trained model 180 (more precisely, an application program utilizing the trained model 180) to provide driving assistance to the driver of the vehicle M2.
[0056] FIG. 9 is a diagram illustrating an example of driving assistance using a trained model 180. FIG. 9 illustrates an example of driving assistance in which a vehicle M2 inputs interior and exterior images captured by a camera mounted thereon while traveling into the trained model 180, and the trained model 180 outputs information to an HMI (human machine interface) to alert the driver to a pedestrian captured in the exterior image, taking into account the driver's line of sight captured in the interior image. As shown in FIG. 9, for example, the HMI displays a risk area RA2 corresponding to a pedestrian P5 captured in the exterior image, and outputs a warning message ("Please be careful not to look away from the road") as text or audio information if the driver's line of sight captured in the interior image is not directed toward the pedestrian P5. This allows for driving assistance that takes into account the driver's state.
[0057] Next, the flow of processing executed by the image processing device 100 will be described with reference to Fig. 10 and Fig. 11. Fig. 10 is a diagram showing an example of the flow of processing executed by the image conversion unit 140. The processing shown in Fig. 10 is executed, for example, when an inside image or an outside image of the vehicle M1 is captured by a camera mounted on the vehicle M1 and processed by the image processing unit 130.
[0058] First, the image conversion unit 140 acquires a captured image included in the captured image data 172 that has been processed by the image processing unit 130 (step S100). Next, the image conversion unit 140 selects one face shown in the acquired captured image (step S102).
[0059] Next, the image conversion unit 140 determines whether the size of the selected face is equal to or larger than a first threshold Th1 (step S104). If it is determined that the size of the selected face is equal to or larger than the first threshold Th1, the image conversion unit 140 converts the face into the face of another person (step S106). On the other hand, if it is determined that the size of the selected face is smaller than the first threshold Th1, the image conversion unit 140 next determines whether the distance of the selected face is equal to or smaller than a second threshold Th2 (step S108).
[0060] If it is determined that the distance of the selected face is equal to or less than the second threshold Th2, the image conversion unit 140 proceeds to step S106 and converts the face into the face of another person. On the other hand, if it is determined that the distance of the selected face is greater than the second threshold Th2, the image conversion unit 140 applies mosaic processing to the face (step S110). Next, the image conversion unit 140 determines whether or not the processing has been performed on all faces appearing in the acquired captured image (step S112).
[0061] If it is determined that the processing has been performed on all faces appearing in the acquired captured image, the image conversion unit 140 acquires the image obtained by performing the processing on all faces as a converted image and stores it in the storage unit 170 as converted image data 174 (step S114). On the other hand, if it is determined that the processing has not been performed on all faces appearing in the acquired captured image, the image conversion unit 140 returns the processing to step S102. This ends the processing of this flowchart.
[0062] Fig. 11 is a diagram showing an example of the flow of processing executed by the image determination unit 150. The processing shown in Fig. 11 is executed, for example, at the timing when a time-series converted image is obtained by applying the above-described conversion processing to a time-series captured image taken in one driving cycle from the start to the stop of the vehicle M1.
[0063] First, the image determination unit 150 acquires time-series converted images (step S200). Next, the image determination unit 150 selects, from the acquired time-series converted images, the faces of people who were tracked as the same person before conversion (step S202).
[0064] Next, the image determination unit 150 extracts feature points from the faces of persons tracked as the same person before conversion from each of the converted images in the time series, and by performing matching, determines whether these faces are the same after conversion (step S204). If it is determined that the faces are the same after conversion, the image determination unit 150 then determines whether the acquired converted images in the time series are images inside a vehicle (step S206). On the other hand, if it is determined that the faces are not the same, the image determination unit 150 causes the image conversion unit 140 to again convert the faces of persons tracked as the same person in the captured images in the time series before conversion (step S208). Thereafter, the image determination unit 150 again performs the process of step S204 on the converted faces.
[0065] If it is determined in step S206 that the acquired time-series converted images are vehicle interior images, the image determination unit 150 determines whether the gaze direction and facial direction of these faces match those of the images before conversion (step S210). On the other hand, if it is determined that the acquired time-series converted images are not vehicle interior images, that is, are vehicle exterior images, the image determination unit 150 determines whether the facial direction of these faces match those of the images before conversion (step S212). If it is determined that they do not match in the processing of step S210 or step S212, the image determination unit 150 proceeds to the processing of step S208.
[0066] If it is determined in the processing of step S210 or step S212 that there is a match, the image determination unit 150 determines that these faces have been converted normally, and determines whether or not the processing has been performed on all faces captured in the time-series converted images (step S214). If it is determined that the processing has been performed on all faces captured in the time-series converted images, the image determination unit 150 acquires these time-series converted images as images for annotation, and causes the transmission / reception control unit 120 to transmit the acquired images for annotation to the terminal device 200 (step S216). On the other hand, if it is determined that the processing has not been performed on all faces captured in the time-series converted images, the image determination unit 150 returns the processing to step S202. This ends the processing of this flowchart.
[0067] According to the present embodiment described above, if it is determined that a plurality of input images that have undergone anonymization processing satisfy predetermined requirements, a predetermined process is performed on the plurality of input images that have undergone anonymization processing. The predetermined process includes changing the faces of persons depicted in the plurality of input images to the faces of different persons. The predetermined requirements include ensuring that the faces of persons who were tracked as the same person and depicted in the plurality of input images that have undergone anonymization processing are the faces of the same person obtained by the anonymization processing. In other words, in this embodiment, faces that belonged to the same person before the anonymization processing are guaranteed to be the faces of the same person even after the anonymization processing, and are used as training data. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of persons depicted in the face images.
[0068] Furthermore, according to this embodiment, the predetermined requirement includes that face orientation information of a person tracked as the same person in multiple input images matches face orientation information of the same person in the multiple input images that have been anonymized. That is, this embodiment ensures that face orientation information of the same person remains unchanged even after anonymization. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of people depicted in face images.
[0069] Furthermore, according to this embodiment, the predetermined requirements are determined according to image attributes, which are the shooting conditions of the multiple input images. That is, in this embodiment, a predetermined process is executed, taking into account the shooting conditions of each of the multiple input images, for example, to store the input images as learning information for generating a behavior prediction model. This makes it possible to generate learning data that is effective for training a machine learning model while protecting the privacy of people depicted in the face images.
[0070] Furthermore, according to this embodiment, it is determined whether to perform anonymization processing using a first method or a second method different from the first method based on the size of the face in each of a plurality of input images or the distance from the shooting point to the face. That is, in this embodiment, the method of anonymization processing to be performed on the face is changed depending on whether it is useful for training a machine learning model. This makes it possible to generate training data that is effective for training a machine learning model while protecting the privacy of people depicted in the face images.
[0071] [Variations] As described above, in this embodiment, an example has been described in which, when the image determination unit 150 determines that a face depicted in a converted image does not satisfy predetermined requirements, the converted image is reconverted or subjected to mosaic processing. However, when the image determination unit 150 determines that the predetermined requirements are not satisfied, the image determination unit 150 may not perform predetermined processing on the converted image, that is, may limit the application of predetermined processing (such as not storing the image or not transmitting it to a server).
[0072] Furthermore, in the present embodiment, an example has been described in which the image processing device 100 is implemented as a server device separate from the vehicle M1. However, as a modification of the present embodiment, the image processing device 100, more specifically, a device having at least the functions of the image processing unit 130, the image conversion unit 140, and the image determination unit 150, may be mounted on the vehicle M1 as an in-vehicle device. In this case, the in-vehicle device performs processing by the image processing unit 130 described above on an image captured by the in-vehicle camera, anonymization by the image conversion unit 140, and determination by the image determination unit 150. Thereafter, the in-vehicle device transmits the anonymized image, the facial continuity and consistency of directional information of which have been confirmed by the image determination unit 150, to an external image server.
[0073] When the image server receives an anonymized image from the vehicle M1, it stores the received anonymized image in a storage unit as annotation image data, and either transmits the annotation image data to the annotator's terminal device 200 or allows the terminal device 200 to access the annotation image data. When the image server receives annotated image data from the terminal device 200, it generates a trained model 180 based on the annotated image data and distributes the generated trained model 180 to the vehicle M2. This method, like the present embodiment, can generate training data effective for training a machine learning model while protecting the privacy of persons depicted in facial images. Furthermore, according to this modification, the in-vehicle device performs an anonymization process on the image before transmitting the anonymized image to the image server, thereby further reliably protecting the privacy of persons depicted in facial images.
[0074] Furthermore, as another aspect, the in-vehicle device may have only some of the functions of the image processing unit 130, the image conversion unit 140, and the image determination unit 150, and the image server may have the remaining functions. For example, the in-vehicle device may have the functions of the image processing unit 130 and the image conversion unit 140, and the image server may have the function of the image determination unit 150, or the in-vehicle device may have the function of the image processing unit 130, and the image server may have the functions of the image conversion unit 140 and the image determination unit 150.
[0075] The above-described embodiment can be expressed as follows. a storage medium for storing computer-readable instructions; a processor connected to the storage medium; The processor executes the computer-readable instructions to: Acquire image attributes, which are the shooting manner of the input image; performing an anonymization process on the input image; determining whether the input image that has been anonymized satisfies predetermined requirements; If it is determined that the input image that has been anonymized satisfies the predetermined requirement, a predetermined process is performed on the input image that has been anonymized; The image processing device is configured so that the predetermined requirement is determined according to the acquired image attribute.
[0076] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]
[0077] 100 Image processing device 110 Communications Department 120 Transmission and reception control section 130 Image processing section 140 Image conversion unit 150 Image Judgment Unit 160 Trained model generation unit 170 Storage section 172 Captured image data 174 Converted Image Data 176 Image data for annotation 178 Annotated Image Data 180 trained models
Claims
1. an image attribute acquisition unit that acquires image attributes, which are the shooting manner of an input image; an image conversion unit that performs an anonymization process on the input image; an image determination unit that determines whether the input image that has been anonymized satisfies predetermined requirements, When the image determination unit determines that the input image that has been anonymized satisfies the predetermined requirement, the image determination unit performs a predetermined process on the input image that has been anonymized, the predetermined requirement is determined in accordance with the acquired image attribute, The predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as learning information for generating a machine learning model. Image processing device.
2. the predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as a target image for annotation work. The image processing device according to claim 1 .
3. the predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as learning information for generating a behavior prediction model that predicts the behavior of a person depicted in the input image. The image processing device according to claim 1 .
4. the predetermined processing is processing of transmitting the input image that has been subjected to the anonymization processing to an image server via a communication means; The image processing device according to claim 1 .
5. The image attribute is information indicating at least whether the input image is an image captured inside a vehicle equipped with a camera that captured the input image, or an image captured outside the vehicle. The image processing device according to claim 1 .
6. the anonymization process includes a process of changing a face of a person shown in the input image to a face of another person, The predetermined requirement includes whether or not a gaze direction of the face of the person coincides with a gaze direction of the face of the other person when the image attribute indicates that the image is an image of the interior of the vehicle. The image processing device according to claim 5 .
7. the predetermined requirements include, when the image attribute indicates that the image is an image of the interior of the vehicle, whether or not a gaze direction of the face of the person and a gaze direction of the face of the other person match, and whether or not a facial direction of the face of the person and a facial direction of the face of the other person match, The predetermined requirement includes whether or not a facial direction of the face of the person matches a facial direction of the face of the other person when the image attribute indicates that the image is an image of the outside of the vehicle. The image processing device according to claim 6 .
8. The predetermined requirement does not include whether the gaze direction of the face of the person coincides with the gaze direction of the face of the other person when the image attribute indicates that the image is an image of the outside of the vehicle. The image processing device according to claim 6 .
9. When the image determination unit determines that the input image to which the anonymization process has been applied does not satisfy the predetermined requirement, the image conversion unit applies the anonymization process to the input image again. The image processing device according to claim 1 .
10. When the image determination unit determines that the input image that has been anonymized does not satisfy the predetermined requirement, the image conversion unit does not perform the predetermined process on the input image that has been anonymized. The image processing device according to claim 1 .
11. an image attribute acquisition unit that acquires image attributes, which are the shooting manner of an input image; an image conversion unit that performs an anonymization process on the input image; an image determination unit that determines whether the input image that has been anonymized satisfies predetermined requirements, When the image determination unit determines that the input image that has been anonymized satisfies the predetermined requirement, the image determination unit performs a predetermined process on the input image that has been anonymized, the predetermined requirement is determined in accordance with the acquired image attribute, The predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as learning information for generating a machine learning model. Image processing system.
12. The computer Acquire image attributes, which are the shooting manner of the input image; performing an anonymization process on the input image; determining whether the input image that has been anonymized satisfies predetermined requirements; If it is determined that the input image that has been anonymized satisfies the predetermined requirement, a predetermined process is performed on the input image that has been anonymized; the predetermined requirement is determined in accordance with the acquired image attribute, The predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as learning information for generating a machine learning model. Image processing methods.
13. On the computer, Acquire image attributes, which are the shooting manner of the input image; performing an anonymization process on the input image; determining whether the input image that has been anonymized satisfies predetermined requirements; If it is determined that the input image that has been anonymized satisfies the predetermined requirement, a predetermined process is performed on the input image that has been anonymized; the predetermined requirement is determined in accordance with the acquired image attribute, The predetermined processing is processing of storing the input image that has been subjected to the anonymization processing as learning information for generating a machine learning model. program.
Citation Information
Patent Citations
Method for producing base material for additive for casting of steel
JP1984030450A
Information processing device and program
JP2014085796A
Image processing system, information processing device, and program
JP2017187850A
Image processor and method for processing image
JP2020061081A
Age privacy protection method and system for face recognition
JP2020170496A